5 months ago

Optimizing Bi-Encoder for Named Entity Recognition via Contrastive Learning

Sheng Zhang; Hao Cheng; Jianfeng Gao; Hoifung Poon

Abstract

We present a bi-encoder framework for named entity recognition (NER), which applies contrastive learning to map candidate text spans and entity types into the same vector representation space. Prior work predominantly approaches NER as sequence labeling or span classification. We instead frame NER as a representation learning problem that maximizes the similarity between the vector representations of an entity mention and its type. This makes it easy to handle nested and flat NER alike, and can better leverage noisy self-supervision signals. A major challenge to this bi-encoder formulation for NER lies in separating non-entity spans from entity mentions. Instead of explicitly labeling all non-entity spans as the same class $\texttt{Outside}$ ($\texttt{O}$) as in most prior methods, we introduce a novel dynamic thresholding loss. Experiments show that our method performs well in both supervised and distantly supervised settings, for nested and flat NER alike, establishing new state of the art across standard datasets in the general domain (e.g., ACE2004, ACE2005) and high-value verticals such as biomedicine (e.g., GENIA, NCBI, BC5CDR, JNLPBA). We release the code at github.com/microsoft/binder.

Code Repositories

microsoft/binder

Official

pytorch

Benchmarks

Benchmark	Methodology	Metrics
named-entity-recognition-ner-on-bc5cdr	BINDER	F1: 91.9
named-entity-recognition-ner-on-jnlpba	BINDER	F1: 80.3
nested-named-entity-recognition-on-ace-2004	BINDER	F1: 88.7
nested-named-entity-recognition-on-ace-2005	BINDER	F1: 89.5
nested-named-entity-recognition-on-genia	BINDER	F1: 80.5

Build AI with AI

From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.

AI Co-coding

Ready-to-use GPUs

Best Pricing

Get Started

Hyper Newsletters

Subscribe to our latest updates

We will deliver the latest updates of the week to your inbox at nine o'clock every Monday morning

Command Palette