Pith. sign in

REVIEW 1 cited by

Strong Heuristics for Named Entity Linking

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2207.02824 v1 pith:QO7LCW4C submitted 2022-07-06 cs.CL cs.IRcs.LG

classification cs.CLcs.IRcs.LG
keywords entityheuristicslinkingmethodsunsupervisedzero-shotemergingentities
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Named entity linking (NEL) in news is a challenging endeavour due to the frequency of unseen and emerging entities, which necessitates the use of unsupervised or zero-shot methods. However, such methods tend to come with caveats, such as no integration of suitable knowledge bases (like Wikidata) for emerging entities, a lack of scalability, and poor interpretability. Here, we consider person disambiguation in Quotebank, a massive corpus of speaker-attributed quotations from the news, and investigate the suitability of intuitive, lightweight, and scalable heuristics for NEL in web-scale corpora. Our best performing heuristic disambiguates 94% and 63% of the mentions on Quotebank and the AIDA-CoNLL benchmark, respectively. Additionally, the proposed heuristics compare favourably to the state-of-the-art unsupervised and zero-shot methods, Eigenthemes and mGENRE, respectively, thereby serving as strong baselines for unsupervised and zero-shot entity linking.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Weakly Supervised Medical Entity Extraction and Linking for Chief Complaints

    cs.CL 2025-09 conditional novelty 5.0 of 10

    A split-and-match weak supervision pipeline trains BERT and BiLSTM models to extract and link medical entities from chief complaints without human annotation, achieving 67.5 F1 on a clinician-labeled test set.

Pith tools