Pith. sign in

REVIEW 4 cited by

SpanBERT: Improving Pre-training by Representing and Predicting Spans

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1907.10529 v3 pith:XTYH5XQ6 submitted 2019-07-24 cs.CL cs.LG

classification cs.CLcs.LG
keywords spanspanbertspansbertcoreferencegainsmodelpre-training
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present SpanBERT, a pre-training method that is designed to better represent and predict spans of text. Our approach extends BERT by (1) masking contiguous random spans, rather than random tokens, and (2) training the span boundary representations to predict the entire content of the masked span, without relying on the individual token representations within it. SpanBERT consistently outperforms BERT and our better-tuned baselines, with substantial gains on span selection tasks such as question answering and coreference resolution. In particular, with the same training data and model size as BERT-large, our single model obtains 94.6% and 88.7% F1 on SQuAD 1.1 and 2.0, respectively. We also achieve a new state of the art on the OntoNotes coreference resolution task (79.6\% F1), strong performance on the TACRED relation extraction benchmark, and even show gains on GLUE.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Structure-Aware Fill-in-the-Middle Pretraining for Code

    cs.CL 2025-05 conditional novelty 7.0 of 10

    AST-FIM masks complete syntax-tree subtrees during fill-in-the-middle pretraining, improving infilling performance on real-world code edits.

  2. Making FETCH! Happen: Finding Emergent Dog Whistles Through Common Habitats

    cs.CL 2024-12 conditional novelty 5.0 of 10

    A new benchmark shows existing NLP systems find almost no novel dog whistles in social media, while the proposed EarShot pipeline raises F0.5 scores to 14.6 on synthetic Reddit, 5.7 on Gab, and 4.6 on Twitter.

  3. Exploring Long-Term Prediction of Type 2 Diabetes Microvascular Complications

    cs.LG 2024-12 conditional novelty 5.0 of 10

    A code-agnostic text representation of EHRs outperformed code-based input for predicting retinopathy, nephropathy and neuropathy at 1, 5, and 10 years in 133,784 UK type 2 diabetes patients.

  4. Masked Diffusion Language Models with Frequency-Informed Training

    cs.CL 2025-09 conditional novelty 4.0 of 10

    Masked diffusion language models trained on 100M words match a hybrid GPT-BERT baseline on BabyLM tests, with a rare-word-focused masking variant.

Pith tools