Pith. sign in

REVIEW 1 cited by

Language ID in the Wild: Unexpected Challenges on the Path to a Thousand-Language Web Text Corpus

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2010.14571 v2 pith:OIDFK2EE submitted 2020-10-27 cs.CL cs.LG

classification cs.CLcs.LG
keywords langidlanguagelanguagesmodelstextcorporacorpusmany
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large text corpora are increasingly important for a wide variety of Natural Language Processing (NLP) tasks, and automatic language identification (LangID) is a core technology needed to collect such datasets in a multilingual context. LangID is largely treated as solved in the literature, with models reported that achieve over 90% average F1 on as many as 1,366 languages. We train LangID models on up to 1,629 languages with comparable quality on held-out test sets, but find that human-judged LangID accuracy for web-crawl text corpora created using these models is only around 5% for many lower-resource languages, suggesting a need for more robust evaluation. Further analysis revealed a variety of error modes, arising from domain mismatch, class imbalance, language similarity, and insufficiently expressive models. We propose two classes of techniques to mitigate these errors: wordlist-based tunable-precision filters (for which we release curated lists in about 500 languages) and transformer-based semi-supervised LangID models, which increase median dataset precision from 5.5% to 71.2%. These techniques enable us to create an initial data set covering 100K or more relatively clean sentences in each of 500+ languages, paving the way towards a 1,000-language web text corpus.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. On the use of Performer and Agent Attention for Spoken Language Identification

    eess.AS 2025-02 conditional novelty 4.0 of 10

    Replacing standard self-attention with performer attention in the pooling layer of a language identification model improves average accuracy on VoxPopuli, FLEURS, and VoxLingua, while agent attention is comparable and...

Pith tools