Pith. sign in

REVIEW 1 cited by

Learning to Recognize Dialect Features

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2010.12707 v3 pith:CG54W5SS submitted 2020-10-23 cs.CL

classification cs.CL
keywords dialectfeaturesdialectsbuildingdetectionfeaturelearningminimal
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Building NLP systems that serve everyone requires accounting for dialect differences. But dialects are not monolithic entities: rather, distinctions between and within dialects are captured by the presence, absence, and frequency of dozens of dialect features in speech and text, such as the deletion of the copula in "He {} running". In this paper, we introduce the task of dialect feature detection, and present two multitask learning approaches, both based on pretrained transformers. For most dialects, large-scale annotated corpora for these features are unavailable, making it difficult to train recognizers. We train our models on a small number of minimal pairs, building on how linguists typically define dialect features. Evaluation on a test set of 22 dialect features of Indian English demonstrates that these models learn to recognize many features with high accuracy, and that a few minimal pairs can be as effective for training as thousands of labeled examples. We also demonstrate the downstream applicability of dialect feature detection both as a measure of dialect density and as a dialect classifier.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LEAP: LLM-powered End-to-end Automatic Library for Processing Social Science Queries on Unstructured Data

    cs.DB 2025-01 conditional novelty 6.0 of 10

    LEAP, an LLM-based library, automatically selects ML functions and writes SQL-like code to answer 92% of 120 social science queries over unstructured data on the first attempt, and 100% within three attempts.

Pith tools