Pith. sign in

REVIEW 1 cited by

Identifying Spurious Correlations for Robust Text Classification

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2010.02458 v1 pith:Z4OKGH66 submitted 2020-10-06 cs.LG cs.CLcs.IR

classification cs.LGcs.CLcs.IR
keywords classificationcorrelationsspurioustextapproachdistinguishevenfeatures
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The predictions of text classifiers are often driven by spurious correlations -- e.g., the term `Spielberg' correlates with positively reviewed movies, even though the term itself does not semantically convey a positive sentiment. In this paper, we propose a method to distinguish spurious and genuine correlations in text classification. We treat this as a supervised classification problem, using features derived from treatment effect estimators to distinguish spurious correlations from "genuine" ones. Due to the generic nature of these features and their small dimensionality, we find that the approach works well even with limited training examples, and that it is possible to transport the word classifier to new domains. Experiments on four datasets (sentiment classification and toxicity detection) suggest that using this approach to inform feature selection also leads to more robust classification, as measured by improved worst-case accuracy on the samples affected by spurious correlations.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CURE: Controlled Unlearning for Robust Embeddings -- Mitigating Conceptual Shortcuts in Pre-Trained Language Models

    cs.CL 2025-09 conditional novelty 5.0 of 10

    CURE removes concept-level spurious correlations from pre-trained embeddings via a content extractor, a reversal network, and a margin-controlled contrastive module, improving OOD sentiment F1 by up to 10 points on IM...

Pith tools