Pith. sign in

REVIEW 1 cited by

Learning to Model and Ignore Dataset Bias with Mixed Capacity Ensembles

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2011.03856 v1 pith:AERT3XDB submitted 2020-11-07 cs.LG cs.CLcs.CV

classification cs.LGcs.CLcs.CV
keywords capacitydatasetsmodelcorrelationsdatasetbiasmethodanswering
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Many datasets have been shown to contain incidental correlations created by idiosyncrasies in the data collection process. For example, sentence entailment datasets can have spurious word-class correlations if nearly all contradiction sentences contain the word "not", and image recognition datasets can have tell-tale object-background correlations if dogs are always indoors. In this paper, we propose a method that can automatically detect and ignore these kinds of dataset-specific patterns, which we call dataset biases. Our method trains a lower capacity model in an ensemble with a higher capacity model. During training, the lower capacity model learns to capture relatively shallow correlations, which we hypothesize are likely to reflect dataset bias. This frees the higher capacity model to focus on patterns that should generalize better. We ensure the models learn non-overlapping approaches by introducing a novel method to make them conditionally independent. Importantly, our approach does not require the bias to be known in advance. We evaluate performance on synthetic datasets, and four datasets built to penalize models that exploit known biases on textual entailment, visual question answering, and image recognition tasks. We show improvement in all settings, including a 10 point gain on the visual question answering dataset.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Multimodal Political Bias Identification and Neutralization

    cs.CY 2025-06 unverdicted novelty 4.0 of 10

    A proposed multimodal pipeline to identify and reduce political bias in news text and images remains unvalidated: the report presents architecture and qualitative examples, without quantitative results for most components.

Pith tools