REVIEW 6 cited by
Dataset Cartography: Mapping and Diagnosing Datasets with Training Dynamics
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large datasets have become commonplace in NLP research. However, the increased emphasis on data quantity has made it challenging to assess the quality of data. We introduce Data Maps---a model-based tool to characterize and diagnose datasets. We leverage a largely ignored source of information: the behavior of the model on individual instances during training (training dynamics) for building data maps. This yields two intuitive measures for each example---the model's confidence in the true class, and the variability of this confidence across epochs---obtained in a single run of training. Experiments across four datasets show that these model-dependent measures reveal three distinct regions in the data map, each with pronounced characteristics. First, our data maps show the presence of "ambiguous" regions with respect to the model, which contribute the most towards out-of-distribution generalization. Second, the most populous regions in the data are "easy to learn" for the model, and play an important role in model optimization. Finally, data maps uncover a region with instances that the model finds "hard to learn"; these often correspond to labeling errors. Our results indicate that a shift in focus from quantity to quality of data could lead to robust models and improved out-of-distribution generalization.
Forward citations
Cited by 6 Pith papers
-
The methodology of Constructing the Large-Scale Dataset for Detecting Presuicidal and Anti-Suicidal Signals in Social Media Texts in Russian
A new methodology and a 57,810-example Russian social media dataset with 33 presuicidal and 12 anti-suicidal classes, plus baseline RuBERT experiments.
-
CDC: Causal Domain Clustering for Multi-Domain Recommendation
CDC clusters domains by combining isolated and interactive transfer-effect measurements, weighted by a causal-distance-based cohesion coefficient, and jointly optimizes target clusters and source training sets.
-
Foundation Model Insights and a Multi-Model Approach for Superior Fine-Grained One-shot Subset Selection
RAM-APL combines distance rankings and pseudo-class label accuracy from two foundation models to select training subsets, outperforming twelve baselines on fine-grained image datasets.
-
Principled Data Selection for Alignment: The Hidden Risks of Difficult Examples
Preference examples vary in difficulty; overly difficult examples degrade DPO alignment, and filtering them out improves AlpacaEval 2 win rates by 9-16 percentage points.
-
Uncertainty-Driven Reliability: Selective Prediction and Trustworthy Deployment in Modern Machine Learning
A training-dynamics abstention method matches deep ensembles at a fraction of the training cost, and a five-term error budget explains why selective classifiers still fall short of the oracle.
-
Beyond Explainable AI (XAI): An Overdue Paradigm Shift and Post-XAI Research Directions
Current XAI methods for DNNs and LLMs rest on paradoxes and false assumptions that demand a paradigm shift to verification protocols, scientific foundations, context-aware design, and faithful model analysis rather th...
Discussion (0). Continue with ORCID to comment.