Pith. sign in

REVIEW 4 cited by

Understanding Dataset Difficulty with $\mathcal{V}$-Usable Information

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2110.08420 v3 pith:NALMLT5X submitted 2021-10-16 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords datasetmathcaldifferentdifficultygiveninformationdifficultmodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Estimating the difficulty of a dataset typically involves comparing state-of-the-art models to humans; the bigger the performance gap, the harder the dataset is said to be. However, this comparison provides little understanding of how difficult each instance in a given distribution is, or what attributes make the dataset difficult for a given model. To address these questions, we frame dataset difficulty -- w.r.t. a model $\mathcal{V}$ -- as the lack of $\mathcal{V}$-$\textit{usable information}$ (Xu et al., 2019), where a lower value indicates a more difficult dataset for $\mathcal{V}$. We further introduce $\textit{pointwise $\mathcal{V}$-information}$ (PVI) for measuring the difficulty of individual instances w.r.t. a given distribution. While standard evaluation metrics typically only compare different models for the same dataset, $\mathcal{V}$-$\textit{usable information}$ and PVI also permit the converse: for a given model $\mathcal{V}$, we can compare different datasets, as well as different instances/slices of the same dataset. Furthermore, our framework allows for the interpretability of different input attributes via transformations of the input, which we use to discover annotation artefacts in widely-used NLP benchmarks.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Rationales Are Not Silver Bullets: Measuring the Impact of Rationales on Model Performance and Reliability

    cs.CL 2025-05 conditional novelty 7.0 of 10

    Rationale-augmented finetuning can hurt accuracy while improving calibration, with the sizes of both effects tied linearly to task difficulty.

  2. LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing

    cs.CL 2025-07 conditional novelty 6.0 of 10

    Small reward models trained on LitBench reach 78% agreement with upvote-derived human preferences in creative writing, beating all zero-shot LLM judges tested.

  3. VisNec: Measuring and Leveraging Visual Necessity for Multimodal Instruction Tuning

    cs.CV 2026-03 conditional novelty 5.0 of 10

    Selecting instruction-tuning samples by the loss difference between text-only and multimodal prediction (VisNec) lets a model match or exceed full-data performance with only 15% of the data.

  4. Efficient and Stealthy Jailbreak Attacks via Adversarial Prompt Distillation from LLMs to SLMs

    cs.CL 2025-05 reject novelty 5.0 of 10

    Distilling a large model's jailbreak-prompt skill into BERT-scale models reportedly yields high attack success at lower compute, but the paper's inconsistent results make the claim unverified.

Pith tools