REVIEW 10 cited by
Dynabench: Rethinking Benchmarking in NLP
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We introduce Dynabench, an open-source platform for dynamic dataset creation and model benchmarking. Dynabench runs in a web browser and supports human-and-model-in-the-loop dataset creation: annotators seek to create examples that a target model will misclassify, but that another person will not. In this paper, we argue that Dynabench addresses a critical need in our community: contemporary models quickly achieve outstanding performance on benchmark tasks but nonetheless fail on simple challenge examples and falter in real-world scenarios. With Dynabench, dataset creation, model development, and model assessment can directly inform each other, leading to more robust and informative benchmarks. We report on four initial NLP tasks, illustrating these concepts and highlighting the promise of the platform, and address potential objections to dynamic benchmarking as a new standard for the field.
Forward citations
Cited by 10 Pith papers
-
EpiQAL: Benchmarking Large Language Models in Epidemiological Question Answering and Reasoning
EpiQAL, a literature-grounded benchmark with three reasoning levels, shows current LLMs handle factual recall well but are substantially weaker at multi-step epidemiological inference.
-
Whose View of Safety? A Deep DIVE Dataset for Pluralistic Alignment of Text-to-Image Models
A demographically diverse annotation dataset shows that safety perceptions for text-to-image outputs vary by rater identity and that conventional safety classifiers under-detect bias harms flagged by minority-group raters.
-
CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal Tasks
CalibForge generates terminal-agent training tasks by revising candidates until a strong solver passes and a weak solver fails, and students trained on the resulting 5,431 tasks gain up to 30 points on held-out benchmarks.
-
Potemkin Understanding in Large Language Models
LLMs frequently pass definition questions yet fail to use the same concepts in classification, generation, and editing tasks, a gap the authors call potemkin understanding.
-
LLM Performance for Code Generation on Noisy Tasks
LLMs solve heavily obfuscated benchmark tasks, and performance decay under obfuscation differs sharply between old and new datasets, which the authors interpret as a signature of training-data contamination.
-
Benchmarks Are Not Monolithic: Sample-Level Auditing and Orchestration for LLM Evaluation
Benchmark items are heterogeneous; the authors annotate them with three LLM judges and use the labels to build filtered subsets, but the validation evidence is weak and the comparison method is flawed.
-
Too long; didn't solve
Prompt length and solution length both rise with LLM failure on expert-authored adversarial math problems, linking structural length to empirical difficulty.
-
Private, Verifiable, and Auditable AI Systems
A thesis demonstrating partial prototypes for zk-verifiable model evaluation and privacy-preserving retrieval, and arguing these pieces can compose into end-to-end auditable AI systems.
-
Improving LLM Safety and Helpfulness using SFT and DPO: A Study on OPT-350M
For OPT-350M on the Anthropic HH-RLHF set, SFT plus DPO gives the highest combined helpfulness/harmlessness score, but not the highest safety score.
-
Agentic Web: Weaving the Next Web with AI Agents
A position paper defines the Agentic Web as the next web era and proposes a three-dimensional conceptual framework for understanding and building it.
Discussion (0). Sign in to comment.