Pith. sign in

REVIEW 10 cited by

Dynabench: Rethinking Benchmarking in NLP

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2104.14337 v1 pith:5JAIX3ZM submitted 2021-04-07 cs.CL cs.AI

classification cs.CLcs.AI
keywords dynabenchmodelbenchmarkingcreationdatasetdynamicexamplesplatform
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce Dynabench, an open-source platform for dynamic dataset creation and model benchmarking. Dynabench runs in a web browser and supports human-and-model-in-the-loop dataset creation: annotators seek to create examples that a target model will misclassify, but that another person will not. In this paper, we argue that Dynabench addresses a critical need in our community: contemporary models quickly achieve outstanding performance on benchmark tasks but nonetheless fail on simple challenge examples and falter in real-world scenarios. With Dynabench, dataset creation, model development, and model assessment can directly inform each other, leading to more robust and informative benchmarks. We report on four initial NLP tasks, illustrating these concepts and highlighting the promise of the platform, and address potential objections to dynamic benchmarking as a new standard for the field.

Discussion (0). Sign in to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EpiQAL: Benchmarking Large Language Models in Epidemiological Question Answering and Reasoning

    cs.CL 2026-01 conditional novelty 7.0 of 10

    EpiQAL, a literature-grounded benchmark with three reasoning levels, shows current LLMs handle factual recall well but are substantially weaker at multi-step epidemiological inference.

  2. Whose View of Safety? A Deep DIVE Dataset for Pluralistic Alignment of Text-to-Image Models

    cs.LG 2025-07 conditional novelty 7.0 of 10

    A demographically diverse annotation dataset shows that safety perceptions for text-to-image outputs vary by rater identity and that conventional safety classifiers under-detect bias harms flagged by minority-group raters.

  3. CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal Tasks

    cs.LG 2026-08 conditional novelty 6.0 of 10

    CalibForge generates terminal-agent training tasks by revising candidates until a strong solver passes and a weak solver fails, and students trained on the resulting 5,431 tasks gain up to 30 points on held-out benchmarks.

  4. Potemkin Understanding in Large Language Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    LLMs frequently pass definition questions yet fail to use the same concepts in classification, generation, and editing tasks, a gap the authors call potemkin understanding.

  5. LLM Performance for Code Generation on Noisy Tasks

    cs.LG 2025-05 conditional novelty 6.0 of 10

    LLMs solve heavily obfuscated benchmark tasks, and performance decay under obfuscation differs sharply between old and new datasets, which the authors interpret as a signature of training-data contamination.

  6. Benchmarks Are Not Monolithic: Sample-Level Auditing and Orchestration for LLM Evaluation

    cs.CL 2026-07 reject novelty 5.0 of 10

    Benchmark items are heterogeneous; the authors annotate them with three LLM judges and use the labels to build filtered subsets, but the validation evidence is weak and the comparison method is flawed.

  7. Too long; didn't solve

    cs.AI 2026-04 unverdicted novelty 5.0 of 10

    Prompt length and solution length both rise with LLM failure on expert-authored adversarial math problems, linking structural length to empirical difficulty.

  8. Private, Verifiable, and Auditable AI Systems

    cs.CR 2025-08 conditional novelty 4.0 of 10

    A thesis demonstrating partial prototypes for zk-verifiable model evaluation and privacy-preserving retrieval, and arguing these pieces can compose into end-to-end auditable AI systems.

  9. Improving LLM Safety and Helpfulness using SFT and DPO: A Study on OPT-350M

    cs.CL 2025-09 reject novelty 3.0 of 10

    For OPT-350M on the Anthropic HH-RLHF set, SFT plus DPO gives the highest combined helpfulness/harmlessness score, but not the highest safety score.

  10. Agentic Web: Weaving the Next Web with AI Agents

    cs.AI 2025-07 conditional novelty 3.0 of 10

    A position paper defines the Agentic Web as the next web era and proposes a three-dimensional conceptual framework for understanding and building it.

Pith tools