Pith. sign in

REVIEW 4 cited by

CommonsenseQA 2.0: Exposing the Limits of AI through Gamification

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2201.05320 v1 pith:AHQTW5TX submitted 2022-01-14 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords gamedatamodelsbenchmarkscommonsenseqademonstrategamificationhuman
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Constructing benchmarks that test the abilities of modern natural language understanding models is difficult - pre-trained language models exploit artifacts in benchmarks to achieve human parity, but still fail on adversarial examples and make errors that demonstrate a lack of common sense. In this work, we propose gamification as a framework for data construction. The goal of players in the game is to compose questions that mislead a rival AI while using specific phrases for extra points. The game environment leads to enhanced user engagement and simultaneously gives the game designer control over the collected data, allowing us to collect high-quality data at scale. Using our method we create CommonsenseQA 2.0, which includes 14,343 yes/no questions, and demonstrate its difficulty for models that are orders-of-magnitude larger than the AI used in the game itself. Our best baseline, the T5-based Unicorn with 11B parameters achieves an accuracy of 70.2%, substantially higher than GPT-3 (52.9%) in a few-shot inference setup. Both score well below human performance which is at 94.1%.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. When Incentives Backfire, Data Stops Being Human

    cs.CY 2025-02 conditional novelty 6.0 of 10

    Incentive-driven crowdwork erodes intrinsic motivation and data quality, so data collection should be redesigned around intrinsic motivation, with games as a promising template.

  2. Debate-on-Graph: Reliable and Adaptive Reasoning of Large Language Model on Uncertain Knowledge Graph

    cs.CL 2026-07 conditional novelty 5.0 of 10

    Confidence-aware subgraph retrieval plus asymmetric multi-agent debate improves LLM QA accuracy on four benchmarks using uncertain knowledge graphs.

  3. A Statistical Physics of Language Model Reasoning

    cs.AI 2025-06 conditional novelty 5.0 of 10

    A switching linear dynamical system on a 40-dimensional projection of LLM hidden states captures about half the variance of reasoning trajectories and predicts belief shifts during adversarial prompts.

  4. Training-free Truthfulness Detection via Sparse MLP Value Vectors

    cs.CL 2025-09 conditional novelty 4.0 of 10

    TruthV detects true answers by majority-voting the argmax/argmin preferences of a sparse set of MLP value vectors selected on 30 labeled examples.

Pith tools