REVIEW 4 cited by
CommonsenseQA 2.0: Exposing the Limits of AI through Gamification
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Constructing benchmarks that test the abilities of modern natural language understanding models is difficult - pre-trained language models exploit artifacts in benchmarks to achieve human parity, but still fail on adversarial examples and make errors that demonstrate a lack of common sense. In this work, we propose gamification as a framework for data construction. The goal of players in the game is to compose questions that mislead a rival AI while using specific phrases for extra points. The game environment leads to enhanced user engagement and simultaneously gives the game designer control over the collected data, allowing us to collect high-quality data at scale. Using our method we create CommonsenseQA 2.0, which includes 14,343 yes/no questions, and demonstrate its difficulty for models that are orders-of-magnitude larger than the AI used in the game itself. Our best baseline, the T5-based Unicorn with 11B parameters achieves an accuracy of 70.2%, substantially higher than GPT-3 (52.9%) in a few-shot inference setup. Both score well below human performance which is at 94.1%.
Forward citations
Cited by 4 Pith papers
-
When Incentives Backfire, Data Stops Being Human
Incentive-driven crowdwork erodes intrinsic motivation and data quality, so data collection should be redesigned around intrinsic motivation, with games as a promising template.
-
Debate-on-Graph: Reliable and Adaptive Reasoning of Large Language Model on Uncertain Knowledge Graph
Confidence-aware subgraph retrieval plus asymmetric multi-agent debate improves LLM QA accuracy on four benchmarks using uncertain knowledge graphs.
-
A Statistical Physics of Language Model Reasoning
A switching linear dynamical system on a 40-dimensional projection of LLM hidden states captures about half the variance of reasoning trajectories and predicts belief shifts during adversarial prompts.
-
Training-free Truthfulness Detection via Sparse MLP Value Vectors
TruthV detects true answers by majority-voting the argmax/argmin preferences of a sparse set of MLP value vectors selected on 30 labeled examples.
Discussion (0). Continue with ORCID to comment.