REVIEW 6 cited by
Cosmos QA: Machine Reading Comprehension with Contextual Commonsense Reasoning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Understanding narratives requires reading between the lines, which in turn, requires interpreting the likely causes and effects of events, even when they are not mentioned explicitly. In this paper, we introduce Cosmos QA, a large-scale dataset of 35,600 problems that require commonsense-based reading comprehension, formulated as multiple-choice questions. In stark contrast to most existing reading comprehension datasets where the questions focus on factual and literal understanding of the context paragraph, our dataset focuses on reading between the lines over a diverse collection of people's everyday narratives, asking such questions as "what might be the possible reason of ...?", or "what would have happened if ..." that require reasoning beyond the exact text spans in the context. To establish baseline performances on Cosmos QA, we experiment with several state-of-the-art neural architectures for reading comprehension, and also propose a new architecture that improves over the competitive baselines. Experimental results demonstrate a significant gap between machine (68.4%) and human performance (94%), pointing to avenues for future research on commonsense machine comprehension. Dataset, code and leaderboard is publicly available at https://wilburone.github.io/cosmos.
Forward citations
Cited by 6 Pith papers
-
Representations Shape Weak-to-Strong Generalization: Theoretical Insights and Empirical Predictions
Weak-to-strong performance is governed by the overlap between the weak model's unlearnable error space and the strong model's principal-representation space, quantified by ||P_s(I-P_w)||.
-
Relating Misfit to Gain in Weak-to-Strong Generalization Beyond the Squared Loss
For convex and approximately convex model classes, the loss gain in weak-to-strong learning is at least the KL misfit between strong and weak models, plus an error term that vanishes as k grows.
-
Debate Helps Weak-to-Strong Generalization
Debate transcripts from two strong models, used as context when training an ensemble of weak models, improve weak-to-strong generalization on four NLP classification benchmarks.
-
Learning Conformal Abstention Policies for Adaptive Risk Management in Large Language and Vision-Language Models
CAP tunes conformal thresholds with RL to switch between single answers, sets, and abstention, but its test-set-fitting undermines the claimed statistical guarantees.
-
Options-Aware Dense Retrieval for Multiple-Choice query Answering
Fine-tuning a sentence transformer so query-plus-options embeddings mimic oracle query-plus-answer embeddings improves evidence retrieval and multiple-choice accuracy on QuALITY.
-
Embodied Spatial Intelligence: from Implicit Scene Modeling to Spatial Reasoning
The thesis demonstrates that combining implicit 3D scene representations with LLM-based reasoning, using text as an interface, yields strong performance on robotic perception and spatial language tasks.
Discussion (0). Continue with ORCID to comment.