Pith. sign in

REVIEW 4 cited by

Think you have Solved Direct-Answer Question Answering? Try ARC-DA, the Direct-Answer AI2 Reasoning Challenge

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2102.03315 v1 pith:AZ55QUY7 submitted 2021-02-05 cs.CL cs.AI

classification cs.CLcs.AI
keywords questionsarc-dadatasetdirect-answerreasoninganswersappropriatechallenge
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present the ARC-DA dataset, a direct-answer ("open response", "freeform") version of the ARC (AI2 Reasoning Challenge) multiple-choice dataset. While ARC has been influential in the community, its multiple-choice format is unrepresentative of real-world questions, and multiple choice formats can be particularly susceptible to artifacts. The ARC-DA dataset addresses these concerns by converting questions to direct-answer format using a combination of crowdsourcing and expert review. The resulting dataset contains 2985 questions with a total of 8436 valid answers (questions typically have more than one valid answer). ARC-DA is one of the first DA datasets of natural questions that often require reasoning, and where appropriate question decompositions are not evident from the questions themselves. We describe the conversion approach taken, appropriate evaluation metrics, and several strong models. Although high, the best scores (81% GENIE, 61.4% F1, 63.2% ROUGE-L) still leave considerable room for improvement. In addition, the dataset provides a natural setting for new research on explanation, as many questions require reasoning to construct answers. We hope the dataset spurs further advances in complex question-answering by the community. ARC-DA is available at https://allenai.org/data/arc-da

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SymStep: Symbolic Step Verification for Logical Reasoning

    cs.AI 2026-07 conditional novelty 6.0 of 10

    SymStep couples atomic LLM deductions to a deterministic constraint propagator with MRV hints, reaching ~97–100% on constraint-dense logic puzzles where CoT scores 0%.

  2. Toward Preference-aligned Large Language Models via Residual-based Model Steering

    cs.CL 2025-09 conditional novelty 5.0 of 10

    Preference signals in LLM residual streams can be distilled into inference-time steering vectors that improve math and code benchmarks using only 100 preference pairs.

  3. Beyond Manually Designed Pruning Policies with Second-Level Performance Prediction: A Pruning Framework for LLMs

    cs.LG 2025-08 conditional novelty 5.0 of 10

    A predictor-based agent generates LLM pruning policies in seconds and reports large perplexity reductions on Llama2-7B and Llama3-8B.

  4. WSM: Decay-Free Learning Rate Schedule via Checkpoint Merging for LLM Pre-training

    cs.CL 2025-07 conditional novelty 5.0 of 10

    Checkpoint merging during constant-LR training can replace LR decay and yields improved LLM benchmark scores over Warmup-Stable-Decay.

Pith tools