Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T20:07:03.928169Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 1 inbound Pith citation observation for arXiv:2508.11252.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T20:07:03.928169Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-09T12:24:33.998226Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-09T12:26:16.710584Z
62 of 62 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 46e93cf4-032b-4c97-b851-e926fb5e3dd1 · outbound
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information van Eemeren, Rob Grootendorst, Sally Jackson, Scott Jacobs, Agnes van Rees, Francisca Snoeck Henkemans, Eveline T
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a76876a4-59e5-4133-aee0-4ae45c21033f · outbound
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Dictionary of philosophy
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 370b6277-7626-4223-a4a4-8ce3fdf716b4 · outbound
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information McCarthy and P.J
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6283a7de-d9c5-4213-92df-a28911622977 · outbound
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Programs with common sense, 1959
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b63c295f-3367-4fc6-8bdf-0d54e7e06fae · outbound
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Newell and H
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7bd09b23-7d48-4ca8-99bf-aa9a9404c831 · outbound
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information OpenAI o1 System Card
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98a098ea-9d90-4b3d-b421-de3ddcb3e9bc · outbound
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e1b6ea8-3138-4a37-98bf-503e9607f5b3 · outbound
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information 2024 aime i
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7a53f98a-62bf-43f9-9a55-9e3c0cc643d0 · outbound
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Let's verify step by step
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation eba0e665-bb2d-4ed3-903f-c5887baf8f86 · outbound
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Omni- MATH : A universal olympiad level mathematic benchmark for large language models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6cf53a6a-4452-4e61-bc37-4641b64d79ce · outbound
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Open r1: A fully open reproduction of deepseek-r1, January 2025
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df75b3d2-5714-4907-8b0e-f846dab3f8c1 · outbound
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Clamber: A benchmark of identifying and clarifying ambiguous information needs in large language models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 386b49d3-e878-442f-84cd-25f37252c31a · outbound
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Rethinking conversational agents in the era of llms: Proactivity, non-collaborativity, and beyond
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fd0d4cff-1480-45c6-b34f-09197e4f2f9a · outbound
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Li, Been Kim
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d640403b-1be5-469e-ae88-de6dbf05c71f · outbound
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Reasoning attack: Inducing llm to never-end thinking, 2025
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7919619a-e362-4901-8965-1a66dc01ce8c · outbound
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Training language models to follow instructions with human feedback
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8ae85c5a-de73-456d-9582-33f2c95668a5 · outbound
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information s1: Simple test-time scaling
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 812c533b-10ff-4c3b-aeff-60ba40608fcd · outbound
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Bespoke-stratos Labs
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0855451a-dbc2-4725-a890-92e160cdfdac · outbound
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Training Verifiers to Solve Math Word Problems
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 925a3393-d92f-4266-af04-2f787ab03547 · outbound
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e340aae-809c-4196-9127-57d204cb4abe · outbound
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Sky-t1: Train your own o1 preview model within \ 450
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1543a353-e86a-4853-b45d-d8927b37cbc1 · outbound
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information RedStar: Does Scaling Long-CoT Data Unlock Better Slow-Reasoning Systems?
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 683f945d-43e1-4346-ae69-9d6eb0b96bd3 · outbound
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Qwen3: Think deeper, act faster
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b74a40ff-1242-4935-853b-dae7802d9374 · outbound
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Claude 3.7 sonnet and claude code
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 717194d9-c392-4cae-a616-9d30993f1c48 · outbound
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Grok 3 beta — the age of reasoning agents
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation feb82631-4813-4c5b-9148-414021f9a9b6 · outbound
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information A Survey on LLM-as-a-Judge
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fde7b267-3a0b-473f-a97e-4fb6b5eca28a · outbound
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information The Lessons of Developing Process Reward Models in Mathematical Reasoning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82397fd9-be3f-4701-be79-d949ae8adcb3 · outbound
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35fa5a0e-da24-46bb-8174-0f646ff6a458 · outbound
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bba2630-f867-4ec9-8945-ab843d9d2d02 · outbound
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f0d8dbf4-632b-46fa-ac4f-dcb6c23728ad · outbound
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Distilling step-by-step! outperforming larger language models with less training data and smaller model sizes
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 12ee1f7d-a8bb-4fec-b038-eca99d9085ee · outbound
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Distilling the Knowledge in a Neural Network
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83059a89-9761-47af-b6aa-e0b209a60029 · outbound
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Math-Verify: Math Verification Library
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 40f39d92-e928-4cdb-8c44-b759ec7efc31 · outbound
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1136c703-50c5-466f-b7ff-b7b6b98446e0 · outbound
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Interference and inhibition in cognition and behavior: Unifying themes for educational psychology
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3bc7e2fb-86f4-4f89-83d1-2ae4594a183b · outbound
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information How students “unpack” the structure of a word problem: Graphic representations and problem solving
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation dac55333-c023-47e4-ad17-4d65f1ebac92 · outbound
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8d60990c-3ea2-4a7e-8f74-550313053840 · outbound
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Metacognition: A literature review
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cca27954-eebb-4e8a-a7d4-d7c8fe423332 · outbound
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Strategies for improving learner metacognition in health professional education
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ef5d890e-0c43-4392-8c98-7ca9a231aada · outbound
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62942cf8-05e8-44b3-9002-0c9956cc902c · outbound
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information A survey of deep active learning
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9bf6058f-c77c-4194-a102-3fd4beb95319 · outbound
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Deep bayesian active learning with image data
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ad39eb8c-7773-4d34-98e3-5954f177a0a4 · outbound
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Reinforcement learning: An introduction , volume 1
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c52c2a08-ece9-4395-8a33-e350216a0313 · outbound
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Partially Observable Task and Motion Planning with Uncertainty and Risk Awareness
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7e5facde-f8c8-4345-b778-7db9d575b6f1 · outbound
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Combined task and motion planning under partial observability: An optimization-based approach
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0b1c7b4b-5a4b-45aa-b921-d6890b2a258b · outbound
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information The communicative function of ambiguity in language
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9e8c9df5-a049-484b-b03c-fb9c513cac9d · outbound
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information The puzzle of ambiguity
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 43c56b16-8780-4f11-87bb-f1f8ff5ba0ad · outbound
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Semantic ambiguity within and across languages: An integrative review
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 996c406a-4b06-4770-86ac-2fc167ea365f · outbound
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information What computers can’t do: The limits of artificial intelligence
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e0f6b72d-cb2b-44a1-a86b-39d5e21a3c38 · outbound
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Foundational Challenges in Assuring Alignment and Safety of Large Language Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bce34754-5435-49c7-85d1-d7eeed121372 · outbound
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information How to Enable Effective Cooperation Between Humans and NLP Models: A Survey of Principles, Formalizations, and Beyond
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 21861de3-46da-4924-a50a-962d68c5d7e7 · outbound
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5f92a559-b612-4dcf-bfbe-2adab9850b45 · outbound
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information AmbigQA: Answering Ambiguous Open-domain Questions
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38ffefd5-7b95-41c5-ab5a-a093a90e3186 · outbound
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information ChatShop: Interactive Information Seeking with Language Agents
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 101e6db0-1c75-4f15-974e-b079df9cc900 · outbound
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Style: Improving domain transferability of asking clarification questions in large language model powered conversational agents
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f12f2614-8b02-43aa-92f8-4ceff6d62045 · outbound
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Prompting and evaluating large language models for proactive dialogues: Clarification, target-guided, and non-collaboration
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6f8c9949-2158-486e-a38a-f15f06fbbb6c · outbound
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information CLAM: Selective Clarification for Ambiguous Questions with Generative Language Models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 558401d8-0e9d-4913-820d-01f876df472e · outbound
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Selectively answering ambiguous questions
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7805dadc-569b-4275-8b6a-582e6b7a74dc · outbound
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Multiwoz-a large-scale multi-domain wizard-of-oz dataset for task-oriented dialogue modelling
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4a57f24e-63db-4e72-9e29-27f8ff31b604 · outbound
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Towards scalable multi-domain conversational agents: The schema-guided dialogue dataset
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9ac5de61-2202-48e8-8946-673f04381adc · outbound
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information We need to consider disagreement in evaluation
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9b8e2b2c-3d43-4e59-8cac-9a8b70c433bb · outbound
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Everyone’s voice matters: Quantifying annotation disagreement using demographic information
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation aa4ce1f3-3f5a-44b5-9695-194bfd0ae40b · inbound
MIRA-Math: A Benchmark for Minimal Information Requesting and Mathematical Reasoning Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.