Pith. sign in

Paper Citation Record · LEDGER

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information

As of 8 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 1 inbound Pith citation observation for arXiv:2508.11252.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.11252 v1

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T20:07:03.928169Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-09T12:24:33.998226Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T12:26:16.710584Z

Reference resolution

62 of 62 outbound references displayed

  • verified exact2
  • verified fuzzy40
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 46e93cf4-032b-4c97-b851-e926fb5e3dd1 · outbound

This paper cites van Eemeren, Rob Grootendorst, Sally Jackson, Scott Jacobs, Agnes van Rees, Francisca Snoeck Henkemans, Eveline T.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information van Eemeren, Rob Grootendorst, Sally Jackson, Scott Jacobs, Agnes van Rees, Francisca Snoeck Henkemans, Eveline T

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:12.728559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:06:58.190648Z digest=sha256:8f8e8d9bb07c76579238306303b1d269e0d1cc16fdb219584ed39f1d02ccb749

Observation a76876a4-59e5-4133-aee0-4ae45c21033f · outbound

This paper cites Dictionary of philosophy.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Dictionary of philosophy

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:12.581522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:06:58.245217Z digest=sha256:a5d05725f2bae7ea7157ae0e74222b454f19cac320c6e5728921b8a90f8c8df3

Observation 370b6277-7626-4223-a4a4-8ce3fdf716b4 · outbound

This paper cites McCarthy and P.J.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information McCarthy and P.J

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:12.484770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:06:58.394610Z digest=sha256:9b83a3232ff79f8e3d785a36cd630e7e9645031536ede32fea5ebef76039b4ab

Observation 6283a7de-d9c5-4213-92df-a28911622977 · outbound

This paper cites Programs with common sense, 1959.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Programs with common sense, 1959

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:12.356361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:06:58.500892Z digest=sha256:ad123b39fa646ae3651a9c74c0c0d805ebe3d2dfee6218ad5c0f1926fe3c2f4e

Observation b63c295f-3367-4fc6-8bdf-0d54e7e06fae · outbound

This paper cites Newell and H.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Newell and H

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:12.262508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:06:58.644951Z digest=sha256:a9b17f215c41dcc1fbabbe96343bae95e5bb26b1e86cae7bc61bb61408fe2a0c

Observation 7bd09b23-7d48-4ca8-99bf-aa9a9404c831 · outbound

This paper cites OpenAI o1 System Card.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information OpenAI o1 System Card

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T20:06:58.787169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:06:58.787169Z digest=sha256:b9cee03e9c978ba9c47c9f559853c160bf6f57810d8c5074328a61fc91430ed0

Observation 98a098ea-9d90-4b3d-b421-de3ddcb3e9bc · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T20:06:58.934923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:06:58.934923Z digest=sha256:489c9a15d44c66cb215fa24852c639caa3a55dd58a3dd67f204bc63ba5daff53

Observation 7e1b6ea8-3138-4a37-98bf-503e9607f5b3 · outbound

This paper cites 2024 aime i.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information 2024 aime i

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:12.165657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:06:59.038089Z digest=sha256:dc714d673643354dd97ab1e5c169a227291585731fbc507e55eb72089eb89d2c

Observation 7a53f98a-62bf-43f9-9a55-9e3c0cc643d0 · outbound

This paper cites Let's verify step by step.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Let's verify step by step

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:12.035357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:06:59.117727Z digest=sha256:b7621182105b71a8db6f8d487c3666af4d604150d3e906f9ec182d7c75873d3a

Observation eba0e665-bb2d-4ed3-903f-c5887baf8f86 · outbound

This paper cites Omni- MATH : A universal olympiad level mathematic benchmark for large language models.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Omni- MATH : A universal olympiad level mathematic benchmark for large language models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:11.949130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:06:59.214909Z digest=sha256:20d470fcc6aa5e3b47f1fe5c891f47ffdab04ca64f278eb2191254bc16326095

Observation 6cf53a6a-4452-4e61-bc37-4641b64d79ce · outbound

This paper cites Open r1: A fully open reproduction of deepseek-r1, January 2025.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Open r1: A fully open reproduction of deepseek-r1, January 2025

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T20:06:59.358545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:06:59.358545Z digest=sha256:fcba9f443c74fc3736810a5bb9b2c0bca307332f00a69ab3c41d4e01567b0da3

Observation df75b3d2-5714-4907-8b0e-f846dab3f8c1 · outbound

This paper cites Clamber: A benchmark of identifying and clarifying ambiguous information needs in large language models.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Clamber: A benchmark of identifying and clarifying ambiguous information needs in large language models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:11.794651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:06:59.465558Z digest=sha256:f3270969c3bd8f8ffb183b6cb0549cd2a34e25826220a800e2e84202d5b23d99

Observation 386b49d3-e878-442f-84cd-25f37252c31a · outbound

This paper cites Rethinking conversational agents in the era of llms: Proactivity, non-collaborativity, and beyond.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Rethinking conversational agents in the era of llms: Proactivity, non-collaborativity, and beyond

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:11.649126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:06:59.574990Z digest=sha256:ef1a925716ba119639ff78bcc7bb63782b2e256187ceada49d66de73df0ee52b

Observation fd0d4cff-1480-45c6-b34f-09197e4f2f9a · outbound

This paper cites Li, Been Kim.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Li, Been Kim

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T20:06:59.670963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:06:59.670963Z digest=sha256:4aeccfe44c645a697052768fc35351070ec0ca993902fc70f77479ed70baf0c6

Observation d640403b-1be5-469e-ae88-de6dbf05c71f · outbound

This paper cites Reasoning attack: Inducing llm to never-end thinking, 2025.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Reasoning attack: Inducing llm to never-end thinking, 2025

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:11.542959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:06:59.756678Z digest=sha256:222b6e9b27d3d0f4374609094ccbd0b775670428d12abcae21e4d4af37bd604e

Observation 7919619a-e362-4901-8965-1a66dc01ce8c · outbound

This paper cites Training language models to follow instructions with human feedback.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Training language models to follow instructions with human feedback

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:11.417146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:06:59.850459Z digest=sha256:8ed7beb2e9ad39f0dbeea889f4559d5761cc49fde0efdff9865616f72d712ff4

Observation 8ae85c5a-de73-456d-9582-33f2c95668a5 · outbound

This paper cites s1: Simple test-time scaling.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information s1: Simple test-time scaling

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T20:06:59.954946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:06:59.954946Z digest=sha256:2df1be87559797cfd33adce9186d784063f678afe3097e115b031c5a44b8e762

Observation 812c533b-10ff-4c3b-aeff-60ba40608fcd · outbound

This paper cites Bespoke-stratos Labs.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Bespoke-stratos Labs

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:11.112782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:07:00.050629Z digest=sha256:1b381c36fc8d2dc93215a28927ce45fd10bf69d173189bfa627492030c729b03

Observation 0855451a-dbc2-4725-a890-92e160cdfdac · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Training Verifiers to Solve Math Word Problems

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T20:07:00.139772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:07:00.139772Z digest=sha256:f75d0abbccb0a2ec390aa76d2f37f0d66a2d381cf0d064ef950eba9c55e74a60

Observation 925a3393-d92f-4266-af04-2f787ab03547 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T20:07:00.246609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:07:00.246609Z digest=sha256:bd40dbf17b66661222cbaf2edf3b0612cac7575a6eb40ac25ddadd1cf5b10775

Observation 2e340aae-809c-4196-9127-57d204cb4abe · outbound

This paper cites Sky-t1: Train your own o1 preview model within \ 450.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Sky-t1: Train your own o1 preview model within \ 450

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:10.674559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:07:00.355608Z digest=sha256:09feedbd61e281700dfd451b0cb795f2654bb3623850df3da248a26a51883566

Observation 1543a353-e86a-4853-b45d-d8927b37cbc1 · outbound

This paper cites RedStar: Does Scaling Long-CoT Data Unlock Better Slow-Reasoning Systems?.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information RedStar: Does Scaling Long-CoT Data Unlock Better Slow-Reasoning Systems?

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T20:07:00.485136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:07:00.485136Z digest=sha256:ef606feb7a9ce4977ebe5edd34c40a9e389014c61b01bbe097d03bd21c5fa2cd

Observation 683f945d-43e1-4346-ae69-9d6eb0b96bd3 · outbound

This paper cites Qwen3: Think deeper, act faster.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Qwen3: Think deeper, act faster

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:10.365162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:07:00.591595Z digest=sha256:fd00f47fd4854beba8c2d4288845a820f3decd873be21eb7844326cce8eb5ab2

Observation b74a40ff-1242-4935-853b-dae7802d9374 · outbound

This paper cites Claude 3.7 sonnet and claude code.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Claude 3.7 sonnet and claude code

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:10.221580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:07:00.699335Z digest=sha256:19deb25a4fd7132ece621a1f4b1ee29dad03ec04c536bb28262531c26d43cadc

Observation 717194d9-c392-4cae-a616-9d30993f1c48 · outbound

This paper cites Grok 3 beta — the age of reasoning agents.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Grok 3 beta — the age of reasoning agents

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:10.033377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:07:00.778952Z digest=sha256:b689543142342974075298b8bde22112d09db8546b6d7aa6cd00539fbf6b5dd2

Observation feb82631-4813-4c5b-9148-414021f9a9b6 · outbound

This paper cites A Survey on LLM-as-a-Judge.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information A Survey on LLM-as-a-Judge

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T20:07:00.871004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:07:00.871004Z digest=sha256:180dfb2a5d325e617c24c0b49fe829839e8f80c1ce1f1e64600ec19a25e942ab

Observation fde7b267-3a0b-473f-a97e-4fb6b5eca28a · outbound

This paper cites The Lessons of Developing Process Reward Models in Mathematical Reasoning.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information The Lessons of Developing Process Reward Models in Mathematical Reasoning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T20:07:00.983996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:07:00.983996Z digest=sha256:5850d4ba05e9c8cd4710c355b4f9f218ff1f87e1df4ff8d3fe4a5923e53b1d2a

Observation 82397fd9-be3f-4701-be79-d949ae8adcb3 · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T20:07:01.076649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:07:01.076649Z digest=sha256:f48895edaab8a5ed1f9e714f89846feefff47d8f1ef060b472c1e97a1586b9cd

Observation 35fa5a0e-da24-46bb-8174-0f646ff6a458 · outbound

This paper cites Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T20:07:01.215034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:07:01.215034Z digest=sha256:813cd0c97261b1bd323e328ef1884bda420ff30cb84943f3b17f5415e2178acf

Observation 3bba2630-f867-4ec9-8945-ab843d9d2d02 · outbound

This paper cites A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:09.852083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:07:01.280601Z digest=sha256:b15cf725b03a654e3d7fd6484543a1af4260a33093bf0f280b182aa16c43e2e6

Observation f0d8dbf4-632b-46fa-ac4f-dcb6c23728ad · outbound

This paper cites Distilling step-by-step! outperforming larger language models with less training data and smaller model sizes.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Distilling step-by-step! outperforming larger language models with less training data and smaller model sizes

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:09.715437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:07:01.363095Z digest=sha256:00c44a4079ef8370eeae09189b3d3b1d3a2a491ed58518856ba532876d7e9dc3

Observation 12ee1f7d-a8bb-4fec-b038-eca99d9085ee · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Distilling the Knowledge in a Neural Network

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T20:07:01.454580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:07:01.454580Z digest=sha256:5d24c16d8323c650c7dbdf2b1f874670d4470863bbcd9a6b5a2d1bdc4805645b

Observation 83059a89-9761-47af-b6aa-e0b209a60029 · outbound

This paper cites Math-Verify: Math Verification Library.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Math-Verify: Math Verification Library

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:09.408545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:07:01.521708Z digest=sha256:2d96c0070647a67694ff15110aabb1c832f4b9f3e3be6170542704c13093ebd7

Observation 40f39d92-e928-4cdb-8c44-b759ec7efc31 · outbound

This paper cites OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T20:07:01.606002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:07:01.606002Z digest=sha256:3826b8c8a149b548d9074ab5a33ec3b93f7a590f61f9bc597ba3abe61eb54ae0

Observation 1136c703-50c5-466f-b7ff-b7b6b98446e0 · outbound

This paper cites Interference and inhibition in cognition and behavior: Unifying themes for educational psychology.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Interference and inhibition in cognition and behavior: Unifying themes for educational psychology

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:09.160064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:07:01.695399Z digest=sha256:74b6e83904726541cb499333906d98938622a3c1d9cf25e1b682c7e749825085

Observation 3bc7e2fb-86f4-4f89-83d1-2ae4594a183b · outbound

This paper cites How students “unpack” the structure of a word problem: Graphic representations and problem solving.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information How students “unpack” the structure of a word problem: Graphic representations and problem solving

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:08.776804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:07:01.788385Z digest=sha256:5d2775634ea73cef1f40d4467761ee7862578a4ad945c1715416183d3da00e5d

Observation dac55333-c023-47e4-ad17-4d65f1ebac92 · outbound

This paper cites an unresolved cited work.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-05T20:07:08.606236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:07:01.852303Z digest=sha256:17af81e82ecdc2b21893ceff1d00a94dd5af70e7b02b41dfd7d77e7e16f69bfc

Observation 8d60990c-3ea2-4a7e-8f74-550313053840 · outbound

This paper cites Metacognition: A literature review.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Metacognition: A literature review

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:08.406532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:07:01.936919Z digest=sha256:3ef8b3843e82f36ffd7ac41228ecd05cc71e125fd419956d734261622d576be6

Observation cca27954-eebb-4e8a-a7d4-d7c8fe423332 · outbound

This paper cites Strategies for improving learner metacognition in health professional education.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Strategies for improving learner metacognition in health professional education

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:08.210450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:07:01.995855Z digest=sha256:a95ed79b9f63e6f46e261ee29e193a1fe4d0343c9ac89ad9ddac2c2b75849b11

Observation ef5d890e-0c43-4392-8c98-7ca9a231aada · outbound

This paper cites Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T20:07:02.080403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:07:02.080403Z digest=sha256:0a0141c1544ef19e22fd4890755a222e26ddf2d976f4f99c4f170bd5030af6a3

Observation 62942cf8-05e8-44b3-9002-0c9956cc902c · outbound

This paper cites A survey of deep active learning.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information A survey of deep active learning

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:07.968999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:07:02.145598Z digest=sha256:2169e947794e14c790421f1edb75afd370fd88ec688af524fa63fc066e3147ef

Observation 9bf6058f-c77c-4194-a102-3fd4beb95319 · outbound

This paper cites Deep bayesian active learning with image data.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Deep bayesian active learning with image data

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:07.796796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:07:02.234196Z digest=sha256:c8e5cf76ba258086d7ec6a9e6263b42b572612bea04d707ad139e62138c1a1bc

Observation ad39eb8c-7773-4d34-98e3-5954f177a0a4 · outbound

This paper cites Reinforcement learning: An introduction , volume 1.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Reinforcement learning: An introduction , volume 1

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:07.562472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:07:02.310110Z digest=sha256:e6e0e6f8cf0c69a4c9e831649b0470d5e2ac9019e4e653b624ddb7360684b8a6

Observation c52c2a08-ece9-4395-8a33-e350216a0313 · outbound

This paper cites Partially Observable Task and Motion Planning with Uncertainty and Risk Awareness.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Partially Observable Task and Motion Planning with Uncertainty and Risk Awareness

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-08-05T20:07:04.307527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:07:02.384922Z digest=sha256:f03bcd3b12c06510a7b128ce432b6c325d972eb964f5e19b523f9d99441b026e

Observation 7e5facde-f8c8-4345-b778-7db9d575b6f1 · outbound

This paper cites Combined task and motion planning under partial observability: An optimization-based approach.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Combined task and motion planning under partial observability: An optimization-based approach

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:07.331095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:07:02.475798Z digest=sha256:f3306e097179ecd8a5d1bcf6386f8bc742d6cba30c70099596fd8ff4e4f49376

Observation 0b1c7b4b-5a4b-45aa-b921-d6890b2a258b · outbound

This paper cites The communicative function of ambiguity in language.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information The communicative function of ambiguity in language

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:07.103378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:07:02.560443Z digest=sha256:e9d00e42ed34bcc47417983730ca716f9283c206921586a93078cd1c5541dfc0

Observation 9e8c9df5-a049-484b-b03c-fb9c513cac9d · outbound

This paper cites The puzzle of ambiguity.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information The puzzle of ambiguity

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:06.910037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:07:02.645566Z digest=sha256:0da5fd14255ed0c17c0680d5efef1ebf20b3497bfdb3cc60a3d9d1f36cae47f9

Observation 43c56b16-8780-4f11-87bb-f1f8ff5ba0ad · outbound

This paper cites Semantic ambiguity within and across languages: An integrative review.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Semantic ambiguity within and across languages: An integrative review

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:06.715372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:07:02.733138Z digest=sha256:b71fd6eb2e31ca7c28bb0952a933a871fb52522deee8f8ada1da65d1952de286

Observation 996c406a-4b06-4770-86ac-2fc167ea365f · outbound

This paper cites What computers can’t do: The limits of artificial intelligence.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information What computers can’t do: The limits of artificial intelligence

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:06.478269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:07:02.816955Z digest=sha256:670d4104b771022bda2bf15f859f541d3fece36c9305c911eacb8010cf7fb66c

Observation e0f6b72d-cb2b-44a1-a86b-39d5e21a3c38 · outbound

This paper cites Foundational Challenges in Assuring Alignment and Safety of Large Language Models.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T20:07:02.902944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:07:02.902944Z digest=sha256:7c2ccf986deeabe689606fc0b263f4f76e79e74eb776baadea3532d732fa4805

Observation bce34754-5435-49c7-85d1-d7eeed121372 · outbound

This paper cites How to Enable Effective Cooperation Between Humans and NLP Models: A Survey of Principles, Formalizations, and Beyond.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information How to Enable Effective Cooperation Between Humans and NLP Models: A Survey of Principles, Formalizations, and Beyond

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-08-05T20:07:04.127404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:07:02.991235Z digest=sha256:90f588080ac4d104af6eef4e26f92151a57be87ce96200727841629e786d38e3

Observation 21861de3-46da-4924-a50a-962d68c5d7e7 · outbound

This paper cites Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:06.214899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:07:03.077429Z digest=sha256:5de6ef544d166e92a91aa1ac2e26b27dda430736af47ed357935d32538db4a6f

Observation 5f92a559-b612-4dcf-bfbe-2adab9850b45 · outbound

This paper cites AmbigQA: Answering Ambiguous Open-domain Questions.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information AmbigQA: Answering Ambiguous Open-domain Questions

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T20:07:03.139186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:07:03.139186Z digest=sha256:057b7665af77e9f7b52947e743c50f26f97ab9a3a78631e393d01b89dc5d6bd8

Observation 38ffefd5-7b95-41c5-ab5a-a093a90e3186 · outbound

This paper cites ChatShop: Interactive Information Seeking with Language Agents.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information ChatShop: Interactive Information Seeking with Language Agents

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-05T20:07:03.226004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:07:03.226004Z digest=sha256:8b1443adcc188e849e1e1ba8df76df13974fc65a258a51b08b2080215f6acc33

Observation 101e6db0-1c75-4f15-974e-b079df9cc900 · outbound

This paper cites Style: Improving domain transferability of asking clarification questions in large language model powered conversational agents.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Style: Improving domain transferability of asking clarification questions in large language model powered conversational agents

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:05.933634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:07:03.313584Z digest=sha256:1dfc8cc739c7624eb945927e4c3481c32634c7d5aaf6fab5f8e907e7924e0fbe

Observation f12f2614-8b02-43aa-92f8-4ceff6d62045 · outbound

This paper cites Prompting and evaluating large language models for proactive dialogues: Clarification, target-guided, and non-collaboration.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Prompting and evaluating large language models for proactive dialogues: Clarification, target-guided, and non-collaboration

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:05.688837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:07:03.400260Z digest=sha256:c1d44030fdec213920ad3e2fc8918334211e1f13713990c9caa3b37cd297953b

Observation 6f8c9949-2158-486e-a38a-f15f06fbbb6c · outbound

This paper cites CLAM: Selective Clarification for Ambiguous Questions with Generative Language Models.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information CLAM: Selective Clarification for Ambiguous Questions with Generative Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-05T20:07:03.485390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:07:03.485390Z digest=sha256:b7182db184ad71cc183548793bbe70d1bb892edecf244b0658c2d094b5e3cd7c

Observation 558401d8-0e9d-4913-820d-01f876df472e · outbound

This paper cites Selectively answering ambiguous questions.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Selectively answering ambiguous questions

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:05.482114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:07:03.576713Z digest=sha256:b818e0599f8bd0d64d55f58fbd4f36889f1576f886db7fd674068ef85dac0892

Observation 7805dadc-569b-4275-8b6a-582e6b7a74dc · outbound

This paper cites Multiwoz-a large-scale multi-domain wizard-of-oz dataset for task-oriented dialogue modelling.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Multiwoz-a large-scale multi-domain wizard-of-oz dataset for task-oriented dialogue modelling

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:05.242301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:07:03.656980Z digest=sha256:523e15377782cd961ceeae5ac8369085e3100211c43216b24d5165c003bc992a

Observation 4a57f24e-63db-4e72-9e29-27f8ff31b604 · outbound

This paper cites Towards scalable multi-domain conversational agents: The schema-guided dialogue dataset.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Towards scalable multi-domain conversational agents: The schema-guided dialogue dataset

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:05.020669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:07:03.743739Z digest=sha256:526b54058e86e08fe1f45bac29d2baf8896dc6017b31ca8d16021df0c73beaec

Observation 9ac5de61-2202-48e8-8946-673f04381adc · outbound

This paper cites We need to consider disagreement in evaluation.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information We need to consider disagreement in evaluation

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:04.803185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:07:03.839908Z digest=sha256:439732a4460e72ecdd9d6e881c669ad2736ff38fd1097391a6af1646f2451f13

Observation 9b8e2b2c-3d43-4e59-8cac-9a8b70c433bb · outbound

This paper cites Everyone’s voice matters: Quantifying annotation disagreement using demographic information.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Everyone’s voice matters: Quantifying annotation disagreement using demographic information

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:04.664550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:07:03.928169Z digest=sha256:9861821c092042c626a4f0e5149123cc0e6ea0ae8559c0bbb52fbcea5a641103

Pith citing papers

Observation aa4ce1f3-3f5a-44b5-9695-194bfd0ae40b · inbound

MIRA-Math: A Benchmark for Minimal Information Requesting and Mathematical Reasoning cites this paper.

MIRA-Math: A Benchmark for Minimal Information Requesting and Mathematical Reasoning Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T12:26:16.711939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-07-09T12:24:33.998226Z digest=sha256:2e30ccc83b7958734be95f52b8e0594bc49771fbb9d105dc2e1c239a503179a9