Pith. sign in

Paper Citation Record · LEDGER

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information

As of 14 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 1 inbound Pith citation observation for arXiv:2508.11252.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.11252 v1

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T20:07:03.928169Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-09T12:24:33.998226Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T12:26:16.710584Z

Reference resolution

62 of 62 outbound references displayed

  • verified exact2
  • verified fuzzy40
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 46e93cf4-032b-4c97-b851-e926fb5e3dd1 · outbound

This paper cites van Eemeren, Rob Grootendorst, Sally Jackson, Scott Jacobs, Agnes van Rees, Francisca Snoeck Henkemans, Eveline T.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information van Eemeren, Rob Grootendorst, Sally Jackson, Scott Jacobs, Agnes van Rees, Francisca Snoeck Henkemans, Eveline T

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:12.728559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T20:06:58.190648Z digest=sha256:283f434684a5c198efb02cc12445e18958bdb2b87431135b97154f025b51dbd7

Observation a76876a4-59e5-4133-aee0-4ae45c21033f · outbound

This paper cites Dictionary of philosophy.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Dictionary of philosophy

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:12.581522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T20:06:58.245217Z digest=sha256:f507eb0553786bcff7a133684d7ef22d5c049f8ac1cfe06e6e954a4bdf4840df

Observation 370b6277-7626-4223-a4a4-8ce3fdf716b4 · outbound

This paper cites McCarthy and P.J.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information McCarthy and P.J

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:12.484770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T20:06:58.394610Z digest=sha256:f5fc24894ceda21883b0fb1e025cb11bae9b34cdd28db269408b1a9ab16ad4db

Observation 6283a7de-d9c5-4213-92df-a28911622977 · outbound

This paper cites Programs with common sense, 1959.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Programs with common sense, 1959

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:12.356361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T20:06:58.500892Z digest=sha256:97d21be7c55ffe04ce91073ca0a10d7af6740b95601494528d72dd346080d000

Observation b63c295f-3367-4fc6-8bdf-0d54e7e06fae · outbound

This paper cites Newell and H.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Newell and H

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:12.262508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T20:06:58.644951Z digest=sha256:8ea402e5c404bcb2c6279d077a56c7ae7bdf1c3381ba13373a5e03c8a6102d40

Observation 7bd09b23-7d48-4ca8-99bf-aa9a9404c831 · outbound

This paper cites OpenAI o1 System Card.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information OpenAI o1 System Card

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T20:06:58.787169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:06:58.787169Z digest=sha256:db0c5cd085742f87cd1106115342f2aeb86c35fc37fe86067845570f2aae1524

Observation 98a098ea-9d90-4b3d-b421-de3ddcb3e9bc · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T20:06:58.934923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:06:58.934923Z digest=sha256:43d69eba09d077d55d1ea284ca0761ce5d1675ea8149c7f3278ba3511a9d1093

Observation 7e1b6ea8-3138-4a37-98bf-503e9607f5b3 · outbound

This paper cites 2024 aime i.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information 2024 aime i

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:12.165657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T20:06:59.038089Z digest=sha256:45add1f0d4b719c6913b50d27e9050d2964750d28376a5d2661e432e71ac3d52

Observation 7a53f98a-62bf-43f9-9a55-9e3c0cc643d0 · outbound

This paper cites Let's verify step by step.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Let's verify step by step

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:12.035357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T20:06:59.117727Z digest=sha256:5c947f52ccc918968d1bc831f1283f8c6bacfcac0816abda39fe9d46b7c0ed41

Observation eba0e665-bb2d-4ed3-903f-c5887baf8f86 · outbound

This paper cites Omni- MATH : A universal olympiad level mathematic benchmark for large language models.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Omni- MATH : A universal olympiad level mathematic benchmark for large language models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:11.949130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T20:06:59.214909Z digest=sha256:ca1e0c5a05bd16165c49435b703fc4bb6d743f90641be072ee9f88c1a81735d7

Observation 6cf53a6a-4452-4e61-bc37-4641b64d79ce · outbound

This paper cites Open r1: A fully open reproduction of deepseek-r1, January 2025.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Open r1: A fully open reproduction of deepseek-r1, January 2025

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T20:06:59.358545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:06:59.358545Z digest=sha256:aa04873ba764574a9e5acb40320255bc1cd59b3433f462bbed3240cffaa05dc4

Observation df75b3d2-5714-4907-8b0e-f846dab3f8c1 · outbound

This paper cites Clamber: A benchmark of identifying and clarifying ambiguous information needs in large language models.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Clamber: A benchmark of identifying and clarifying ambiguous information needs in large language models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:11.794651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T20:06:59.465558Z digest=sha256:531f893de43761f78a365b420be4203106a2071e1d761d13de2c7d2c385dee4b

Observation 386b49d3-e878-442f-84cd-25f37252c31a · outbound

This paper cites Rethinking conversational agents in the era of llms: Proactivity, non-collaborativity, and beyond.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Rethinking conversational agents in the era of llms: Proactivity, non-collaborativity, and beyond

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:11.649126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T20:06:59.574990Z digest=sha256:4cac864682f325c02cd9f1945327a0dbc95d642ec07a30566859e153c4f04efe

Observation fd0d4cff-1480-45c6-b34f-09197e4f2f9a · outbound

This paper cites Li, Been Kim.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Li, Been Kim

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T20:06:59.670963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:06:59.670963Z digest=sha256:c958a7322a66ed8f6f0309a951654ed75793d4aa83157bdaa4e03b7a7156ddf6

Observation d640403b-1be5-469e-ae88-de6dbf05c71f · outbound

This paper cites Reasoning attack: Inducing llm to never-end thinking, 2025.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Reasoning attack: Inducing llm to never-end thinking, 2025

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:11.542959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T20:06:59.756678Z digest=sha256:1ce04d78218c22bb8676241a627c2f815546316e86492a02c51c3d68676d0912

Observation 7919619a-e362-4901-8965-1a66dc01ce8c · outbound

This paper cites Training language models to follow instructions with human feedback.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Training language models to follow instructions with human feedback

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:11.417146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T20:06:59.850459Z digest=sha256:70ef95843660330d5ad6c2157d5a69c9fdb52104c2e5f02ebc7cda260dbf3890

Observation 8ae85c5a-de73-456d-9582-33f2c95668a5 · outbound

This paper cites s1: Simple test-time scaling.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information s1: Simple test-time scaling

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T20:06:59.954946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:06:59.954946Z digest=sha256:0815c7d5430ad7142dcc3dd6838bfe07147f7ae13b9b8abc1c54cea67ca91cce

Observation 812c533b-10ff-4c3b-aeff-60ba40608fcd · outbound

This paper cites Bespoke-stratos Labs.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Bespoke-stratos Labs

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:11.112782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T20:07:00.050629Z digest=sha256:72c472a03a2657748c06afade9ac7e3a67aaa22a2078e94acafb588239349af7

Observation 0855451a-dbc2-4725-a890-92e160cdfdac · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Training Verifiers to Solve Math Word Problems

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T20:07:00.139772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:07:00.139772Z digest=sha256:6d296e454ead57b37e36317d17e15c53e1a9bac572343d17dcc4104673092ff4

Observation 925a3393-d92f-4266-af04-2f787ab03547 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T20:07:00.246609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:07:00.246609Z digest=sha256:7b1f27924a33995a47c84e8c6eceea1fff8fe98c951d45ec97a7cb8726187331

Observation 2e340aae-809c-4196-9127-57d204cb4abe · outbound

This paper cites Sky-t1: Train your own o1 preview model within \ 450.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Sky-t1: Train your own o1 preview model within \ 450

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:10.674559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T20:07:00.355608Z digest=sha256:3809fccee72c028e86be4ccc0a9c01346feb99a05f3a0b6446ff57db6272b3ca

Observation 1543a353-e86a-4853-b45d-d8927b37cbc1 · outbound

This paper cites RedStar: Does Scaling Long-CoT Data Unlock Better Slow-Reasoning Systems?.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information RedStar: Does Scaling Long-CoT Data Unlock Better Slow-Reasoning Systems?

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T20:07:00.485136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:07:00.485136Z digest=sha256:240daf2f9324138199fe3982d625ee916e8151164242f595cc39faa9b8c47c54

Observation 683f945d-43e1-4346-ae69-9d6eb0b96bd3 · outbound

This paper cites Qwen3: Think deeper, act faster.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Qwen3: Think deeper, act faster

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:10.365162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T20:07:00.591595Z digest=sha256:652f3caa2029cf7d7229cc63e32435154ebb4a227864d9d6dde55c488cc0ffc4

Observation b74a40ff-1242-4935-853b-dae7802d9374 · outbound

This paper cites Claude 3.7 sonnet and claude code.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Claude 3.7 sonnet and claude code

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:10.221580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T20:07:00.699335Z digest=sha256:8781892ba1ed030ab114485fc701ed844588e3af92048022fffc06482cdc3903

Observation 717194d9-c392-4cae-a616-9d30993f1c48 · outbound

This paper cites Grok 3 beta — the age of reasoning agents.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Grok 3 beta — the age of reasoning agents

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:10.033377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T20:07:00.778952Z digest=sha256:69232ad5636bba4bb74d4f3ab8da43611e634ca519f414194ecbe8d9003b878e

Observation feb82631-4813-4c5b-9148-414021f9a9b6 · outbound

This paper cites A Survey on LLM-as-a-Judge.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information A Survey on LLM-as-a-Judge

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T20:07:00.871004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:07:00.871004Z digest=sha256:ffb28f8e22e3fb5ebf437f7414abf4938725526bd1ec7e0eb3f10eeabb206d9b

Observation fde7b267-3a0b-473f-a97e-4fb6b5eca28a · outbound

This paper cites The Lessons of Developing Process Reward Models in Mathematical Reasoning.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information The Lessons of Developing Process Reward Models in Mathematical Reasoning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T20:07:00.983996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:07:00.983996Z digest=sha256:bb21f4c7d1d7504d15ebd3f908c271d9d5136f3f6f2abeb6b6cfae574b4132fd

Observation 82397fd9-be3f-4701-be79-d949ae8adcb3 · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T20:07:01.076649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:07:01.076649Z digest=sha256:5b642d8c9f967738e89ef68e1d438c318e2229a5963764cd7e592e28eca5d8b2

Observation 35fa5a0e-da24-46bb-8174-0f646ff6a458 · outbound

This paper cites Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T20:07:01.215034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:07:01.215034Z digest=sha256:1ead18d9eb375aa34e97ac528bc8f0a26b524e7ddfc3e6b5e5e86ad20e212339

Observation 3bba2630-f867-4ec9-8945-ab843d9d2d02 · outbound

This paper cites A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:09.852083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T20:07:01.280601Z digest=sha256:4f23aa5dc882ecdf70e7ba820473e1c0b1d73e65e0b9f23727cbf660981de588

Observation f0d8dbf4-632b-46fa-ac4f-dcb6c23728ad · outbound

This paper cites Distilling step-by-step! outperforming larger language models with less training data and smaller model sizes.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Distilling step-by-step! outperforming larger language models with less training data and smaller model sizes

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:09.715437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T20:07:01.363095Z digest=sha256:8e5247ddd0d9d25cf25182cbfe76d7d1fa5c1b9050a1f79d2f35a4514f069888

Observation 12ee1f7d-a8bb-4fec-b038-eca99d9085ee · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Distilling the Knowledge in a Neural Network

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T20:07:01.454580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:07:01.454580Z digest=sha256:9b4215c63d3762ad9df7f5283a96693760b06be3796c9a2d05d0b776458a6319

Observation 83059a89-9761-47af-b6aa-e0b209a60029 · outbound

This paper cites Math-Verify: Math Verification Library.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Math-Verify: Math Verification Library

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:09.408545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T20:07:01.521708Z digest=sha256:e39e5f80917b98650b47f6a59bd0709f27a7d6d7f599aa037446f2db072504eb

Observation 40f39d92-e928-4cdb-8c44-b759ec7efc31 · outbound

This paper cites OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T20:07:01.606002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:07:01.606002Z digest=sha256:078b41623c76d3acc57d0153cced3ab25165c9b5b925bacc425486f75423464e

Observation 1136c703-50c5-466f-b7ff-b7b6b98446e0 · outbound

This paper cites Interference and inhibition in cognition and behavior: Unifying themes for educational psychology.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Interference and inhibition in cognition and behavior: Unifying themes for educational psychology

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:09.160064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T20:07:01.695399Z digest=sha256:500baad58d3494772d63c4eb03d7409dbbbc011be261ea8318a82e601222ceb0

Observation 3bc7e2fb-86f4-4f89-83d1-2ae4594a183b · outbound

This paper cites How students “unpack” the structure of a word problem: Graphic representations and problem solving.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information How students “unpack” the structure of a word problem: Graphic representations and problem solving

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:08.776804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T20:07:01.788385Z digest=sha256:5028a24ca31f9e97a9deb75182f723fe2693b9514a785417b50fff4ce2c70a5d

Observation dac55333-c023-47e4-ad17-4d65f1ebac92 · outbound

This paper cites an unresolved cited work.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-05T20:07:08.606236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T20:07:01.852303Z digest=sha256:b55bcad51823c1758bcbf80c20009a4bb6071120ca3af509cbfdc49d5dc874c0

Observation 8d60990c-3ea2-4a7e-8f74-550313053840 · outbound

This paper cites Metacognition: A literature review.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Metacognition: A literature review

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:08.406532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T20:07:01.936919Z digest=sha256:06eb3630bf367b0746db9edd2aadbc1413a319156ed2c77017094b03f8ad5112

Observation cca27954-eebb-4e8a-a7d4-d7c8fe423332 · outbound

This paper cites Strategies for improving learner metacognition in health professional education.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Strategies for improving learner metacognition in health professional education

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:08.210450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T20:07:01.995855Z digest=sha256:2a9bb03d19f597a8923b5e1d3a4c4eee82266147c5e90ecf56cce71def3520a3

Observation ef5d890e-0c43-4392-8c98-7ca9a231aada · outbound

This paper cites Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T20:07:02.080403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:07:02.080403Z digest=sha256:992f1634dd9904910b7b7819403de332006bb965c90fdcb81b4863a976421590

Observation 62942cf8-05e8-44b3-9002-0c9956cc902c · outbound

This paper cites A survey of deep active learning.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information A survey of deep active learning

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:07.968999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T20:07:02.145598Z digest=sha256:9b033610a07ba06e8c33878e1a73293786f05b8541250302e5f7547fc67c890a

Observation 9bf6058f-c77c-4194-a102-3fd4beb95319 · outbound

This paper cites Deep bayesian active learning with image data.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Deep bayesian active learning with image data

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:07.796796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T20:07:02.234196Z digest=sha256:a969eb9e3e11ad664eb99d60110b7a021a2d303e12b4abed09119431621044f9

Observation ad39eb8c-7773-4d34-98e3-5954f177a0a4 · outbound

This paper cites Reinforcement learning: An introduction , volume 1.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Reinforcement learning: An introduction , volume 1

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:07.562472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T20:07:02.310110Z digest=sha256:596c91b1852c59967672455b3134395c61affc30e71074817e0d455b2d0a40c2

Observation c52c2a08-ece9-4395-8a33-e350216a0313 · outbound

This paper cites Partially Observable Task and Motion Planning with Uncertainty and Risk Awareness.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Partially Observable Task and Motion Planning with Uncertainty and Risk Awareness

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-08-05T20:07:04.307527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T20:07:02.384922Z digest=sha256:1a3d78e32aee1fa86349576caa95077a23d73f3754fb99a491ede5261139d76c

Observation 7e5facde-f8c8-4345-b778-7db9d575b6f1 · outbound

This paper cites Combined task and motion planning under partial observability: An optimization-based approach.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Combined task and motion planning under partial observability: An optimization-based approach

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:07.331095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T20:07:02.475798Z digest=sha256:744e72bd64ee28177ec24e2165e27883e6ed6975f434cdaa83ebabe8c40f2ea3

Observation 0b1c7b4b-5a4b-45aa-b921-d6890b2a258b · outbound

This paper cites The communicative function of ambiguity in language.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information The communicative function of ambiguity in language

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:07.103378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T20:07:02.560443Z digest=sha256:347e78484be35fb8689c66b424b296dd4be936c3d8f0ad36bf53d0aa5bd07f39

Observation 9e8c9df5-a049-484b-b03c-fb9c513cac9d · outbound

This paper cites The puzzle of ambiguity.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information The puzzle of ambiguity

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:06.910037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T20:07:02.645566Z digest=sha256:8880c45b09a1efc2f880f4f8e64ff16217f12c4246517278c399db505530fd54

Observation 43c56b16-8780-4f11-87bb-f1f8ff5ba0ad · outbound

This paper cites Semantic ambiguity within and across languages: An integrative review.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Semantic ambiguity within and across languages: An integrative review

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:06.715372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T20:07:02.733138Z digest=sha256:f10508a262b36d2575ef57d2c8a9346fea0c4bd97383597e0cde635445abff76

Observation 996c406a-4b06-4770-86ac-2fc167ea365f · outbound

This paper cites What computers can’t do: The limits of artificial intelligence.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information What computers can’t do: The limits of artificial intelligence

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:06.478269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T20:07:02.816955Z digest=sha256:6ce2041f152a8bd51308958f752ef9475e548cb4789180940a7ee24a8097f88d

Observation e0f6b72d-cb2b-44a1-a86b-39d5e21a3c38 · outbound

This paper cites Foundational Challenges in Assuring Alignment and Safety of Large Language Models.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T20:07:02.902944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:07:02.902944Z digest=sha256:8c1ec0b505b76a455df47cbcc3479ebb4264f9a08a1084f09d2ed6b230cc527e

Observation bce34754-5435-49c7-85d1-d7eeed121372 · outbound

This paper cites How to Enable Effective Cooperation Between Humans and NLP Models: A Survey of Principles, Formalizations, and Beyond.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information How to Enable Effective Cooperation Between Humans and NLP Models: A Survey of Principles, Formalizations, and Beyond

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-08-05T20:07:04.127404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T20:07:02.991235Z digest=sha256:5fbbe8c88c0343d9b0ae2a489619f37d59572fa2c0b4c2b789e32dcfe9be7c0d

Observation 21861de3-46da-4924-a50a-962d68c5d7e7 · outbound

This paper cites Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:06.214899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T20:07:03.077429Z digest=sha256:385676b9ad1775cf7b3120ee872be52cd354319c77417e8f92b6d69f17134625

Observation 5f92a559-b612-4dcf-bfbe-2adab9850b45 · outbound

This paper cites AmbigQA: Answering Ambiguous Open-domain Questions.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information AmbigQA: Answering Ambiguous Open-domain Questions

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T20:07:03.139186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:07:03.139186Z digest=sha256:9c29f206908991b1b2f75a36f82f68b5c90b49333ae534055dc86d3b1041bb18

Observation 38ffefd5-7b95-41c5-ab5a-a093a90e3186 · outbound

This paper cites ChatShop: Interactive Information Seeking with Language Agents.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information ChatShop: Interactive Information Seeking with Language Agents

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-05T20:07:03.226004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:07:03.226004Z digest=sha256:b67fa6ac16dde51211547eda896323c472f20ff4d7ffc22a8b67765de5b211bc

Observation 101e6db0-1c75-4f15-974e-b079df9cc900 · outbound

This paper cites Style: Improving domain transferability of asking clarification questions in large language model powered conversational agents.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Style: Improving domain transferability of asking clarification questions in large language model powered conversational agents

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:05.933634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T20:07:03.313584Z digest=sha256:f7f52c8fa142dfaec95ee62a4f13babd9fea88c8e004983aceff09647bb661d2

Observation f12f2614-8b02-43aa-92f8-4ceff6d62045 · outbound

This paper cites Prompting and evaluating large language models for proactive dialogues: Clarification, target-guided, and non-collaboration.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Prompting and evaluating large language models for proactive dialogues: Clarification, target-guided, and non-collaboration

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:05.688837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T20:07:03.400260Z digest=sha256:f295020d7931d0c967953b7a9828d5ea74bba78be5ef96a09a84b9aa7f786474

Observation 6f8c9949-2158-486e-a38a-f15f06fbbb6c · outbound

This paper cites CLAM: Selective Clarification for Ambiguous Questions with Generative Language Models.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information CLAM: Selective Clarification for Ambiguous Questions with Generative Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-05T20:07:03.485390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:07:03.485390Z digest=sha256:f28f92559f806270d4e6cbbcd9fd1ec17beff333dc8619ae0661e3f262af68fd

Observation 558401d8-0e9d-4913-820d-01f876df472e · outbound

This paper cites Selectively answering ambiguous questions.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Selectively answering ambiguous questions

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:05.482114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T20:07:03.576713Z digest=sha256:c9cc86e429aea04e884d83e7a54a90dac027c68acd6d6083b3ba3fce0ba3bc7b

Observation 7805dadc-569b-4275-8b6a-582e6b7a74dc · outbound

This paper cites Multiwoz-a large-scale multi-domain wizard-of-oz dataset for task-oriented dialogue modelling.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Multiwoz-a large-scale multi-domain wizard-of-oz dataset for task-oriented dialogue modelling

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:05.242301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T20:07:03.656980Z digest=sha256:d31ba3dbcd62f4141721dd24c2cbec837a159ca6f2962172a16f2d3b3d66e4a3

Observation 4a57f24e-63db-4e72-9e29-27f8ff31b604 · outbound

This paper cites Towards scalable multi-domain conversational agents: The schema-guided dialogue dataset.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Towards scalable multi-domain conversational agents: The schema-guided dialogue dataset

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:05.020669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T20:07:03.743739Z digest=sha256:f8e600ff941274c71ad034de5c4db8aa22075262556526adc3eb8a0cd3043cea

Observation 9ac5de61-2202-48e8-8946-673f04381adc · outbound

This paper cites We need to consider disagreement in evaluation.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information We need to consider disagreement in evaluation

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:04.803185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T20:07:03.839908Z digest=sha256:51b4138cc5a0b4cd7dc56a27954056327e939f0036867c38ce4213fd0764cc6e

Observation 9b8e2b2c-3d43-4e59-8cac-9a8b70c433bb · outbound

This paper cites Everyone’s voice matters: Quantifying annotation disagreement using demographic information.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Everyone’s voice matters: Quantifying annotation disagreement using demographic information

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:07:04.664550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T20:07:03.928169Z digest=sha256:b07f36b7a370d6e02fb5227e59b248bdafee3423311a558408cde8d0f865dc39

Pith citing papers

Observation aa4ce1f3-3f5a-44b5-9695-194bfd0ae40b · inbound

MIRA-Math: A Benchmark for Minimal Information Requesting and Mathematical Reasoning cites this paper.

MIRA-Math: A Benchmark for Minimal Information Requesting and Mathematical Reasoning Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T12:26:16.711939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-07-09T12:24:33.998226Z digest=sha256:7e95c91d51e9e756a87aa25d8279799a78e84da342c5f9749d6eb5d4c6edf6f7