Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T12:14:48.950737Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 5 inbound Pith citation observations for arXiv:2504.13367.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T12:14:48.950737Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:35:43.292037Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
31 of 31 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 43c2f155-1b99-4149-bfaa-030f7a739ab3 · outbound
THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4dc3494-2e7f-460b-89c0-2fc42f74b2ae · outbound
THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 131b31e2-6268-414d-a5fa-76286ca826cb · outbound
THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models Over-reasoning and redundant calculation of large language models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8e826a7d-29de-4795-90b3-c6956d153280 · outbound
THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models Training Verifiers to Solve Math Word Problems
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40b657e6-77ab-432e-bf09-ce7f7a9b9f7f · outbound
THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89397f45-798e-4cc4-b659-dea66024f808 · outbound
THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by Step
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8138f147-c3ab-43d0-b091-4447a71f61e3 · outbound
THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models Mercury: A code efficiency benchmark for code large language models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ebc2cf72-08f3-4931-b578-8a19db43dc35 · outbound
THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models The Llama 3 Herd of Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 890f2dd9-ab3c-4901-85ab-af273b6b27af · outbound
THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models Efficiently Scaling LLM Reasoning with Certaindex
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14891e74-26e4-459a-90ac-3eb649d0eb48 · outbound
THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models Measuring mathematical problem solving with the math dataset
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9687924-55a1-4f17-88e8-0144c8d3e4a6 · outbound
THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models OpenAI o1 System Card
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ea386ff-7360-4b18-b448-7968903ce8cb · outbound
THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08a3d596-66ee-4b56-83f2-0580513257f2 · outbound
THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models Let's Verify Step by Step
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18416b25-db0e-4c2a-846a-818b13c5e27d · outbound
THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models Zebralogic: Benchmarking the logical reasoning ability of language models, 2024
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b0895ca6-2b2a-4fea-99eb-56b6826f249b · outbound
THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models Can Language Models Learn to Skip Steps?
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18ef791c-9136-41eb-bb59-eeeaf2c62948 · outbound
THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models Gemma: Open Models Based on Gemini Research and Technology
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05040b97-1084-40e0-9fa5-d14a24de9d5d · outbound
THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models s1: Simple test-time scaling
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8da46fcd-b93a-47b4-bc08-1c3ed408f3ac · outbound
THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models Automatically correcting large language models: Surveying the landscape of diverse automated correction strategies
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 09fe61dd-0ab5-4525-816d-8cbe1ee31d2f · outbound
THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models Qwq-32b: Embracing the power of reinforcement learning, March 2025
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f16e2643-33ed-428c-9448-276d20d50424 · outbound
THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e839c21f-89ab-424d-abcb-1bda4628b216 · outbound
THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 558dea1c-e9b8-4584-8e69-fd9aac865849 · outbound
THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac3de784-0977-4b60-8761-e7545f859678 · outbound
THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba55d97d-57b5-48b0-b4ac-e7e288568540 · outbound
THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models Inference scaling laws: An empirical analysis of compute-optimal inference for problem-solving with language models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation bb2eac34-96ff-4878-8cbb-57175fa68ac2 · outbound
THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models Qwen2.5 Technical Report
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f3c8050-3965-40e9-931c-3b6c50b10716 · outbound
THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models Towards thinking-optimal scaling of test-time compute for llm reasoning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f6f807f-5a06-4f50-a79f-b2d089b54981 · outbound
THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models Think Smarter not Harder: Adaptive Reasoning with Inference Aware Optimization
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cdeee7e-923d-4e0c-a08e-82d9643bf942 · outbound
THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52afbd10-403f-48a1-ae6b-0685b53e8fc9 · outbound
THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models @esa (Ref
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40f9b3f5-fa2a-4150-b323-baef6c1e8ba7 · outbound
THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models Unresolved cited work
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddfa9544-8400-4079-bfc6-3dddd38848c9 · outbound
THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models JH cBP ]D aW Z *+b tW77wϙ 9w;Ri @ O
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61691de3-89af-489a-b581-36ef3b0a6fab · inbound
AlphaOne: Reasoning Models Thinking Slow and Fast at Test Time THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30c26e1a-4e14-4d49-b0b5-ba9f08e06c7d · inbound
How Far Are We from Optimal Reasoning Efficiency? THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 965ee1d5-4697-42e8-951f-79d218f4d34c · inbound
Do LLMs Overthink Basic Math Reasoning? Benchmarking the Accuracy-Efficiency Tradeoff in Language Models THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 7efa6940-8689-4eaa-abf6-800d893ff98e · inbound
BudgetThinker: Empowering Budget-aware LLM Reasoning with Control Tokens THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14627de6-5ed9-4683-92de-60bac275179f · inbound
Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.