Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:37:03.196744Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 8 inbound Pith citation observations for arXiv:2505.18298.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:37:03.196744Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T21:45:10.270039Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T09:19:43.854814Z
35 of 35 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 442ab1f7-9fc6-484f-bc92-061aefa9a772 · outbound
Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f53ce4f-0aef-4a8a-8e9f-3731ce5fe379 · outbound
Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Training language models to reason efficiently.arXiv preprint arXiv:2502.04463, 2025
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72a9f013-a6a3-47ba-a1df-020837f1e845 · outbound
Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3663651-1041-422a-858d-3edb87b86a86 · outbound
Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Training Verifiers to Solve Math Word Problems
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51308257-6ebd-4863-a821-4f87adf34c01 · outbound
Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards competitive programming platform.https://codeforces.com/, 2025
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c970b6b3-d015-41fc-ba95-5c2ab9399b22 · outbound
Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Efficient reasoning models: A survey.arXiv preprint arXiv:2504.10903, 2025
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ee06ac6-e9c8-4669-a56a-d1e83aa8f32b · outbound
Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38674f00-a0af-40af-b4f9-bd9e328a6162 · outbound
Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Scaling laws for reward model overoptimization
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8832b3e4-85ff-4a45-9030-0a56a9b4c4ce · outbound
Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b90ca9c3-e075-477f-9356-016a9a549240 · outbound
Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Token-Budget-Aware LLM Reasoning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be7cfda0-09a0-4455-874b-d240ddc9084d · outbound
Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcef9361-7a71-4c42-88a1-def13466c23d · outbound
Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Measuring Mathematical Problem Solving With the MATH Dataset
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 664a0074-5786-45a7-a606-0116e12b066e · outbound
Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5856107e-8d78-40cd-af8d-bca9abfc7d6f · outbound
Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards C3ot: Generating shorter chain-of- thought without compromising effectiveness
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ab5e3d9-f52c-45d2-b9af-3ab106d5cec7 · outbound
Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Training Language Models to Self-Correct via Reinforcement Learning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31e03873-e1f0-4215-94d9-6e52937bf939 · outbound
Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Let’s verify step by step
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be441b8d-1204-4559-8db8-fb50648a4f2a · outbound
Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1076a454-ec4d-4c8c-b4cb-24c664984679 · outbound
Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Tianjun Zhang, Li Erran Li, Raluca Ada Popa, and Ion Stoica
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 505ccb87-69d6-4fee-9cd0-1cc0f6d00f67 · outbound
Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e1c8307-a8e8-4f91-836a-98a20af4269f · outbound
Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Self-Training Elicits Concise Reasoning in Large Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abf27e81-4979-46ec-bfe9-c90a609f7654 · outbound
Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Concise Thoughts: Impact of Output Length on LLM Reasoning and Cost
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b90efee-0b7f-4959-addf-336cbe67c12a · outbound
Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Learning to reason with llms.https://openai.com/index/learning-to-reason-with-llms/, 2024
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 60242fa7-0d39-445e-a632-bc67c3386528 · outbound
Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Gpqa: A graduate-level google-proof q&a benchmark
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08aef909-c430-4afe-9729-9393712e7ad7 · outbound
Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Code Llama: Open Foundation Models for Code
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 158a1cca-3fa1-41c1-9c31-f17173af4c38 · outbound
Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f4b5653-1520-4686-9390-b95c16d4e23b · outbound
Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Between Underthinking and Overthinking: An Empirical Study of Reasoning Length and correctness in LLMs
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c121087-817b-458e-85a1-5ec109884a4b · outbound
Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f35abcbc-a4b1-4f97-bd70-f2bc9589c7f4 · outbound
Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Qwq-32b-preview.https://qwenlm.github.io/blog/qwq-32b-preview/, 2025
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8ee41af2-69d6-45d3-bbe2-f7952706419e · outbound
Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Harnessing the Reasoning Economy: A Survey of Efficient Reasoning for Large Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91d8a4e6-f946-45b6-a219-4417b9846b2b · outbound
Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Self-Consistency Improves Chain of Thought Reasoning in Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d748111-dfab-45c2-b6bb-fd5012ffa94b · outbound
Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Chain-of-thought prompting elicits reasoning in large language models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3f68f8b-fa53-436f-93dc-151615958a5f · outbound
Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Unlocking Efficient Long-to-Short LLM Reasoning with Model Merging
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bd1b620-86a1-489e-bec5-64558b313831 · outbound
Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Tokenskip: Controllable chain-of-thought compression in llms.arXiv preprint arXiv:2502.12067, 2025
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56280a21-99c1-43fe-94e3-d410c4ac50c2 · outbound
Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Towards thinking-optimal scaling of test-time compute for llm reasoning.arXiv preprint arXiv:2502.18080, 2025
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68153ac2-28be-4582-9590-89a703dc1273 · outbound
Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Shorterbetter: Guiding reasoning models to find optimal inference length for efficient reasoning.arXiv preprint arXiv:2504.21370, 2025
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d52c04d0-32e9-4730-b36a-609bd4a0a2d3 · inbound
Do Thinking Tokens Help or Trap? Towards More Efficient Large Reasoning Model Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02d91076-4512-4bc6-ad22-2425111373bc · inbound
Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards
Reference 168
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e800aca3-effb-423a-8206-db125821c2c5 · inbound
Learning to Reason Efficiently with Discounted Reinforcement Learning Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards
Reference 2010
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56fffa5f-907e-49f8-a5be-9506a292801d · inbound
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4369648a-efbc-44ca-b123-55488db6a1dc · inbound
Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f1e542d4-bdda-4a54-853a-dcf3a992cc5b · inbound
CLORE: Content-Level Optimization for Reasoning Efficiency Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation bb273491-ab63-42ba-8e1b-99c1e0dfdacf · inbound
CARE: Competence-Aware Reward Shaping for Adaptive Reasoning Length in Video-MLLMs Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3782feab-adce-4d78-806e-cc9f1647c702 · inbound
Beyond Penalizing Mistakes: Stabilizing Efficiency Training in Large Reasoning Models via Adaptive Correct-Only Rewards Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.