Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T10:31:37.980968Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 1 inbound Pith citation observation for arXiv:2509.04027.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T10:31:37.980968Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-18T00:02:24.352947Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-18T00:02:25.420148Z
53 of 53 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 653d22f5-53a2-4c5d-bd2d-c921f972abc5 · outbound
Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0ece3d5-b6d2-4ecf-b2e1-086f710063dc · outbound
Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a99df52-5d06-4985-9c17-1f152740cc5b · outbound
Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Language models are few-shot learners
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7eadecf7-832c-4bab-b6de-39a6d38e2ae3 · outbound
Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d8dd784-010b-452d-ac80-2bda6dbdfb2a · outbound
Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f1e3433-c858-49cf-852f-f44ed4fbd56d · outbound
Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought An Empirical Study on Eliciting and Improving R1-like Reasoning Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0469f239-9c70-4708-bb4d-0844b17ddcdc · outbound
Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Training Verifiers to Solve Math Word Problems
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 022a69a8-5b54-430b-bcf8-f3e022ec8531 · outbound
Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b0b2129-bb5a-4ea7-92e5-788a2ffe8717 · outbound
Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Towards a Theoretical Understanding of Synthetic Data in LLM Post-Training: A Reverse-Bottleneck Perspective
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa980435-17e7-40da-a5ea-0905e43f7941 · outbound
Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Rethinking External Slow-Thinking: From Snowball Errors to Probability of Correct Reasoning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bec0efd-2831-4438-9c99-49d7773451f0 · outbound
Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought On distances in uniformly random networks
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 758aa007-4123-474a-a24a-2b735e7b363f · outbound
Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Measuring Mathematical Problem Solving With the MATH Dataset
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6dd9e2e9-cabd-40dd-9a55-44c83f72e0ef · outbound
Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 295b6ad4-dade-4ed1-8e1a-2143e6727bdb · outbound
Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Survey of hallucination in natural language generation
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0a8439d8-5404-4d31-926f-f181fcc433c4 · outbound
Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Enhancing LLM Reasoning with Reward-guided Tree Search
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ada20603-706a-4c55-be5c-3159391cb56d · outbound
Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0270234c-1444-4e64-a0a9-8a94c501a99b · outbound
Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Decoupled Weight Decay Regularization
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc6366ea-b12b-4c5c-807e-a77ef077744a · outbound
Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Some pac-bayesian theorems
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a7d29796-62ed-4555-b21d-81aadb7bb521 · outbound
Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought The Llama 3 Herd of Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99b43959-d9de-4c31-8bb8-631a3602bb7c · outbound
Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Nearest neighbor distance in three-dimensional space
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 18994fd2-58ee-43b0-8eaf-e72069020449 · outbound
Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Learning to reason with llms, 2024
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a947c17b-463c-415e-bf3c-2bae4b76115f · outbound
Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Introducing openai o3 and o4-mini, 2025
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a54cd083-fb2f-4c11-8b4a-c49510dbb830 · outbound
Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Qwq: Reflect deeply on the boundaries of the unknown, November 2024
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6e0006be-46fd-4878-b922-e522d5a14d00 · outbound
Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Qwen2.5 Technical Report
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f30b8c27-7c1d-4049-a9f9-9ec2465ce29f · outbound
Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Qwen3 Technical Report
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed7d4204-0de5-4b55-ae19-7960911b9773 · outbound
Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Benchmarking prompt sensitivity in large language models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2571ea31-6b53-4441-9afc-b0e244fadcb5 · outbound
Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought How much does your data exploration overfit? controlling bias via information usage
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e3cad880-d189-4e67-bcf5-def90fd559b8 · outbound
Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Proximal Policy Optimization Algorithms
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3121a39-00b2-46a5-a1ab-654f921f914f · outbound
Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Scaling Test-Time Compute Without Verification or RL is Suboptimal
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70d53b55-92e9-418f-8b2e-2d0617b2bff0 · outbound
Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Spurious Rewards: Rethinking Training Signals in RLVR
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddd1d032-1732-4740-a8a4-bf1983309fac · outbound
Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a67f407a-8b36-4419-adf6-781640bcc753 · outbound
Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought A bayesian perspective on generalization and stochastic gradient descent
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21dc2f48-eff6-498b-8b74-69cc9287bd3b · outbound
Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b671102f-5d38-4cf1-801e-336e7c3bd280 · outbound
Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5bd238d-19cf-4f4f-afa0-29c1ecc4ad53 · outbound
Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Understanding Chain-of-Thought in LLMs through Information Theory
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcb68915-e810-4f07-9ce4-755df89ed3d5 · outbound
Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Alphazero-like tree-search can guide large language model decoding and training
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f9dfb13e-8c84-4749-abd2-807892557a2e · outbound
Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Self-Consistency Improves Chain of Thought Reasoning in Language Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f90a21e-332b-47fe-8db8-77670c608173 · outbound
Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Emergent Abilities of Large Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f0d9a22-a300-49fa-b531-53970135c7ae · outbound
Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Chain-of-thought prompting elicits reasoning in large language models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d155c5de-8c81-4195-a9f0-170c5792b7e6 · outbound
Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 295e2ff3-2aef-472a-9343-23a1c135d222 · outbound
Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Information-theoretic analysis of generalization capability of learning algorithms
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f010506b-53ab-4da0-8a04-855bf8c5800b · outbound
Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Tree of thoughts: Deliberate problem solving with large language models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25fff4de-bd51-413e-97d0-775fd7de3d63 · outbound
Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c9667f0-812e-4e67-8d41-0e8f7003c117 · outbound
Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7f1001c-2bab-4f60-8465-b46f3d0069de · outbound
Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97bdd282-cc34-4964-b295-0a99b8da6273 · outbound
Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d10c2c0-d047-4c30-b3ad-a1a73d9121db · outbound
Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6174444d-a37b-45bc-95a6-2838c75b832c · outbound
Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Rest-mcts*: Llm self-training via process reward guided tree search
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbe4cbfe-156e-46e5-874d-4a5c2542018f · outbound
Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 238330f5-2f6a-4340-aa8c-b4321bac6f44 · outbound
Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Prosa: Assessing and understanding the prompt sensitivity of llms
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fa14debb-98ee-4397-88ac-456b940528a8 · outbound
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c1cc768-96eb-40c5-9962-4f4079d4f000 · outbound
Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Unresolved cited work
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2db63415-3ea4-479d-9bbd-00fd28ff3b2b · outbound
Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Unresolved cited work
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8bb9d9c-cda0-4a4f-8558-1a01fd930ed4 · inbound
A Survey of Reinforcement Learning for Large Reasoning Models Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought
Reference 152
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.