Pith. sign in

Paper Citation Record · LEDGER

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought

As of 8 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 1 inbound Pith citation observation for arXiv:2509.04027.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.04027 v4

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T10:31:37.980968Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-18T00:02:24.352947Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T00:02:25.420148Z

Reference resolution

53 of 53 outbound references displayed

  • verified exact0
  • verified fuzzy12
  • unresolved41
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 653d22f5-53a2-4c5d-bd2d-c921f972abc5 · outbound

This paper cites write newline.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.734744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.734744Z digest=sha256:ded06e819699561a631b91e14d0491b96e83553ca1a9bf3319eaea05ff2d12d1

Observation e0ece3d5-b6d2-4ecf-b2e1-086f710063dc · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.745578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.745578Z digest=sha256:0267e92acd3411dd2b70f136f60b55b0b9f6d9e3c5ff3130d563807d2ef5b69a

Observation 3a99df52-5d06-4985-9c17-1f152740cc5b · outbound

This paper cites Language models are few-shot learners.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Language models are few-shot learners

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.751394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.751394Z digest=sha256:b3de53ff524f3db788e3f9e743d5331a7977332cc215bd8e8346cbd3cd8a755f

Observation 7eadecf7-832c-4bab-b6de-39a6d38e2ae3 · outbound

This paper cites Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.755360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.755360Z digest=sha256:4d0104672c7e6d22de4926921d5cbebea1a350d202c891953704f7b812566d25

Observation 1d8dd784-010b-452d-ac80-2bda6dbdfb2a · outbound

This paper cites Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.759551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.759551Z digest=sha256:f3afd65d9f443d46d735a3359d19ab5def7bfbece8682e2772bcd3f52c8f68a4

Observation 0f1e3433-c858-49cf-852f-f44ed4fbd56d · outbound

This paper cites An Empirical Study on Eliciting and Improving R1-like Reasoning Models.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought An Empirical Study on Eliciting and Improving R1-like Reasoning Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.763822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.763822Z digest=sha256:6e34d3ec0aef8ca7d956cf8c0578e9426fadf7c462d483a0e21ab3a5163a3126

Observation 0469f239-9c70-4708-bb4d-0844b17ddcdc · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Training Verifiers to Solve Math Word Problems

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.768029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.768029Z digest=sha256:fdaf0373cc25d4640658e098271745dd1971c1a68466040818239a896677ba16

Observation 022a69a8-5b54-430b-bcf8-f3e022ec8531 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.772049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.772049Z digest=sha256:37bc573190a723071750ebc924b02c13c3797f46b17208160eac8acce7647622

Observation 8b0b2129-bb5a-4ea7-92e5-788a2ffe8717 · outbound

This paper cites Towards a Theoretical Understanding of Synthetic Data in LLM Post-Training: A Reverse-Bottleneck Perspective.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Towards a Theoretical Understanding of Synthetic Data in LLM Post-Training: A Reverse-Bottleneck Perspective

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.775535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.775535Z digest=sha256:2a61bc5e500a04c425b799c126626d4a97d88bbe58a3a33ae1067325a5c8304e

Observation fa980435-17e7-40da-a5ea-0905e43f7941 · outbound

This paper cites Rethinking External Slow-Thinking: From Snowball Errors to Probability of Correct Reasoning.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Rethinking External Slow-Thinking: From Snowball Errors to Probability of Correct Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.779402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.779402Z digest=sha256:975f3231463dd9d533015e4ce61037ad861257aa611d4a4505a67e6e1b90b41a

Observation 9bec0efd-2831-4438-9c99-49d7773451f0 · outbound

This paper cites On distances in uniformly random networks.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought On distances in uniformly random networks

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:38.536573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T10:31:37.783759Z digest=sha256:5a047eb1d49165f201cb1fe8b33c240c16b32de78d16d7009d37762cfe856e11

Observation 758aa007-4123-474a-a24a-2b735e7b363f · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Measuring Mathematical Problem Solving With the MATH Dataset

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.787916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.787916Z digest=sha256:71601831562e22002b7470d695e572268b5422f6946b6c3d7d4131fafbf97012

Observation 6dd9e2e9-cabd-40dd-9a55-44c83f72e0ef · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.792123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.792123Z digest=sha256:47f0818e31e8804c502dfbc1ace6b5e44bc170250a19e3cbbac4e39b7ad9c43d

Observation 295b6ad4-dade-4ed1-8e1a-2143e6727bdb · outbound

This paper cites Survey of hallucination in natural language generation.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Survey of hallucination in natural language generation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:38.525058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T10:31:37.796260Z digest=sha256:3085c3bfe6d7faa65b921627054d5e1db3e6736509c0aa80a03de7db765631ef

Observation 0a8439d8-5404-4d31-926f-f181fcc433c4 · outbound

This paper cites Enhancing LLM Reasoning with Reward-guided Tree Search.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Enhancing LLM Reasoning with Reward-guided Tree Search

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.800280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.800280Z digest=sha256:eb79165956e8761bad511bb9ae18eb4a8b0ff584def0f44d2eefab7cbc25710e

Observation ada20603-706a-4c55-be5c-3159391cb56d · outbound

This paper cites On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.804582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.804582Z digest=sha256:8d2d5c9a7a67a3ac16c196d1e0d1b53adf4aaf4d98ea5116d92afc42d0855702

Observation 0270234c-1444-4e64-a0a9-8a94c501a99b · outbound

This paper cites Decoupled Weight Decay Regularization.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Decoupled Weight Decay Regularization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.808968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.808968Z digest=sha256:f787c208233fa15f10ef60d2e03e774b13e0b1ed4ec586548cb3f8574a936063

Observation cc6366ea-b12b-4c5c-807e-a77ef077744a · outbound

This paper cites Some pac-bayesian theorems.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Some pac-bayesian theorems

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:38.514186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T10:31:37.813305Z digest=sha256:a984029780a555b9ce61b5be080509d7c5804d96da86ecdd73eabfe5b6f4f298

Observation a7d29796-62ed-4555-b21d-81aadb7bb521 · outbound

This paper cites The Llama 3 Herd of Models.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought The Llama 3 Herd of Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.819065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.819065Z digest=sha256:4ad503563cbb8d7c2ec0eac7fc54b70790e626aab294e60e98647750493a4e7c

Observation 99b43959-d9de-4c31-8bb8-631a3602bb7c · outbound

This paper cites Nearest neighbor distance in three-dimensional space.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Nearest neighbor distance in three-dimensional space

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:38.502802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T10:31:37.823520Z digest=sha256:92241222c60d5d7d1d25098ee09efb20361eca00a6b0cc7e799e674e7c6fcaea

Observation 18994fd2-58ee-43b0-8eaf-e72069020449 · outbound

This paper cites Learning to reason with llms, 2024.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Learning to reason with llms, 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:38.490571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T10:31:37.827688Z digest=sha256:9a034149d468e6fdcb476b49cd766b3b5f74c01ae5581703d233af1fdbcc6e8d

Observation a947c17b-463c-415e-bf3c-2bae4b76115f · outbound

This paper cites Introducing openai o3 and o4-mini, 2025.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Introducing openai o3 and o4-mini, 2025

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:38.477629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T10:31:37.831809Z digest=sha256:65c4e6c6b14159212cc93ac05c79087e790edc2a700fbe7d3643a8df1e126970

Observation a54cd083-fb2f-4c11-8b4a-c49510dbb830 · outbound

This paper cites Qwq: Reflect deeply on the boundaries of the unknown, November 2024.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Qwq: Reflect deeply on the boundaries of the unknown, November 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:38.464167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T10:31:37.836030Z digest=sha256:08d54240c9f0d0313240650cd96a6270f88773c08ccb58dd5ba592551c245560

Observation 6e0006be-46fd-4878-b922-e522d5a14d00 · outbound

This paper cites Qwen2.5 Technical Report.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Qwen2.5 Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.840224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.840224Z digest=sha256:bff7b2015b14a33df6fa28790ce70ec5eed32aeca2ffd3b6e4a08892b3a6782f

Observation f30b8c27-7c1d-4049-a9f9-9ec2465ce29f · outbound

This paper cites Qwen3 Technical Report.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Qwen3 Technical Report

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.844935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.844935Z digest=sha256:2ab1f7e5002c7bb42c8a1ffb0f4db88231eaacbaede42731aa5f4ea01090b072

Observation ed7d4204-0de5-4b55-ae19-7960911b9773 · outbound

This paper cites Benchmarking prompt sensitivity in large language models.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Benchmarking prompt sensitivity in large language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:38.451294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T10:31:37.849362Z digest=sha256:b38c5b64754b398513e3afe480a0ea7cd2d1b3d5f227d2865465f06f804a0c5c

Observation 2571ea31-6b53-4441-9afc-b0e244fadcb5 · outbound

This paper cites How much does your data exploration overfit? controlling bias via information usage.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought How much does your data exploration overfit? controlling bias via information usage

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:38.438143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T10:31:37.854081Z digest=sha256:7b0b256545eb309c3d147e115ff5bff493fc47a89ed7a9c012dcc3d4e070010b

Observation e3cad880-d189-4e67-bcf5-def90fd559b8 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Proximal Policy Optimization Algorithms

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.858445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.858445Z digest=sha256:c104a9257cecb59c42a8c053338e50354c44aec16cf340aac1550e0921273700

Observation c3121a39-00b2-46a5-a1ab-654f921f914f · outbound

This paper cites Scaling Test-Time Compute Without Verification or RL is Suboptimal.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Scaling Test-Time Compute Without Verification or RL is Suboptimal

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.862981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.862981Z digest=sha256:97898055cc91b3f9ec747316746adf1b38720e9970cb6333ae44839b7f8f5434

Observation 70d53b55-92e9-418f-8b2e-2d0617b2bff0 · outbound

This paper cites Spurious Rewards: Rethinking Training Signals in RLVR.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Spurious Rewards: Rethinking Training Signals in RLVR

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.869115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.869115Z digest=sha256:465fb3aeb1244a8aa642e1902be5c633f0d1650ac46b54f034bed90749ec401f

Observation ddd1d032-1732-4740-a8a4-bf1983309fac · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.873877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.873877Z digest=sha256:30b9b12e2b7bba7497c1b0a2a58ac3a9eb9d9632a1b4c89f2f72a398ba9334f7

Observation a67f407a-8b36-4419-adf6-781640bcc753 · outbound

This paper cites A bayesian perspective on generalization and stochastic gradient descent.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought A bayesian perspective on generalization and stochastic gradient descent

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.879593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.879593Z digest=sha256:79da8bae6f67d632fc55570026a1bc4999c233c4b0cc279d7ff856be93492020

Observation 21dc2f48-eff6-498b-8b74-69cc9287bd3b · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.884342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.884342Z digest=sha256:bf5cae8ee798e544dec7ce7c4237db1a94212d465f32851929630984e256ea17

Observation b671102f-5d38-4cf1-801e-336e7c3bd280 · outbound

This paper cites Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.889298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.889298Z digest=sha256:b04901ccfa87c51912b7c00abaa3d86314d0576aa7512889e77df4836637e1ba

Observation b5bd238d-19cf-4f4f-afa0-29c1ecc4ad53 · outbound

This paper cites Understanding Chain-of-Thought in LLMs through Information Theory.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Understanding Chain-of-Thought in LLMs through Information Theory

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.894683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.894683Z digest=sha256:492850ba948e2f88685d2eb22bbb8ca4a3f39dbbe35792c62bb530b766af31c5

Observation fcb68915-e810-4f07-9ce4-755df89ed3d5 · outbound

This paper cites Alphazero-like tree-search can guide large language model decoding and training.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Alphazero-like tree-search can guide large language model decoding and training

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:38.418602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T10:31:37.899762Z digest=sha256:834973f191f6346be7a7ce2d20d2221216f18cbcc561e2a96002637a0ef8a9f9

Observation f9dfb13e-8c84-4749-abd2-807892557a2e · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.904517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.904517Z digest=sha256:26d5629b5714c06889467a2d34937291866a356d5a9cb7c5b6b9e79939b55e7a

Observation 4f90a21e-332b-47fe-8db8-77670c608173 · outbound

This paper cites Emergent Abilities of Large Language Models.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Emergent Abilities of Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.909588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.909588Z digest=sha256:109e32e07e0e6266c78411e934dbff9c90909dbf1bf34410d2df8764bef804f1

Observation 5f0d9a22-a300-49fa-b531-53970135c7ae · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Chain-of-thought prompting elicits reasoning in large language models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.914448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.914448Z digest=sha256:d73e08a861ab9800bf503f3d34ee473059928de5ac3156a79809c61100997493

Observation d155c5de-8c81-4195-a9f0-170c5792b7e6 · outbound

This paper cites Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.919359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.919359Z digest=sha256:9d698a1aac56fa9ee7ad3f9f3edc38bbe73e556327d1ef4cef98e355e3dd90af

Observation 295e2ff3-2aef-472a-9343-23a1c135d222 · outbound

This paper cites Information-theoretic analysis of generalization capability of learning algorithms.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Information-theoretic analysis of generalization capability of learning algorithms

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:38.398678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T10:31:37.924011Z digest=sha256:808a4d2a33ff806953a8ba143f928dfb3dda040425b25b4060a99f17f423dd09

Observation f010506b-53ab-4da0-8a04-855bf8c5800b · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Tree of thoughts: Deliberate problem solving with large language models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.928512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.928512Z digest=sha256:3cdd183619a9237d2da3583995dd5bed166a0ff566012089a81a17cd67a790e8

Observation 25fff4de-bd51-413e-97d0-775fd7de3d63 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.933115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.933115Z digest=sha256:3a21c151a55899bebb2aa53537c1d3c5d2668bcca98f87f17e1846c8b53422f8

Observation 5c9667f0-812e-4e67-8d41-0e8f7003c117 · outbound

This paper cites What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.938067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.938067Z digest=sha256:8cdd7f28fc7c4cf707a4f15f46c71eac57746e8d655297959679454910d2bf15

Observation c7f1001c-2bab-4f60-8465-b46f3d0069de · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.942705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.942705Z digest=sha256:0449abd5a54149cb3a88b30bb066cb3e985a7f8d9c41a76e641f2be664373a25

Observation 97bdd282-cc34-4964-b295-0a99b8da6273 · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.947268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.947268Z digest=sha256:7f606f96a9b4a98cc3ffaf02c1a353d091e45b21abb28ebfec2bf7ba56c75bf6

Observation 1d10c2c0-d047-4c30-b3ad-a1a73d9121db · outbound

This paper cites Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.951774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.951774Z digest=sha256:05f6b3c228da8be22657649d5adee053b5eb973488086eba8db1cd50b5a51c34

Observation 6174444d-a37b-45bc-95a6-2838c75b832c · outbound

This paper cites Rest-mcts*: Llm self-training via process reward guided tree search.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Rest-mcts*: Llm self-training via process reward guided tree search

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.956608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.956608Z digest=sha256:a9d61893225ca49f5f303c6130c58be29a27f43ffe31cdacb3bc444dc45ce2ef

Observation cbe4cbfe-156e-46e5-874d-4a5c2542018f · outbound

This paper cites Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.961065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.961065Z digest=sha256:72a12cb6b7c7d6ae8a64c2fe5f690c2cebcb8e67a45c988b18b2bca68dfadf5d

Observation 238330f5-2f6a-4340-aa8c-b4321bac6f44 · outbound

This paper cites Prosa: Assessing and understanding the prompt sensitivity of llms.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Prosa: Assessing and understanding the prompt sensitivity of llms

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:38.367597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T10:31:37.966101Z digest=sha256:f0b77d06ee7f228a70c2ea9f917524dd65f92b7de5f03c0f37aa3daa2fb4b1ad

Observation fa14debb-98ee-4397-88ac-456b940528a8 · outbound

This paper cites @esa (Ref.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought @esa (Ref

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.971079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.971079Z digest=sha256:09fb7122210b2576e51e8f21495846a57d0f5c2ac607763037080c9c27ac18d8

Observation 9c1cc768-96eb-40c5-9962-4f4079d4f000 · outbound

This paper cites an unresolved cited work.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.976071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.976071Z digest=sha256:752488cdce63ebb56eaaf533f6b81b8b1a465efb4119ec6730b166c6026d6ef0

Observation 2db63415-3ea4-479d-9bbd-00fd28ff3b2b · outbound

This paper cites an unresolved cited work.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Unresolved cited work

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.980968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.980968Z digest=sha256:195e1c39fba632e4695b3bfa5dccd5e3261e60b3678816a0c468b40b6c78c927

Pith citing papers

Observation f8bb9d9c-cda0-4a4f-8558-1a01fd930ed4 · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought

Reference 152

Resolution
verified exact
arxiv_id, observed 2026-06-05T02:16:19.319482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:f1675eb45103749b2f16ea0c0a1e62b40edcba0292771484c085cb2c0eacc34a