Pith. sign in

Paper Citation Record · LEDGER

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought

As of 10 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 1 inbound Pith citation observation for arXiv:2509.04027.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.04027 v4

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T10:31:37.980968Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-18T00:02:24.352947Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T00:02:25.420148Z

Reference resolution

53 of 53 outbound references displayed

  • verified exact0
  • verified fuzzy12
  • unresolved41
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 653d22f5-53a2-4c5d-bd2d-c921f972abc5 · outbound

This paper cites write newline.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.734744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.734744Z digest=sha256:728dea0450076a33fbc9a7ae71c474447d2b70ecedee894ff8f86e9fbc1405e5

Observation e0ece3d5-b6d2-4ecf-b2e1-086f710063dc · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.745578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.745578Z digest=sha256:871a37efb86777f0e1d550dbfb2b2ae6c2d02668805d5de27be4dca148a72b92

Observation 3a99df52-5d06-4985-9c17-1f152740cc5b · outbound

This paper cites Language models are few-shot learners.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Language models are few-shot learners

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.751394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.751394Z digest=sha256:f3c4b50c4ac2592da4e286ea48a2f22ff50ecca78b124c15cf9072992382f87c

Observation 7eadecf7-832c-4bab-b6de-39a6d38e2ae3 · outbound

This paper cites Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.755360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.755360Z digest=sha256:20276cc67cf6e7486e8605c48a0de4adb4802f58c60fb47fd3644d4fb7102def

Observation 1d8dd784-010b-452d-ac80-2bda6dbdfb2a · outbound

This paper cites Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.759551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.759551Z digest=sha256:a992a3d7508ee2a9b45a4446e095988447d1eac25155e7b937da6716b477d9d7

Observation 0f1e3433-c858-49cf-852f-f44ed4fbd56d · outbound

This paper cites An Empirical Study on Eliciting and Improving R1-like Reasoning Models.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought An Empirical Study on Eliciting and Improving R1-like Reasoning Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.763822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.763822Z digest=sha256:3e71763bb414a3240569de8d6f4887158dac89bea23589a335139fc7c8d90d34

Observation 0469f239-9c70-4708-bb4d-0844b17ddcdc · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Training Verifiers to Solve Math Word Problems

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.768029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.768029Z digest=sha256:c1928e765b1be5ff53a2e5f0f1fbf05d1c8bc8e74b35fde9d1f3188f477df1c2

Observation 022a69a8-5b54-430b-bcf8-f3e022ec8531 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.772049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.772049Z digest=sha256:3a1d30ef3175699cfa4d273ede05c44ed7ca03b0858f69cc0d3e94a307c1bcfa

Observation 8b0b2129-bb5a-4ea7-92e5-788a2ffe8717 · outbound

This paper cites Towards a Theoretical Understanding of Synthetic Data in LLM Post-Training: A Reverse-Bottleneck Perspective.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Towards a Theoretical Understanding of Synthetic Data in LLM Post-Training: A Reverse-Bottleneck Perspective

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.775535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.775535Z digest=sha256:54e32d13685a0135d78714c2126fed407a0244fa2b0408521fae3ef3ed6c82bb

Observation fa980435-17e7-40da-a5ea-0905e43f7941 · outbound

This paper cites Rethinking External Slow-Thinking: From Snowball Errors to Probability of Correct Reasoning.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Rethinking External Slow-Thinking: From Snowball Errors to Probability of Correct Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.779402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.779402Z digest=sha256:5eaf32b240cf3f857e57ed3861135f860fcf132f5e8a3ed1a171a81ff199a71b

Observation 9bec0efd-2831-4438-9c99-49d7773451f0 · outbound

This paper cites On distances in uniformly random networks.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought On distances in uniformly random networks

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:38.536573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T10:31:37.783759Z digest=sha256:044566155f31ba10ecf15a6916ff28ecb32142757ae81f3bd368dc56e1374ae9

Observation 758aa007-4123-474a-a24a-2b735e7b363f · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Measuring Mathematical Problem Solving With the MATH Dataset

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.787916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.787916Z digest=sha256:aa129c4bff45e36019bf90bbd82a053401071eb29ce7e590cb0873c233fe5887

Observation 6dd9e2e9-cabd-40dd-9a55-44c83f72e0ef · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.792123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.792123Z digest=sha256:e24a18084e9cee1675062296803dd44dfac36cafa082efdf10b052e135f89cc9

Observation 295b6ad4-dade-4ed1-8e1a-2143e6727bdb · outbound

This paper cites Survey of hallucination in natural language generation.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Survey of hallucination in natural language generation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:38.525058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T10:31:37.796260Z digest=sha256:5798eebe010bfdf7cd53f37dc11cd7a5d438ade2038418e144cc93796452cc23

Observation 0a8439d8-5404-4d31-926f-f181fcc433c4 · outbound

This paper cites Enhancing LLM Reasoning with Reward-guided Tree Search.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Enhancing LLM Reasoning with Reward-guided Tree Search

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.800280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.800280Z digest=sha256:889f8a328854ffa3f492957b1633bd4aed458679db2522218d61b5b0aa1133de

Observation ada20603-706a-4c55-be5c-3159391cb56d · outbound

This paper cites On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.804582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.804582Z digest=sha256:c3bc98d13916459161798efd071014084f0ae2c7ffc3728590dbe8996d964d2a

Observation 0270234c-1444-4e64-a0a9-8a94c501a99b · outbound

This paper cites Decoupled Weight Decay Regularization.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Decoupled Weight Decay Regularization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.808968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.808968Z digest=sha256:f2e12e22bd318e30e75e752f73e8c024d0fd2eb961137aa5e1728d6d3c70be71

Observation cc6366ea-b12b-4c5c-807e-a77ef077744a · outbound

This paper cites Some pac-bayesian theorems.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Some pac-bayesian theorems

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:38.514186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T10:31:37.813305Z digest=sha256:25addbe78955e3511c3b0979ce471ae1e33ac083ed6f4b4dda815032478ce719

Observation a7d29796-62ed-4555-b21d-81aadb7bb521 · outbound

This paper cites The Llama 3 Herd of Models.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought The Llama 3 Herd of Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.819065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.819065Z digest=sha256:ab149f7513477d3112a6678e39a9f7b403af62c685a57bec77f43a9cc9ff4ec7

Observation 99b43959-d9de-4c31-8bb8-631a3602bb7c · outbound

This paper cites Nearest neighbor distance in three-dimensional space.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Nearest neighbor distance in three-dimensional space

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:38.502802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T10:31:37.823520Z digest=sha256:36b9c76ff9524fc18939e49b807c0b576c7c6248da4dfd79975facb4d231db6a

Observation 18994fd2-58ee-43b0-8eaf-e72069020449 · outbound

This paper cites Learning to reason with llms, 2024.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Learning to reason with llms, 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:38.490571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T10:31:37.827688Z digest=sha256:a97a784d131a822bfe64a2ec976d356a02559e24e5ae125f0ff3770500460864

Observation a947c17b-463c-415e-bf3c-2bae4b76115f · outbound

This paper cites Introducing openai o3 and o4-mini, 2025.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Introducing openai o3 and o4-mini, 2025

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:38.477629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T10:31:37.831809Z digest=sha256:d9e155511d6260ed919af6e9dfdd8e3e586dd94749ea6fa8d0f66f4fe06ea921

Observation a54cd083-fb2f-4c11-8b4a-c49510dbb830 · outbound

This paper cites Qwq: Reflect deeply on the boundaries of the unknown, November 2024.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Qwq: Reflect deeply on the boundaries of the unknown, November 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:38.464167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T10:31:37.836030Z digest=sha256:5c9424e737011f7d411987c29c27c04707702c5d90d7a13a351b7cff3c42255c

Observation 6e0006be-46fd-4878-b922-e522d5a14d00 · outbound

This paper cites Qwen2.5 Technical Report.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Qwen2.5 Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.840224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.840224Z digest=sha256:1393320cefa0862b58891614a62603b66aba90533d45befb92250cffda6b0e53

Observation f30b8c27-7c1d-4049-a9f9-9ec2465ce29f · outbound

This paper cites Qwen3 Technical Report.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Qwen3 Technical Report

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.844935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.844935Z digest=sha256:db0f9bf06f969928b6832a550f210674d7d03aa231e75e63ee0e98d46bcbcffa

Observation ed7d4204-0de5-4b55-ae19-7960911b9773 · outbound

This paper cites Benchmarking prompt sensitivity in large language models.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Benchmarking prompt sensitivity in large language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:38.451294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T10:31:37.849362Z digest=sha256:9a17877b495c95d686d9c2fef6f43e580e22b769a0d1a1d8b6fffc15ca203e5d

Observation 2571ea31-6b53-4441-9afc-b0e244fadcb5 · outbound

This paper cites How much does your data exploration overfit? controlling bias via information usage.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought How much does your data exploration overfit? controlling bias via information usage

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:38.438143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T10:31:37.854081Z digest=sha256:f59074d40d4c3161cf5bd4264ddb754ae1c8d6f66f95998a221a355834326e2f

Observation e3cad880-d189-4e67-bcf5-def90fd559b8 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Proximal Policy Optimization Algorithms

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.858445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.858445Z digest=sha256:159d8af1e6326bfa721f4b7ccdb3b92b59233fa5aa01be9c00dfa1a8e6a00d4f

Observation c3121a39-00b2-46a5-a1ab-654f921f914f · outbound

This paper cites Scaling Test-Time Compute Without Verification or RL is Suboptimal.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Scaling Test-Time Compute Without Verification or RL is Suboptimal

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.862981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.862981Z digest=sha256:30f09d0ebfd0ca19ccbf2fb7ea0597871501da7fec6681ad1b46b46333b645d0

Observation 70d53b55-92e9-418f-8b2e-2d0617b2bff0 · outbound

This paper cites Spurious Rewards: Rethinking Training Signals in RLVR.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Spurious Rewards: Rethinking Training Signals in RLVR

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.869115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.869115Z digest=sha256:615ed7c9ce5e1522240aa3cdbf843ecc61f6963440199f449c39154fb0215201

Observation ddd1d032-1732-4740-a8a4-bf1983309fac · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.873877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.873877Z digest=sha256:00e52ecb5041590111b2001a95dce37e9bc39c928cf8927756e13467592f122c

Observation a67f407a-8b36-4419-adf6-781640bcc753 · outbound

This paper cites A bayesian perspective on generalization and stochastic gradient descent.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought A bayesian perspective on generalization and stochastic gradient descent

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.879593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.879593Z digest=sha256:4d366ce9b6708721c1a559f616d46bc3bc3bae253e27984a8f3ef0de3d75b6da

Observation 21dc2f48-eff6-498b-8b74-69cc9287bd3b · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.884342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.884342Z digest=sha256:8d3c32ec54a42c317e82b4aa3156446218ebacb51d67eeb33514bcf63ed3ce9f

Observation b671102f-5d38-4cf1-801e-336e7c3bd280 · outbound

This paper cites Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.889298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.889298Z digest=sha256:46fe77c26fe15d3893a1e9ae1c386786e757c3fe2101b76e08f0e835855c95bc

Observation b5bd238d-19cf-4f4f-afa0-29c1ecc4ad53 · outbound

This paper cites Understanding Chain-of-Thought in LLMs through Information Theory.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Understanding Chain-of-Thought in LLMs through Information Theory

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.894683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.894683Z digest=sha256:00b20eb8e1df9c4f7aa7cd11048e5e7192f69474faadfece06f2162a1963d290

Observation fcb68915-e810-4f07-9ce4-755df89ed3d5 · outbound

This paper cites Alphazero-like tree-search can guide large language model decoding and training.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Alphazero-like tree-search can guide large language model decoding and training

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:38.418602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T10:31:37.899762Z digest=sha256:9368826e9bc5270120fbb2fdfeca3536ad64a102c4c0951a448467802ee57598

Observation f9dfb13e-8c84-4749-abd2-807892557a2e · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.904517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.904517Z digest=sha256:83ba628b3f40a6612b9aca83c3a5a0b43cf1374b7caf998db0cbd16bf139a128

Observation 4f90a21e-332b-47fe-8db8-77670c608173 · outbound

This paper cites Emergent Abilities of Large Language Models.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Emergent Abilities of Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.909588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.909588Z digest=sha256:398ee1c27cd9792fd254c539893d11354d99c5baa059b33c157ccec7da205581

Observation 5f0d9a22-a300-49fa-b531-53970135c7ae · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Chain-of-thought prompting elicits reasoning in large language models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.914448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.914448Z digest=sha256:bbf2e8e84b637acddf986fbce6d95af45351c39b00308cdba0c606510af7f5e3

Observation d155c5de-8c81-4195-a9f0-170c5792b7e6 · outbound

This paper cites Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.919359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.919359Z digest=sha256:ac9007b597affa976373632efddc4963ac833e04074a51ee6c00e5952dbe52ad

Observation 295e2ff3-2aef-472a-9343-23a1c135d222 · outbound

This paper cites Information-theoretic analysis of generalization capability of learning algorithms.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Information-theoretic analysis of generalization capability of learning algorithms

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:38.398678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T10:31:37.924011Z digest=sha256:feecea4534303abc476931ab5f066a94c9dd3af41d63ba9735682acea18c1035

Observation f010506b-53ab-4da0-8a04-855bf8c5800b · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Tree of thoughts: Deliberate problem solving with large language models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.928512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.928512Z digest=sha256:afdf6bb949e39a887a01a18e38efb90603b94b2b5fddd863afe9e091413d3566

Observation 25fff4de-bd51-413e-97d0-775fd7de3d63 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.933115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.933115Z digest=sha256:0a17ff84e253ac52d914347ddbb410a85aa33af3ce634bb0cbc4c1eef5bceba0

Observation 5c9667f0-812e-4e67-8d41-0e8f7003c117 · outbound

This paper cites What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.938067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.938067Z digest=sha256:6434fe9ecb6b8a09a94b1eb189b066fe4dc717918b65320a3f39819185882653

Observation c7f1001c-2bab-4f60-8465-b46f3d0069de · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.942705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.942705Z digest=sha256:c112b53a2cdf0264c62f591418801810bfca9c7ed2035520c8842a97d715fbed

Observation 97bdd282-cc34-4964-b295-0a99b8da6273 · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.947268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.947268Z digest=sha256:946194458b0f69d5c0f031c39e4ef0695db9e610cd371f5a93099d5faac7a251

Observation 1d10c2c0-d047-4c30-b3ad-a1a73d9121db · outbound

This paper cites Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.951774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.951774Z digest=sha256:d17859164b6d89c53e2df62a62dccf2867453f717a4751bd0c9d6fbe61f1ea15

Observation 6174444d-a37b-45bc-95a6-2838c75b832c · outbound

This paper cites Rest-mcts*: Llm self-training via process reward guided tree search.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Rest-mcts*: Llm self-training via process reward guided tree search

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.956608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.956608Z digest=sha256:b93796b08afcab3a5bb05f5cb0161be1959a15bc5c998dec04bbdb68ebf9fbb7

Observation cbe4cbfe-156e-46e5-874d-4a5c2542018f · outbound

This paper cites Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.961065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.961065Z digest=sha256:c5b974c0c87ea6601d12b8a1871a344e98388cb270338dc4b2ec23b38b53632d

Observation 238330f5-2f6a-4340-aa8c-b4321bac6f44 · outbound

This paper cites Prosa: Assessing and understanding the prompt sensitivity of llms.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Prosa: Assessing and understanding the prompt sensitivity of llms

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:38.367597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T10:31:37.966101Z digest=sha256:7182bddbf568eb790561c0e1109a44496befc4c4fe6eaae0fc37d7cee3f3443d

Observation fa14debb-98ee-4397-88ac-456b940528a8 · outbound

This paper cites @esa (Ref.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought @esa (Ref

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.971079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.971079Z digest=sha256:450b3f2ba433ada62b1179bcecb166581518a06f5e9eaaf4981859be7c6b995c

Observation 9c1cc768-96eb-40c5-9962-4f4079d4f000 · outbound

This paper cites an unresolved cited work.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.976071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.976071Z digest=sha256:ae5bfaed42213d4c80df7d6cd1420879b739378dae851783674cebc18efcb076

Observation 2db63415-3ea4-479d-9bbd-00fd28ff3b2b · outbound

This paper cites an unresolved cited work.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Unresolved cited work

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.980968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.980968Z digest=sha256:88b5da5ddc263db0097d398133ee98b9bb6f52d192834b79e76e6021f4b27135

Pith citing papers

Observation f8bb9d9c-cda0-4a4f-8558-1a01fd930ed4 · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought

Reference 152

Resolution
verified exact
arxiv_id, observed 2026-06-05T02:16:19.319482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:cd5b37ddc69ec78d9e44437a6094f9ff44e3e1ce486cb66ee238b5d31403b835