Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:28:31.536924Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 28 inbound Pith citation observations for arXiv:2506.05256.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:28:31.536924Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T17:54:17.890896Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T18:40:03.297902Z
17 of 17 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 1c2e65a2-a26b-454f-a04d-17ba24728abd · outbound
Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning Xingyu Chen, Jiahao Xu, Tian Liang, Zhiwei He, Jianhui Pang, Dian Yu, Linfeng Song, Qiuzhi Liu, Mengfei Zhou, Zhuosheng Zhang, Rui Wang, Zhaopeng Tu, Haitao Mi, and Dong Yu
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21756724-5271-4197-82a1-57ed4f1d5c21 · outbound
Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7fd6e85-6f0c-4ed1-88e5-0df5bdedf482 · outbound
Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c17b38e5-69a3-4707-bee7-f1d9912d501c · outbound
Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning Measuring Mathematical Problem Solving With the MATH Dataset
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 506a3c43-7cb5-4e58-9910-0f3a7236edaf · outbound
Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6266d389-f09d-486d-90b0-605c95de981e · outbound
Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b9d6cb8-82a1-4aef-85bd-e8ee20e8affa · outbound
Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning s1: Simple test-time scaling
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc219677-9c0c-4254-ba77-9456d1e1314d · outbound
Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa4c12c2-f0d5-4d45-a420-c80e7d78f85f · outbound
Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning HybridFlow: A Flexible and Efficient RLHF Framework
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90f3d839-d2ed-47a8-b0f0-16b8717c79e3 · outbound
Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73b43e58-27e3-456e-afad-2512004f8b5c · outbound
Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49c3822b-ec35-401e-a7dc-565487c5fc36 · outbound
Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79876e20-bc1d-4a1e-be92-fa16de0f8d84 · outbound
Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning Towards System 2 Reasoning in LLMs: Learning How to Think With Meta Chain-of-Thought
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f65d5de8-86c7-4087-8689-c8d1d2d00d61 · outbound
Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58daa4dc-cb8d-4f60-915a-c311ae2b4f9c · outbound
Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning Backtracking
Reference 2023
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 419cdef7-efd5-4303-9654-5891a7af65a2 · outbound
Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99687c21-cddc-4122-ad45-a82abd54e362 · outbound
Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c090e580-0bf6-4349-920f-77f28307ce44 · inbound
Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning
Reference 208
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16630b78-1342-41b9-9ed6-dc4aacfcc365 · inbound
Learning to Reason Efficiently with Discounted Reinforcement Learning Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e8c6e8f-b5e7-42ac-8320-38c74b5bfdf1 · inbound
Schoenfeld's Anatomy of Mathematical Reasoning by Language Models Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4200b067-bc3b-474c-96d4-a8b2e435ba8a · inbound
CASHEW: Stabilizing Multimodal Reasoning via Iterative Trajectory Aggregation Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39b95d97-b81b-4681-bedc-cdd7eba02034 · inbound
Neural Chain-of-Thought Search: Searching the Optimal Reasoning Path to Enhance Large Language Models Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 713381ab-65ec-40bf-a8ce-7c0fa3ade253 · inbound
On the Optimal Reasoning Length for RL-Trained Language Models Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4a96d6c-7bae-4897-91f2-4885d260985a · inbound
ATTNPO: Attention-Guided Process Supervision for Efficient Reasoning Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 886e77fa-05e3-4eae-8cca-de6899a4bc07 · inbound
CODA: Difficulty-Aware Compute Allocation for Adaptive Reasoning Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 84601f54-4850-4575-a90c-73210dcff0f1 · inbound
Length Value Model: Scalable Value Pretraining for Token-Level Length Modeling Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 344b53f1-e67c-445a-b1a3-61e04d7501c6 · inbound
Length Value Model: Scalable Value Pretraining for Token-Level Length Modeling Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 724ea3d7-2db4-4f6c-b903-a7e256621619 · inbound
ZAYA1-8B Technical Report Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning
Reference 227
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0933af8a-9e79-4afc-921b-197b5fea19cc · inbound
DUET: Optimize Token-Budget Allocation for Reinforcement Learning with Verifiable Rewards Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation df06833a-2720-4f97-a639-0fa242ebb1a6 · inbound
LEAD: Length-Efficient Adaptive and Dynamic Reasoning for Large Language Models Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation deecccc9-3022-4352-bddd-5df8ab3f24d5 · inbound
Taming the Thinker: Conditional Entropy Shaping for Adaptive LLM Reasoning Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 07878212-db95-4bed-acb5-058a6488d63c · inbound
CLORE: Content-Level Optimization for Reasoning Efficiency Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b7ff19b3-3818-4609-aad2-73facda3bc15 · inbound
Trust Region On-Policy Distillation Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning
Reference 243
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bae9feb9-53fe-44eb-96f0-cc4d667df1c3 · inbound
Libra: Efficient Resource Management for Agentic RL Post-Training Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 57c65ba3-1f45-48aa-acd3-0038a3bd3b17 · inbound
From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 14dfd15c-1b8c-4bd1-bbbb-bea699947083 · inbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3e778285-b681-44bd-9652-4e0890b5ae65 · inbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8dac336-9aea-4217-932d-7a2c3366b668 · inbound
Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning
Reference 231
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 124265fb-a462-4763-ba7b-0bd2a6e40fdf · inbound
Beyond Penalizing Mistakes: Stabilizing Efficiency Training in Large Reasoning Models via Adaptive Correct-Only Rewards Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3742ff11-eeed-4114-8244-565d834dc2c9 · inbound
ZONOS2 Technical Report Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning
Reference 268
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 59b6edb1-7561-488e-8233-1e658a261cf1 · inbound
ZONOS2 Technical Report Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning
Reference 268
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 25ff8107-92f1-4eb6-a063-61c8481c02df · inbound
Length Penalties Make Chain-of-Thought Less Monitorable Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72a9e867-bdbe-411f-8fa7-12cecb22928b · inbound
Length Penalties Make Chain-of-Thought Less Monitorable Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c6f5859-2199-4f0b-a129-13d2f68e69a2 · inbound
Length Penalties Make Chain-of-Thought Less Monitorable Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22d43118-764f-4480-89c7-77bb223199be · inbound
ZUNA1.1: A more flexible EEG foundation model for Denoising and Super-resolution Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning
Reference 237
Source-reported events for the cited work
Unavailable: canonical work link unavailable.