Pith. sign in

Paper Citation Record · LEDGER

Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

As of 18 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 30 inbound Pith citation observations for arXiv:2506.05256.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.05256 v2

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:28:31.536924Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 30 of 30 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:50:30.707660Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T18:40:03.297902Z

Reference resolution

17 of 17 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1c2e65a2-a26b-454f-a04d-17ba24728abd · outbound

This paper cites Xingyu Chen, Jiahao Xu, Tian Liang, Zhiwei He, Jianhui Pang, Dian Yu, Linfeng Song, Qiuzhi Liu, Mengfei Zhou, Zhuosheng Zhang, Rui Wang, Zhaopeng Tu, Haitao Mi, and Dong Yu.

Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning Xingyu Chen, Jiahao Xu, Tian Liang, Zhiwei He, Jianhui Pang, Dian Yu, Linfeng Song, Qiuzhi Liu, Mengfei Zhou, Zhuosheng Zhang, Rui Wang, Zhaopeng Tu, Haitao Mi, and Dong Yu

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:30.048003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:30.048003Z digest=sha256:19623d33e689e94033fad848bcb050d76ac4107af23764cc089c91f68748504e

Observation 21756724-5271-4197-82a1-57ed4f1d5c21 · outbound

This paper cites Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs.

Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:30.149754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:30.149754Z digest=sha256:56993d063b110a385c2df56451581f0d2b68ecc70123c7874d0ba298c8d53280

Observation d7fd6e85-6f0c-4ed1-88e5-0df5bdedf482 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:30.270494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:30.270494Z digest=sha256:92db4810d1b62794dcd5d73d4d12d88f96c5a6faa6bc6275f832bdf6c9032ef5

Observation c17b38e5-69a3-4707-bee7-f1d9912d501c · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning Measuring Mathematical Problem Solving With the MATH Dataset

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:30.467637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:30.467637Z digest=sha256:2e400b4bb1c493a76e81bba34de844c66ab5def43d3fdd9423eadcfc8cbe0c89

Observation 506a3c43-7cb5-4e58-9910-0f3a7236edaf · outbound

This paper cites ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning.

Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:30.568819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:30.568819Z digest=sha256:ab88ef61c2a530b6dc52f8e98ee313fed9f9f8a9df65cc54eec5c3ffa5b48655

Observation 6266d389-f09d-486d-90b0-605c95de981e · outbound

This paper cites Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems.

Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:30.642928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:30.642928Z digest=sha256:ebc033627624a6409bce770ebb2cd04de079843a1a6564567ce005553d3ad34f

Observation 3b9d6cb8-82a1-4aef-85bd-e8ee20e8affa · outbound

This paper cites s1: Simple test-time scaling.

Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning s1: Simple test-time scaling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:30.721513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:30.721513Z digest=sha256:7c72f8e35b9b8ea27fe1d7f0b63d3c2bfa36cc83a1b7d007d91bd943e816ea08

Observation fc219677-9c0c-4254-ba77-9456d1e1314d · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:30.799790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:30.799790Z digest=sha256:25c5908cb3237f64a356ee2630de3b163807252f87a368aff0e7cfac77a761ea

Observation aa4c12c2-f0d5-4d45-a420-c80e7d78f85f · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning HybridFlow: A Flexible and Efficient RLHF Framework

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:30.901007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:30.901007Z digest=sha256:0c356229abd72672eed8d05f80cb0573712d544133cc0bc4e48ead5bee73471b

Observation 90f3d839-d2ed-47a8-b0f0-16b8717c79e3 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:30.991793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:30.991793Z digest=sha256:b77a921c759fe57dd26f057bb41fdd26b034be5ffce28d145238f05237a915ac

Observation 73b43e58-27e3-456e-afad-2512004f8b5c · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:31.092526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:31.092526Z digest=sha256:459b71ec4dc66356645a76dfbf2a14d0e694ce38eee25b368ae12e326b56553b

Observation 49c3822b-ec35-401e-a7dc-565487c5fc36 · outbound

This paper cites Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models.

Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:31.175370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:31.175370Z digest=sha256:9bf5f36f37720133e1afa4809b0f9475206a4f1b1a5dc4da061cb2bcdeed1a13

Observation 79876e20-bc1d-4a1e-be92-fa16de0f8d84 · outbound

This paper cites Towards System 2 Reasoning in LLMs: Learning How to Think With Meta Chain-of-Thought.

Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning Towards System 2 Reasoning in LLMs: Learning How to Think With Meta Chain-of-Thought

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:31.296126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:31.296126Z digest=sha256:6826ac22066239078b3066f4b463bea4d061407d12089a9e42223389df5875df

Observation f65d5de8-86c7-4087-8689-c8d1d2d00d61 · outbound

This paper cites Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning.

Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:31.403312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:31.403312Z digest=sha256:74f4096dcf326540e69e971cd6abc6d6ab6a3506236b15872e4c2e7cc9bebe2d

Observation 58daa4dc-cb8d-4f60-915a-c311ae2b4f9c · outbound

This paper cites Backtracking.

Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning Backtracking

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:32.170425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:28:31.536924Z digest=sha256:66229735c76326178f3f3a30cd9d5e9e36734825d95ed54756325cbddfc1223d

Observation 419cdef7-efd5-4303-9654-5891a7af65a2 · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:30.367732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:30.367732Z digest=sha256:00c26ca329c5c5c3ee6a6fb8a4ab7838f8b0b32e1b2bdb917f844f722ce0a5ed

Observation 99687c21-cddc-4122-ad45-a82abd54e362 · outbound

This paper cites L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning.

Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:29.980350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:29.980350Z digest=sha256:b5e7739da912e81e0bd4187af0aa0d5c2d35639fa3c0054e6d5b120bc862be6c

Pith citing papers

Observation c090e580-0bf6-4349-920f-77f28307ce44 · inbound

Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey cites this paper.

Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

Reference 208

Resolution
unresolved
no resolver link, observed 2026-08-06T17:54:17.890896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:54:17.890896Z digest=sha256:a8d8766bb865387a20bafbf3bc261b9afdea72567d931e52de6df13bb601fe58

Observation 84067768-e4ab-44df-a498-07d6b4bca3d6 · inbound

Verifying Meta-Awareness via Predictive Rewards in Reasoning Models cites this paper.

Verifying Meta-Awareness via Predictive Rewards in Reasoning Models Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T15:50:30.707660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:50:30.707660Z digest=sha256:20d1dd5625de3b9dcab17ca94a0c4cd7f4173d197d6c1e3fb0cbc34adff65be2

Observation 16630b78-1342-41b9-9ed6-dc4aacfcc365 · inbound

Learning to Reason Efficiently with Discounted Reinforcement Learning cites this paper.

Learning to Reason Efficiently with Discounted Reinforcement Learning Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T07:59:54.120284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:59:54.120284Z digest=sha256:61922be1184d36e6b1fbc076a3772c2a90dcd01bad08f95c9ff3be90111bef27

Observation 4e8c6e8f-b5e7-42ac-8320-38c74b5bfdf1 · inbound

Schoenfeld's Anatomy of Mathematical Reasoning by Language Models cites this paper.

Schoenfeld's Anatomy of Mathematical Reasoning by Language Models Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T20:41:15.359772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T20:39:34.858032Z digest=sha256:5013c4775e2b96961675bbaf089e869e87efdb37ade417c8cca1f9ca55b7d1d0

Observation 4200b067-bc3b-474c-96d4-a8b2e435ba8a · inbound

CASHEW: Stabilizing Multimodal Reasoning via Iterative Trajectory Aggregation cites this paper.

CASHEW: Stabilizing Multimodal Reasoning via Iterative Trajectory Aggregation Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T11:02:10.268276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:02:10.268276Z digest=sha256:b045730698ef999a091b9d1fbe90b6b218d1d228c583f97874d79ed054311d83

Observation 39b95d97-b81b-4681-bedc-cdd7eba02034 · inbound

Neural Chain-of-Thought Search: Searching the Optimal Reasoning Path to Enhance Large Language Models cites this paper.

Neural Chain-of-Thought Search: Searching the Optimal Reasoning Path to Enhance Large Language Models Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T13:22:55.264368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T13:21:36.606855Z digest=sha256:9632ec26e2cfcfdb53e3bbe1096ae581c26036a63a695b84ae8ef8125902403d

Observation 713381ab-65ec-40bf-a8ce-7c0fa3ade253 · inbound

On the Optimal Reasoning Length for RL-Trained Language Models cites this paper.

On the Optimal Reasoning Length for RL-Trained Language Models Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T02:53:07.791956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:53:07.791956Z digest=sha256:1659435dd40696f429cb3acafc522644d38fa1acdc4f47ac7fd0eaa74908c9e2

Observation a4a96d6c-7bae-4897-91f2-4885d260985a · inbound

ATTNPO: Attention-Guided Process Supervision for Efficient Reasoning cites this paper.

ATTNPO: Attention-Guided Process Supervision for Efficient Reasoning Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:47:10.890611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T02:44:18.723713Z digest=sha256:e54d883c5029c616adb396f49ea507897ae6122502af65a4e7245cb506c337ca

Observation 886e77fa-05e3-4eae-8cca-de6899a4bc07 · inbound

CODA: Difficulty-Aware Compute Allocation for Adaptive Reasoning cites this paper.

CODA: Difficulty-Aware Compute Allocation for Adaptive Reasoning Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:25:55.461203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T14:23:00.793443Z digest=sha256:58a34af6c36783cc45c8621710d9711a18d034cf3d86dc57210012b61c672db8

Observation 84601f54-4850-4575-a90c-73210dcff0f1 · inbound

Length Value Model: Scalable Value Pretraining for Token-Level Length Modeling cites this paper.

Length Value Model: Scalable Value Pretraining for Token-Level Length Modeling Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:31:26.152468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-07T10:42:27.644514Z digest=sha256:daa5c1637264fb9aea15773dfaf42a37662331a634123d7b86bc0622d9ae3ce5

Observation 344b53f1-e67c-445a-b1a3-61e04d7501c6 · inbound

Length Value Model: Scalable Value Pretraining for Token-Level Length Modeling cites this paper.

Length Value Model: Scalable Value Pretraining for Token-Level Length Modeling Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T15:28:13.138448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T15:28:13.138448Z digest=sha256:56a8e3dbe5f9b50789eda6fc66efa0812acc0ea9ff1eddd93ac352b261587c09

Observation 724ea3d7-2db4-4f6c-b903-a7e256621619 · inbound

ZAYA1-8B Technical Report cites this paper.

ZAYA1-8B Technical Report Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

Reference 227

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:26:04.968571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-08T17:36:37.182196Z digest=sha256:097d61b38bf4fca88809819bfe0db01657f39d3921970b8ce7a4f7da59ac6668

Observation 0933af8a-9e79-4afc-921b-197b5fea19cc · inbound

DUET: Optimize Token-Budget Allocation for Reinforcement Learning with Verifiable Rewards cites this paper.

DUET: Optimize Token-Budget Allocation for Reinforcement Learning with Verifiable Rewards Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:46:28.454063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T01:57:11.065744Z digest=sha256:9f1c27844d669ab4b84143148f17385fadaaf2017d2475b6340bc9112ec2cdf5

Observation df06833a-2720-4f97-a639-0fa242ebb1a6 · inbound

LEAD: Length-Efficient Adaptive and Dynamic Reasoning for Large Language Models cites this paper.

LEAD: Length-Efficient Adaptive and Dynamic Reasoning for Large Language Models Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:31:16.760115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T02:30:42.407934Z digest=sha256:71d4c8a89367da79c57978969e382746ca43731799ac7c4c344048719b4f559c

Observation deecccc9-3022-4352-bddd-5df8ab3f24d5 · inbound

Taming the Thinker: Conditional Entropy Shaping for Adaptive LLM Reasoning cites this paper.

Taming the Thinker: Conditional Entropy Shaping for Adaptive LLM Reasoning Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T06:43:05.955153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-20T06:40:06.103206Z digest=sha256:292348fe4e3ba51865d03073c33b5b93a950c64eb3492eb895b3098b9c1849bd

Observation 07878212-db95-4bed-acb5-058a6488d63c · inbound

CLORE: Content-Level Optimization for Reasoning Efficiency cites this paper.

CLORE: Content-Level Optimization for Reasoning Efficiency Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:51:08.348019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T05:50:23.111591Z digest=sha256:15c1666b8ad22312b025a0c3e433feaf438de6b6658909993678b5cf7d8d08d7

Observation b7ff19b3-3818-4609-aad2-73facda3bc15 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

Reference 243

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T20:56:13.590565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:42a448aa65110ce7736c7d6d4f4d3d78c0e7bd39eceac40ed51485e7575eb5f3

Observation bae9feb9-53fe-44eb-96f0-cc4d667df1c3 · inbound

Libra: Efficient Resource Management for Agentic RL Post-Training cites this paper.

Libra: Efficient Resource Management for Agentic RL Post-Training Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:16:27.163309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T11:02:00.385932Z digest=sha256:85156cc0bfa20ca9aaf8dd7210661c1f52c5274c163b244a54abf02128df2b3f

Observation 57c65ba3-1f45-48aa-acd3-0038a3bd3b17 · inbound

From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning cites this paper.

From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

Reference 82

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:38:56.069868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-27T01:13:11.483599Z digest=sha256:dcd4da39d437e7f5e694e57e3b8f02f09c8f225c7b298555dcf48c0b13523c23

Observation 14dfd15c-1b8c-4bd1-bbbb-bea699947083 · inbound

Token-Operations-Oriented Inference Optimization Techniques for Large Models cites this paper.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

Reference 88

Resolution
verified exact
arxiv_id, observed 2026-07-04T05:09:36.866760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T16:15:22.543601Z digest=sha256:0a0c1a5e85d31b45255d3f76f192eb517d0e23896159e76e4d3e95143b3ea0aa

Observation 3e778285-b681-44bd-9652-4e0890b5ae65 · inbound

Token-Operations-Oriented Inference Optimization Techniques for Large Models cites this paper.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:07.680058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:07.680058Z digest=sha256:9aada7a376ace7c2150cd71c3ef157ca68ea9194fe8fa26a21f0aafb20797ac3

Observation a8dac336-9aea-4217-932d-7a2c3366b668 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

Reference 231

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:59:40.698593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:b8316e21abb4d5babffaad43f5175a2b3eb2bbacf6ac0282fdb877e29198986b

Observation 124265fb-a462-4763-ba7b-0bd2a6e40fdf · inbound

Beyond Penalizing Mistakes: Stabilizing Efficiency Training in Large Reasoning Models via Adaptive Correct-Only Rewards cites this paper.

Beyond Penalizing Mistakes: Stabilizing Efficiency Training in Large Reasoning Models via Adaptive Correct-Only Rewards Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T09:19:43.879965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-26T10:12:30.692295Z digest=sha256:ed33caf9958c2f7b61b0fa8fea8c7bd76a3083fead6c5db27ea7e7e1203ce3af

Observation 3742ff11-eeed-4114-8244-565d834dc2c9 · inbound

ZONOS2 Technical Report cites this paper.

ZONOS2 Technical Report Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

Reference 268

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T18:40:03.299162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-25T22:37:15.072758Z digest=sha256:c9fbfdd131612f8972711d498e5b37cec2dcaa42e8fbb5f274af35a85a2f6dc8

Observation 59b6edb1-7561-488e-8233-1e658a261cf1 · inbound

ZONOS2 Technical Report cites this paper.

ZONOS2 Technical Report Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

Reference 268

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T18:15:59.049118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-29T02:07:31.791835Z digest=sha256:678ab366582a103fd9b0933babb72484a1e4e66af1c3809c856de2f032487190

Observation 25ff8107-92f1-4eb6-a063-61c8481c02df · inbound

Length Penalties Make Chain-of-Thought Less Monitorable cites this paper.

Length Penalties Make Chain-of-Thought Less Monitorable Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

Reference 80

Resolution
unresolved
no resolver link, observed 2026-07-14T15:45:54.532529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T15:45:54.532529Z digest=sha256:bca7c1fd05fc2e3c665f1f4b94d6d49ed88b6a90e0b1a015e5836377f7aba237

Observation 72a9e867-bdbe-411f-8fa7-12cecb22928b · inbound

Length Penalties Make Chain-of-Thought Less Monitorable cites this paper.

Length Penalties Make Chain-of-Thought Less Monitorable Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-02T08:06:10.856964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:06:10.856964Z digest=sha256:bf4a97f5e5920c95fce5e45d348841cd49f7edbbc25739c281ae31cc7f621dae

Observation 1c6f5859-2199-4f0b-a129-13d2f68e69a2 · inbound

Length Penalties Make Chain-of-Thought Less Monitorable cites this paper.

Length Penalties Make Chain-of-Thought Less Monitorable Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T04:30:30.198277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:30:30.198277Z digest=sha256:3000a29dc1600bd22686aef2572565b62be0b645a0660878299818b716502704

Observation 22d43118-764f-4480-89c7-77bb223199be · inbound

ZUNA1.1: A more flexible EEG foundation model for Denoising and Super-resolution cites this paper.

ZUNA1.1: A more flexible EEG foundation model for Denoising and Super-resolution Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

Reference 237

Resolution
unresolved
no resolver link, observed 2026-08-01T09:52:11.853588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:52:11.853588Z digest=sha256:243cfd8c6ee0c15f05e7c15dcc7a2176b38919044db14afb759487e063ebe7f3

Observation db95060a-dc02-4e79-81fc-8f0d19a2f530 · inbound

BAP-SQL: Budget-Aware Observation Planning for Agentic Text-to-SQL cites this paper.

BAP-SQL: Budget-Aware Observation Planning for Agentic Text-to-SQL Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T15:01:59.546960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:01:59.546960Z digest=sha256:3d2dfa9e2d9649569d920218cb61a0f8085de0249f84cf993e9792382327a3c9