Pith. sign in

Paper Citation Record · LEDGER

LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization

As of 18 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 3 inbound Pith citation observations for arXiv:2507.15758.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.15758 v2

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:29:44.311377Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T17:04:38.759385Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T19:58:53.662557Z

Reference resolution

28 of 28 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 63efb1d1-3055-422a-bacc-3f2a525f4774 · outbound

This paper cites L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning.

LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:42.182628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:42.182628Z digest=sha256:035639a79b1280f168267004c6aacafa903e1e06a85ace5ecbe0306628f43a25

Observation 0c2ae177-ebe1-478c-a172-115c2560f6d9 · outbound

This paper cites Thinkless: LLM Learns When to Think.

LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization Thinkless: LLM Learns When to Think

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:42.379745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:42.379745Z digest=sha256:b0c7833fbb0e4346c982817bc6d157a376c5a0e73f3de5a1712c83380e9994a3

Observation fc79b765-1550-4f56-949e-23c0b32a0c8c · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:42.446317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:42.446317Z digest=sha256:6695bbbd997bf39a26fa170cb5a6b7283b3736bc6d5b05c914f00b61534d5a2f

Observation 615c17f0-de06-40d5-aa00-90439ad8af3b · outbound

This paper cites Token-Budget-Aware LLM Reasoning.

LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization Token-Budget-Aware LLM Reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:42.490719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:42.490719Z digest=sha256:eb6fe836a1eaf3ad0c29e9b4d2c29e0ddb1795e1986188d6af26e792def05040

Observation 3b866691-815f-4263-bdb7-5c74eda232b8 · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:42.574658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:42.574658Z digest=sha256:b02d61a62c05e6055f26df605bef3d9d73c35f86a5977c2083c2ba052af49493

Observation ffc127bb-2f9d-4e59-97fc-161b990fc451 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization Measuring Massive Multitask Language Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:42.642768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:42.642768Z digest=sha256:13f61894fbd2d1a908eb73b7f217df1e4269b6422fd271ba8755abc2060f6a4f

Observation 3c80df40-f92a-4d48-9390-0400681285f8 · outbound

This paper cites Hapo: Training language models to reason concisely via history-aware policy optimization.

LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization Hapo: Training language models to reason concisely via history-aware policy optimization

Reference 10

Resolution
verified exact
doi, observed 2026-08-06T15:29:44.660422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T15:29:42.818923Z digest=sha256:4633cd8d2eb523301cde4e736bb90b3871574837975206eb512173bf3bb95650

Observation 5628c428-fb7c-4db6-843d-608cd7a3b9a0 · outbound

This paper cites Overthink: Slowdown attacks on reasoning llms.

LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization Overthink: Slowdown attacks on reasoning llms

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:42.907541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:42.907541Z digest=sha256:0f5eaf0cd9414bd16f9bb994f36c208bc190e6a0ac3e585b875e8e5c0d1f2f17

Observation 40826666-312f-482b-9b61-0a3ca07a52ce · outbound

This paper cites AdaCoT: Pareto-Optimal Adaptive Chain-of-Thought Triggering via Reinforcement Learning.

LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization AdaCoT: Pareto-Optimal Adaptive Chain-of-Thought Triggering via Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:42.965888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:42.965888Z digest=sha256:77968ad48a80eeec3244e73ed2b78896992f05617fc7871dc12fc3512f34d313

Observation 4e7f71cc-cee9-4d46-84ad-e4eb1a61b23c · outbound

This paper cites O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning.

LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:43.034909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:43.034909Z digest=sha256:3239dde5db4ff17c374974e760223b0445d04e006fa3b08da8091d12f147e77f

Observation 031ce5b9-3f8a-48e5-a7ff-64797412ca69 · outbound

This paper cites Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems.

LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:43.103227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:43.103227Z digest=sha256:b343f921ccda7f0833be82afb5bf1805f309e9563ab1fe7b293a8289593d4b56

Observation fe021649-68fc-4f31-969b-6cb2d99e2b6e · outbound

This paper cites s1: Simple test-time scaling.

LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization s1: Simple test-time scaling

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:43.152143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:43.152143Z digest=sha256:7abb6742b37879537b3859030274697701480a74f905960a40a911c28588730f

Observation 1ca9559b-1e6c-4fd5-938d-e6263e64c308 · outbound

This paper cites RouteLLM: Learning to Route LLMs with Preference Data.

LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization RouteLLM: Learning to Route LLMs with Preference Data

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:43.237643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:43.237643Z digest=sha256:6d50460a8277a81199e29da0732397f8d671d13123cd867cea947c56679be3bd

Observation bd09387e-8315-438e-b042-0a555fa54d84 · outbound

This paper cites Concise: Confidence-guided compression in step-by-step efficient reasoning.

LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization Concise: Confidence-guided compression in step-by-step efficient reasoning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:43.338210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:43.338210Z digest=sha256:2ca321a63db7b282386772ae0a9ba6716f0e50dc772bafd429e0907143481efa

Observation 64582e5a-def5-485d-9a17-831dec3742b0 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:43.402800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:43.402800Z digest=sha256:d51a1f80d598ff2d48549216a3c2f9ba5efaf3365ddb69253a4a4d5b64942410

Observation 16de9413-fb86-45b2-9678-08447fd5fd0d · outbound

This paper cites Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models.

LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:43.470138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:43.470138Z digest=sha256:8597e2d015bacd9c86d54b72759c203218dfb64cdaa5dc23a744a6ca6dd1c58c

Observation 970106e8-2575-4f0f-b7b5-89e0db836bf9 · outbound

This paper cites Learning when to think: Shaping adaptive reasoning in r1-style models via multi-stage rl.

LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization Learning when to think: Shaping adaptive reasoning in r1-style models via multi-stage rl

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:43.549175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:43.549175Z digest=sha256:d4e830510e62dd007448126df96c164589bb39f0e8db1ca39a67be691cb01dad

Observation d5d660cd-42a0-4987-9d75-5cd43b6ef237 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:43.658272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:43.658272Z digest=sha256:9050f2dd841f5e5a05e292dbda5b96c511289c903310db64dadb0c66fa072d83

Observation 88529ffc-21a4-4fa3-a5bc-debeba036124 · outbound

This paper cites From Decoding to Meta-Generation: Inference-time Algorithms for Large Language Models.

LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization From Decoding to Meta-Generation: Inference-time Algorithms for Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:43.802572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:43.802572Z digest=sha256:b27a1d5ce098fc4d197fb87c692bc65286265573ff202363fec58e012cc3b30b

Observation 2f248839-594e-4aa9-a5cc-a8be95e04fe2 · outbound

This paper cites Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models.

LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:43.891579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:43.891579Z digest=sha256:f354d5ec81cd2549bc51624af093ce7ef3175b994165e00dc08424e183c40f39

Observation 63e39cad-253f-40a9-9eae-35936d284e34 · outbound

This paper cites Tokenskip: Controllable chain-of-thought compression in llms.

LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization Tokenskip: Controllable chain-of-thought compression in llms

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:43.955617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:43.955617Z digest=sha256:46378787a8b19dc5683f9781ce1e80338b3537adaa4e400765e80fda97b9e65b

Observation fdba8db0-4a46-44b5-93ec-e6e4ed21c265 · outbound

This paper cites Chain of Draft: Thinking Faster by Writing Less.

LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization Chain of Draft: Thinking Faster by Writing Less

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:44.051838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:44.051838Z digest=sha256:065b004ce6c6ceea585811b3d2ea1337c6ee9a0bc5e6312deb0d2a6025934352

Observation 1a6ea9b0-bdc5-43fc-a0ff-04b2443f1d8b · outbound

This paper cites AdaptThink: Reasoning Models Can Learn When to Think.

LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization AdaptThink: Reasoning Models Can Learn When to Think

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:44.311377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:44.311377Z digest=sha256:3b23c731809fac8f185f607adb9b04d1f847e9a7965a0fc691df59a65e7d3958

Observation 429ce97d-9577-489b-b6e0-91645c229703 · outbound

This paper cites ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning.

LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:42.735481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:42.735481Z digest=sha256:019f06410f32a42bc08bb2ea78b4f9f441392cc267fa9603cd09fc9242a39485

Observation e228e5ed-4276-4971-b943-3801b8bf9c36 · outbound

This paper cites Thoughts Are All Over the Place: On the Underthinking of o1-Like LLMs.

LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization Thoughts Are All Over the Place: On the Underthinking of o1-Like LLMs

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:43.738162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:43.738162Z digest=sha256:934a2ecd4fd662d683abdbc869250a790d081955c8da3eef913de0ad3201e338

Observation 4002c37a-0be0-465f-ae43-b3bb56358af2 · outbound

This paper cites Demystifying Long Chain-of-Thought Reasoning in LLMs.

LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization Demystifying Long Chain-of-Thought Reasoning in LLMs

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:44.153200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:44.153200Z digest=sha256:09236ed3e891f12d924309fe946b1ab788c021d82ce9562231a2977d6582b8d4

Observation 68ca42b8-4c01-42fa-b60b-a0580f4d084d · outbound

This paper cites Learning to Route LLMs with Confidence Tokens.

LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization Learning to Route LLMs with Confidence Tokens

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:42.324127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:42.324127Z digest=sha256:307217f3aeb06a8746ddf7dd3fbb6bbbdea8e8359089b6208307ec63a2a7dcd5

Observation 7167743f-36ee-4791-80e2-9212c2d9a929 · outbound

This paper cites Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs.

LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:42.254233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:42.254233Z digest=sha256:23c0d76fb0ea8a604e57009ec570ca060f4c39c85e53ecf8debd073ea41b60ef

Pith citing papers

Observation 2fe9cc27-c640-41c8-b2d0-218cc82279ba · inbound

BudgetThinker: Empowering Budget-aware LLM Reasoning with Control Tokens cites this paper.

BudgetThinker: Empowering Budget-aware LLM Reasoning with Control Tokens LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T17:04:38.759385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:04:38.759385Z digest=sha256:d576f9403022d862ecc11f9da28c485c90446e68c45fac31e7840652981b1f28

Observation 79333e2d-250d-4ace-aee3-0a39cbf8206b · inbound

StaRPO: Stability-Augmented Reinforcement Policy Optimization cites this paper.

StaRPO: Stability-Augmented Reinforcement Policy Optimization LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:21:00.059202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T18:11:49.056805Z digest=sha256:1a1c83de915ca790974aec2ca2c67117d8c070e0f4df14a197310bf8917069c6

Observation 66042c5c-5030-4a85-830e-16688a93f121 · inbound

Overthink-Triggered Slowdown Attacks on LVLM-Based Robotic Systems cites this paper.

Overthink-Triggered Slowdown Attacks on LVLM-Based Robotic Systems LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-03T19:58:53.664207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-03T19:52:11.018335Z digest=sha256:2a5396ed0e7c7d27babb70f4a144a606e95aaeee22b01b0a9bb3b59a93ed3ab7