Pith. sign in

Paper Citation Record · LEDGER

LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization

As of 9 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 3 inbound Pith citation observations for arXiv:2507.15758.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.15758 v2

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:29:44.311377Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T17:04:38.759385Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T19:58:53.662557Z

Reference resolution

28 of 28 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 63efb1d1-3055-422a-bacc-3f2a525f4774 · outbound

This paper cites L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning.

LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:42.182628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:42.182628Z digest=sha256:028a2014236a74bf40570201bad19f479314c776585e98db40386e431c0c8a53

Observation 0c2ae177-ebe1-478c-a172-115c2560f6d9 · outbound

This paper cites Thinkless: LLM Learns When to Think.

LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization Thinkless: LLM Learns When to Think

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:42.379745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:42.379745Z digest=sha256:0fea0eef1349af663677568b28df5a2019238ac16b5c64f475fefb4142e235ac

Observation fc79b765-1550-4f56-949e-23c0b32a0c8c · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:42.446317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:42.446317Z digest=sha256:5d74b3704704d9c91d304bdea303c240ce3c91a9464550b3c5bffe352eaa244d

Observation 615c17f0-de06-40d5-aa00-90439ad8af3b · outbound

This paper cites Token-Budget-Aware LLM Reasoning.

LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization Token-Budget-Aware LLM Reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:42.490719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:42.490719Z digest=sha256:00641ce0e012b344b506193bdbdc365a6f2e256981f08cc1cbbaf0b6839e54fa

Observation 3b866691-815f-4263-bdb7-5c74eda232b8 · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:42.574658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:42.574658Z digest=sha256:fda41b289eeb99a27667e2a854c1501dd5f1336878adfd98bd17b91b9ec75c09

Observation ffc127bb-2f9d-4e59-97fc-161b990fc451 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization Measuring Massive Multitask Language Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:42.642768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:42.642768Z digest=sha256:0980452e0c4e47e7e619899ca47eda0678b395d191cf1311939d54adf6c398f0

Observation 3c80df40-f92a-4d48-9390-0400681285f8 · outbound

This paper cites Hapo: Training language models to reason concisely via history-aware policy optimization.

LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization Hapo: Training language models to reason concisely via history-aware policy optimization

Reference 10

Resolution
verified exact
doi, observed 2026-08-06T15:29:44.660422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:29:42.818923Z digest=sha256:c95f05e7abe3205e74e352801d4b624ccd9fa61196d6ce72e6dbca2881b4880b

Observation 5628c428-fb7c-4db6-843d-608cd7a3b9a0 · outbound

This paper cites Overthink: Slowdown attacks on reasoning llms.

LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization Overthink: Slowdown attacks on reasoning llms

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:42.907541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:42.907541Z digest=sha256:ec645f9c454755c37ee0942652ca3b8767b9bc497a9d1e067721b4eb04d02e04

Observation 40826666-312f-482b-9b61-0a3ca07a52ce · outbound

This paper cites AdaCoT: Pareto-Optimal Adaptive Chain-of-Thought Triggering via Reinforcement Learning.

LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization AdaCoT: Pareto-Optimal Adaptive Chain-of-Thought Triggering via Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:42.965888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:42.965888Z digest=sha256:7f9abd859137969b4c7f3a92c41a4109ecbe44d3f0cae81b7399def02bd50be1

Observation 4e7f71cc-cee9-4d46-84ad-e4eb1a61b23c · outbound

This paper cites O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning.

LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:43.034909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:43.034909Z digest=sha256:679b9f03fd3de589d9c3d30722b924cb541c585df967b1eaa401812b8d28f606

Observation 031ce5b9-3f8a-48e5-a7ff-64797412ca69 · outbound

This paper cites Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems.

LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:43.103227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:43.103227Z digest=sha256:a70e986b291607bce7d1773641c20169bd6ff9acf9a902fe4c262445d7d9aca9

Observation fe021649-68fc-4f31-969b-6cb2d99e2b6e · outbound

This paper cites s1: Simple test-time scaling.

LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization s1: Simple test-time scaling

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:43.152143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:43.152143Z digest=sha256:4644ed6302e8fec0bf22a170729e4087a45777d8e2cd9e23f17687a44e7fb556

Observation 1ca9559b-1e6c-4fd5-938d-e6263e64c308 · outbound

This paper cites RouteLLM: Learning to Route LLMs with Preference Data.

LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization RouteLLM: Learning to Route LLMs with Preference Data

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:43.237643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:43.237643Z digest=sha256:2a190cff31c0978fbe50147fc9b6c70a75485c8f9f4c0c4173f8601351ad62ac

Observation bd09387e-8315-438e-b042-0a555fa54d84 · outbound

This paper cites Concise: Confidence-guided compression in step-by-step efficient reasoning.

LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization Concise: Confidence-guided compression in step-by-step efficient reasoning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:43.338210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:43.338210Z digest=sha256:50379a62de14188d89218a1b900ef8a6afbfc9b8eca8a61a7521fb93f3f363c9

Observation 64582e5a-def5-485d-9a17-831dec3742b0 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:43.402800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:43.402800Z digest=sha256:dd2b8a03bc38db0fa62ad0f016661c761645ef78adbde8b1b9183ccdfeda5c1f

Observation 16de9413-fb86-45b2-9678-08447fd5fd0d · outbound

This paper cites Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models.

LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:43.470138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:43.470138Z digest=sha256:5c5a7cd37b7cb4297d71335208ffbb50981e8c0b84a51b0476909c5c9330a411

Observation 970106e8-2575-4f0f-b7b5-89e0db836bf9 · outbound

This paper cites Learning when to think: Shaping adaptive reasoning in r1-style models via multi-stage rl.

LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization Learning when to think: Shaping adaptive reasoning in r1-style models via multi-stage rl

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:43.549175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:43.549175Z digest=sha256:ba8bc1b45acd78480a5358991776679d1b021de93b9f84b9af9e891afedbbd3d

Observation d5d660cd-42a0-4987-9d75-5cd43b6ef237 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:43.658272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:43.658272Z digest=sha256:3bebddff10ca9a62650aca5b3a6864f50ec6bb62e45961f7f2a1851cb809ef78

Observation 88529ffc-21a4-4fa3-a5bc-debeba036124 · outbound

This paper cites From Decoding to Meta-Generation: Inference-time Algorithms for Large Language Models.

LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization From Decoding to Meta-Generation: Inference-time Algorithms for Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:43.802572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:43.802572Z digest=sha256:5e2251f793d20164445bf1a6ae2893ca0d0293c38803373c34da60227e8ec566

Observation 2f248839-594e-4aa9-a5cc-a8be95e04fe2 · outbound

This paper cites Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models.

LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:43.891579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:43.891579Z digest=sha256:025d672e2664c814dacbe28423db01dcfac37032a34616a0818e2dc18edb6f9a

Observation 63e39cad-253f-40a9-9eae-35936d284e34 · outbound

This paper cites Tokenskip: Controllable chain-of-thought compression in llms.

LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization Tokenskip: Controllable chain-of-thought compression in llms

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:43.955617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:43.955617Z digest=sha256:4345ac78ec8197bd8931fd0f1ae7ed2ac2efc1d8ad6e140c41753291fe8d571f

Observation fdba8db0-4a46-44b5-93ec-e6e4ed21c265 · outbound

This paper cites Chain of Draft: Thinking Faster by Writing Less.

LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization Chain of Draft: Thinking Faster by Writing Less

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:44.051838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:44.051838Z digest=sha256:ffc4b30c39f0c7650b38f46f504b8f1d9289192ced6a6fff3f695a3503a11606

Observation 1a6ea9b0-bdc5-43fc-a0ff-04b2443f1d8b · outbound

This paper cites AdaptThink: Reasoning Models Can Learn When to Think.

LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization AdaptThink: Reasoning Models Can Learn When to Think

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:44.311377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:44.311377Z digest=sha256:cebd827fab605f0eb73e09dbe92ce4c61165d0f20914f4a63a1faf19df76ed50

Observation 429ce97d-9577-489b-b6e0-91645c229703 · outbound

This paper cites ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning.

LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:42.735481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:42.735481Z digest=sha256:85366560a7532b67369b1f37458b0742e0574775280bdfe3873a301640e44521

Observation e228e5ed-4276-4971-b943-3801b8bf9c36 · outbound

This paper cites Thoughts Are All Over the Place: On the Underthinking of o1-Like LLMs.

LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization Thoughts Are All Over the Place: On the Underthinking of o1-Like LLMs

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:43.738162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:43.738162Z digest=sha256:e963ebc699ae9e4c9e57955e1af016bb724ab450a7e9b52f5c1bfcfb78a4bf0b

Observation 4002c37a-0be0-465f-ae43-b3bb56358af2 · outbound

This paper cites Demystifying Long Chain-of-Thought Reasoning in LLMs.

LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization Demystifying Long Chain-of-Thought Reasoning in LLMs

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:44.153200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:44.153200Z digest=sha256:1beb2f7548911b634c6a89d424f80156f37c8cabc8af5b22ed7b8dc0b81e0baa

Observation 68ca42b8-4c01-42fa-b60b-a0580f4d084d · outbound

This paper cites Learning to Route LLMs with Confidence Tokens.

LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization Learning to Route LLMs with Confidence Tokens

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:42.324127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:42.324127Z digest=sha256:3866ffcf9e7266edb335eb01441fb5f2979dbd4118dff5e808dc0575a99d6fe6

Observation 7167743f-36ee-4791-80e2-9212c2d9a929 · outbound

This paper cites Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs.

LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:42.254233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:42.254233Z digest=sha256:8adb66d10051ee72ff9fc1381ef43ac763ec21de298171589600d18afad726c1

Pith citing papers

Observation 2fe9cc27-c640-41c8-b2d0-218cc82279ba · inbound

BudgetThinker: Empowering Budget-aware LLM Reasoning with Control Tokens cites this paper.

BudgetThinker: Empowering Budget-aware LLM Reasoning with Control Tokens LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T17:04:38.759385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:04:38.759385Z digest=sha256:ec07dbefc441c5758844e30521d24d4d0c32bb433ed6a593e98a215a936cf96c

Observation 79333e2d-250d-4ace-aee3-0a39cbf8206b · inbound

StaRPO: Stability-Augmented Reinforcement Policy Optimization cites this paper.

StaRPO: Stability-Augmented Reinforcement Policy Optimization LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:21:00.059202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T18:11:49.056805Z digest=sha256:a24418d857c74c7f8a7a46546ee7fef90b4ab9d9dad76047eae8ecb61bdac61f

Observation 66042c5c-5030-4a85-830e-16688a93f121 · inbound

Overthink-Triggered Slowdown Attacks on LVLM-Based Robotic Systems cites this paper.

Overthink-Triggered Slowdown Attacks on LVLM-Based Robotic Systems LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-03T19:58:53.664207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-03T19:52:11.018335Z digest=sha256:e66304e36c41561412e07d216782603ef215aad63189784f1ea5c98aaad6c5d9