Pith. sign in

Paper Citation Record · LEDGER

Hierarchical Budget Policy Optimization for Adaptive Reasoning

As of 10 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2507.15844.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.15844 v3

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:31:28.767651Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact3
  • verified fuzzy0
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 282d0983-76a0-4e7d-9269-597ef0fabc77 · outbound

This paper cites L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning.

Hierarchical Budget Policy Optimization for Adaptive Reasoning L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:26.600626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:26.600626Z digest=sha256:c00ece7c3655eb77e9e6b61aa4c464b2f370e96d6df6477ab8a33d23401fc531

Observation bc1f98b9-2226-48e4-a27a-b217e678a2d8 · outbound

This paper cites URL https://doi.org/10.

Hierarchical Budget Policy Optimization for Adaptive Reasoning URL https://doi.org/10

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:26.692180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:26.692180Z digest=sha256:234697fc2cfa1033d2e462e2527294c5bbb21de863e436e0f913799140e21ace

Observation 2a4f7ba3-08d8-4a34-9586-a2519ef98720 · outbound

This paper cites Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models.

Hierarchical Budget Policy Optimization for Adaptive Reasoning Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:26.868293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:26.868293Z digest=sha256:c4225695f58f43129187b104960f87d899aa292217164c08bba8e76e5f8b531a

Observation 4fc7dd44-8184-4f22-b1a9-b918be534d1f · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Hierarchical Budget Policy Optimization for Adaptive Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:27.319209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:27.319209Z digest=sha256:2e9a7fd80a05421785249a097402e1cf9fa55b2b8b7aa7f0f9ea18fb22a97942

Observation 6c0aaae6-8018-48ee-8579-0d863a965da8 · outbound

This paper cites Thinkless: LLM Learns When to Think.

Hierarchical Budget Policy Optimization for Adaptive Reasoning Thinkless: LLM Learns When to Think

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:27.442287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:27.442287Z digest=sha256:bc53e8499fde40bf7a3eb76d8fa91b842fc1ddbf5c1100a3a3edd7d2bc018031

Observation 87229f48-e3d9-4439-b38f-18c0427e85b8 · outbound

This paper cites ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning.

Hierarchical Budget Policy Optimization for Adaptive Reasoning ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:27.609818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:27.609818Z digest=sha256:f6219a35ea172c8bf26097a9a55bf8347e81c23a6677c42ce32a2cfc657d0ee9

Observation d8c3f112-3400-466c-8a33-b1f78825ba81 · outbound

This paper cites ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning.

Hierarchical Budget Policy Optimization for Adaptive Reasoning ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:27.748148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:27.748148Z digest=sha256:f9baf1b32d2cdbfe3b51f005bb379531899f050e4c64568f66de660c60203f0d

Observation c534d5d1-4d77-4b2e-b298-c0cf235de80c · outbound

This paper cites URLhttps://doi.org/10.48550/arXiv.2505.11225.

Hierarchical Budget Policy Optimization for Adaptive Reasoning URLhttps://doi.org/10.48550/arXiv.2505.11225

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:27.958656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:27.958656Z digest=sha256:f0f5f9b5597816ee7b4d250064190120c187552ca33c6c3c902102cacc25b9ef

Observation e0714b9d-f4a4-46c2-9aee-ba6de4f8bcca · outbound

This paper cites Think Only When You Need with Large Hybrid-Reasoning Models.

Hierarchical Budget Policy Optimization for Adaptive Reasoning Think Only When You Need with Large Hybrid-Reasoning Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:28.102020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:28.102020Z digest=sha256:a8d997207e43d1c8b2f6419af125257a266785ba5167cf009433259ba40ad120

Observation 02bf1114-67cd-4d06-8c9d-f67a7d202e2f · outbound

This paper cites SelfBudgeter: Adaptive Token Allocation for Efficient LLM Reasoning.

Hierarchical Budget Policy Optimization for Adaptive Reasoning SelfBudgeter: Adaptive Token Allocation for Efficient LLM Reasoning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:28.228889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:28.228889Z digest=sha256:7823898b7fafe22c64c2979deacee92cfd022e5aacd63388ffacb96d7a5a3f1c

Observation 3754db1e-dab9-47d9-9389-9b7d56f7f4d3 · outbound

This paper cites ThinkSwitcher: When to Think Hard, When to Think Fast.

Hierarchical Budget Policy Optimization for Adaptive Reasoning ThinkSwitcher: When to Think Hard, When to Think Fast

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:28.312222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:28.312222Z digest=sha256:64a688dc0c3bdcbb7278c0fb0a9f9fbaafe81c4b5a7d8be5315551bc1ec418c0

Observation 06e9ff18-fa4c-457d-ab68-213df15f9805 · outbound

This paper cites AdaCoT: Pareto-Optimal Adaptive Chain-of-Thought Triggering via Reinforcement Learning.

Hierarchical Budget Policy Optimization for Adaptive Reasoning AdaCoT: Pareto-Optimal Adaptive Chain-of-Thought Triggering via Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:28.441783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:28.441783Z digest=sha256:dbfe7f23ea2a30099e34c72e23ba4e95ab90f58ebc810b64cc2ae6bf861cb568

Observation 8c52a92a-bdbe-41f5-8298-c7fb10f3f786 · outbound

This paper cites Reasoning Models Can Be Effective Without Thinking.

Hierarchical Budget Policy Optimization for Adaptive Reasoning Reasoning Models Can Be Effective Without Thinking

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:28.526529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:28.526529Z digest=sha256:3f356f75f169f17b2e6bdb74499c0488f000dd7614c23c5aa5d39c80f26708ec

Observation 040d2d23-8868-452a-8f41-466916fccb9b · outbound

This paper cites Reasoning Models Can Be Effective Without Thinking.

Hierarchical Budget Policy Optimization for Adaptive Reasoning Reasoning Models Can Be Effective Without Thinking

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:28.578878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:28.578878Z digest=sha256:09e27ce87951f4f1128e1875858ca9c286af4c444c20f8d376a9018c9d8ab186

Observation 35312d21-cb31-4f3e-8bda-dc2b7aae9586 · outbound

This paper cites Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems.

Hierarchical Budget Policy Optimization for Adaptive Reasoning Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:28.703711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:28.703711Z digest=sha256:e45c293ece7f6b2c87f1e664f9a62f443bd8544cd13722bd6a13b9e976b40498

Observation 5968c0c8-97cd-47c5-a1f7-f0cf736660e6 · outbound

This paper cites s1: Simple test-time scaling.

Hierarchical Budget Policy Optimization for Adaptive Reasoning s1: Simple test-time scaling

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:28.722892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:28.722892Z digest=sha256:c746d8f2b8dd58cf24d73bbbed3999d8fa0563241b7c8566c8402012c0a06d11

Observation 08e6a92c-a9df-4c47-970e-d32541be5277 · outbound

This paper cites Accessed: 2025-07-22.

Hierarchical Budget Policy Optimization for Adaptive Reasoning Accessed: 2025-07-22

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:28.728133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:28.728133Z digest=sha256:47717dceb2eeefab382b8833cbf9890cc53587dc5c76eef8c4e4a4c2497b27e6

Observation b823124c-5b7f-4289-abb1-10c120f67e88 · outbound

This paper cites URL https://doi.org/ 10.48550/arXiv.2505.04881.

Hierarchical Budget Policy Optimization for Adaptive Reasoning URL https://doi.org/ 10.48550/arXiv.2505.04881

Reference 22

Resolution
verified exact
doi, observed 2026-08-06T15:31:29.181641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:31:28.733587Z digest=sha256:858c75fb83d4b4a3cae09d34fc7c188d773bf8dc656788b6bdd5d188bcf3c6e7

Observation a077caf7-ac67-47b6-8c0f-5df45d927f99 · outbound

This paper cites Learning when to think: Shaping adaptive reasoning in r1-style models via multi-stage RL.

Hierarchical Budget Policy Optimization for Adaptive Reasoning Learning when to think: Shaping adaptive reasoning in r1-style models via multi-stage RL

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:28.738590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:28.738590Z digest=sha256:0dffd7b778e4de8065b7e297b263412f13eebba289ee3d59fc71c48ecef698f5

Observation 5c5a568d-8990-4f95-a77f-857af800e004 · outbound

This paper cites URL https://doi.org/ 10.48550/arXiv.2505.10832.

Hierarchical Budget Policy Optimization for Adaptive Reasoning URL https://doi.org/ 10.48550/arXiv.2505.10832

Reference 24

Resolution
verified exact
doi, observed 2026-08-06T15:31:29.109749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:31:28.743999Z digest=sha256:0eb2403249a5fa1dc54fc983875e5756c2849191512b54b8e8fb6b2232979461

Observation 3a5e8ee1-c570-490f-991a-e77582d4a762 · outbound

This paper cites PATS: Process-Level Adaptive Thinking Mode Switching.

Hierarchical Budget Policy Optimization for Adaptive Reasoning PATS: Process-Level Adaptive Thinking Mode Switching

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:28.748477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:28.748477Z digest=sha256:50734f3e54517ac68130a78a2a6b7ce75ba3402a2b83d5089e06ae6debbeb0f6

Observation 44edbed9-a921-4666-a661-fa77e446d855 · outbound

This paper cites URL https://doi.org/10.48550/arXiv.2505.20258.

Hierarchical Budget Policy Optimization for Adaptive Reasoning URL https://doi.org/10.48550/arXiv.2505.20258

Reference 26

Resolution
verified exact
doi, observed 2026-08-06T15:31:29.005466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:31:28.753494Z digest=sha256:863d06bd0a4045ea4caf8c807b8299f625867d5c42db20fafa1061517c2bb2d8

Observation d4571c65-7566-4e78-958d-fde3db65fa6b · outbound

This paper cites Scalable Chain of Thoughts via Elastic Reasoning.

Hierarchical Budget Policy Optimization for Adaptive Reasoning Scalable Chain of Thoughts via Elastic Reasoning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:28.757943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:28.757943Z digest=sha256:930f67f6c6a3de9706c0a07fb0253ad6c92843837f45db501d2ba97f8ebc0761

Observation 092df8c9-b248-4eca-91ce-1631970e42fb · outbound

This paper cites URLhttps://doi.org/10.48550/arXiv.2504.15895.

Hierarchical Budget Policy Optimization for Adaptive Reasoning URLhttps://doi.org/10.48550/arXiv.2504.15895

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:28.763066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:28.763066Z digest=sha256:663abe4de801df5f1d582b5b6589dc8a61383891a3786a796d43f4c3914ecd14

Observation a884b321-ad3e-4dd1-8214-a6eead52d5c3 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Hierarchical Budget Policy Optimization for Adaptive Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:28.767651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:28.767651Z digest=sha256:ff1ae226dda00ea95e900e0c83a867cc610f39887a8cdfd8ef099f5950726434

Observation 68b378fb-239c-496d-a120-d98241007ce4 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Hierarchical Budget Policy Optimization for Adaptive Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:27.109371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:27.109371Z digest=sha256:e2696a8886f77fe6577894796e2cd43884777bd7beecac88f3191e7637c40e77

Observation 4b63b925-88cd-4732-a3ce-000e546a4658 · outbound

This paper cites AdaCoT: Pareto-Optimal Adaptive Chain-of-Thought Triggering via Reinforcement Learning.

Hierarchical Budget Policy Optimization for Adaptive Reasoning AdaCoT: Pareto-Optimal Adaptive Chain-of-Thought Triggering via Reinforcement Learning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:28.425196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:28.425196Z digest=sha256:17a1046ca148f7f087ef9c519d69c6802e57b9ff614699f70889ffb67fb9c38b

Observation c7fbf56c-13cb-4258-ac3f-e4ae51ea39cc · outbound

This paper cites Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs.

Hierarchical Budget Policy Optimization for Adaptive Reasoning Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:26.956879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:26.956879Z digest=sha256:321debc0271e7e2722be724cd4f9f6b74a03e81779e8673436a52553ef31f8ba

Observation 4af2d2de-57bc-4c3d-b51b-d6bed7cff138 · outbound

This paper cites L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning.

Hierarchical Budget Policy Optimization for Adaptive Reasoning L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:26.625005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:26.625005Z digest=sha256:72bf91b070eff576d4b781ed56d37a96f5c521d0bcf9bdd1e351a6549bfbea85

Pith citing papers

No inbound Pith citation observations are available.