Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T05:53:52.797095Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 0 inbound Pith citation observations for arXiv:2608.05080.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T05:53:52.797095Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
33 of 33 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a1297d76-745d-4ed0-b42b-eb4cdfa6d4eb · outbound
Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning Assumption C.6 is the stability condition connecting the controller analysis to policy optimization
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c5f83f90-9fa2-41f6-820a-a27d5847c104 · outbound
Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning EvolveRouter: Co-Evolving Routing and Prompt for Multi-Agent Question Answering
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1ee55ea-e882-4ff7-b926-c5183599e476 · outbound
Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning Tree search for llm agent reinforcement learning.arXiv preprint arXiv:2509.21240,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b314d96-44a0-40c2-aeb1-33fb9af16d44 · outbound
Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning Understanding R1-Zero-Like Training: A Critical Perspective
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b770c64f-c6f7-468a-9cce-d88cb13e5802 · outbound
Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning Adaptive rollout allocation for online reinforcement learning with verifiable rewards.arXiv preprint arXiv:2602.01601,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 824626db-cf8b-4f7e-b9ee-082df4fc9507 · outbound
Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning Group distribu- tionally robust optimization-driven reinforcement learning for llm reasoning.arXiv preprint arXiv:2601.19280,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae3eb28e-29f0-4e35-8350-1a307876aa15 · outbound
Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning SAGE: Answer-Conditioned Uncertainty Targets for Verbal Uncertainty Alignment
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation edbb19bf-4184-4a46-a168-220e7f27a0d8 · outbound
Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning RAGEN-2: Reasoning Collapse in Agentic RL
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc6ba695-5295-4b4a-98c2-932ca0d4c436 · outbound
Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning Entropy-tree: Tree-based decoding with entropy-guided exploration.arXiv preprint arXiv:2601.15296,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51a9fe03-06bf-4c44-97a5-a18e719450d7 · outbound
Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning Scaling search-augmented llm reasoning via adaptive information control
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f92a6c1b-c9bd-48bf-adf1-9cb27978882d · outbound
Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning Wei Xiong, Chenlu Ye, Baohao Liao, Hanze Dong, Xinxing Xu, Christof Monz, Jiang Bian, Nan Jiang, and Tong Zhang
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ce64937-e1b1-4807-9fe5-06d152e115f9 · outbound
Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning Llms4all: A review of large language models across academic disciplines.arXiv preprint arXiv:2509.19580,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 454f7f37-0091-4243-aba9-34b22322ce7d · outbound
Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ae2f86d9-a61b-4345-84a8-11d2420ec17a · outbound
Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning Why Semantic Entropy Fails: Geometry-Aware and Calibrated Uncertainty for Policy Optimization
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8114cd00-53cd-4efb-baa9-9557b6121c75 · outbound
Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning First Return, Entropy-Eliciting Explore
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c837d4d8-8b63-4426-836d-dc3c6a75e150 · outbound
Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6285bcc4-50ff-498e-ba5a-b8f4509aca7d · outbound
Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning APPENDIXCONTENTTABLE A Related Work 14 B Implementation Details and Additional Experiments 14 B.1 Benchmarks
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 98c29281-dc7e-47d9-bd71-bd471f03e792 · outbound
Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning While effective, these methods rely on final outcomes and thus overlook where uncertainty arises within multi-step reasoning
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 179f783e-53a7-48e4-a01b-05041af08e16 · outbound
Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning IGRPO (Zhang et al., 2026a) instead expands trajectory trees according to information gain and derives an induced teacher distribution for policy optimization
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5579dc2f-5b03-407a-aa6f-6c999df98cb7 · outbound
Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning The agent interacts with the environment through search, page navigation, item inspection, and purchase decisions
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f2e637b2-b2c6-4298-b857-8367df30363e · outbound
Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning 1 n nX i=1 Acent i 2 # = n−1 n σ2 R,E
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7821f00e-40cf-4b90-87a1-756e48fedc62 · outbound
Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning VIP provides a complementary gradient-level justification: under standard conditional i.i.d
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a362d1fe-9974-4424-9999-0d0a8dc5fddb · outbound
Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning The variation-budget literature motivates modeling policy-induced non-stationarity through the drift term DT (Besbes et al., 2014; 2015)
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation bfbe4bae-8ba9-4e69-84e3-2c9c2da9f545 · outbound
Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning 42" rather than
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b3529ecf-c5f6-4dd1-9817-c09442526626 · outbound
Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning name":"sql_query
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ed38e50f-78f7-4050-ac0c-0ba23b2eb645 · outbound
Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Men, Kejuan Yang, et al
Reference 2007
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b4b3596d-a445-4a55-a7dc-1030a185020b · outbound
Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning 3SPO: State-Score-Supervised Policy Optimization for LLM Agents
Reference 2011
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5835ec8-5613-4908-abb1-6a43d8d84e55 · outbound
Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning Temperature as a meta-policy: Adaptive temperature in llm reinforcement learning.arXiv preprint arXiv:2602.11779,
Reference 2015
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 36e3614f-6bda-4940-aed6-f89bc3d6f519 · outbound
Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning Coba-rl: Capability-oriented budget allocation for reinforcement learning in llms.arXiv preprint arXiv:2602.03048,
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation edb736c8-4c8a-407f-a756-b069625ef69e · outbound
Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6b3308e-4fc6-48d8-8f73-733b20fd7814 · outbound
Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning Finite-time analysis of the multiarmed bandit problem.Machine learning, 47(2):235–256, 2002a
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ef582f9-9263-4486-b582-b33393401d0b · outbound
Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning How to Allocate, How to Learn? Dynamic Rollout Allocation and Advantage Modulation for Policy Optimization
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a96c60cf-83ce-48c8-a059-c68faa26629b · outbound
Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning Agentic entropy-balanced policy optimization
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.