Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-08T13:01:45.049174Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 2 inbound Pith citation observations for arXiv:2607.06223.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-08T13:01:45.049174Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T05:53:52.763644Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-06T05:53:52.854862Z
37 of 37 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 653bad21-17df-4070-bd47-d3474750dee1 · outbound
Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901, 2020
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 32b1664c-bd0d-4432-91aa-a452b6374cef · outbound
Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Palm: Scaling language modeling with pathways.Journal of machine learning research, 24(240): 1–113, 2023
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f1d3c3ec-21cd-4711-8191-2f83006a4011 · outbound
Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents LLaMA: Open and Efficient Foundation Language Models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation eecfc933-8c46-4f11-83ac-80370982ca02 · outbound
Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents ReAct: Synergizing Reasoning and Acting in Language Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6eb9376e-9dd2-4705-9c20-335efced767d · outbound
Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Toolformer: Language models can teach themselves to use tools.Advances in neural information processing systems, 36: 68539–68551
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 59fbeac2-527b-467f-85be-ae71a3a1a923 · outbound
Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1a918ce6-6f2d-4f62-853f-055961e0244e · outbound
Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Deep Research Agents: A Systematic Examination And Roadmap
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3555b2ba-7936-4a1e-b79c-7e9d3df39f71 · outbound
Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Searching for best practices in retrieval- augmented generation
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b38c324e-277e-4cc6-846a-2aab5023bafa · outbound
Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Search-o1: Agentic search-enhanced large reasoning models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6d71ee68-0fae-4634-be81-d148ace315d6 · outbound
Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation cc264f8a-82bf-408a-8192-331fc975c8ec · outbound
Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents ToolRL: Reward is All Tool Learning Needs
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8f58128f-f53e-43de-9396-3f476fc3f513 · outbound
Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Proximal Policy Optimization Algorithms
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ea69393f-50df-47b8-9222-cda18a03251d · outbound
Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 5db26127-3237-46a2-8a55-3a47aa4ea8f9 · outbound
Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 924d65f0-371c-4ed0-9a18-57c48f7c031f · outbound
Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Group-in-Group Policy Optimization for LLM Agent Training
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 5517b29e-903c-4223-89a1-f84079b4f418 · outbound
Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Information gain-based policy optimization: A simple and effective approach for multi-turn llm agents
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 957ff1d1-e873-46b1-b8d2-06258fc57d2f · outbound
Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Tree search for llm agent reinforcement learning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f41b96ba-06e1-48cb-af03-3eb52abc5429 · outbound
Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Agentic entropy-balanced policy optimization
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 746a0598-b11f-487d-9302-53595c7fd4f5 · outbound
Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 533705bc-d6b6-4971-afa2-0eb313595668 · outbound
Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Back to basics: Revisiting reinforce-style optimization for learning from human feedback in llms
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1ace522c-3a58-4e76-a26b-5012f4e794d6 · outbound
Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c40d23b7-eece-4928-99e9-19d3e0d190cf · outbound
Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Group Sequence Policy Optimization
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 76a4e2d2-bb8f-4746-9e5b-1a996670f524 · outbound
Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents ZeroSearch: Incentivize the Search Capability of LLMs without Searching
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation fbd51867-5c1e-427a-970d-6080a2d3ed18 · outbound
Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 40f437aa-3049-4bf9-8921-3c39e18ee3c7 · outbound
Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Smart-searcher: Incentivizing the dynamic knowledge acquisition of llms via reinforcement learning
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 733fe686-4d61-4648-8494-88de7018c776 · outbound
Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9641d324-553f-4063-8ade-d0947055e109 · outbound
Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Agentic Reinforced Policy Optimization
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b46cd4e0-7718-4392-829a-35b24d6e2297 · outbound
Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Natural questions: a benchmark for question answering research.Transactions of the Association for Computational Linguistics, 7:453–466
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7a7ae40f-1efd-4503-abf7-7c5332196998 · outbound
Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 71e13e9f-c272-4e1e-a83f-32eee4a60b52 · outbound
Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents When not to trust language models: Investigating effectiveness of parametric and non-parametric memories
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 588a4f76-4c1f-496c-8435-803fdfe81971 · outbound
Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Hotpotqa: A dataset for diverse, explainable multi-hop question answering
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ca273a76-9e6c-4df2-abef-9dd2c4417ae1 · outbound
Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Constructing a multi-hop qa dataset for comprehensive evaluation of reasoning steps
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d7a05ff2-e2b9-45ef-bf4f-34290dfa4c1b · outbound
Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Musique: Multihop questions via single-hop question composition.Transactions of the Association for Computational Linguistics, 10:539–554
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 95cbde69-9b32-4dbe-af78-721da3d53d77 · outbound
Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Measuring and narrowing the compositionality gap in language models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e92f7f51-20c9-40bc-81e6-3d8406a88775 · outbound
Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 76e253ca-14b7-44a0-ab5e-51ed8aa1c835 · outbound
Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Text Embeddings by Weakly-Supervised Contrastive Pre-training
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7b48ea87-c656-4864-8393-c46097a58a1b · outbound
Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Asymptotic evaluation of certain markov process expectations for large time, i.Communications on pure and applied mathematics, 28 (1):1–47, 1975
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 601c01fc-f827-46be-80ab-f4ef28a8bd95 · inbound
Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 454f7f37-0091-4243-aba9-34b22322ce7d · inbound
Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.