Pith. sign in

Paper Citation Record · LEDGER

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents

As of 15 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 2 inbound Pith citation observations for arXiv:2607.06223.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.06223 v1

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-08T13:01:45.049174Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T05:53:52.763644Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T05:53:52.854862Z

Reference resolution

37 of 37 outbound references displayed

  • verified exact20
  • verified fuzzy16
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 653bad21-17df-4070-bd47-d3474750dee1 · outbound

This paper cites Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901, 2020.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901, 2020

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T13:04:57.061259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:d0450cf4a3dc7368ba00fe15481385ea1d409503273a39258fb80cdbf2953242

Observation 32b1664c-bd0d-4432-91aa-a452b6374cef · outbound

This paper cites Palm: Scaling language modeling with pathways.Journal of machine learning research, 24(240): 1–113, 2023.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Palm: Scaling language modeling with pathways.Journal of machine learning research, 24(240): 1–113, 2023

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T13:04:57.083108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:3bed213e69fb943bb5f1b97771e933cc8a5166c956c6f2b12c2c9557329012c4

Observation f1d3c3ec-21cd-4711-8191-2f83006a4011 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents LLaMA: Open and Efficient Foundation Language Models

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-08T13:04:56.738338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:9636c92221ae7f53d7b576fb2c49f76d63ccee33c683ce8876879c3f302aa115

Observation eecfc933-8c46-4f11-83ac-80370982ca02 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents ReAct: Synergizing Reasoning and Acting in Language Models

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-08T13:04:56.701696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:1b19a099e49269a2a5add24c15063aef32efd00f60c9dbcad43885aa9db2b6f8

Observation 6eb9376e-9dd2-4705-9c20-335efced767d · outbound

This paper cites Toolformer: Language models can teach themselves to use tools.Advances in neural information processing systems, 36: 68539–68551.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Toolformer: Language models can teach themselves to use tools.Advances in neural information processing systems, 36: 68539–68551

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T13:04:57.094150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:59fcbaa6600e02ecfc9bde88056035b4b53ab69db930d20ad4ec42e2e79f1e1c

Observation 59fbeac2-527b-467f-85be-ae71a3a1a923 · outbound

This paper cites The Landscape of Agentic Reinforcement Learning for LLMs: A Survey.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-08T13:04:56.742129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:30cf0ad77b54af99356e5976ce650462ac853041624e4ef7ded6984d8d365df3

Observation 1a918ce6-6f2d-4f62-853f-055961e0244e · outbound

This paper cites Deep Research Agents: A Systematic Examination And Roadmap.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Deep Research Agents: A Systematic Examination And Roadmap

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-08T13:04:56.716778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:c0bc9914d948f661b6ef75cb03fa04bbcf662a331f3230c70d9eb8b0daa5570b

Observation 3555b2ba-7936-4a1e-b79c-7e9d3df39f71 · outbound

This paper cites Searching for best practices in retrieval- augmented generation.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Searching for best practices in retrieval- augmented generation

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T13:04:57.091978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:6d8fa837232a59e4e3c756c194a22802aeb6aec3ffc7abc6ae571a97ef2aac65

Observation b38c324e-277e-4cc6-846a-2aab5023bafa · outbound

This paper cites Search-o1: Agentic search-enhanced large reasoning models.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Search-o1: Agentic search-enhanced large reasoning models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T13:04:57.096554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:d7851f68404c1b773dc457f58c3581bc6f376157a8daabaf628b250421c8ffa0

Observation 6d71ee68-0fae-4634-be81-d148ace315d6 · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-08T13:04:56.721618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:0586c54fb7d688d2dc8129cd48c850ac8b607131b91e2b0bf6cc226701b25af7

Observation cc264f8a-82bf-408a-8192-331fc975c8ec · outbound

This paper cites ToolRL: Reward is All Tool Learning Needs.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents ToolRL: Reward is All Tool Learning Needs

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-08T13:04:56.719157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:604e7eaf8d17bf59d717f29006faa84c48d418a8f751a61ace10385f83ef7cbb

Observation 8f58128f-f53e-43de-9396-3f476fc3f513 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Proximal Policy Optimization Algorithms

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-08T13:04:56.711606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:c088e9ff46e7bac01bab038bc215a04dc35ed912ba60b85e06b24878c27d92ca

Observation ea69393f-50df-47b8-9222-cda18a03251d · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T13:04:57.085272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:2e8b026d2cb225bb304ab59b6a768bb0891b80cb7d150205451307db5123626d

Observation 5db26127-3237-46a2-8a55-3a47aa4ea8f9 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-08T13:04:56.723999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:37a8a72aa999f9386f280aefe7c8a32983e61d2159696e22f063b3f0b71cbc96

Observation 924d65f0-371c-4ed0-9a18-57c48f7c031f · outbound

This paper cites Group-in-Group Policy Optimization for LLM Agent Training.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Group-in-Group Policy Optimization for LLM Agent Training

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-08T13:04:56.714142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:f610747a962924a38584e6ee43bf6c0c204599cbc37990f7a5e94a808b9c2af4

Observation 5517b29e-903c-4223-89a1-f84079b4f418 · outbound

This paper cites Information gain-based policy optimization: A simple and effective approach for multi-turn llm agents.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Information gain-based policy optimization: A simple and effective approach for multi-turn llm agents

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-08T13:04:56.696136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:9cd3dca5c5bdd8aee7d438df72c2e620a960eaebbb61080aaccf3cb8e19bd625

Observation 957ff1d1-e873-46b1-b8d2-06258fc57d2f · outbound

This paper cites Tree search for llm agent reinforcement learning.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Tree search for llm agent reinforcement learning

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-08T13:04:56.744901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:c54cdee9fb9d1ce0d55c2125b4ee0f8722340a51c66a958494727dad1c328e78

Observation f41b96ba-06e1-48cb-af03-3eb52abc5429 · outbound

This paper cites Agentic entropy-balanced policy optimization.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Agentic entropy-balanced policy optimization

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-08T13:04:56.704280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:09f987fd1f29d88ad096d7b9c22ba7863c76012a0432827b22caadc05f572e99

Observation 746a0598-b11f-487d-9302-53595c7fd4f5 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-08T13:04:56.709206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:b5f7ee946a7cf056911f8c9d39b1c4c30fa9a2f82461057769f281d2335fdadc

Observation 533705bc-d6b6-4971-afa2-0eb313595668 · outbound

This paper cites Back to basics: Revisiting reinforce-style optimization for learning from human feedback in llms.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Back to basics: Revisiting reinforce-style optimization for learning from human feedback in llms

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T13:04:57.087498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:9a8dcd135b647f458fe294bb5c3fbfe035990c094c75d2d194c419e734441a86

Observation 1ace522c-3a58-4e76-a26b-5012f4e794d6 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-08T13:04:56.735738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:b23e2ecf045df4230b402ea7ca5ac982a8bf9587896fce4ce51c21ea802f8b8f

Observation c40d23b7-eece-4928-99e9-19d3e0d190cf · outbound

This paper cites Group Sequence Policy Optimization.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Group Sequence Policy Optimization

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-08T13:04:56.706723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:3dc4a0bef8d0602ef2cd9d5cdf0a2018c35af3b9286bdf4438dac0784f2d823f

Observation 76a4e2d2-bb8f-4746-9e5b-1a996670f524 · outbound

This paper cites ZeroSearch: Incentivize the Search Capability of LLMs without Searching.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents ZeroSearch: Incentivize the Search Capability of LLMs without Searching

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-08T13:04:56.726732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:0c9f81f951d4598ca5efe129272a9d2d331dc4839e277a5900d4ef6a566d1200

Observation fbd51867-5c1e-427a-970d-6080a2d3ed18 · outbound

This paper cites R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-08T13:04:56.732119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:cbeeca6b394e3e8ad9b5d988efbe42671666e5123a8daf4ce8dd0c93962e5f94

Observation 40f437aa-3049-4bf9-8921-3c39e18ee3c7 · outbound

This paper cites Smart-searcher: Incentivizing the dynamic knowledge acquisition of llms via reinforcement learning.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Smart-searcher: Incentivizing the dynamic knowledge acquisition of llms via reinforcement learning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T13:04:57.073965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:1fff31cf3b4b404c2c636bdabc45c1db5af257eb24c4092134953547f2244b6b

Observation 733fe686-4d61-4648-8494-88de7018c776 · outbound

This paper cites StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-07-08T13:04:56.729725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:76eb1da74252d0d6b77f4956aa83b01817f7d226076e5f7e2f9bcefec5b67e62

Observation 9641d324-553f-4063-8ade-d0947055e109 · outbound

This paper cites Agentic Reinforced Policy Optimization.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Agentic Reinforced Policy Optimization

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-08T13:04:56.693393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:e03b7bdd01588b12ddf79166868f51f24aa0cf5280c8691be11724039ab22312

Observation b46cd4e0-7718-4392-829a-35b24d6e2297 · outbound

This paper cites Natural questions: a benchmark for question answering research.Transactions of the Association for Computational Linguistics, 7:453–466.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Natural questions: a benchmark for question answering research.Transactions of the Association for Computational Linguistics, 7:453–466

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T13:04:57.089868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:8383df39f6131a7b4cf8182850252d2b8c01e1c9099f8988f973ee7848987569

Observation 7a7ae40f-1efd-4503-abf7-7c5332196998 · outbound

This paper cites Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T13:04:57.076213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:30008b7aca662945e3d3a376029a45b81251b5d38873ea8e048171e6eb8a138e

Observation 71e13e9f-c272-4e1e-a83f-32eee4a60b52 · outbound

This paper cites When not to trust language models: Investigating effectiveness of parametric and non-parametric memories.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents When not to trust language models: Investigating effectiveness of parametric and non-parametric memories

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T13:04:57.081101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:5bb0ed8b94d2b582c9ecd63f8cd01b8c07f4611ea3fac26f1aa6907b3e96a190

Observation 588a4f76-4c1f-496c-8435-803fdfe81971 · outbound

This paper cites Hotpotqa: A dataset for diverse, explainable multi-hop question answering.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Hotpotqa: A dataset for diverse, explainable multi-hop question answering

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T13:04:57.072005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:50353c654f9a1268bf4e1c9b341bbd9d5b2e6c44480dbba6c0eb710b874b44b2

Observation ca273a76-9e6c-4df2-abef-9dd2c4417ae1 · outbound

This paper cites Constructing a multi-hop qa dataset for comprehensive evaluation of reasoning steps.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Constructing a multi-hop qa dataset for comprehensive evaluation of reasoning steps

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T13:04:57.069900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:db2337633df4002c119b3b05c60e63321889a3859977c38895a7d36051e50539

Observation d7a05ff2-e2b9-45ef-bf4f-34290dfa4c1b · outbound

This paper cites Musique: Multihop questions via single-hop question composition.Transactions of the Association for Computational Linguistics, 10:539–554.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Musique: Multihop questions via single-hop question composition.Transactions of the Association for Computational Linguistics, 10:539–554

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T13:04:57.063647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:665ee65cc0bc21b94ee95e533d2c23b861f943c3c9928e35611627182f5c425e

Observation 95cbde69-9b32-4dbe-af78-721da3d53d77 · outbound

This paper cites Measuring and narrowing the compositionality gap in language models.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Measuring and narrowing the compositionality gap in language models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T13:04:57.066034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:777a1032ffd62cadf824758011d1c71b9a0fd435755fe67b0bdd5aa91653727a

Observation e92f7f51-20c9-40bc-81e6-3d8406a88775 · outbound

This paper cites an unresolved cited work.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-07-08T13:04:57.067705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:73acf0819563b20a9631ed3c761858b2b37a77a27a0ab3bd7e1bde6f6a590704

Observation 76e253ca-14b7-44a0-ab5e-51ed8aa1c835 · outbound

This paper cites Text Embeddings by Weakly-Supervised Contrastive Pre-training.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Text Embeddings by Weakly-Supervised Contrastive Pre-training

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-07-08T13:04:56.699304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:b28c7b18b1c668626143859af138faf885ca442963d858cfa7ef47f40d371b25

Observation 7b48ea87-c656-4864-8393-c46097a58a1b · outbound

This paper cites Asymptotic evaluation of certain markov process expectations for large time, i.Communications on pure and applied mathematics, 28 (1):1–47, 1975.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Asymptotic evaluation of certain markov process expectations for large time, i.Communications on pure and applied mathematics, 28 (1):1–47, 1975

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T13:04:57.079064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:74aa94cb2ba70b1d40e7e29af38cbdad3d14f0acf6e6123f0038e3d9f11f56e0

Pith citing papers

Observation 601c01fc-f827-46be-80ab-f4ef28a8bd95 · inbound

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning cites this paper.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:42.254343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:42.254343Z digest=sha256:c435fe80f13cca47f98387a6436ae48f69193ce43505a1febe132325643e4839

Observation 454f7f37-0091-4243-aba9-34b22322ce7d · inbound

Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning cites this paper.

Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-06T05:53:52.858284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T05:53:52.763644Z digest=sha256:83b93b9d27d027c7885646f71825c086c1b6328169b5d3e8cf08300dec4b93e7