Pith. sign in

Paper Citation Record · LEDGER

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents

As of 17 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 2 inbound Pith citation observations for arXiv:2607.06223.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.06223 v1

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-08T13:01:45.049174Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T05:53:52.763644Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T05:53:52.854862Z

Reference resolution

37 of 37 outbound references displayed

  • verified exact20
  • verified fuzzy16
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 653bad21-17df-4070-bd47-d3474750dee1 · outbound

This paper cites Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901, 2020.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901, 2020

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T13:04:57.061259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:eb95285794240b5ce5fa76aa237ef6eecd3808c630f1aa64c16cda10cc02b51d

Observation 32b1664c-bd0d-4432-91aa-a452b6374cef · outbound

This paper cites Palm: Scaling language modeling with pathways.Journal of machine learning research, 24(240): 1–113, 2023.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Palm: Scaling language modeling with pathways.Journal of machine learning research, 24(240): 1–113, 2023

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T13:04:57.083108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:6aa61667207416f24f84f4c3880278db0be8c085810ffd71e91744f85010d57f

Observation f1d3c3ec-21cd-4711-8191-2f83006a4011 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents LLaMA: Open and Efficient Foundation Language Models

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-08T13:04:56.738338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:1f01cd919161b07b9d9f80db599c8c03d39dcfe1ab09dbe4a97e5a1e25f53478

Observation eecfc933-8c46-4f11-83ac-80370982ca02 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents ReAct: Synergizing Reasoning and Acting in Language Models

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-08T13:04:56.701696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:f4470f2f47e5c6403d4d58d4af51a0b8ab0382a60a31c4cdc5ef003c07e455be

Observation 6eb9376e-9dd2-4705-9c20-335efced767d · outbound

This paper cites Toolformer: Language models can teach themselves to use tools.Advances in neural information processing systems, 36: 68539–68551.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Toolformer: Language models can teach themselves to use tools.Advances in neural information processing systems, 36: 68539–68551

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T13:04:57.094150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:377d2092853ecf5a2379ffd9f0eed8a10e1249ffc172c65bff7cdc693dafbb3e

Observation 59fbeac2-527b-467f-85be-ae71a3a1a923 · outbound

This paper cites The Landscape of Agentic Reinforcement Learning for LLMs: A Survey.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-08T13:04:56.742129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:1aeb61611596e641f18a50871adddf37d7c6fcea11a9d13ca44d786bc7236e45

Observation 1a918ce6-6f2d-4f62-853f-055961e0244e · outbound

This paper cites Deep Research Agents: A Systematic Examination And Roadmap.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Deep Research Agents: A Systematic Examination And Roadmap

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-08T13:04:56.716778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:10fd7fd8d57f357f0880c60eb0e37381ac642eca525d8f4a786b28c55b196aac

Observation 3555b2ba-7936-4a1e-b79c-7e9d3df39f71 · outbound

This paper cites Searching for best practices in retrieval- augmented generation.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Searching for best practices in retrieval- augmented generation

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T13:04:57.091978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:a46858bb1e34abe33c7cc43c7714c04c41cf069ff1476d1f10ef9f38ce85ff0d

Observation b38c324e-277e-4cc6-846a-2aab5023bafa · outbound

This paper cites Search-o1: Agentic search-enhanced large reasoning models.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Search-o1: Agentic search-enhanced large reasoning models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T13:04:57.096554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:d6a54b7f4f5d3ad72ba5398b44b2cb45570c475e69d8522877794cfc7cc1625f

Observation 6d71ee68-0fae-4634-be81-d148ace315d6 · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-08T13:04:56.721618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:e139c878014cdf229fd578a8de2e7d0c2b3ca27ef3ac69c8ec6460a3d3a48b9b

Observation cc264f8a-82bf-408a-8192-331fc975c8ec · outbound

This paper cites ToolRL: Reward is All Tool Learning Needs.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents ToolRL: Reward is All Tool Learning Needs

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-08T13:04:56.719157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:98359739e289a6c4d812fe32d7b871cb70429e9cfb996f7ed0d509a7a4fcb936

Observation 8f58128f-f53e-43de-9396-3f476fc3f513 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Proximal Policy Optimization Algorithms

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-08T13:04:56.711606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:1684f2938b548ae1fc503147ec8eedb877f7aff889887ec7a2f6de42eb776e31

Observation ea69393f-50df-47b8-9222-cda18a03251d · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T13:04:57.085272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:b3f472c9ef09ce1479e4863cc6ed16e8490f8c492c11b8f55200ba7bb713411a

Observation 5db26127-3237-46a2-8a55-3a47aa4ea8f9 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-08T13:04:56.723999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:73689369b4164a715dc6def28d174d50a18dfe085ee78ab2ee40a6872337082f

Observation 924d65f0-371c-4ed0-9a18-57c48f7c031f · outbound

This paper cites Group-in-Group Policy Optimization for LLM Agent Training.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Group-in-Group Policy Optimization for LLM Agent Training

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-08T13:04:56.714142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:563783529dcd4ef2cffcbe6d192fbb767212c88a7ca755351cfabb774b83593c

Observation 5517b29e-903c-4223-89a1-f84079b4f418 · outbound

This paper cites Information gain-based policy optimization: A simple and effective approach for multi-turn llm agents.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Information gain-based policy optimization: A simple and effective approach for multi-turn llm agents

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-08T13:04:56.696136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:84d0f8b60469a0fcad398e2578a00e52d8cf3b943b57184c06c37fb75954cf76

Observation 957ff1d1-e873-46b1-b8d2-06258fc57d2f · outbound

This paper cites Tree search for llm agent reinforcement learning.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Tree search for llm agent reinforcement learning

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-08T13:04:56.744901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:e855fab05e6ce53adf2fa4685869bb3b63733bc1e64023ceb28539b0c87f9efc

Observation f41b96ba-06e1-48cb-af03-3eb52abc5429 · outbound

This paper cites Agentic entropy-balanced policy optimization.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Agentic entropy-balanced policy optimization

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-08T13:04:56.704280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:61d337514198daf0a7ae8d52e60a9ebc603f5d9bfec0dfb87bedb12d61c3fb6a

Observation 746a0598-b11f-487d-9302-53595c7fd4f5 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-08T13:04:56.709206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:b8081732a6c650e3f9e86d96c5e7ae68d9feb69c6bf39beee7097d138745af74

Observation 533705bc-d6b6-4971-afa2-0eb313595668 · outbound

This paper cites Back to basics: Revisiting reinforce-style optimization for learning from human feedback in llms.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Back to basics: Revisiting reinforce-style optimization for learning from human feedback in llms

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T13:04:57.087498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:1032c05aa18215b713ec075ef8f16f6ed23158d86c8e511738f4ea73080a9f47

Observation 1ace522c-3a58-4e76-a26b-5012f4e794d6 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-08T13:04:56.735738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:4697d210545d97d410962ea9ab6ae17db978698f9c2646e55c239f7ea6fccf8a

Observation c40d23b7-eece-4928-99e9-19d3e0d190cf · outbound

This paper cites Group Sequence Policy Optimization.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Group Sequence Policy Optimization

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-08T13:04:56.706723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:858b8fc346743733d3674fc17fa39b19c9f8949dcae3d387f1102a48306a7ae9

Observation 76a4e2d2-bb8f-4746-9e5b-1a996670f524 · outbound

This paper cites ZeroSearch: Incentivize the Search Capability of LLMs without Searching.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents ZeroSearch: Incentivize the Search Capability of LLMs without Searching

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-08T13:04:56.726732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:f67feddfc0efc7f670b1ae1d59842edea3911f1b3148e7bee18ca5da526606a2

Observation fbd51867-5c1e-427a-970d-6080a2d3ed18 · outbound

This paper cites R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-08T13:04:56.732119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:383bbabfd79f8121adea6d66c4aedab6f8f8abe45e94595df4abbac16079dee1

Observation 40f437aa-3049-4bf9-8921-3c39e18ee3c7 · outbound

This paper cites Smart-searcher: Incentivizing the dynamic knowledge acquisition of llms via reinforcement learning.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Smart-searcher: Incentivizing the dynamic knowledge acquisition of llms via reinforcement learning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T13:04:57.073965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:9c94b00ec341a11d7ccd354a370241c5037af04e07bdd045ac3636c17d38f384

Observation 733fe686-4d61-4648-8494-88de7018c776 · outbound

This paper cites StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-07-08T13:04:56.729725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:1fb29cdc0b279feb2a6a909839049dc8a379c6acdb72ea1806b1c5709d0b00f1

Observation 9641d324-553f-4063-8ade-d0947055e109 · outbound

This paper cites Agentic Reinforced Policy Optimization.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Agentic Reinforced Policy Optimization

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-08T13:04:56.693393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:b61c4ff7056cab87a739f311a3fe2012bcea60db0b07cdf72559e5a0b34cd582

Observation b46cd4e0-7718-4392-829a-35b24d6e2297 · outbound

This paper cites Natural questions: a benchmark for question answering research.Transactions of the Association for Computational Linguistics, 7:453–466.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Natural questions: a benchmark for question answering research.Transactions of the Association for Computational Linguistics, 7:453–466

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T13:04:57.089868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:687d064a616695bf766d7bb0f2bd2e3a9e910ff5312cacb07ef6d494775b5d44

Observation 7a7ae40f-1efd-4503-abf7-7c5332196998 · outbound

This paper cites Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T13:04:57.076213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:d1a11d4c4c4f691955d55228b52af886c40553f597b0de22c59b299a3049ef8f

Observation 71e13e9f-c272-4e1e-a83f-32eee4a60b52 · outbound

This paper cites When not to trust language models: Investigating effectiveness of parametric and non-parametric memories.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents When not to trust language models: Investigating effectiveness of parametric and non-parametric memories

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T13:04:57.081101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:4b8d1f55154cbaee5efb8b9f9f00390607d4f4052f912e45fc6026aa5f8d4bde

Observation 588a4f76-4c1f-496c-8435-803fdfe81971 · outbound

This paper cites Hotpotqa: A dataset for diverse, explainable multi-hop question answering.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Hotpotqa: A dataset for diverse, explainable multi-hop question answering

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T13:04:57.072005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:7e1362d1010a7be6989b08ab2a9844e8eb1e37b2bcbc775e4ae93abe5b3887a5

Observation ca273a76-9e6c-4df2-abef-9dd2c4417ae1 · outbound

This paper cites Constructing a multi-hop qa dataset for comprehensive evaluation of reasoning steps.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Constructing a multi-hop qa dataset for comprehensive evaluation of reasoning steps

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T13:04:57.069900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:f056ea321c8c03a21db12b90c3435668b6cf2e5ed34fa616b418bac6f0061124

Observation d7a05ff2-e2b9-45ef-bf4f-34290dfa4c1b · outbound

This paper cites Musique: Multihop questions via single-hop question composition.Transactions of the Association for Computational Linguistics, 10:539–554.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Musique: Multihop questions via single-hop question composition.Transactions of the Association for Computational Linguistics, 10:539–554

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T13:04:57.063647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:5ca0dca27265b07d3fbe5e783f4937cf0cf7aaad131a71907b4f03c4a57e415d

Observation 95cbde69-9b32-4dbe-af78-721da3d53d77 · outbound

This paper cites Measuring and narrowing the compositionality gap in language models.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Measuring and narrowing the compositionality gap in language models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T13:04:57.066034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:767afb6f0d3e26ac5ab5a8fdb1b67728c6754a21b6468bb014d4bdba81b2a2e5

Observation e92f7f51-20c9-40bc-81e6-3d8406a88775 · outbound

This paper cites an unresolved cited work.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-07-08T13:04:57.067705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:9102a686d8355c1fa670dc4989251ec3b69dd3af112f1ecf149578cd5515dacb

Observation 76e253ca-14b7-44a0-ab5e-51ed8aa1c835 · outbound

This paper cites Text Embeddings by Weakly-Supervised Contrastive Pre-training.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Text Embeddings by Weakly-Supervised Contrastive Pre-training

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-07-08T13:04:56.699304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:c3ecca06553eccc6cf7663ce07338bf7092c8aa60dd105ed6b25bc1218ed5946

Observation 7b48ea87-c656-4864-8393-c46097a58a1b · outbound

This paper cites Asymptotic evaluation of certain markov process expectations for large time, i.Communications on pure and applied mathematics, 28 (1):1–47, 1975.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Asymptotic evaluation of certain markov process expectations for large time, i.Communications on pure and applied mathematics, 28 (1):1–47, 1975

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T13:04:57.079064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:fab79ef4ba5182fb1f7d7e1814da21f42b0cfa9e89e277c69d7f9d22c231bcbd

Pith citing papers

Observation 601c01fc-f827-46be-80ab-f4ef28a8bd95 · inbound

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning cites this paper.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:42.254343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:42.254343Z digest=sha256:72cad8ce928c7db16ff6d673d7ee95eb35ee08722ea4f891efcfe9210c4bfb98

Observation 454f7f37-0091-4243-aba9-34b22322ce7d · inbound

Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning cites this paper.

Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-06T05:53:52.858284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T05:53:52.763644Z digest=sha256:57c60899a0e4989e66e547954b70fa5fc44a1528fee282fa2cd8a5be554b7c2b