Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T00:37:25.304636Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 1 inbound Pith citation observation for arXiv:2608.00969.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T00:37:25.304636Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T00:37:22.965089Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-06T00:37:26.010355Z
30 of 30 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 54373557-c971-46ed-9353-69796bbe3beb · outbound
PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3fb9add0-3062-4433-bdf0-1c54e8dc20ea · outbound
PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent It has enabled open- domain question answering [3] by LLMs through the retrieval of evidence from large corpora and the generation of answers based on the retrieved context
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 435cc06d-50a3-40dd-ac41-ede8b629232d · outbound
PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent Then we provide details on our proposed coverage reward with teacher-guided training framework
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a0463496-3f85-4695-8e62-ed67c9fcb8c1 · outbound
PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d1fca82b-1a51-448a-a05d-20c18d89ff59 · outbound
PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0456a6a4-4db1-4e42-b685-43be9bbeb345 · outbound
PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation dc596767-7e36-40ec-9e0f-5150a5e5f9cd · outbound
PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent React: Synergizing reasoning and acting in lan- guage models,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9b1eb094-aace-4fff-a20c-40390a6b0152 · outbound
PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7ffc96a-a5c4-4c25-95c9-7060a7b281c3 · outbound
PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent Deep- researcher: Scaling deep research via reinforcement learning in real-world environments,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 64ff9fca-eef5-48d8-918e-39fcd395f3de · outbound
PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent Reading wikipedia to answer open-domain questions,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation bcccccfc-dc1b-4079-b3b5-a75eeba4d86c · outbound
PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent Natural questions: a benchmark for question answering re- search,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8f4edd40-8e4e-440b-9d8e-7bf609785314 · outbound
PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent Hotpotqa: A dataset for diverse, explain- able multi-hop question answering,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 664e026b-47a3-4c85-9345-a7b6aab7a057 · outbound
PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent Interleaving retrieval with chain-of-thought reasoning for knowledge-intensive multi-step questions,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 42d91b26-5ed7-48e4-ab77-65c6c60556c6 · outbound
PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent In parallel, Ze- roSearch [16] addresses the high cost and instability of RL by training with simulated retrieval during RL
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation dde2cfa6-3482-4b1b-9008-de6eb1cd687d · outbound
PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent Chain-of-retrieval augmented generation,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a09846f1-9385-43e4-9eb6-7e153ec07bd4 · outbound
PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent DeepRAG: Thinking to Retrieve Step by Step for Large Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12b88403-6e6b-49d1-a9fa-03a354bda57c · outbound
PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent Self- rag: Learning to retrieve, generate, and critique through self- reflection,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7c021b4d-6adc-46d6-8dcd-28b0f7a9678c · outbound
PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent HiPRAG: Hierarchical Process Rewards for Efficient Agentic Retrieval Augmented Generation
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbf7f50f-2bf9-4d17-9be9-06154b315522 · outbound
PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent Beyond the limitation of a single query: Train your llm for query expansion with reinforcement learning,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c39cdff7-14a1-448e-923c-1e6dbc1b3a2c · outbound
PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent Frugalrag: Less is more in rl finetuning for multi-hop question answering,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2381ac15-4ae2-4160-822a-36d73801b58d · outbound
PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0edeed5-dc2f-4829-b781-6d3531d018e5 · outbound
PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent Search wisely: Mitigating sub-optimal agentic searches by reducing uncertainty,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2c1bf993-9c97-46b7-a545-892b5d0e4681 · outbound
PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent ZeroSearch: Incentivize the Search Capability of LLMs without Searching
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc26b024-8e89-41d8-a5a1-cd82da4b0c6e · outbound
PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent Proximal Policy Optimization Algorithms
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0c35b74-3571-4c96-848e-98407ef02022 · outbound
PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4355908-5b51-401c-9410-c6ab1fbab9d4 · outbound
PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent When not to trust language models: Investigating effec- tiveness of parametric and non-parametric memories,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b61f4cc7-fb6b-47d7-b93b-0d4dd29fb68f · outbound
PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent Con- structing a multi-hop qa dataset for comprehensive evaluation of reasoning steps,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 83f4963e-bad6-4ddf-af4b-6b0a7d284845 · outbound
PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent Musique: Multihop questions via single-hop question composi- tion,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fb31c455-c342-4e19-8fb5-5bc8dfe86381 · outbound
PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent Dense passage retrieval for open-domain question answering,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5187a253-79c8-4d32-a87c-c7495542da9e · outbound
PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent Text Embeddings by Weakly-Supervised Contrastive Pre-training
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54373557-c971-46ed-9353-69796bbe3beb · inbound
PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.