Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:54:30.541560Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 14 inbound Pith citation observations for arXiv:2506.04185.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:54:30.541560Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T21:12:12.793725Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T06:19:38.663842Z
44 of 44 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation fbdc3536-4226-4620-aa14-1d80f49ea013 · outbound
R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning Unresolved cited work
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c155b641-b8d7-4924-8234-4a2ea31e74bf · outbound
R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning Unresolved cited work
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 728e8402-de64-4bb0-9f16-18a2cbdb3567 · outbound
R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f14615f3-bf2f-4d9f-a4a3-038dbcb03f25 · outbound
R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning RQ-RAG: Learning to Refine Queries for Retrieval Augmented Generation
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0458c6c1-fb1a-4acc-99d6-97775c449014 · outbound
R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ea1ee574-c118-4222-b4dd-7a7a6fd9f773 · outbound
R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b5b447c7-48d1-455a-bb56-c6389374d78c · outbound
R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4d27694-6d90-4eab-b4c9-3dfd99f42d05 · outbound
R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning Retrieval-Augmented Generation for Large Language Models: A Survey
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aefe27e4-2382-442a-85b2-ee35c1f8f2b1 · outbound
R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning DeepRAG: Thinking to Retrieve Step by Step for Large Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3dae0aa0-caff-4bfd-a084-1803c332fe30 · outbound
R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2117c3f-9435-42da-923f-427d56c31482 · outbound
R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning REALM: Retrieval-Augmented Language Model Pre-Training
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8cb34d6-2812-41e1-bfc2-cf3f690c9db6 · outbound
R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning Unresolved cited work
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f6099d7-897f-4297-ab19-3e0e65e07534 · outbound
R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning OpenAI o1 System Card
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 408e75e1-1756-4072-9b89-bcf8d9d25e38 · outbound
R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning Unresolved cited work
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52555175-f859-48d1-9c71-24b6822bb064 · outbound
R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning Xu, Luyu Gao, Zhiqing Sun, Qian Liu, Jane Dwivedi - Yu, Yiming Yang, Jamie Callan, and Graham Neubig
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c274fb56-4c2b-4edd-af2c-0b8fb4f8b6e6 · outbound
R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c9dacb0-3a3c-45a9-854f-ae11ede5480e · outbound
R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning FlashRAG: A Modular Toolkit for Efficient Retrieval-Augmented Generation Research
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 627dcc3a-6959-49c1-b1bf-f05d5267f601 · outbound
R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning Weld, and Luke Zettlemoyer
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20bd6a40-5b96-4d90-a4ef-8a52550e4607 · outbound
R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning Unresolved cited work
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c522333-8a80-4eb4-8078-4b876b79c84a · outbound
R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning Unresolved cited work
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17d2ab95-057d-402f-9bee-7ed50560ea97 · outbound
R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning Training Language Models to Self-Correct via Reinforcement Learning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca83f98c-38fb-4070-ab48-be9fd457c959 · outbound
R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning Unresolved cited work
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a196208-39d9-4970-9db7-a4b97e708d13 · outbound
R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning u ttler, Mike Lewis, Wen - tau Yih, Tim Rockt \
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46fda872-3ffb-4558-a455-605f44d7e1c4 · outbound
R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning Unresolved cited work
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c433b464-ea9f-4427-b624-05adb958da6f · outbound
R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c24b190f-2bfa-4b03-aadd-3d8b7aed4419 · outbound
R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f8bb9a67-7d7d-4f46-b83f-0bd423b7d5cf · outbound
R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning GPT-4 Technical Report
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d686e0a-dd7c-4f91-acc8-61c7629fb771 · outbound
R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation dc63a20b-8944-4eb4-bc6d-1ece1d2221cb · outbound
R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning Smith, and Mike Lewis
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcf82ccf-3bca-47ba-9c8a-d3e1b1016088 · outbound
R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning Hamilton, Chris Dyer, and Dani Yogatama
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bae75213-a5d1-42b0-86b0-f506496e4722 · outbound
R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning Proximal Policy Optimization Algorithms
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c30857ee-2823-4424-89db-3c9bbfc19591 · outbound
R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning Unresolved cited work
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8032718-9c9e-405e-a0c7-87b19773f16c · outbound
R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18bd1d35-7962-433c-85db-c9a5f62d1ff6 · outbound
R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning Unresolved cited work
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd59f18c-0d90-4f7e-a453-cdf0ff7fbca9 · outbound
R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning Unresolved cited work
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af55d101-5b80-4f5e-ac71-85e70df80ed6 · outbound
R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning Unresolved cited work
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c931d564-86c3-4854-8096-736595ec00c7 · outbound
R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning Large Language Models are Better Reasoners with Self-Verification
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a9be03b-020e-41a3-9b34-421110aa12fb · outbound
R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning Unresolved cited work
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 82a47b6f-2838-4763-95c5-5382dddf54a7 · outbound
R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning Qwen2.5 Technical Report
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9a73950-9e81-4df9-97c0-d657ab55938b · outbound
R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning Cohen, Ruslan Salakhutdinov, and Christopher D
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 274306c3-fb9d-4eaf-9901-3a7d595e016f · outbound
R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning Unresolved cited work
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2f4f4220-882a-42d8-b558-e67ea766ce81 · outbound
R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb7b1617-4827-4063-a3b5-eef1198cf550 · outbound
R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49404bab-a40d-4d0f-8bac-a4fa4024a915 · outbound
R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning Unresolved cited work
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a035ada7-ce11-4bbb-8836-820e9b7b309f · inbound
ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1aa069b-9b7b-46a1-aba8-ddf010ed2d6b · inbound
Erase to Improve: Erasable Reinforcement Learning for Search-Augmented LLMs R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a70854fa-5586-4614-b096-b05e0ba3d765 · inbound
Learning to Trust: Dynamic Utilization of Retrieval-Augmented Generation for E-commerce Search Relevance R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6f35d705-251a-4281-a182-3841362f4f59 · inbound
Beyond Correctness: Rewarding Faithful Reasoning in Retrieval-Augmented Generation R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84164662-8da3-46e0-b4b5-be266237471d · inbound
LatentRAG: Latent Reasoning and Retrieval for Efficient Agentic RAG R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7f9edbd2-78df-4063-88ba-511dbfaa9a57 · inbound
Dynamic Skill Lifecycle Management for Agentic Reinforcement Learning R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a8f2398a-4187-4976-af04-f952f24d8fdf · inbound
Dynamic Skill Lifecycle Management for Agentic Reinforcement Learning R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8c269676-9853-4254-ae24-6187fab04f09 · inbound
C-MIG: Multi-view Information Gain-based Retrieval-Augmented Generation for Clinical Diagnosis Reasoning R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 48e413c0-87bf-409f-9bd6-e0dcad6295ec · inbound
C-MIG: Multi-view Information Gain-based Retrieval-Augmented Generation for Clinical Diagnosis Reasoning R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff0c6177-c78e-43ad-b05b-23eb7ce5d77b · inbound
ARCO: Adaptive Rubrics with Co-Evolution for Multi-Step LLM-Based Agents R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 018f5e59-856c-457f-8186-0402c437d89f · inbound
ARCO: Adaptive Rubrics with Co-Evolution for Multi-Step LLM-Based Agents R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ee5da9b-6551-4441-a027-aeaa9de77755 · inbound
DocArena: Turning Raw Documents into Controllable Training Environments for Document Search Agents R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4edeba9a-bc41-4575-8bbc-4c477388e225 · inbound
Fishing Out Free Riders: Shapley-Based Reward Attribution for Parallel Reasoning via Reinforcement Learning R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e83d5df-5c7a-45e8-9376-d24c1ef67354 · inbound
Fetch-then-Explore: Decoupling Selection from Extraction over a Persistent Workspace for Search Agents R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning
Reference 104
Source-reported events for the cited work
Unavailable: canonical work link unavailable.