Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T20:17:14.224235Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 72 of 72 outbound references and 5 inbound Pith citation observations for arXiv:2508.10874.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T20:17:14.224235Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T16:18:44.717716Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T17:09:59.246273Z
72 of 72 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e660644d-2afa-4c06-bfb3-b2ace5a8519e · outbound
SSRL: Self-Search Reinforcement Learning Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5174f725-50ae-492a-a825-3a6c242c5328 · outbound
SSRL: Self-Search Reinforcement Learning Language models are few-shot learners
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81b68b3e-8c5a-4e6e-9f3b-686b6d3fa244 · outbound
SSRL: Self-Search Reinforcement Learning ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 066bb555-cc25-4e86-8aa1-0bcaf8648914 · outbound
SSRL: Self-Search Reinforcement Learning Process Reinforcement through Implicit Rewards
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 704fe00b-b2eb-4fd2-baee-79c97318e82e · outbound
SSRL: Self-Search Reinforcement Learning The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82eed66f-a399-4593-9a7b-cbc81c49c252 · outbound
SSRL: Self-Search Reinforcement Learning A Survey of Sim-to-Real Methods in RL: Progress, Prospects and Challenges with Foundation Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25d1b355-7e9d-44b4-b808-93831202c15f · outbound
SSRL: Self-Search Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fcc3ac4-0e5b-497c-a10c-6ff4f05bb5ef · outbound
SSRL: Self-Search Reinforcement Learning Citations and trust in llm generated responses
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 42c01a12-0b8f-491c-bbc5-b0c84ccbab26 · outbound
SSRL: Self-Search Reinforcement Learning Competitive Programming with Large Reasoning Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f631ddf0-5609-442e-b5b5-5a9908f3b02d · outbound
SSRL: Self-Search Reinforcement Learning ReTool: Reinforcement Learning for Strategic Tool Use in LLMs
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1635789a-eb37-4632-98e2-5d0c265d6d45 · outbound
SSRL: Self-Search Reinforcement Learning Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac084a07-791c-4efc-9c54-bb54ff541e1b · outbound
SSRL: Self-Search Reinforcement Learning A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5bbfe28-e7c3-489a-86d2-017851b123fb · outbound
SSRL: Self-Search Reinforcement Learning Enabling Large Language Models to Generate Text with Citations
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d62b9b14-86a4-4930-9d3e-b97ab0a2f297 · outbound
SSRL: Self-Search Reinforcement Learning The Llama 3 Herd of Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3aaea7a4-83e8-4d98-933a-e9d41467abcc · outbound
SSRL: Self-Search Reinforcement Learning Is Your LLM Secretly a World Model of the Internet? Model-Based Planning for Web Agents
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f829c9f-5fbd-4ad7-bc12-8da44b93d272 · outbound
SSRL: Self-Search Reinforcement Learning Reasoning with Language Model is Planning with World Model
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e79197a4-5345-4fe4-8716-3f031433b673 · outbound
SSRL: Self-Search Reinforcement Learning Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 669820f7-1974-43f8-85c3-eb94427b164f · outbound
SSRL: Self-Search Reinforcement Learning REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e06a672-a89a-4390-a2df-320dc727fe33 · outbound
SSRL: Self-Search Reinforcement Learning An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a52d4d0-b364-4586-b9d8-587a9ba961d1 · outbound
SSRL: Self-Search Reinforcement Learning Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3cff19dd-0333-4be6-91eb-5e001564d9a5 · outbound
SSRL: Self-Search Reinforcement Learning TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1c5cf19-f396-4075-800d-3440538a65b6 · outbound
SSRL: Self-Search Reinforcement Learning Sim2real transfer for reinforcement learning without dynamics randomization
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 64d4c93a-bbf1-4918-8e8a-76fe19db7a7a · outbound
SSRL: Self-Search Reinforcement Learning Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71bf551d-680e-43a1-a83b-17835cde46b4 · outbound
SSRL: Self-Search Reinforcement Learning Scalable agent alignment via reward modeling: a research direction
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cae576c8-daad-4d64-87ee-56724e58b22c · outbound
SSRL: Self-Search Reinforcement Learning S*: Test Time Scaling for Code Generation
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0c76f63-b0dc-4add-bdbb-cefe3781db74 · outbound
SSRL: Self-Search Reinforcement Learning Emergent world representations: Exploring a sequence model trained on a synthetic task
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 34037597-400a-4e64-99f6-751bc4f6b000 · outbound
SSRL: Self-Search Reinforcement Learning CorpusLM: Towards a Unified Language Model on Corpus for Knowledge-Intensive Tasks
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 25160b53-2978-4454-8697-0e4ca2084107 · outbound
SSRL: Self-Search Reinforcement Learning Search-o1: Agentic Search-Enhanced Large Reasoning Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe433099-c955-48a3-be81-58857651866a · outbound
SSRL: Self-Search Reinforcement Learning From matching to generation: A survey on generative information retrieval
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b65de2a5-3dfe-497a-805a-14bbe9d2336c · outbound
SSRL: Self-Search Reinforcement Learning A Survey of Generative Search and Recommendation in the Era of Large Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9cbfc218-ed6b-42a1-b648-4e03b3fb8028 · outbound
SSRL: Self-Search Reinforcement Learning Learning to rank in generative retrieval
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2abfb961-6d12-41ba-b824-58ff7ed19bc7 · outbound
SSRL: Self-Search Reinforcement Learning Multi-Agent Verification: Scaling Test-Time Compute with Multiple Verifiers
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 981d5a87-31a4-489d-b30b-7edf2723ae84 · outbound
SSRL: Self-Search Reinforcement Learning DeepSeek-V3 Technical Report
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0972242c-07d0-4677-81ae-206435b91c3f · outbound
SSRL: Self-Search Reinforcement Learning Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34cd607f-393e-46b8-b686-fedebd29a29c · outbound
SSRL: Self-Search Reinforcement Learning Understanding R1-Zero-Like Training: A Critical Perspective
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cc82a03-e7dc-4e1f-9ff4-38693d2f6e13 · outbound
SSRL: Self-Search Reinforcement Learning Generative multi-modal knowledge retrieval with large language models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a09b9e17-e2c0-4bc8-926a-6aa6d86d27ef · outbound
SSRL: Self-Search Reinforcement Learning OpenAI o1 System Card
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbf8efaf-4312-4d5e-984a-68622d4d12cc · outbound
SSRL: Self-Search Reinforcement Learning Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42ffeb56-5f9b-421a-b8a4-120aebb67274 · outbound
SSRL: Self-Search Reinforcement Learning Measuring and Narrowing the Compositionality Gap in Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e82d6f1e-83ba-4f7a-9dd9-24b42ab33272 · outbound
SSRL: Self-Search Reinforcement Learning ToolRL: Reward is All Tool Learning Needs
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4275843b-7c40-4f5d-91ff-4e5dfaa92949 · outbound
SSRL: Self-Search Reinforcement Learning WebCPM: Interactive Web Search for Chinese Long-form Question Answering
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fb3313db-e19b-4ff9-b1ad-dbfe0c1f15ce · outbound
SSRL: Self-Search Reinforcement Learning TreeBoN: Enhancing Inference-Time Alignment with Speculative Tree-Search and Best-of-N Sampling
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 166ef671-64e5-4e50-9dab-f72d13048071 · outbound
SSRL: Self-Search Reinforcement Learning Qwen2.5 Technical Report
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e27dd7cf-c73d-43e3-9281-374eed56989c · outbound
SSRL: Self-Search Reinforcement Learning Proximal Policy Optimization Algorithms
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a7746c4-18cf-4309-beb6-d3d9beed30a0 · outbound
SSRL: Self-Search Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78856db5-fe51-411c-96bf-e5a93e222d04 · outbound
SSRL: Self-Search Reinforcement Learning Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9038f3a-fd3a-454f-9deb-93ccf45879b9 · outbound
SSRL: Self-Search Reinforcement Learning HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39bd7740-f962-4252-8546-8abd5810261a · outbound
SSRL: Self-Search Reinforcement Learning Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8f7c3d2-1a78-4583-892c-4749ddd15efa · outbound
SSRL: Self-Search Reinforcement Learning ZeroSearch: Incentivize the Search Capability of LLMs without Searching
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97d453e7-df30-4c77-9a58-058311b88095 · outbound
SSRL: Self-Search Reinforcement Learning Worldcoder, a model-based llm agent: Building world models by writing code and interacting with the environment
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2221edb6-7c2d-4f56-9f77-b4833d2fdb2b · outbound
SSRL: Self-Search Reinforcement Learning Transformer memory as a differentiable search index
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19cb28b6-ecbe-45e0-9a22-0be3e0b0946d · outbound
SSRL: Self-Search Reinforcement Learning Kimi K2: Open Agentic Intelligence
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e34e3623-2212-45eb-996d-5d4a39a593ff · outbound
SSRL: Self-Search Reinforcement Learning MuSiQue: Multihop Questions via Single-hop Question Composition
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84542cda-df2e-4f4b-ad73-aa6fa116ac58 · outbound
SSRL: Self-Search Reinforcement Learning A neural corpus indexer for document retrieval
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef6e6e4c-790c-4b24-8c58-4716de72ec9d · outbound
SSRL: Self-Search Reinforcement Learning OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac04a7b9-7d04-41c4-be28-29de60856bbc · outbound
SSRL: Self-Search Reinforcement Learning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a8f45fd-17ae-43a2-88d7-41a99c1bec28 · outbound
SSRL: Self-Search Reinforcement Learning Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 914e4dc4-33d4-4f3d-9017-87ea06285be2 · outbound
SSRL: Self-Search Reinforcement Learning Measuring short-form factuality in large language models
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65abb7e1-0218-4aa4-80d4-013c5f0659fc · outbound
SSRL: Self-Search Reinforcement Learning BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27d9c8ff-0767-4307-923d-b64106e9eaf9 · outbound
SSRL: Self-Search Reinforcement Learning Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46d3f3ca-4670-4aad-9895-4aff26ec06c8 · outbound
SSRL: Self-Search Reinforcement Learning Qwen3 Technical Report
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6745111-7cb8-4e4f-9809-50c5f97366c1 · outbound
SSRL: Self-Search Reinforcement Learning GTA1: GUI Test-time Scaling Agent
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d32414e7-31d1-4337-9f87-0812da702845 · outbound
SSRL: Self-Search Reinforcement Learning HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86181fa9-55ff-4a86-b63b-b2e198330cfd · outbound
SSRL: Self-Search Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1541f30a-c8b1-4377-9cba-48384864dc1f · outbound
SSRL: Self-Search Reinforcement Learning Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eaee824b-2e7b-430b-a2b6-27158195c2c0 · outbound
SSRL: Self-Search Reinforcement Learning DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c74b1a72-ace6-4ffb-8732-1be03af50989 · outbound
SSRL: Self-Search Reinforcement Learning Scaling Test-time Compute for LLM Agents
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5c57250-75b1-47c9-a70b-708ea24a4f42 · outbound
SSRL: Self-Search Reinforcement Learning TTRL: Test-Time Reinforcement Learning
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c01981d2-1ba7-461f-8d22-5c30ae39d937 · outbound
SSRL: Self-Search Reinforcement Learning write newline
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8133ea35-7fd3-4484-bca2-6b690abcd79c · outbound
SSRL: Self-Search Reinforcement Learning @esa (Ref
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7510aa87-e88a-49dd-b9bd-7891ea4830a5 · outbound
SSRL: Self-Search Reinforcement Learning Unresolved cited work
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69291813-5d61-4668-b016-ea021efba23c · outbound
SSRL: Self-Search Reinforcement Learning Unresolved cited work
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b4875cd-9bb5-40da-a5e3-985edbd95183 · inbound
Dynamic Collaboration of Multi-Language Models based on Minimal Complete Semantic Units SSRL: Self-Search Reinforcement Learning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f591b89f-408e-4b86-9ffb-52c78f9a0b70 · inbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey SSRL: Self-Search Reinforcement Learning
Reference 292
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 997be28b-ee60-4cac-a025-0ed4cfbcf684 · inbound
A Survey of Reinforcement Learning for Large Reasoning Models SSRL: Self-Search Reinforcement Learning
Reference 132
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c7c41de9-4fe9-4c31-8288-f52bc1be63bb · inbound
Erase to Improve: Erasable Reinforcement Learning for Search-Augmented LLMs SSRL: Self-Search Reinforcement Learning
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9a8e4261-7149-48d1-8b52-f4ca13a5f9c2 · inbound
Qwen-AgentWorld: Language World Models for General Agents SSRL: Self-Search Reinforcement Learning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.