Pith. sign in

Paper Citation Record · LEDGER

SSRL: Self-Search Reinforcement Learning

As of 19 August 2026, this Paper Citation Record lists 72 of 72 outbound references and 5 inbound Pith citation observations for arXiv:2508.10874.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.10874 v1

Coverage vector

measured 72 of 72 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T20:17:14.224235Z

measured 77 of 77 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T16:18:44.717716Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T17:09:59.246273Z

Reference resolution

72 of 72 outbound references displayed

  • verified exact2
  • verified fuzzy6
  • unresolved64
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e660644d-2afa-4c06-bfb3-b2ace5a8519e · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

SSRL: Self-Search Reinforcement Learning Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:07.082566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:07.082566Z digest=sha256:b99c0295afe0a894e2c7525ec6dd860837839e88d0f212f0ebba0d21d14b7715

Observation 5174f725-50ae-492a-a825-3a6c242c5328 · outbound

This paper cites Language models are few-shot learners.

SSRL: Self-Search Reinforcement Learning Language models are few-shot learners

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:07.140938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:07.140938Z digest=sha256:dd5f26afb9b66dede24a33eee7282bf25e377cea2d44fe42d998cbecd7bad432

Observation 81b68b3e-8c5a-4e6e-9f3b-686b6d3fa244 · outbound

This paper cites ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning.

SSRL: Self-Search Reinforcement Learning ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:07.244058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:07.244058Z digest=sha256:33e2f58e1f4620bfbfbf5f83952e4e75a152a8f6a6d7fe7cb69a05573a9ea62c

Observation 066bb555-cc25-4e86-8aa1-0bcaf8648914 · outbound

This paper cites Process Reinforcement through Implicit Rewards.

SSRL: Self-Search Reinforcement Learning Process Reinforcement through Implicit Rewards

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:07.379560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:07.379560Z digest=sha256:a95427c7ed027d4124a6eb58590060be02e5a2aa37c00264c4e95d63d7911c9f

Observation 704fe00b-b2eb-4fd2-baee-79c97318e82e · outbound

This paper cites The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models.

SSRL: Self-Search Reinforcement Learning The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:07.516675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:07.516675Z digest=sha256:ccdf769ae3edc3997d3708ff50b335210e6bfcd6480eee8a62efa2fff0009f94

Observation 82eed66f-a399-4593-9a7b-cbc81c49c252 · outbound

This paper cites A Survey of Sim-to-Real Methods in RL: Progress, Prospects and Challenges with Foundation Models.

SSRL: Self-Search Reinforcement Learning A Survey of Sim-to-Real Methods in RL: Progress, Prospects and Challenges with Foundation Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:07.610161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:07.610161Z digest=sha256:40bded4c9c769c2ff7b290f4bd94d5e15c702b18f23d55f41f8d402c9c7444c8

Observation 25d1b355-7e9d-44b4-b808-93831202c15f · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

SSRL: Self-Search Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:07.780237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:07.780237Z digest=sha256:a48fffe34b82134984bb022ed6acc353b3619c34b79208b6b662303b14e59dd6

Observation 9fcc3ac4-0e5b-497c-a10c-6ff4f05bb5ef · outbound

This paper cites Citations and trust in llm generated responses.

SSRL: Self-Search Reinforcement Learning Citations and trust in llm generated responses

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:17:16.055399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-05T20:17:07.896751Z digest=sha256:08e7acedd7a859abe37a1298155e79670cd8865096b966c4f98f970d9a71d611

Observation 42c01a12-0b8f-491c-bbc5-b0c84ccbab26 · outbound

This paper cites Competitive Programming with Large Reasoning Models.

SSRL: Self-Search Reinforcement Learning Competitive Programming with Large Reasoning Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:08.011061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:08.011061Z digest=sha256:b8406690279df44c47407578cf35f8e1eff928628b2bb2cbb47506b3c02b083e

Observation f631ddf0-5609-442e-b5b5-5a9908f3b02d · outbound

This paper cites ReTool: Reinforcement Learning for Strategic Tool Use in LLMs.

SSRL: Self-Search Reinforcement Learning ReTool: Reinforcement Learning for Strategic Tool Use in LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:08.153794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:08.153794Z digest=sha256:58248f110f94196068eb2daa3b38c4c2e467a9ec293b1a95ee87a8a2adc390d7

Observation 1635789a-eb37-4632-98e2-5d0c265d6d45 · outbound

This paper cites Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs.

SSRL: Self-Search Reinforcement Learning Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:08.299466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:08.299466Z digest=sha256:2be1afb0eef12d507bb904ef225593dc2d6061525c4848e14b5b95d019c959ec

Observation ac084a07-791c-4efc-9c54-bb54ff541e1b · outbound

This paper cites A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence.

SSRL: Self-Search Reinforcement Learning A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:08.416770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:08.416770Z digest=sha256:7565a82e0bc38d700772edbe8f0ea345771c56be8211a31b3a0e6a7dc7b7ab48

Observation e5bbfe28-e7c3-489a-86d2-017851b123fb · outbound

This paper cites Enabling Large Language Models to Generate Text with Citations.

SSRL: Self-Search Reinforcement Learning Enabling Large Language Models to Generate Text with Citations

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:08.576057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:08.576057Z digest=sha256:1847cf5860959500e0996813ba56363755fc55bf65aaa2fa5962afe14c08c76a

Observation d62b9b14-86a4-4930-9d3e-b97ab0a2f297 · outbound

This paper cites The Llama 3 Herd of Models.

SSRL: Self-Search Reinforcement Learning The Llama 3 Herd of Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:08.686291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:08.686291Z digest=sha256:de404f8e8e084d99acc8cac9646425ef0b4c0020ffbdacddc03f4996ef7034df

Observation 3aaea7a4-83e8-4d98-933a-e9d41467abcc · outbound

This paper cites Is Your LLM Secretly a World Model of the Internet? Model-Based Planning for Web Agents.

SSRL: Self-Search Reinforcement Learning Is Your LLM Secretly a World Model of the Internet? Model-Based Planning for Web Agents

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:08.840484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:08.840484Z digest=sha256:4ec794255fc005e77c1935df6f41b23499db0a8902b52b66824235cb427cc1ff

Observation 7f829c9f-5fbd-4ad7-bc12-8da44b93d272 · outbound

This paper cites Reasoning with Language Model is Planning with World Model.

SSRL: Self-Search Reinforcement Learning Reasoning with Language Model is Planning with World Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:09.037240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:09.037240Z digest=sha256:cf98668b76615cbec9c0f793d7f6ae78981e344327cb9b8c382c96c9bad9490b

Observation e79197a4-5345-4fe4-8716-3f031433b673 · outbound

This paper cites Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps.

SSRL: Self-Search Reinforcement Learning Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:09.215321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:09.215321Z digest=sha256:f3c7fed9086ea1b1839b9bda2e0a7bc8d0089a99c373235e08b600e0ba5ba61b

Observation 669820f7-1974-43f8-85c3-eb94427b164f · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

SSRL: Self-Search Reinforcement Learning REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:09.323202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:09.323202Z digest=sha256:f95d891506ae65ad4f076888c08140ea109c72d7fb32bf50302499ecf21a85ba

Observation 3e06a672-a89a-4390-a2df-320dc727fe33 · outbound

This paper cites An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents.

SSRL: Self-Search Reinforcement Learning An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:09.393341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:09.393341Z digest=sha256:4b6c19b252ccb601af34e1313065a684a753c8e8e6be1b6754cade3a842b629a

Observation 1a52d4d0-b364-4586-b9d8-587a9ba961d1 · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

SSRL: Self-Search Reinforcement Learning Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:09.409539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:09.409539Z digest=sha256:ef2671bbbcf29a4f379466af1465ea43af42d865df520db5e2c8558110b5f0be

Observation 3cff19dd-0333-4be6-91eb-5e001564d9a5 · outbound

This paper cites TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension.

SSRL: Self-Search Reinforcement Learning TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:09.457321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:09.457321Z digest=sha256:1cae30cd64085c0fb964a0cf3ec200f13a7d1faf07d76e6d0e1c4af32c764584

Observation d1c5cf19-f396-4075-800d-3440538a65b6 · outbound

This paper cites Sim2real transfer for reinforcement learning without dynamics randomization.

SSRL: Self-Search Reinforcement Learning Sim2real transfer for reinforcement learning without dynamics randomization

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:17:15.944894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-05T20:17:09.570035Z digest=sha256:4a6404d2f0f69d4382a1ffd0c3b7effcd6a79ffeadc7f4e5c9fde95f36a29cd2

Observation 64d4c93a-bbf1-4918-8e8a-76fe19db7a7a · outbound

This paper cites Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov.

SSRL: Self-Search Reinforcement Learning Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:09.639320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:09.639320Z digest=sha256:ada76608e42d48903a2931e363cf288b4dd3cd14410054cae81e878769208a59

Observation 71bf551d-680e-43a1-a83b-17835cde46b4 · outbound

This paper cites Scalable agent alignment via reward modeling: a research direction.

SSRL: Self-Search Reinforcement Learning Scalable agent alignment via reward modeling: a research direction

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:09.710424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:09.710424Z digest=sha256:8b2e15bda03f3fddcfd5876002fae5490fa1e8df01bca97d2d3f42ce2ada1790

Observation cae576c8-daad-4d64-87ee-56724e58b22c · outbound

This paper cites S*: Test Time Scaling for Code Generation.

SSRL: Self-Search Reinforcement Learning S*: Test Time Scaling for Code Generation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:09.862906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:09.862906Z digest=sha256:5f5527c318777b23dc67dd569badefbf305132cceec1c6ba9ac1c4eb138c4ab4

Observation c0c76f63-b0dc-4add-bdbb-cefe3781db74 · outbound

This paper cites Emergent world representations: Exploring a sequence model trained on a synthetic task.

SSRL: Self-Search Reinforcement Learning Emergent world representations: Exploring a sequence model trained on a synthetic task

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:17:15.829090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-05T20:17:09.970482Z digest=sha256:b5473177cfb07f1c27117005c2d6b0b18444f8b94a4cbde15bfc7f89d6c57755

Observation 34037597-400a-4e64-99f6-751bc4f6b000 · outbound

This paper cites CorpusLM: Towards a Unified Language Model on Corpus for Knowledge-Intensive Tasks.

SSRL: Self-Search Reinforcement Learning CorpusLM: Towards a Unified Language Model on Corpus for Knowledge-Intensive Tasks

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-05T20:17:15.004781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-05T20:17:10.055596Z digest=sha256:0ba0fd6946f237fa1fc18673852800d5140b2f9d3bafe694ca742a84138d28bf

Observation 25160b53-2978-4454-8697-0e4ca2084107 · outbound

This paper cites Search-o1: Agentic Search-Enhanced Large Reasoning Models.

SSRL: Self-Search Reinforcement Learning Search-o1: Agentic Search-Enhanced Large Reasoning Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:10.120457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:10.120457Z digest=sha256:6ab8753ec2abc3297941ed8265ba58fa6d7d7813b7d2b6365470a51b8c3fbd5f

Observation fe433099-c955-48a3-be81-58857651866a · outbound

This paper cites From matching to generation: A survey on generative information retrieval.

SSRL: Self-Search Reinforcement Learning From matching to generation: A survey on generative information retrieval

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:17:15.701341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-05T20:17:10.230492Z digest=sha256:a0cde49aef6b69445199fbbf5d5dc1a3bd021fda6a2407b4215d23fbd6da2a91

Observation b65de2a5-3dfe-497a-805a-14bbe9d2336c · outbound

This paper cites A Survey of Generative Search and Recommendation in the Era of Large Language Models.

SSRL: Self-Search Reinforcement Learning A Survey of Generative Search and Recommendation in the Era of Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:10.345154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:10.345154Z digest=sha256:343fe4b03ba04a0669cff1f2374c8753bdb810bdb99695aa47cbf635c99f38e6

Observation 9cbfc218-ed6b-42a1-b648-4e03b3fb8028 · outbound

This paper cites Learning to rank in generative retrieval.

SSRL: Self-Search Reinforcement Learning Learning to rank in generative retrieval

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:17:15.572763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-05T20:17:10.430071Z digest=sha256:9a6ec839c90c3fdec13253403375256fad3156de4a6b83eb4cf916910bd25629

Observation 2abfb961-6d12-41ba-b824-58ff7ed19bc7 · outbound

This paper cites Multi-Agent Verification: Scaling Test-Time Compute with Multiple Verifiers.

SSRL: Self-Search Reinforcement Learning Multi-Agent Verification: Scaling Test-Time Compute with Multiple Verifiers

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:10.502885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:10.502885Z digest=sha256:4908d4c5598a00542db1cd6bdd9a3b16bb0f69c52b91855fa19281e0e20b4808

Observation 981d5a87-31a4-489d-b30b-7edf2723ae84 · outbound

This paper cites DeepSeek-V3 Technical Report.

SSRL: Self-Search Reinforcement Learning DeepSeek-V3 Technical Report

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:10.603582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:10.603582Z digest=sha256:b1aa29f4dfcb676d0052da3dc7b896ddac4368abd8ec02f6513ca11aa6d53b3c

Observation 0972242c-07d0-4677-81ae-206435b91c3f · outbound

This paper cites Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling.

SSRL: Self-Search Reinforcement Learning Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:10.690804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:10.690804Z digest=sha256:303341291271b43274d2977ccd1191a5a8d597dcd98a7f8e25a6f2b72359b189

Observation 34cd607f-393e-46b8-b686-fedebd29a29c · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

SSRL: Self-Search Reinforcement Learning Understanding R1-Zero-Like Training: A Critical Perspective

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:10.754730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:10.754730Z digest=sha256:c331ce0a5bc793a90ebb5f3fc09ecca2e2974d5f6b1bb363bdeacf09ad5304d5

Observation 2cc82a03-e7dc-4e1f-9ff4-38693d2f6e13 · outbound

This paper cites Generative multi-modal knowledge retrieval with large language models.

SSRL: Self-Search Reinforcement Learning Generative multi-modal knowledge retrieval with large language models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:10.833429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:10.833429Z digest=sha256:04478fb592917f1e7f4c3eee6cd698a472ac1ae8b4a39500b2a4ca612cf7c973

Observation a09b9e17-e2c0-4bc8-926a-6aa6d86d27ef · outbound

This paper cites OpenAI o1 System Card.

SSRL: Self-Search Reinforcement Learning OpenAI o1 System Card

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:10.945859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:10.945859Z digest=sha256:7bedfc1e427f79667c201dd19419e22c1194944d2902ff11e390a6fe47d63565

Observation bbf8efaf-4312-4d5e-984a-68622d4d12cc · outbound

This paper cites Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad.

SSRL: Self-Search Reinforcement Learning Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:11.057856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:11.057856Z digest=sha256:f8eb639b0e9964999c83b5eaabd5e521351143d6c01135019b2c1d33832bb174

Observation 42ffeb56-5f9b-421a-b8a4-120aebb67274 · outbound

This paper cites Measuring and Narrowing the Compositionality Gap in Language Models.

SSRL: Self-Search Reinforcement Learning Measuring and Narrowing the Compositionality Gap in Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:11.138260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:11.138260Z digest=sha256:1cec42ce9b564bed061ddb37bbf565e249c3a85c71ab5cb943719470b72fdbad

Observation e82d6f1e-83ba-4f7a-9dd9-24b42ab33272 · outbound

This paper cites ToolRL: Reward is All Tool Learning Needs.

SSRL: Self-Search Reinforcement Learning ToolRL: Reward is All Tool Learning Needs

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:11.220482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:11.220482Z digest=sha256:eb9c2ecbce5e08b1738fdea6bf84a7aad617ef69ece5e6bc457384769f92ab24

Observation 4275843b-7c40-4f5d-91ff-4e5dfaa92949 · outbound

This paper cites WebCPM: Interactive Web Search for Chinese Long-form Question Answering.

SSRL: Self-Search Reinforcement Learning WebCPM: Interactive Web Search for Chinese Long-form Question Answering

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-08-05T20:17:14.717417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-05T20:17:11.334274Z digest=sha256:f5e158062ca488b531a7c7d0c7471efce1c9e3715a2877e603c885a3a2330157

Observation fb3313db-e19b-4ff9-b1ad-dbfe0c1f15ce · outbound

This paper cites TreeBoN: Enhancing Inference-Time Alignment with Speculative Tree-Search and Best-of-N Sampling.

SSRL: Self-Search Reinforcement Learning TreeBoN: Enhancing Inference-Time Alignment with Speculative Tree-Search and Best-of-N Sampling

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:11.414943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:11.414943Z digest=sha256:4a6c21ab1bb782ee79feeb452d1148170d69807fdf91aa458a4607120f724654

Observation 166ef671-64e5-4e50-9dab-f72d13048071 · outbound

This paper cites Qwen2.5 Technical Report.

SSRL: Self-Search Reinforcement Learning Qwen2.5 Technical Report

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:11.513187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:11.513187Z digest=sha256:0619093179afed57d5790a15e8e917736826ce097c48ec028b50dda25e6465be

Observation e27dd7cf-c73d-43e3-9281-374eed56989c · outbound

This paper cites Proximal Policy Optimization Algorithms.

SSRL: Self-Search Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:11.582880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:11.582880Z digest=sha256:3a304c36fb297e6703c1bc950af4c0d19ed30436dd0d2ffe1631c4a4f907ed28

Observation 4a7746c4-18cf-4309-beb6-d3d9beed30a0 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

SSRL: Self-Search Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:11.674940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:11.674940Z digest=sha256:879eb2e26d900c6d72af5ca6fdd01dfb1aa06887e0c15b670cf25cefb0ae0617

Observation 78856db5-fe51-411c-96bf-e5a93e222d04 · outbound

This paper cites Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction.

SSRL: Self-Search Reinforcement Learning Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:11.770737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:11.770737Z digest=sha256:52fa0f1b0eb0e76fd8e8eb952f2712e26a5bd86be8d834197c53a5ff0d9f823e

Observation e9038f3a-fd3a-454f-9deb-93ccf45879b9 · outbound

This paper cites HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face.

SSRL: Self-Search Reinforcement Learning HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:11.877617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:11.877617Z digest=sha256:d6479996819d121f8dcff8306404e159565a2fa2b219f9083df1079c67e629ce

Observation 39bd7740-f962-4252-8546-8abd5810261a · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

SSRL: Self-Search Reinforcement Learning Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:11.983940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:11.983940Z digest=sha256:169f10cdadf5e8dfc16d02d184e03dde851056fc8601ac991c249ea5a90267f0

Observation b8f7c3d2-1a78-4583-892c-4749ddd15efa · outbound

This paper cites ZeroSearch: Incentivize the Search Capability of LLMs without Searching.

SSRL: Self-Search Reinforcement Learning ZeroSearch: Incentivize the Search Capability of LLMs without Searching

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:12.061269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:12.061269Z digest=sha256:70887e1910aa3f9026294faed5135920cf20ac3c9b22e6f62738a0d545a957e7

Observation 97d453e7-df30-4c77-9a58-058311b88095 · outbound

This paper cites Worldcoder, a model-based llm agent: Building world models by writing code and interacting with the environment.

SSRL: Self-Search Reinforcement Learning Worldcoder, a model-based llm agent: Building world models by writing code and interacting with the environment

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:17:15.423680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-05T20:17:12.144246Z digest=sha256:54512b044086a6b2c3d2ffabe25e59fb45cf442fdd3f489020b2cabb4b3faf7a

Observation 2221edb6-7c2d-4f56-9f77-b4833d2fdb2b · outbound

This paper cites Transformer memory as a differentiable search index.

SSRL: Self-Search Reinforcement Learning Transformer memory as a differentiable search index

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:12.247238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:12.247238Z digest=sha256:e9d8e39675292b907dad0e37ee2bb223da9eef2ba45012388f4bacdb73fa4baf

Observation 19cb28b6-ecbe-45e0-9a22-0be3e0b0946d · outbound

This paper cites Kimi K2: Open Agentic Intelligence.

SSRL: Self-Search Reinforcement Learning Kimi K2: Open Agentic Intelligence

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:12.345692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:12.345692Z digest=sha256:d47b7cc9f1dde63d412fdfccd29507db6ecefd22d327c93c3fe3c6a9d0c56f9d

Observation e34e3623-2212-45eb-996d-5d4a39a593ff · outbound

This paper cites MuSiQue: Multihop Questions via Single-hop Question Composition.

SSRL: Self-Search Reinforcement Learning MuSiQue: Multihop Questions via Single-hop Question Composition

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:12.419079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:12.419079Z digest=sha256:bf7975e4eefb1829c8e1ab91b326a207faafab8d118fd2f3c39f974781761e8d

Observation 84542cda-df2e-4f4b-ad73-aa6fa116ac58 · outbound

This paper cites A neural corpus indexer for document retrieval.

SSRL: Self-Search Reinforcement Learning A neural corpus indexer for document retrieval

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:12.524672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:12.524672Z digest=sha256:d05f080c26ec6e265ad56787956523145dd29c9097350a217e2325b92909f4dd

Observation ef6e6e4c-790c-4b24-8c58-4716de72ec9d · outbound

This paper cites OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling.

SSRL: Self-Search Reinforcement Learning OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:12.589663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:12.589663Z digest=sha256:10d154c5e5bd0c45dbd4e8fe16d5d402130ed434abc83ff54a034dd87206e751

Observation ac04a7b9-7d04-41c4-be28-29de60856bbc · outbound

This paper cites RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning.

SSRL: Self-Search Reinforcement Learning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:12.728962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:12.728962Z digest=sha256:aea48218c328ca46aa06f7fcd63a81605238376b0a4cb176a850ce5a43392d6a

Observation 7a8f45fd-17ae-43a2-88d7-41a99c1bec28 · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

SSRL: Self-Search Reinforcement Learning Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:12.825794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:12.825794Z digest=sha256:32f98734b1cb03afbcd7df48675481972d9d7fa68d4f487c66231cf9e675a14b

Observation 914e4dc4-33d4-4f3d-9017-87ea06285be2 · outbound

This paper cites Measuring short-form factuality in large language models.

SSRL: Self-Search Reinforcement Learning Measuring short-form factuality in large language models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:12.901494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:12.901494Z digest=sha256:34cd5f9cec73e4dea570605e05cc9890090369b27f90aa7e05818ccccccfa506

Observation 65abb7e1-0218-4aa4-80d4-013c5f0659fc · outbound

This paper cites BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents.

SSRL: Self-Search Reinforcement Learning BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:13.008368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:13.008368Z digest=sha256:b22cd377829c409c5d227f4e62017a08fb97f84977976ae4c678c8bddc32fcf5

Observation 27d9c8ff-0767-4307-923d-b64106e9eaf9 · outbound

This paper cites Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models.

SSRL: Self-Search Reinforcement Learning Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:13.081546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:13.081546Z digest=sha256:ecdeee8be5acd072d035f07d815ce7d2b72923f29de6b32d70c1a508b9ae645e

Observation 46d3f3ca-4670-4aad-9895-4aff26ec06c8 · outbound

This paper cites Qwen3 Technical Report.

SSRL: Self-Search Reinforcement Learning Qwen3 Technical Report

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:13.289705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:13.289705Z digest=sha256:05174349e3be7853b8a67cabfc5e5a9744a8966aac891945e08b278faffb8e20

Observation b6745111-7cb8-4e4f-9809-50c5f97366c1 · outbound

This paper cites GTA1: GUI Test-time Scaling Agent.

SSRL: Self-Search Reinforcement Learning GTA1: GUI Test-time Scaling Agent

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:13.398421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:13.398421Z digest=sha256:41e0f2adf61b3edaa4406446ff2744c0a8ef0385c035345da52bbb7f41f9ee0c

Observation d32414e7-31d1-4337-9f87-0812da702845 · outbound

This paper cites HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering.

SSRL: Self-Search Reinforcement Learning HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:13.464960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:13.464960Z digest=sha256:63560023141953f8cc1e7be0a8bf37cdc9e9fb61b7992632c12f0d2791dcb303

Observation 86181fa9-55ff-4a86-b63b-b2e198330cfd · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

SSRL: Self-Search Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:13.572346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:13.572346Z digest=sha256:3e25cb42c83c896a3faa7e2432fe2bf05264e7831b47535e8a51b325d7acf9fb

Observation 1541f30a-c8b1-4377-9cba-48384864dc1f · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

SSRL: Self-Search Reinforcement Learning Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:13.673603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:13.673603Z digest=sha256:262762c3e6deb75bb76f69c01969e80f359746d9c0c1330ac4ad532611d476b6

Observation eaee824b-2e7b-430b-a2b6-27158195c2c0 · outbound

This paper cites DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments.

SSRL: Self-Search Reinforcement Learning DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:13.787573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:13.787573Z digest=sha256:735ac6ad6d205e9cdab58648ce927837e89faef5535ea68041f1d924b0a291de

Observation c74b1a72-ace6-4ffb-8732-1be03af50989 · outbound

This paper cites Scaling Test-time Compute for LLM Agents.

SSRL: Self-Search Reinforcement Learning Scaling Test-time Compute for LLM Agents

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:13.855017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:13.855017Z digest=sha256:4f3cb5f1cd0e799cf36c209e28bd8a17127234e6dd1816c2845d6420a0df66b7

Observation e5c57250-75b1-47c9-a70b-708ea24a4f42 · outbound

This paper cites TTRL: Test-Time Reinforcement Learning.

SSRL: Self-Search Reinforcement Learning TTRL: Test-Time Reinforcement Learning

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:13.935000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:13.935000Z digest=sha256:2e4bac2aba1c82f62b0a282c6c7c635518a00c5027ec7d951002673c62d99a19

Observation c01981d2-1ba7-461f-8d22-5c30ae39d937 · outbound

This paper cites write newline.

SSRL: Self-Search Reinforcement Learning write newline

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:14.035628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:14.035628Z digest=sha256:24538ef23232525ae042fc08a399735ee23a25aaf1bc1108ceb8b74daf919149

Observation 8133ea35-7fd3-4484-bca2-6b690abcd79c · outbound

This paper cites @esa (Ref.

SSRL: Self-Search Reinforcement Learning @esa (Ref

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:14.094327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:14.094327Z digest=sha256:318cbbe8bee73b07d3dce41178c858817a2c57c30e44a61992df3bfec4d7c8b2

Observation 7510aa87-e88a-49dd-b9bd-7891ea4830a5 · outbound

This paper cites an unresolved cited work.

SSRL: Self-Search Reinforcement Learning Unresolved cited work

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:14.138015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:14.138015Z digest=sha256:38af47bb708c4c6b5a11fc805336ad3349eb500b64c50110e801873834e6ddb5

Observation 69291813-5d61-4668-b016-ea021efba23c · outbound

This paper cites an unresolved cited work.

SSRL: Self-Search Reinforcement Learning Unresolved cited work

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:14.224235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:14.224235Z digest=sha256:22b287d927a86033f8cfa97d60da8a64a4d64589fa8d889da827061d62762bfd

Pith citing papers

Observation 1b4875cd-9bb5-40da-a5e3-985edbd95183 · inbound

Dynamic Collaboration of Multi-Language Models based on Minimal Complete Semantic Units cites this paper.

Dynamic Collaboration of Multi-Language Models based on Minimal Complete Semantic Units SSRL: Self-Search Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T16:18:44.717716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:18:44.717716Z digest=sha256:42fe0665b2eb98bd47bfa7e8cdb758110ec1e790824777610cfd4468e2a10173

Observation f591b89f-408e-4b86-9ffb-52c78f9a0b70 · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey SSRL: Self-Search Reinforcement Learning

Reference 292

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.095055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:499a18e3333f27809d5266df2e22e165d8a3fe2e399a85d5644b9509be56aff9

Observation 997be28b-ee60-4cac-a025-0ed4cfbcf684 · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models SSRL: Self-Search Reinforcement Learning

Reference 132

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:02:25.347982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:7a3e8e8c873e94dbff9e9f868e7652c40decd2f10075cb6a5985592bb825d580

Observation c7c41de9-4fe9-4c31-8288-f52bc1be63bb · inbound

Erase to Improve: Erasable Reinforcement Learning for Search-Augmented LLMs cites this paper.

Erase to Improve: Erasable Reinforcement Learning for Search-Augmented LLMs SSRL: Self-Search Reinforcement Learning

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:11:18.267234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T11:06:20.058342Z digest=sha256:3277fb638aad8a2401f597b3e2037c9e406c106eb97de30f42641307589ec2ec

Observation 9a8e4261-7149-48d1-8b52-f4ca13a5f9c2 · inbound

Qwen-AgentWorld: Language World Models for General Agents cites this paper.

Qwen-AgentWorld: Language World Models for General Agents SSRL: Self-Search Reinforcement Learning

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-04T17:09:59.247985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-25T23:52:31.403419Z digest=sha256:d4e78e24020b3f8801b9d77ef0ef6a5dc64ab4049fc86a1daf66e9523610236a