Pith. sign in

Paper Citation Record · LEDGER

SSRL: Self-Search Reinforcement Learning

As of 7 August 2026, this Paper Citation Record lists 72 of 72 outbound references and 5 inbound Pith citation observations for arXiv:2508.10874.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.10874 v1

Coverage vector

measured 72 of 72 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T20:17:14.224235Z

measured 77 of 77 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T16:18:44.717716Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T17:09:59.246273Z

Reference resolution

72 of 72 outbound references displayed

  • verified exact2
  • verified fuzzy6
  • unresolved64
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e660644d-2afa-4c06-bfb3-b2ace5a8519e · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

SSRL: Self-Search Reinforcement Learning Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:07.082566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:07.082566Z digest=sha256:d4d2739a50278e05ddd4b74be4bfff3196ca9d5e2a2b523d110506d9f36d0996

Observation 5174f725-50ae-492a-a825-3a6c242c5328 · outbound

This paper cites Language models are few-shot learners.

SSRL: Self-Search Reinforcement Learning Language models are few-shot learners

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:07.140938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:07.140938Z digest=sha256:dd5f26afb9b66dede24a33eee7282bf25e377cea2d44fe42d998cbecd7bad432

Observation 81b68b3e-8c5a-4e6e-9f3b-686b6d3fa244 · outbound

This paper cites ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning.

SSRL: Self-Search Reinforcement Learning ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:07.244058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:07.244058Z digest=sha256:a8c905f1970f2a31522e2ffa40d7908a0464ce4cc3d2549fc24af8d334ffb791

Observation 066bb555-cc25-4e86-8aa1-0bcaf8648914 · outbound

This paper cites Process Reinforcement through Implicit Rewards.

SSRL: Self-Search Reinforcement Learning Process Reinforcement through Implicit Rewards

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:07.379560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:07.379560Z digest=sha256:9282075bbc37d93ee4276a3279ff8ede037cd91fefac87cc40c90a3c38cbac3e

Observation 704fe00b-b2eb-4fd2-baee-79c97318e82e · outbound

This paper cites The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models.

SSRL: Self-Search Reinforcement Learning The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:07.516675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:07.516675Z digest=sha256:c2f7f3f93630e0d6d8cf0978ace84aa154a2551d49a723add4fc8b239620b203

Observation 82eed66f-a399-4593-9a7b-cbc81c49c252 · outbound

This paper cites A Survey of Sim-to-Real Methods in RL: Progress, Prospects and Challenges with Foundation Models.

SSRL: Self-Search Reinforcement Learning A Survey of Sim-to-Real Methods in RL: Progress, Prospects and Challenges with Foundation Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:07.610161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:07.610161Z digest=sha256:98a77f1276f183219038ecd6578c28bc0806c8f9a8c750a0b59c278412ab9ef1

Observation 25d1b355-7e9d-44b4-b808-93831202c15f · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

SSRL: Self-Search Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:07.780237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:07.780237Z digest=sha256:577d8cf2cfe613fa5bbc6e23c91294ddc442826acc1d99238008e25f59005b99

Observation 9fcc3ac4-0e5b-497c-a10c-6ff4f05bb5ef · outbound

This paper cites Citations and trust in llm generated responses.

SSRL: Self-Search Reinforcement Learning Citations and trust in llm generated responses

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:17:16.055399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-05T20:17:07.896751Z digest=sha256:2842e830c4cfb9925eb645aa08a7e89faaf6de770c46f175d7e7621c91926e1e

Observation 42c01a12-0b8f-491c-bbc5-b0c84ccbab26 · outbound

This paper cites Competitive Programming with Large Reasoning Models.

SSRL: Self-Search Reinforcement Learning Competitive Programming with Large Reasoning Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:08.011061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:08.011061Z digest=sha256:f4fef693d5c0652dad700a7eb29e73cc14127cfb1bcb98b5f632e5310d3e5fd3

Observation f631ddf0-5609-442e-b5b5-5a9908f3b02d · outbound

This paper cites ReTool: Reinforcement Learning for Strategic Tool Use in LLMs.

SSRL: Self-Search Reinforcement Learning ReTool: Reinforcement Learning for Strategic Tool Use in LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:08.153794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:08.153794Z digest=sha256:58248f110f94196068eb2daa3b38c4c2e467a9ec293b1a95ee87a8a2adc390d7

Observation 1635789a-eb37-4632-98e2-5d0c265d6d45 · outbound

This paper cites Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs.

SSRL: Self-Search Reinforcement Learning Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:08.299466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:08.299466Z digest=sha256:4e739a105b0572c0dbea563cacace312fab8e224bc9760481303deb6f02810af

Observation ac084a07-791c-4efc-9c54-bb54ff541e1b · outbound

This paper cites A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence.

SSRL: Self-Search Reinforcement Learning A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:08.416770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:08.416770Z digest=sha256:7565a82e0bc38d700772edbe8f0ea345771c56be8211a31b3a0e6a7dc7b7ab48

Observation e5bbfe28-e7c3-489a-86d2-017851b123fb · outbound

This paper cites Enabling Large Language Models to Generate Text with Citations.

SSRL: Self-Search Reinforcement Learning Enabling Large Language Models to Generate Text with Citations

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:08.576057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:08.576057Z digest=sha256:5ada23643e7ed1aaf9b1ed98df1a322a0505a052882e1606c93e39d05fc05b52

Observation d62b9b14-86a4-4930-9d3e-b97ab0a2f297 · outbound

This paper cites The Llama 3 Herd of Models.

SSRL: Self-Search Reinforcement Learning The Llama 3 Herd of Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:08.686291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:08.686291Z digest=sha256:bad17e559e7154cb2c005e25e553dec84c1ab9946822caad99d0ba7ee02d9e6e

Observation 3aaea7a4-83e8-4d98-933a-e9d41467abcc · outbound

This paper cites Is Your LLM Secretly a World Model of the Internet? Model-Based Planning for Web Agents.

SSRL: Self-Search Reinforcement Learning Is Your LLM Secretly a World Model of the Internet? Model-Based Planning for Web Agents

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:08.840484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:08.840484Z digest=sha256:796461ce1ca47a74d5e95b70b744cad38fcb07faa19d87fb139821829fb88557

Observation 7f829c9f-5fbd-4ad7-bc12-8da44b93d272 · outbound

This paper cites Reasoning with Language Model is Planning with World Model.

SSRL: Self-Search Reinforcement Learning Reasoning with Language Model is Planning with World Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:09.037240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:09.037240Z digest=sha256:64754637dc7a9c9c3793fdda913177ae711a14093609206ff4f649f8b37cb403

Observation e79197a4-5345-4fe4-8716-3f031433b673 · outbound

This paper cites Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps.

SSRL: Self-Search Reinforcement Learning Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:09.215321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:09.215321Z digest=sha256:fc587d8f8e462b035d341c0975b7691a93c33c31d2a1f04eadeb887ecad4ba80

Observation 669820f7-1974-43f8-85c3-eb94427b164f · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

SSRL: Self-Search Reinforcement Learning REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:09.323202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:09.323202Z digest=sha256:20150f61405306a7871b21dae468013fa3a4eee704a53f29cf30a182cf613e91

Observation 3e06a672-a89a-4390-a2df-320dc727fe33 · outbound

This paper cites An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents.

SSRL: Self-Search Reinforcement Learning An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:09.393341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:09.393341Z digest=sha256:f6bb55cc494c46abe8693d77f27127b06ba90ca0feb79c9e84aff30a21ad5058

Observation 1a52d4d0-b364-4586-b9d8-587a9ba961d1 · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

SSRL: Self-Search Reinforcement Learning Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:09.409539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:09.409539Z digest=sha256:18f5dffc04488fea7d308f0bf92323e833358706c87eb183ab1d6c79b0f10ea3

Observation 3cff19dd-0333-4be6-91eb-5e001564d9a5 · outbound

This paper cites TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension.

SSRL: Self-Search Reinforcement Learning TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:09.457321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:09.457321Z digest=sha256:d03dba1ad4cee4b496c00e6c0e1ad46ba7c5ff6d0142df59ebadd28f0c2ee6c4

Observation d1c5cf19-f396-4075-800d-3440538a65b6 · outbound

This paper cites Sim2real transfer for reinforcement learning without dynamics randomization.

SSRL: Self-Search Reinforcement Learning Sim2real transfer for reinforcement learning without dynamics randomization

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:17:15.944894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-05T20:17:09.570035Z digest=sha256:557268fbf466211829f29875d99b7b434d46039698da34d40cca49eb2ed1f772

Observation 64d4c93a-bbf1-4918-8e8a-76fe19db7a7a · outbound

This paper cites Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov.

SSRL: Self-Search Reinforcement Learning Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:09.639320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:09.639320Z digest=sha256:ada76608e42d48903a2931e363cf288b4dd3cd14410054cae81e878769208a59

Observation 71bf551d-680e-43a1-a83b-17835cde46b4 · outbound

This paper cites Scalable agent alignment via reward modeling: a research direction.

SSRL: Self-Search Reinforcement Learning Scalable agent alignment via reward modeling: a research direction

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:09.710424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:09.710424Z digest=sha256:a3c2614d4b6b1eb2b56f42ad94196a95977e88f3702968745724e39e23443e88

Observation cae576c8-daad-4d64-87ee-56724e58b22c · outbound

This paper cites S*: Test Time Scaling for Code Generation.

SSRL: Self-Search Reinforcement Learning S*: Test Time Scaling for Code Generation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:09.862906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:09.862906Z digest=sha256:4bf1899fa357bad610cc52de07086e68eba15fa9313a52499f3251e0f9193437

Observation c0c76f63-b0dc-4add-bdbb-cefe3781db74 · outbound

This paper cites Emergent world representations: Exploring a sequence model trained on a synthetic task.

SSRL: Self-Search Reinforcement Learning Emergent world representations: Exploring a sequence model trained on a synthetic task

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:17:15.829090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-05T20:17:09.970482Z digest=sha256:c80ae88d62065624d595fb94823928098cc3cba6b197b2d01f3666fe0fdfd57a

Observation 34037597-400a-4e64-99f6-751bc4f6b000 · outbound

This paper cites CorpusLM: Towards a Unified Language Model on Corpus for Knowledge-Intensive Tasks.

SSRL: Self-Search Reinforcement Learning CorpusLM: Towards a Unified Language Model on Corpus for Knowledge-Intensive Tasks

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-05T20:17:15.004781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-05T20:17:10.055596Z digest=sha256:ade410da3dff3094b8158fbde23403ff1aa85bc4de6145c8204bd88e51e6f330

Observation 25160b53-2978-4454-8697-0e4ca2084107 · outbound

This paper cites Search-o1: Agentic Search-Enhanced Large Reasoning Models.

SSRL: Self-Search Reinforcement Learning Search-o1: Agentic Search-Enhanced Large Reasoning Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:10.120457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:10.120457Z digest=sha256:707cc225f4ec7fa8c984cf6b8e1d986be1c8eb5fd25ea7142270b6beadc3695b

Observation fe433099-c955-48a3-be81-58857651866a · outbound

This paper cites From matching to generation: A survey on generative information retrieval.

SSRL: Self-Search Reinforcement Learning From matching to generation: A survey on generative information retrieval

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:17:15.701341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-05T20:17:10.230492Z digest=sha256:f06e73ceedae35853ecfffe32f4829f2e22d84900acc6a67c4d4caa35f5a998c

Observation b65de2a5-3dfe-497a-805a-14bbe9d2336c · outbound

This paper cites A Survey of Generative Search and Recommendation in the Era of Large Language Models.

SSRL: Self-Search Reinforcement Learning A Survey of Generative Search and Recommendation in the Era of Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:10.345154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:10.345154Z digest=sha256:33fdc509c860c7ff579fc2ee49f1231b9feebe22b56c2727b88b1cf0ba0a69cf

Observation 9cbfc218-ed6b-42a1-b648-4e03b3fb8028 · outbound

This paper cites Learning to rank in generative retrieval.

SSRL: Self-Search Reinforcement Learning Learning to rank in generative retrieval

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:17:15.572763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-05T20:17:10.430071Z digest=sha256:7c11ca3d6fd94995f23e1a8a59cf0dc38b0b5f0f0c969abd5f8d07e48eabf386

Observation 2abfb961-6d12-41ba-b824-58ff7ed19bc7 · outbound

This paper cites Multi-Agent Verification: Scaling Test-Time Compute with Multiple Verifiers.

SSRL: Self-Search Reinforcement Learning Multi-Agent Verification: Scaling Test-Time Compute with Multiple Verifiers

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:10.502885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:10.502885Z digest=sha256:44190e60db27cd9f21ed3931631d677c06f4f9c3638a528e78724aa046e6b0e5

Observation 981d5a87-31a4-489d-b30b-7edf2723ae84 · outbound

This paper cites DeepSeek-V3 Technical Report.

SSRL: Self-Search Reinforcement Learning DeepSeek-V3 Technical Report

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:10.603582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:10.603582Z digest=sha256:2f2a0a20c8d3e1d6abdf50268450673fa8d62d4ec8732d559d420ded93d015d0

Observation 0972242c-07d0-4677-81ae-206435b91c3f · outbound

This paper cites Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling.

SSRL: Self-Search Reinforcement Learning Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:10.690804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:10.690804Z digest=sha256:43e2e35296f2013f13ed549cdf8698e1a3e39dae1e04419cecf28ff640a52da3

Observation 34cd607f-393e-46b8-b686-fedebd29a29c · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

SSRL: Self-Search Reinforcement Learning Understanding R1-Zero-Like Training: A Critical Perspective

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:10.754730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:10.754730Z digest=sha256:a938da176a570e4f1f18843b1231faf4ca80484becd479d08955cedb6fd454a2

Observation 2cc82a03-e7dc-4e1f-9ff4-38693d2f6e13 · outbound

This paper cites Generative multi-modal knowledge retrieval with large language models.

SSRL: Self-Search Reinforcement Learning Generative multi-modal knowledge retrieval with large language models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:10.833429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:10.833429Z digest=sha256:04478fb592917f1e7f4c3eee6cd698a472ac1ae8b4a39500b2a4ca612cf7c973

Observation a09b9e17-e2c0-4bc8-926a-6aa6d86d27ef · outbound

This paper cites OpenAI o1 System Card.

SSRL: Self-Search Reinforcement Learning OpenAI o1 System Card

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:10.945859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:10.945859Z digest=sha256:e74968e0d1e155b583aaed9a10ce64c1f2fa2b87b9c7c0b3d7e2004309d92979

Observation bbf8efaf-4312-4d5e-984a-68622d4d12cc · outbound

This paper cites Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad.

SSRL: Self-Search Reinforcement Learning Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:11.057856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:11.057856Z digest=sha256:4acbc6a636be36925cf7fe3f385d4e6c77ded91292a744e62c0440ddce4139d2

Observation 42ffeb56-5f9b-421a-b8a4-120aebb67274 · outbound

This paper cites Measuring and Narrowing the Compositionality Gap in Language Models.

SSRL: Self-Search Reinforcement Learning Measuring and Narrowing the Compositionality Gap in Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:11.138260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:11.138260Z digest=sha256:6e26d92fffb0d87c4368d104693a49eeb4656d040667559bef900601005b2e1f

Observation e82d6f1e-83ba-4f7a-9dd9-24b42ab33272 · outbound

This paper cites ToolRL: Reward is All Tool Learning Needs.

SSRL: Self-Search Reinforcement Learning ToolRL: Reward is All Tool Learning Needs

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:11.220482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:11.220482Z digest=sha256:367ae64d6c26f96357ffdde987ad89fbd891df9a7d09b207ae4aca9054ca2c2d

Observation 4275843b-7c40-4f5d-91ff-4e5dfaa92949 · outbound

This paper cites WebCPM: Interactive Web Search for Chinese Long-form Question Answering.

SSRL: Self-Search Reinforcement Learning WebCPM: Interactive Web Search for Chinese Long-form Question Answering

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-08-05T20:17:14.717417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-05T20:17:11.334274Z digest=sha256:a1c7ae41c17b22a4722f3f750956cd119f61fdabc4a1d85baf93ca12b8e4401d

Observation fb3313db-e19b-4ff9-b1ad-dbfe0c1f15ce · outbound

This paper cites TreeBoN: Enhancing Inference-Time Alignment with Speculative Tree-Search and Best-of-N Sampling.

SSRL: Self-Search Reinforcement Learning TreeBoN: Enhancing Inference-Time Alignment with Speculative Tree-Search and Best-of-N Sampling

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:11.414943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:11.414943Z digest=sha256:3c507a9cd3b7a0fc14b2130c2178a77af9bf8155f6ee74ddb84695d88fcaf77e

Observation 166ef671-64e5-4e50-9dab-f72d13048071 · outbound

This paper cites Qwen2.5 Technical Report.

SSRL: Self-Search Reinforcement Learning Qwen2.5 Technical Report

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:11.513187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:11.513187Z digest=sha256:3d8e8959136c2980e9dc312540766eeb35cd685933fbf23189647d7cd1271988

Observation e27dd7cf-c73d-43e3-9281-374eed56989c · outbound

This paper cites Proximal Policy Optimization Algorithms.

SSRL: Self-Search Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:11.582880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:11.582880Z digest=sha256:1cd2c51f126ad3d6c4c1c84d6a853041a10846c790e38f60255f9d99847f4e71

Observation 4a7746c4-18cf-4309-beb6-d3d9beed30a0 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

SSRL: Self-Search Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:11.674940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:11.674940Z digest=sha256:879eb2e26d900c6d72af5ca6fdd01dfb1aa06887e0c15b670cf25cefb0ae0617

Observation 78856db5-fe51-411c-96bf-e5a93e222d04 · outbound

This paper cites Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction.

SSRL: Self-Search Reinforcement Learning Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:11.770737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:11.770737Z digest=sha256:bd6819741995e4ad9349be46269c48da159cee7a2bdbe26d053b19b86ba40f11

Observation e9038f3a-fd3a-454f-9deb-93ccf45879b9 · outbound

This paper cites HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face.

SSRL: Self-Search Reinforcement Learning HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:11.877617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:11.877617Z digest=sha256:99e66e71cffec739835361fbdb552720b681acf444597e4f2938b9caaf51ada8

Observation 39bd7740-f962-4252-8546-8abd5810261a · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

SSRL: Self-Search Reinforcement Learning Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:11.983940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:11.983940Z digest=sha256:169f10cdadf5e8dfc16d02d184e03dde851056fc8601ac991c249ea5a90267f0

Observation b8f7c3d2-1a78-4583-892c-4749ddd15efa · outbound

This paper cites ZeroSearch: Incentivize the Search Capability of LLMs without Searching.

SSRL: Self-Search Reinforcement Learning ZeroSearch: Incentivize the Search Capability of LLMs without Searching

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:12.061269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:12.061269Z digest=sha256:70887e1910aa3f9026294faed5135920cf20ac3c9b22e6f62738a0d545a957e7

Observation 97d453e7-df30-4c77-9a58-058311b88095 · outbound

This paper cites Worldcoder, a model-based llm agent: Building world models by writing code and interacting with the environment.

SSRL: Self-Search Reinforcement Learning Worldcoder, a model-based llm agent: Building world models by writing code and interacting with the environment

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:17:15.423680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-05T20:17:12.144246Z digest=sha256:401f764461ec50bb9d28cfb8bd3c11d4614d7dd017d6b095802021e6ba83512b

Observation 2221edb6-7c2d-4f56-9f77-b4833d2fdb2b · outbound

This paper cites Transformer memory as a differentiable search index.

SSRL: Self-Search Reinforcement Learning Transformer memory as a differentiable search index

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:12.247238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:12.247238Z digest=sha256:e9d8e39675292b907dad0e37ee2bb223da9eef2ba45012388f4bacdb73fa4baf

Observation 19cb28b6-ecbe-45e0-9a22-0be3e0b0946d · outbound

This paper cites Kimi K2: Open Agentic Intelligence.

SSRL: Self-Search Reinforcement Learning Kimi K2: Open Agentic Intelligence

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:12.345692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:12.345692Z digest=sha256:cbf237ebe913420af656f7735d8b27673e573c0385042a1e4b4e9954dd9cdcd3

Observation e34e3623-2212-45eb-996d-5d4a39a593ff · outbound

This paper cites MuSiQue: Multihop Questions via Single-hop Question Composition.

SSRL: Self-Search Reinforcement Learning MuSiQue: Multihop Questions via Single-hop Question Composition

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:12.419079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:12.419079Z digest=sha256:405d28faeec71e4814bddff12215ae013db4f97f57e1f60bc32c6085cf8d036c

Observation 84542cda-df2e-4f4b-ad73-aa6fa116ac58 · outbound

This paper cites A neural corpus indexer for document retrieval.

SSRL: Self-Search Reinforcement Learning A neural corpus indexer for document retrieval

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:12.524672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:12.524672Z digest=sha256:d05f080c26ec6e265ad56787956523145dd29c9097350a217e2325b92909f4dd

Observation ef6e6e4c-790c-4b24-8c58-4716de72ec9d · outbound

This paper cites OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling.

SSRL: Self-Search Reinforcement Learning OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:12.589663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:12.589663Z digest=sha256:f4720aa90bb7397299579945256d9a76976dfc4db467472402b0a1c0000d9300

Observation ac04a7b9-7d04-41c4-be28-29de60856bbc · outbound

This paper cites RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning.

SSRL: Self-Search Reinforcement Learning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:12.728962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:12.728962Z digest=sha256:e6341a4eef774670dd85113c194e0b2228ebe8e5bd9e68950e3f7d77de726c73

Observation 7a8f45fd-17ae-43a2-88d7-41a99c1bec28 · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

SSRL: Self-Search Reinforcement Learning Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:12.825794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:12.825794Z digest=sha256:697504342a58c0eda9e9498e6833c48691a16b574ed0b586c995fe56ccf28f6e

Observation 914e4dc4-33d4-4f3d-9017-87ea06285be2 · outbound

This paper cites Measuring short-form factuality in large language models.

SSRL: Self-Search Reinforcement Learning Measuring short-form factuality in large language models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:12.901494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:12.901494Z digest=sha256:62f87b81bb6a9781cddc39b08a21d05a630465028ba7e0e26390c92b4ea9be23

Observation 65abb7e1-0218-4aa4-80d4-013c5f0659fc · outbound

This paper cites BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents.

SSRL: Self-Search Reinforcement Learning BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:13.008368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:13.008368Z digest=sha256:39bf368d5beb8de678592b4ddcfcd7b605efc8ffc86d7043a3c8501e1df8862e

Observation 27d9c8ff-0767-4307-923d-b64106e9eaf9 · outbound

This paper cites Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models.

SSRL: Self-Search Reinforcement Learning Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:13.081546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:13.081546Z digest=sha256:ecdeee8be5acd072d035f07d815ce7d2b72923f29de6b32d70c1a508b9ae645e

Observation 46d3f3ca-4670-4aad-9895-4aff26ec06c8 · outbound

This paper cites Qwen3 Technical Report.

SSRL: Self-Search Reinforcement Learning Qwen3 Technical Report

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:13.289705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:13.289705Z digest=sha256:05174349e3be7853b8a67cabfc5e5a9744a8966aac891945e08b278faffb8e20

Observation b6745111-7cb8-4e4f-9809-50c5f97366c1 · outbound

This paper cites GTA1: GUI Test-time Scaling Agent.

SSRL: Self-Search Reinforcement Learning GTA1: GUI Test-time Scaling Agent

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:13.398421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:13.398421Z digest=sha256:5170392eeede94f65634dc1e2c8cb05097c35fc410886fb011952a8514794cc7

Observation d32414e7-31d1-4337-9f87-0812da702845 · outbound

This paper cites HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering.

SSRL: Self-Search Reinforcement Learning HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:13.464960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:13.464960Z digest=sha256:2c4906860dd87b8975f51447aff684831367c57c3ce6d36edbbdf7eb7acefffd

Observation 86181fa9-55ff-4a86-b63b-b2e198330cfd · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

SSRL: Self-Search Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:13.572346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:13.572346Z digest=sha256:3bfc776e2bf2524dfab706951351b30ca890d0faade462a851b9626fb10f477a

Observation 1541f30a-c8b1-4377-9cba-48384864dc1f · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

SSRL: Self-Search Reinforcement Learning Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:13.673603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:13.673603Z digest=sha256:dcdcd75df06b614e70377823e7a98210e85002efe94c724329f04f216d2a21b6

Observation eaee824b-2e7b-430b-a2b6-27158195c2c0 · outbound

This paper cites DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments.

SSRL: Self-Search Reinforcement Learning DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:13.787573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:13.787573Z digest=sha256:1ddb07180ce4d550c3d4d2f0b674a00bd6fa20541cc0dbee78d77761a1dad23c

Observation c74b1a72-ace6-4ffb-8732-1be03af50989 · outbound

This paper cites Scaling Test-time Compute for LLM Agents.

SSRL: Self-Search Reinforcement Learning Scaling Test-time Compute for LLM Agents

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:13.855017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:13.855017Z digest=sha256:314fafe0bf29523bb29dc4cfe1c79bc8faa1344f922777ea0cb0957d732ee72e

Observation e5c57250-75b1-47c9-a70b-708ea24a4f42 · outbound

This paper cites TTRL: Test-Time Reinforcement Learning.

SSRL: Self-Search Reinforcement Learning TTRL: Test-Time Reinforcement Learning

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:13.935000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:13.935000Z digest=sha256:490e3acd47549ec696be882940adeb77b8f3f8d4798e63afe126848cb9a5f7e7

Observation c01981d2-1ba7-461f-8d22-5c30ae39d937 · outbound

This paper cites write newline.

SSRL: Self-Search Reinforcement Learning write newline

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:14.035628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:14.035628Z digest=sha256:24538ef23232525ae042fc08a399735ee23a25aaf1bc1108ceb8b74daf919149

Observation 8133ea35-7fd3-4484-bca2-6b690abcd79c · outbound

This paper cites @esa (Ref.

SSRL: Self-Search Reinforcement Learning @esa (Ref

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:14.094327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:14.094327Z digest=sha256:318cbbe8bee73b07d3dce41178c858817a2c57c30e44a61992df3bfec4d7c8b2

Observation 7510aa87-e88a-49dd-b9bd-7891ea4830a5 · outbound

This paper cites an unresolved cited work.

SSRL: Self-Search Reinforcement Learning Unresolved cited work

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:14.138015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:14.138015Z digest=sha256:38af47bb708c4c6b5a11fc805336ad3349eb500b64c50110e801873834e6ddb5

Observation 69291813-5d61-4668-b016-ea021efba23c · outbound

This paper cites an unresolved cited work.

SSRL: Self-Search Reinforcement Learning Unresolved cited work

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:14.224235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:14.224235Z digest=sha256:22b287d927a86033f8cfa97d60da8a64a4d64589fa8d889da827061d62762bfd

Pith citing papers

Observation 1b4875cd-9bb5-40da-a5e3-985edbd95183 · inbound

Dynamic Collaboration of Multi-Language Models based on Minimal Complete Semantic Units cites this paper.

Dynamic Collaboration of Multi-Language Models based on Minimal Complete Semantic Units SSRL: Self-Search Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T16:18:44.717716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:18:44.717716Z digest=sha256:3263d70eeda17be6c097ea6a2716c9db5c76b1ca3477bae147c59fbb9f7d5ecc

Observation f591b89f-408e-4b86-9ffb-52c78f9a0b70 · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey SSRL: Self-Search Reinforcement Learning

Reference 292

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.095055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:edbcd4b40ba25151a03e20e057df28f7e14bbe885e0a735af4f348cdc8d24f45

Observation 997be28b-ee60-4cac-a025-0ed4cfbcf684 · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models SSRL: Self-Search Reinforcement Learning

Reference 132

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:02:25.347982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:7724d281c1fcbfb34be5c1360345ffe1b6a5900cca037bd5f9b845612c1fe733

Observation c7c41de9-4fe9-4c31-8288-f52bc1be63bb · inbound

Erase to Improve: Erasable Reinforcement Learning for Search-Augmented LLMs cites this paper.

Erase to Improve: Erasable Reinforcement Learning for Search-Augmented LLMs SSRL: Self-Search Reinforcement Learning

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:11:18.267234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T11:06:20.058342Z digest=sha256:1e8201ad83d83be484a637c3c599252d0d1d8694856f6ef7f32a3712c9ff2cf8

Observation 9a8e4261-7149-48d1-8b52-f4ca13a5f9c2 · inbound

Qwen-AgentWorld: Language World Models for General Agents cites this paper.

Qwen-AgentWorld: Language World Models for General Agents SSRL: Self-Search Reinforcement Learning

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-04T17:09:59.247985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-25T23:52:31.403419Z digest=sha256:b2074b7822ae912bfc31e06915e7fe5fae41fc8649818b86c6b1f807163e97bf