Pith. sign in

Paper Citation Record · LEDGER

L0: Reinforcement Learning to Become General Agents

As of 18 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 2 inbound Pith citation observations for arXiv:2506.23667.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23667 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:45:00.896630Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-15T21:51:10.972744Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T21:51:40.816217Z

Reference resolution

18 of 18 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bef74cd5-63f2-434e-a12e-c43090396fe9 · outbound

This paper cites Group-in-Group Policy Optimization for LLM Agent Training.

L0: Reinforcement Learning to Become General Agents Group-in-Group Policy Optimization for LLM Agent Training

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:59.706854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:44:59.706854Z digest=sha256:3ff8c4ef0b53516137368dc3cbb1f14e52d3255ad3908432f5b1fafed23bd131

Observation 8608f17b-f4a3-47d9-a94b-ba8cc8b87810 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

L0: Reinforcement Learning to Become General Agents DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:59.778161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:44:59.778161Z digest=sha256:68573655697d8082d909245a6b1c4bd3098aac7f027a57e3bad428eaebea099b

Observation 1ff586b0-bb82-4514-b5a2-ad612c94b987 · outbound

This paper cites an unresolved cited work.

L0: Reinforcement Learning to Become General Agents Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:45:02.773356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T21:44:59.884891Z digest=sha256:4c2c9ccdb0c1b32707fd508cce11319f7d42205cce373b1cb97449a44f6f33c1

Observation 899f4bf4-0f05-44c5-bd4e-30cd78ef346f · outbound

This paper cites an unresolved cited work.

L0: Reinforcement Learning to Become General Agents Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:45:02.597391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T21:44:59.933455Z digest=sha256:de2d0b00a9e7c1b80cd37515836623f7f0fbf433b46fa31eafcb83010bd87b5e

Observation 72fa6a89-92e5-439a-b933-439aa24f0a6a · outbound

This paper cites RAGent: Retrieval-based Access Control Policy Generation.

L0: Reinforcement Learning to Become General Agents RAGent: Retrieval-based Access Control Policy Generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:00.007761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:00.007761Z digest=sha256:b1532ac7285ae7b85daf81c9bf50f13d8e4d8904e8348c1254c3e5adadef4054

Observation 7acf283a-68bb-4849-8f3f-ee404e39634e · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

L0: Reinforcement Learning to Become General Agents Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:00.084961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:00.084961Z digest=sha256:0baa85abf9e8cbb9767218b4d087546f1bdce7b11011d0f6969ec431040aa3d5

Observation fa7a83f0-e58a-4379-86a5-f54943bd6301 · outbound

This paper cites an unresolved cited work.

L0: Reinforcement Learning to Become General Agents Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:45:02.398920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T21:45:00.190054Z digest=sha256:03a16a9a32749e3963b7299ce40fa99d213eb6113dec1fcc1a9366e990af7757

Observation 0371d9fe-77f0-4a83-8b55-39be7cf64e4a · outbound

This paper cites an unresolved cited work.

L0: Reinforcement Learning to Become General Agents Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:45:02.224600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T21:45:00.264906Z digest=sha256:f934e5686babe40094952aaef4fa4740901db63d192b4becfdd807cb0a2434ae

Observation 9f9a0550-db7d-461b-b08a-5dabcfa900b1 · outbound

This paper cites Search-o1: Agentic Search-Enhanced Large Reasoning Models.

L0: Reinforcement Learning to Become General Agents Search-o1: Agentic Search-Enhanced Large Reasoning Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:00.349407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:00.349407Z digest=sha256:12ecb7e2f2a6881df59b65f7e5f158de514f2eecc890095fd651c49274b896a8

Observation 8063a002-7b35-4e2b-bd16-40aca3581f7f · outbound

This paper cites an unresolved cited work.

L0: Reinforcement Learning to Become General Agents Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:45:02.072346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T21:45:00.408556Z digest=sha256:52b244a27c2418fec9a7c4f7a378119c960f1401a6fefb76cca339946ddf5f29

Observation 10349242-37f1-4f2e-b007-f0b20d7ab42a · outbound

This paper cites ZeroSearch: Incentivize the Search Capability of LLMs without Searching.

L0: Reinforcement Learning to Become General Agents ZeroSearch: Incentivize the Search Capability of LLMs without Searching

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:00.471283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:00.471283Z digest=sha256:4a535184af6fa0673d5ad4edbfefb94ea49d9c5d2c61101733dea65f75216f48

Observation 6394dd20-57cb-42e8-b990-3c0703ba3894 · outbound

This paper cites an unresolved cited work.

L0: Reinforcement Learning to Become General Agents Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:45:01.893036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T21:45:00.537973Z digest=sha256:4ee5ecf0274c674678714b454b5d3ad59feacf2fcb15a543990751435098db3a

Observation b553d063-6f8b-41cf-a758-fb45b1b0bfd6 · outbound

This paper cites an unresolved cited work.

L0: Reinforcement Learning to Become General Agents Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:45:01.720030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T21:45:00.588082Z digest=sha256:134361bdb3de559e41665b80a3d5c21d708b0d8445962c7e71bd015be1bb1710

Observation f7fbddeb-c4cf-4e61-afac-b28e34b7a3a4 · outbound

This paper cites an unresolved cited work.

L0: Reinforcement Learning to Become General Agents Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:45:01.551841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T21:45:00.635927Z digest=sha256:5617c548a291809f7316f941893e1c65ea0b9202ccf650d9855ccf519f9b2c03

Observation e091268f-41eb-4008-bc86-6a39117b640e · outbound

This paper cites Measuring short-form factuality in large language models.

L0: Reinforcement Learning to Become General Agents Measuring short-form factuality in large language models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:00.718465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:00.718465Z digest=sha256:96abdf00658f39cb31313b1832c972913a215c97dc33744d1b52f0ecf36fc711

Observation 8a587171-c7c8-4a4d-b714-114a72425a2e · outbound

This paper cites Qwen3 Technical Report.

L0: Reinforcement Learning to Become General Agents Qwen3 Technical Report

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:00.782069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:00.782069Z digest=sha256:cda114b07a870126ce0bcbd1af7c7d27dab86150e241b1725d215759059e0844

Observation 51f948a6-9df1-4632-a191-5819d4b31832 · outbound

This paper cites an unresolved cited work.

L0: Reinforcement Learning to Become General Agents Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:45:01.382091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T21:45:00.834498Z digest=sha256:00a711a06900a4fab0dda3ed9eea9c07510bc2baafdad6912aef5337cb8eda8b

Observation 9caf3835-4d4d-403e-ba6c-4f0e747b91a4 · outbound

This paper cites an unresolved cited work.

L0: Reinforcement Learning to Become General Agents Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:45:01.188222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T21:45:00.896630Z digest=sha256:d3116143bf2c6a4df94fb065cc71c3d71cbbd7186a6fa75ca7d87b5a7f8916f7

Pith citing papers

Observation 7f6b2a6f-fb70-41df-906c-053903ef2397 · inbound

Multi-Modal Multi-Agent Reinforcement Learning for Radiology Report Generation cites this paper.

Multi-Modal Multi-Agent Reinforcement Learning for Radiology Report Generation L0: Reinforcement Learning to Become General Agents

Reference 140

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:51:40.820434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-15T21:51:10.972744Z digest=sha256:551c0714be09946e27891286d39aa798ffd2d19c4e56c02f3a37fe8a1ba13dba

Observation 1309c335-5c4e-4801-9ba6-78bcd1cdb231 · inbound

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning cites this paper.

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning L0: Reinforcement Learning to Become General Agents

Reference 173

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:15:49.725122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-10T19:15:27.406778Z digest=sha256:067e5decca704bf16e28d60dea8b98bbd934a1fd6313137ceb8f9050b84f1804