Pith. sign in

Paper Citation Record · LEDGER

L0: Reinforcement Learning to Become General Agents

As of 10 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 2 inbound Pith citation observations for arXiv:2506.23667.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23667 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:45:00.896630Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-15T21:51:10.972744Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T21:51:40.816217Z

Reference resolution

18 of 18 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bef74cd5-63f2-434e-a12e-c43090396fe9 · outbound

This paper cites Group-in-Group Policy Optimization for LLM Agent Training.

L0: Reinforcement Learning to Become General Agents Group-in-Group Policy Optimization for LLM Agent Training

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:59.706854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:44:59.706854Z digest=sha256:a31fdf407936071d6094cee12573d8191ce950bc951164772e44014ecd880841

Observation 8608f17b-f4a3-47d9-a94b-ba8cc8b87810 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

L0: Reinforcement Learning to Become General Agents DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:59.778161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:44:59.778161Z digest=sha256:263a7fe7e2a86d326e666d55363172baba50fdcd3e610dfd6db72d51066c212b

Observation 1ff586b0-bb82-4514-b5a2-ad612c94b987 · outbound

This paper cites an unresolved cited work.

L0: Reinforcement Learning to Become General Agents Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:45:02.773356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T21:44:59.884891Z digest=sha256:e457c2cc6982dd54a00fa56cbe51944d4638f2e5ba879f5a31944981eaa1300e

Observation 899f4bf4-0f05-44c5-bd4e-30cd78ef346f · outbound

This paper cites an unresolved cited work.

L0: Reinforcement Learning to Become General Agents Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:45:02.597391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T21:44:59.933455Z digest=sha256:e5f0b4325fad3c9a08149b96efcd99d6d2ef2b9c502bd12f0d481ebe11359325

Observation 72fa6a89-92e5-439a-b933-439aa24f0a6a · outbound

This paper cites RAGent: Retrieval-based Access Control Policy Generation.

L0: Reinforcement Learning to Become General Agents RAGent: Retrieval-based Access Control Policy Generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:00.007761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:00.007761Z digest=sha256:f136cbc5d6ba30a7ad80745decbb2d4de8698b1c1e6005aa7e4a90a85e8432b8

Observation 7acf283a-68bb-4849-8f3f-ee404e39634e · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

L0: Reinforcement Learning to Become General Agents Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:00.084961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:00.084961Z digest=sha256:b852d3cd9dd6ec311428217f63032a129473fa56f09d30ed0f15200ca016bc29

Observation fa7a83f0-e58a-4379-86a5-f54943bd6301 · outbound

This paper cites an unresolved cited work.

L0: Reinforcement Learning to Become General Agents Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:45:02.398920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T21:45:00.190054Z digest=sha256:2581a689c8a6e076ce39251a652ff5403b6aae29f40d82b73511c85afde9e359

Observation 0371d9fe-77f0-4a83-8b55-39be7cf64e4a · outbound

This paper cites an unresolved cited work.

L0: Reinforcement Learning to Become General Agents Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:45:02.224600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T21:45:00.264906Z digest=sha256:32564ccfd19ba2ae0f7218b6434503a28e5e7475d12412fd685558340b0854bc

Observation 9f9a0550-db7d-461b-b08a-5dabcfa900b1 · outbound

This paper cites Search-o1: Agentic Search-Enhanced Large Reasoning Models.

L0: Reinforcement Learning to Become General Agents Search-o1: Agentic Search-Enhanced Large Reasoning Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:00.349407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:00.349407Z digest=sha256:ba37a56bf74d10429f7c2b4ac9764e1edd59b3743c1f7aae8523e99656a49336

Observation 8063a002-7b35-4e2b-bd16-40aca3581f7f · outbound

This paper cites an unresolved cited work.

L0: Reinforcement Learning to Become General Agents Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:45:02.072346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T21:45:00.408556Z digest=sha256:9530369f7f47201698ad3099444a34d6fdd33a1c6ddb9f2b3cc206100a183d60

Observation 10349242-37f1-4f2e-b007-f0b20d7ab42a · outbound

This paper cites ZeroSearch: Incentivize the Search Capability of LLMs without Searching.

L0: Reinforcement Learning to Become General Agents ZeroSearch: Incentivize the Search Capability of LLMs without Searching

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:00.471283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:00.471283Z digest=sha256:025692f344ef97e70c40aba8ed672d4d3057deaeb6b39a088347f3bfa1f18cf0

Observation 6394dd20-57cb-42e8-b990-3c0703ba3894 · outbound

This paper cites an unresolved cited work.

L0: Reinforcement Learning to Become General Agents Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:45:01.893036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T21:45:00.537973Z digest=sha256:1f1c5e311787895a6a8ec9ea085a49e6b89f3ee91b07d0c67a1b959b31b7cab1

Observation b553d063-6f8b-41cf-a758-fb45b1b0bfd6 · outbound

This paper cites an unresolved cited work.

L0: Reinforcement Learning to Become General Agents Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:45:01.720030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T21:45:00.588082Z digest=sha256:3fc73f2de62bcdaec1e4e88df43ebb028ac23bceadf30cb524fccf16dc058212

Observation f7fbddeb-c4cf-4e61-afac-b28e34b7a3a4 · outbound

This paper cites an unresolved cited work.

L0: Reinforcement Learning to Become General Agents Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:45:01.551841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T21:45:00.635927Z digest=sha256:bd4bf5fd9ae84559ea82febee344f7f82a00cdd66576abadcc48feadc4283dbc

Observation e091268f-41eb-4008-bc86-6a39117b640e · outbound

This paper cites Measuring short-form factuality in large language models.

L0: Reinforcement Learning to Become General Agents Measuring short-form factuality in large language models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:00.718465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:00.718465Z digest=sha256:165283309c972ba8dfb3df8ee2ba7bd96fe67d3fb1d951a760a413cdc94570e7

Observation 8a587171-c7c8-4a4d-b714-114a72425a2e · outbound

This paper cites Qwen3 Technical Report.

L0: Reinforcement Learning to Become General Agents Qwen3 Technical Report

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:00.782069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:00.782069Z digest=sha256:7e1218b4e119d0f354e8815b93b166eb888878dcc5f0f85d874f01376221c4b6

Observation 51f948a6-9df1-4632-a191-5819d4b31832 · outbound

This paper cites an unresolved cited work.

L0: Reinforcement Learning to Become General Agents Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:45:01.382091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T21:45:00.834498Z digest=sha256:e98378262af864c7a8caa19de7d8cd351595626f2936c785238bfea0bef6cb0a

Observation 9caf3835-4d4d-403e-ba6c-4f0e747b91a4 · outbound

This paper cites an unresolved cited work.

L0: Reinforcement Learning to Become General Agents Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:45:01.188222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T21:45:00.896630Z digest=sha256:d89801104e44fd0b01eb35257c7e3a1e97393be2a4bf8101596d0e144d3c4555

Pith citing papers

Observation 7f6b2a6f-fb70-41df-906c-053903ef2397 · inbound

Multi-Modal Multi-Agent Reinforcement Learning for Radiology Report Generation cites this paper.

Multi-Modal Multi-Agent Reinforcement Learning for Radiology Report Generation L0: Reinforcement Learning to Become General Agents

Reference 140

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:51:40.820434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T21:51:10.972744Z digest=sha256:8c840e67fcdc604c6ee4132c36d6ceb600fb0540e0db301b526dacd6c6f95a7b

Observation 1309c335-5c4e-4801-9ba6-78bcd1cdb231 · inbound

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning cites this paper.

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning L0: Reinforcement Learning to Become General Agents

Reference 173

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:15:49.725122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T19:15:27.406778Z digest=sha256:15bb7ef8a7c4d31a5ab10818c049b42ec616c50e25626826e4ce4ab7e355cbb7