Pith. sign in

Paper Citation Record · LEDGER

Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry

As of 7 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 0 inbound Pith citation observations for arXiv:2607.03702.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.03702 v1

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-12T00:33:06.488657Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

19 of 19 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 959d28d3-e55a-4719-a13c-ac9a3d57164b · outbound

This paper cites Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning.

Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-12T00:33:06.488657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:33:06.488657Z digest=sha256:53a0988c5c33796d190abf54aada12e1d0d8a5a40c868e493ec39f2f83202299

Observation 3262efab-1047-4a2e-b3bc-7cf45fe30e57 · outbound

This paper cites Soft Adaptive Policy Optimization.

Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry Soft Adaptive Policy Optimization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-12T00:33:06.488657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:33:06.488657Z digest=sha256:12e32bb421a026823e91015e6e85db3951706b566642344c21d5ffb1730c8268

Observation 70a98fff-27e1-4117-97ec-a411d8f9b5ec · outbound

This paper cites Reinforcement Learning via Self-Distillation.

Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry Reinforcement Learning via Self-Distillation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-12T00:33:06.488657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:33:06.488657Z digest=sha256:a3151205b16a6deb7adb08c57eaf368ac183f8e5e37809d615d7f23fc8256a6b

Observation 7cda9da9-3e62-49fd-baa4-61f36708cd5b · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-12T00:33:06.488657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:33:06.488657Z digest=sha256:2ae6f70560935c9b7b225cbeff2c183ad6150a1e3d5c998062a290cc6c3ec186

Observation f566925c-a7b8-44e2-b50f-3deb4b74c3eb · outbound

This paper cites Yinghao Li, Haorui Wang, and Chao Zhang.

Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry Yinghao Li, Haorui Wang, and Chao Zhang

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-12T00:33:06.488657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:33:06.488657Z digest=sha256:1295c716c27fe32f9dc71cda7a87f99d14732100eb05ac3d6807393f4eb0754b

Observation ac97ecc7-39de-4fcb-8b8f-afd960b0042f · outbound

This paper cites SimpleMem: Efficient Lifelong Memory for LLM Agents.

Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry SimpleMem: Efficient Lifelong Memory for LLM Agents

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-12T00:33:06.488657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:33:06.488657Z digest=sha256:1eb502400bdf56e527c3cf65a71a3899282279469eebc1bbb39658ed03890bf8

Observation eceb5f1b-bf95-4805-91b6-f7f185cea1a5 · outbound

This paper cites Self-Distilled Agentic Reinforcement Learning.

Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry Self-Distilled Agentic Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-12T00:33:06.488657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:33:06.488657Z digest=sha256:0d1773fe36181b4c91a1fd303e2bb0ec0cb29b80529f40b340aa1f542dccbdb3

Observation dd46e826-4f93-4f9f-b525-5d9c1352e1b4 · outbound

This paper cites MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention.

Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-12T00:33:06.488657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:33:06.488657Z digest=sha256:8bacc8188feb6784ecd9f3932a23c1bdad6079835b351c22fe66a18d09ca0022

Observation 9bc99ff2-bf4f-4d2b-aa4e-5bfd34644a46 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry Proximal Policy Optimization Algorithms

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-12T00:33:06.488657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:33:06.488657Z digest=sha256:75cbb3330834670daad042d4ac814c2e81f84d4a0743334f8ded8b7666312d5d

Observation 990de9bf-fb05-4282-a4ac-c72396371ae6 · outbound

This paper cites Experiential reinforcement learning.arXiv preprint arXiv:2602.13949, 2026a.

Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry Experiential reinforcement learning.arXiv preprint arXiv:2602.13949, 2026a

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-12T00:33:06.488657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:33:06.488657Z digest=sha256:9cc4267785b23994e01c6dfe9741755e0553929f4728929c329cd13e56e409f7

Observation ab6ddb41-db83-439c-9ce9-d545c92e1ad1 · outbound

This paper cites Andrew Bagnell, Aarti Singh, and Andrea Zanette.

Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry Andrew Bagnell, Aarti Singh, and Andrea Zanette

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-12T00:33:06.488657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:33:06.488657Z digest=sha256:c7fa553cc3edd40de7b2cef044d9e1cb74f03240c97cc2efee0489aa03a03126

Observation 5c063ef8-7448-4456-9872-05c7a28ec0d3 · outbound

This paper cites ZeroSearch: Incentivize the Search Capability of LLMs without Searching.

Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry ZeroSearch: Incentivize the Search Capability of LLMs without Searching

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-12T00:33:06.488657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:33:06.488657Z digest=sha256:a976a97a8a905a8027dcc037bd8aed1bddceed177e0e752324a9d793c5d50345

Observation 42fbcc21-3a76-4e7c-9dc0-234b73e2a021 · outbound

This paper cites RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning.

Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-12T00:33:06.488657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:33:06.488657Z digest=sha256:7d6caa1c5e5cbe831c0d1aae17593d0cc51bbc7c49456b8fd02f4333f4ad0b57

Observation e2305126-e912-4ecb-8e5c-b2975f885fa4 · outbound

This paper cites RAGEN-2: Reasoning Collapse in Agentic RL.

Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry RAGEN-2: Reasoning Collapse in Agentic RL

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-12T00:33:06.488657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:33:06.488657Z digest=sha256:25076e78efc0b6e95568ed5adac15f65e2a239ccf2cb298ee20061706e4e294c

Observation f1cde520-3c61-4568-80ea-b2f3ad259532 · outbound

This paper cites Meta-reinforcement learning with self-reflection for agentic search.arXiv preprint arXiv:2603.11327,.

Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry Meta-reinforcement learning with self-reflection for agentic search.arXiv preprint arXiv:2603.11327,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-12T00:33:06.488657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:33:06.488657Z digest=sha256:b807052e17e4909172a3a002fc951dd7130d729d4d6598799b30f6bbfc055870

Observation ca7fe6f0-bbc7-4df2-8c8b-0f2721de1539 · outbound

This paper cites MAGE: Meta-reinforcement learning for language agents toward strategic exploration and exploitation.arXiv preprint arXiv:2603.03680,.

Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry MAGE: Meta-reinforcement learning for language agents toward strategic exploration and exploitation.arXiv preprint arXiv:2603.03680,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-12T00:33:06.488657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:33:06.488657Z digest=sha256:2fe216ae6addb0a79a45bed63ca332193ba259c49fedd9e6ef6043468d0f0395

Observation 53c62c7a-e8fd-4950-84fb-20e44709a45f · outbound

This paper cites The landscape of agentic reinforcement learning for llms: A survey.Trans.

Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry The landscape of agentic reinforcement learning for llms: A survey.Trans

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-12T00:33:06.488657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:33:06.488657Z digest=sha256:9502e17a4e345695be4717447bc546bf19aef1463ddcc5b92179946ade54798c

Observation c34e7e5f-a8a7-40f4-b444-3fab060e95ae · outbound

This paper cites Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback.

Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-12T00:33:06.488657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:33:06.488657Z digest=sha256:09a2dbad27e8b0bbafc1acff7f3e4791d9707299aff534a991aa75cdf3eacd11

Observation 76a272b9-981f-4a40-bc55-c531dd7208f9 · outbound

This paper cites Group Sequence Policy Optimization.

Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry Group Sequence Policy Optimization

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-12T00:33:06.488657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:33:06.488657Z digest=sha256:d9d379a3a1a7a88bac649eb9aacfb3076c1b6b5628fe9c2aae6e90579612bf15

Pith citing papers

No inbound Pith citation observations are available.