Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-12T00:33:06.488657Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 0 inbound Pith citation observations for arXiv:2607.03702.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-12T00:33:06.488657Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
19 of 19 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 959d28d3-e55a-4719-a13c-ac9a3d57164b · outbound
Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3262efab-1047-4a2e-b3bc-7cf45fe30e57 · outbound
Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry Soft Adaptive Policy Optimization
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70a98fff-27e1-4117-97ec-a411d8f9b5ec · outbound
Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry Reinforcement Learning via Self-Distillation
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cda9da9-3e62-49fd-baa4-61f36708cd5b · outbound
Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f566925c-a7b8-44e2-b50f-3deb4b74c3eb · outbound
Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry Yinghao Li, Haorui Wang, and Chao Zhang
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac97ecc7-39de-4fcb-8b8f-afd960b0042f · outbound
Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry SimpleMem: Efficient Lifelong Memory for LLM Agents
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eceb5f1b-bf95-4805-91b6-f7f185cea1a5 · outbound
Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry Self-Distilled Agentic Reinforcement Learning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd46e826-4f93-4f9f-b525-5d9c1352e1b4 · outbound
Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bc99ff2-bf4f-4d2b-aa4e-5bfd34644a46 · outbound
Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry Proximal Policy Optimization Algorithms
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 990de9bf-fb05-4282-a4ac-c72396371ae6 · outbound
Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry Experiential reinforcement learning.arXiv preprint arXiv:2602.13949, 2026a
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab6ddb41-db83-439c-9ce9-d545c92e1ad1 · outbound
Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry Andrew Bagnell, Aarti Singh, and Andrea Zanette
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c063ef8-7448-4456-9872-05c7a28ec0d3 · outbound
Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry ZeroSearch: Incentivize the Search Capability of LLMs without Searching
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42fbcc21-3a76-4e7c-9dc0-234b73e2a021 · outbound
Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2305126-e912-4ecb-8e5c-b2975f885fa4 · outbound
Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry RAGEN-2: Reasoning Collapse in Agentic RL
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1cde520-3c61-4568-80ea-b2f3ad259532 · outbound
Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry Meta-reinforcement learning with self-reflection for agentic search.arXiv preprint arXiv:2603.11327,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca7fe6f0-bbc7-4df2-8c8b-0f2721de1539 · outbound
Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry MAGE: Meta-reinforcement learning for language agents toward strategic exploration and exploitation.arXiv preprint arXiv:2603.03680,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53c62c7a-e8fd-4950-84fb-20e44709a45f · outbound
Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry The landscape of agentic reinforcement learning for llms: A survey.Trans
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c34e7e5f-a8a7-40f4-b444-3fab060e95ae · outbound
Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76a272b9-981f-4a40-bc55-c531dd7208f9 · outbound
Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry Group Sequence Policy Optimization
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.