Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T22:41:04.563390Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 6 inbound Pith citation observations for arXiv:2502.01652.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T22:41:04.563390Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:20:10.552803Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-01T15:15:47.315708Z
14 of 14 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 2127b893-b24d-4cd7-9845-ba120ba0d9b5 · outbound
Hybrid Group Relative Policy Optimization: A Multi-Sample Approach to Enhancing Policy Optimization Proximal Policy Optimization Algorithms
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4315583f-a384-4d71-8374-b398ba438e36 · outbound
Hybrid Group Relative Policy Optimization: A Multi-Sample Approach to Enhancing Policy Optimization Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f41338f1-11a4-47c4-86fe-dcc021e27a3a · outbound
Hybrid Group Relative Policy Optimization: A Multi-Sample Approach to Enhancing Policy Optimization Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ad30ddcb-aef3-41bb-9eba-6c4bff633a0f · outbound
Hybrid Group Relative Policy Optimization: A Multi-Sample Approach to Enhancing Policy Optimization Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 13df8e24-49a2-42e7-8a3a-44d7a1a4061d · outbound
Hybrid Group Relative Policy Optimization: A Multi-Sample Approach to Enhancing Policy Optimization S., and Barto, A
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1cb49423-ef76-41ad-8d3c-8ad9c65fd9f7 · outbound
Hybrid Group Relative Policy Optimization: A Multi-Sample Approach to Enhancing Policy Optimization P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 783b9cbb-85ed-4a2d-b019-7adaabb2421f · outbound
Hybrid Group Relative Policy Optimization: A Multi-Sample Approach to Enhancing Policy Optimization Data-efficient Deep Reinforcement Learning for Dexterous Manipulation
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ee1b285-6723-4e72-bd28-f54cea488763 · outbound
Hybrid Group Relative Policy Optimization: A Multi-Sample Approach to Enhancing Policy Optimization Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8c6b52ad-e4cd-47f2-9f56-631f3a6bfdf5 · outbound
Hybrid Group Relative Policy Optimization: A Multi-Sample Approach to Enhancing Policy Optimization J., Guez, A., Sifre, L., Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., and Lanctot, M
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 24261d01-3c06-479e-8c75-894c1965c93c · outbound
Hybrid Group Relative Policy Optimization: A Multi-Sample Approach to Enhancing Policy Optimization Prioritized Experience Replay
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd277b80-0b3e-4913-8454-2c925ba2ef64 · outbound
Hybrid Group Relative Policy Optimization: A Multi-Sample Approach to Enhancing Policy Optimization Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 207a0b77-7332-495e-808f-eba08502c127 · outbound
Hybrid Group Relative Policy Optimization: A Multi-Sample Approach to Enhancing Policy Optimization Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e196a6eb-6410-4ef2-a40a-bba0f60594b4 · outbound
Hybrid Group Relative Policy Optimization: A Multi-Sample Approach to Enhancing Policy Optimization Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5681b994-1d1f-44fc-b216-b4d54bab2baa · outbound
Hybrid Group Relative Policy Optimization: A Multi-Sample Approach to Enhancing Policy Optimization Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ae8c1636-467b-47d2-9de2-31f898d0f24a · inbound
AM-PPO: (Advantage) Alpha-Modulation with Proximal Policy Optimization Hybrid Group Relative Policy Optimization: A Multi-Sample Approach to Enhancing Policy Optimization
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7818414-6f38-458f-88c9-6ff2bb6c12b0 · inbound
Activation Control for Efficiently Eliciting Long Chain-of-thought Ability of Language Models Hybrid Group Relative Policy Optimization: A Multi-Sample Approach to Enhancing Policy Optimization
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26f352d2-bb61-4b59-9368-03abe5d0fc8b · inbound
DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Hybrid Group Relative Policy Optimization: A Multi-Sample Approach to Enhancing Policy Optimization
Reference 121
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82dad115-11ff-4a67-b5ab-5acc990cc05b · inbound
CogDual: Enhancing Dual Cognition of LLMs via Reinforcement Learning with Implicit Rule-Based Rewards Hybrid Group Relative Policy Optimization: A Multi-Sample Approach to Enhancing Policy Optimization
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a77361da-145c-4ce6-951c-ceb4e777e034 · inbound
Advantage Collapse in Group Relative Policy Optimization: Diagnosis and Mitigation Hybrid Group Relative Policy Optimization: A Multi-Sample Approach to Enhancing Policy Optimization
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e5fcd091-6d8a-4b24-85f8-3451dce56e07 · inbound
Advantage Collapse in Group Relative Policy Optimization: Diagnosis and Mitigation Hybrid Group Relative Policy Optimization: A Multi-Sample Approach to Enhancing Policy Optimization
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.