Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 25 inbound Pith citation observations for arXiv:2005.12729.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:15:46.288418Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T12:09:48.784830Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation d1e461ff-bee6-4969-857b-70a087d36390 · inbound
RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6320458f-5f51-4b32-b4d5-70cbfc97563a · inbound
Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO
Reference 208
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation dc0ec6a0-0fe7-4e1f-8560-d83ff9f57534 · inbound
Shallow Preference Signals: Large Language Model Aligns Even Better with Truncated Data? Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19324a54-6076-42a7-9974-d420050cd73f · inbound
Thompson Sampling in Online RLHF with General Function Approximation Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca9784f2-02ca-4061-ad86-be0546e474fc · inbound
On the Effect of Regularization in Policy Mirror Descent Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b798d612-57e9-4f25-9a09-a704bf31f366 · inbound
Shared Control of Holonomic Wheelchairs through Reinforcement Learning Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96838120-0c29-41d2-b247-d6365188a7d3 · inbound
Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO
Reference 216
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44e20981-aeb7-4aa2-b449-20b1aaf520ec · inbound
SERA: Soft-Verified Efficient Repository Agents Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43518c25-4e3d-4773-8eda-17dedb3c887c · inbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9d43959-881d-43c0-9145-44e658aec8e0 · inbound
Bounded Ratio Reinforcement Learning Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 27c12819-cd74-4918-8b56-c044b86bbd6f · inbound
Application of Deep Reinforcement Learning to Event-Triggered Control for Networked Artificial Pancreas Systems Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 07e7e605-660a-40c8-beec-f7d10bbdfcb8 · inbound
Application of Deep Reinforcement Learning to Event-Triggered Control for Networked Artificial Pancreas Systems Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4a49f367-01c8-4aa0-abed-8c07c071b344 · inbound
ANO: A Principled Approach to Robust Policy Optimization Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 52dd2441-8816-4fb8-af6a-aacbb3fc3bea · inbound
Does Synthetic Data Help? Empirical Evidence from Deep Learning Time Series Forecasters Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO
Reference 254
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bb983171-0d8d-45a4-be46-b146e61802ae · inbound
Response Time Enhances Alignment with Heterogeneous Preferences Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5cdcc69b-7f56-4ea8-828b-6657dc5f04f9 · inbound
TOPPO: Rethinking PPO for Multi-Task Reinforcement Learning with Critic Balancing Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 34485d47-4743-4400-a09d-9c895fef8c0d · inbound
Ratio-Variance Regularized Policy Optimization Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f6af7c43-22eb-45ff-a1b1-9f8337fb0d18 · inbound
Personalized Observation Normalization for Federated Reinforcement Learning in Simulation Environments with Heterogeneity Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf90a442-801c-4282-9a38-7337e7bc4607 · inbound
Dynamic Multi-Pair Trading Strategy in Cryptocurrency Markets with Deep Reinforcement Learning Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 316ce026-771d-4e4e-b3c9-692453b5f899 · inbound
Distribution-Agnostic Robust Trajectory Optimization via Chance-Constrained Reinforcement Learning Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c43b365a-919c-42ee-9c54-d17a1c1fdd92 · inbound
LOLLA: Deep Reinforcement Learning for Closed-Loop Link Adaptation Towards a GPU-Accelerated AI-RAN Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fe19167d-f4a3-4b44-9437-b41ed7f7b555 · inbound
Understanding electricity consumption behaviour through Inverse Reinforcement Learning Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47e40277-b8b8-483d-a3a0-2a33dbbe64d8 · inbound
Multi-Turn On-Policy Distillation with Prefix Replay Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO
Reference 147
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0253a4f8-8e1d-4aab-85cf-946b64d8577d · inbound
Multi-Turn On-Policy Distillation with Prefix Replay Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO
Reference 148
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40ce625e-4d6b-4884-b8bc-3a84d544f609 · inbound
ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.