Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2309.00754.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:13:18.743392Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-13T06:17:23.152582Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation beff939c-dcde-4dc1-a8e3-497cdbd1638e · inbound
HybridFlow: A Flexible and Efficient RLHF Framework Efficient RLHF: Reducing the Memory Usage of PPO
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9d141c4d-440b-4d20-8b30-a8982e54cb9d · inbound
Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Efficient RLHF: Reducing the Memory Usage of PPO
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb625c15-5994-4293-911c-4a98e22c47b9 · inbound
Aligning Large Language Models with Implicit Preferences from User-Generated Content Efficient RLHF: Reducing the Memory Usage of PPO
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a86b0c55-969d-4b60-92a1-e5a81a318488 · inbound
A Technical Survey of Reinforcement Learning Techniques for Large Language Models Efficient RLHF: Reducing the Memory Usage of PPO
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50bc0b6a-c2dd-4d41-a0cf-da219e2fa9ae · inbound
Representation-Guided Parameter-Efficient LLM Unlearning Efficient RLHF: Reducing the Memory Usage of PPO
Reference 202
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9964e064-08c2-48bf-b2b2-1a7de21bdffe · inbound
Reference-Sampled Boltzmann Projection for KL-Regularized RLVR: Target-Matched Weighted SFT, Finite One-Shot Gaps, and Policy Mirror Descent Efficient RLHF: Reducing the Memory Usage of PPO
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 786db28a-13ad-4303-b908-504c09691404 · inbound
Reinforcement Learning for Scalable and Trustworthy Intelligent Systems Efficient RLHF: Reducing the Memory Usage of PPO
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d5f26f3d-2527-426f-bff9-191c97966a0b · inbound
MARLaaS: Multi-Tenant Asynchronous Reinforcement Learning as a Service Efficient RLHF: Reducing the Memory Usage of PPO
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a70cb779-6bbd-47c5-b285-e32cf2ec6730 · inbound
Not How Many, But Which: Parameter Placement in Low-Rank Adaptation Efficient RLHF: Reducing the Memory Usage of PPO
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.