Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2012.05862.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-08T13:32:44.291764Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-24T05:56:01.759090Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation e96acce3-4665-415c-9753-e3a6fbe10e02 · inbound
RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment Understanding Learned Reward Functions
Reference 127
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8cb72ddc-32ef-4ce5-8ef5-6d952a873e5c · inbound
Active teacher selection for reward learning Understanding Learned Reward Functions
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 80e47365-ef4f-4799-a910-e53086418e9a · inbound
Improve the Training Efficiency of DRL for Wireless Communication Resource Allocation: The Role of Generative Diffusion Models Understanding Learned Reward Functions
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6a1849f-7f15-4fbf-93f7-6ee490ec18ea · inbound
Beyond Correctness: Harmonizing Process and Outcome Rewards through RL Training Understanding Learned Reward Functions
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4b3466da-ceb1-46c1-be74-f78d98521990 · inbound
Multi-Turn On-Policy Distillation with Prefix Replay Understanding Learned Reward Functions
Reference 174
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation befe4177-5313-4b75-bd0f-69694ed8f6a9 · inbound
Multi-Turn On-Policy Distillation with Prefix Replay Understanding Learned Reward Functions
Reference 175
Source-reported events for the cited work
Unavailable: canonical work link unavailable.