Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T19:38:51.086103Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 6 inbound Pith citation observations for arXiv:2506.15421.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T19:38:51.086103Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-27T18:08:19.797150Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T13:29:50.941267Z
36 of 36 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 49a2ba99-8f48-418d-a2b3-2d472e72612c · outbound
Reward Models in Deep Reinforcement Learning: A Survey Vision-Language Models as a Source of Rewards
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b81b121-5b55-402d-b1e7-7a72d76bf2d1 · outbound
Reward Models in Deep Reinforcement Learning: A Survey Diversity is All You Need: Learning Skills without a Reward Function
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f002d8c3-1e2c-423d-8e3e-a11ad8568f40 · outbound
Reward Models in Deep Reinforcement Learning: A Survey Quantifying Differences in Reward Functions
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0bff0728-f114-4b9d-8edc-c33ad3690e8d · outbound
Reward Models in Deep Reinforcement Learning: A Survey Preprocessing Reward Functions for Interpretability
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4a188ef-28ba-45b9-b6d5-9f3237367050 · outbound
Reward Models in Deep Reinforcement Learning: A Survey Regularized Inverse Reinforcement Learning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9270cc22-8d05-44ab-bba5-0b97572f10d2 · outbound
Reward Models in Deep Reinforcement Learning: A Survey A survey of reinforcement learning from human feedback.arXiv preprint arXiv:2312.14925, 10,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6570f42-3b80-4095-b1cc-01223c13b89b · outbound
Reward Models in Deep Reinforcement Learning: A Survey Preference Transformer: Modeling Human Preferences using Transformers for RL
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc08f9b9-f160-4338-a95e-482f090b333c · outbound
Reward Models in Deep Reinforcement Learning: A Survey Empowerment: A universal agent-centric measure of control
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation fb6b9eb2-7d20-4713-b7cb-c33da20962a7 · outbound
Reward Models in Deep Reinforcement Learning: A Survey Goal-Conditioned Reinforcement Learning: Problems and Solutions
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5144f0ef-b109-4920-a17b-a2d97aab60bc · outbound
Reward Models in Deep Reinforcement Learning: A Survey Reward Modeling with Ordinal Feedback: Wisdom of the Crowd
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a6529e8c-1975-4c4e-86d9-64d2cd39f9a1 · outbound
Reward Models in Deep Reinforcement Learning: A Survey ReFT: Reasoning with Reinforced Fine-Tuning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4795901f-7eab-45bf-ab6e-6c5ea4aec581 · outbound
Reward Models in Deep Reinforcement Learning: A Survey Choreographer: Learning and Adapting Skills in Imagination
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61063ea3-cb44-4b2e-9e8c-769bdd690a8a · outbound
Reward Models in Deep Reinforcement Learning: A Survey Learning to Assist Humans without Inferring Rewards
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 598ec36a-b1d1-4312-86f5-6eeb6144bcfb · outbound
Reward Models in Deep Reinforcement Learning: A Survey METRA: Scalable Unsupervised RL with Metric-Aware Abstraction
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9c43ce1-7ec8-434d-99cf-8d878f4a37cb · outbound
Reward Models in Deep Reinforcement Learning: A Survey Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26e7ff2b-2682-45f3-8248-ac6aaa0779a6 · outbound
Reward Models in Deep Reinforcement Learning: A Survey STARC: A General Framework For Quantifying Differences Between Reward Functions
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd93ee8b-7900-4168-838c-4bfeb2a8f248 · outbound
Reward Models in Deep Reinforcement Learning: A Survey Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b80eeca1-34af-449e-9d5f-de88fe30a17a · outbound
Reward Models in Deep Reinforcement Learning: A Survey Gymnasium: A Standard Interface for Reinforcement Learning Environments
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 199389b0-881b-4c00-9e30-6e75e105ec4a · outbound
Reward Models in Deep Reinforcement Learning: A Survey Hindsight PRIORs for Reward Learning from Human Preferences
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation bea44e5a-ebf5-4d1a-a72c-bffd2ba218a8 · outbound
Reward Models in Deep Reinforcement Learning: A Survey Text2Reward: Reward Shaping with Language Models for Reinforcement Learning
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81ce2191-54d4-409a-8c40-53578b8eaa83 · outbound
Reward Models in Deep Reinforcement Learning: A Survey DeepSeek-Prover-V1.5: Harnessing Proof Assistant Feedback for Reinforcement Learning and Monte-Carlo Tree Search
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9dbf7c04-eef8-4c1e-83be-43a1cee1455f · outbound
Reward Models in Deep Reinforcement Learning: A Survey Generative Verifiers: Reward Modeling as Next-Token Prediction
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82014311-ddbe-4e41-8db7-7b212b576b57 · outbound
Reward Models in Deep Reinforcement Learning: A Survey Maximum entropy inverse reinforcement learning
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b88bbc5f-199d-4b32-961a-8cb31b19f2c7 · outbound
Reward Models in Deep Reinforcement Learning: A Survey Dynamical Distance Learning for Semi-Supervised and Unsupervised Skill Discovery
Reference 1950
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4c7d4f9-2d29-479c-9d14-f6b860535c79 · outbound
Reward Models in Deep Reinforcement Learning: A Survey Exploration by Random Network Distillation
Reference 1952
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebb72092-549a-4dec-84b7-b23e74c8330b · outbound
Reward Models in Deep Reinforcement Learning: A Survey Improving Multimodal Interactive Agents with Reinforcement Learning from Human Feedback
Reference 2004
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation baff5cc1-6d0e-4bf8-adc7-3bcd9acfebef · outbound
Reward Models in Deep Reinforcement Learning: A Survey Models of human preference for learning reward functions
Reference 2005
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8b3e62f-02a8-49eb-b66c-a8dc62d1b0ee · outbound
Reward Models in Deep Reinforcement Learning: A Survey DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 2011
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 875992f6-9b86-4dba-a3c3-852304b4c120 · outbound
Reward Models in Deep Reinforcement Learning: A Survey Learning Robust Rewards with Adversarial Inverse Reinforcement Learning
Reference 2016
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5b5aa4f-bf1a-44fe-8b45-54f78a626e19 · outbound
Reward Models in Deep Reinforcement Learning: A Survey Safe RLHF: Safe Reinforcement Learning from Human Feedback
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53580f90-e529-44f7-a411-577256bbb502 · outbound
Reward Models in Deep Reinforcement Learning: A Survey Constitutional AI: Harmlessness from AI Feedback
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64e3677c-cd60-4840-8c23-e9dacf260f04 · outbound
Reward Models in Deep Reinforcement Learning: A Survey Never Give Up: Learning Directed Exploration Strategies
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b420d97-7414-4ab5-a1da-b8a05c7f42bc · outbound
Reward Models in Deep Reinforcement Learning: A Survey A Connection between Generative Adversarial Networks, Inverse Reinforcement Learning, and Energy-Based Models
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1be62d4a-6672-4734-b03a-985e95e42ca8 · outbound
Reward Models in Deep Reinforcement Learning: A Survey Motif: Intrinsic Motivation from Artificial Intelligence Feedback
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fe16626-45d2-4e09-a520-a493d26b41fe · outbound
Reward Models in Deep Reinforcement Learning: A Survey LiPO: Listwise Preference Optimization through Learning-to-Rank
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f0cf442-4de2-42dd-a4a1-b515064d66f3 · outbound
Reward Models in Deep Reinforcement Learning: A Survey Dynamics-Aware Comparison of Learned Reward Functions
Reference 2025
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 9c30b6bc-215c-4a92-8b91-d9a93b451a39 · inbound
Multi-objective Reinforcement Learning With Augmented States Requires Rewards After Deployment Reward Models in Deep Reinforcement Learning: A Survey
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e53582ca-beba-4ed5-90ca-36d1efeb6910 · inbound
Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Reward Models in Deep Reinforcement Learning: A Survey
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 79cea2e5-be26-43e8-9720-51304d27cddc · inbound
D-VLA: A High-Concurrency Distributed Asynchronous Reinforcement Learning Framework for Vision-Language-Action Models Reward Models in Deep Reinforcement Learning: A Survey
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ee58a84b-58e0-4b21-80ad-9659fcf303f4 · inbound
D-VLA: A High-Concurrency Distributed Asynchronous Reinforcement Learning Framework for Vision-Language-Action Models Reward Models in Deep Reinforcement Learning: A Survey
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6c8a3fc2-9855-4232-b2f2-9bf323ba31e0 · inbound
AudioProcessBench: Benchmark for Identifying Process Errors in Audio-Grounded Reasoning Reward Models in Deep Reinforcement Learning: A Survey
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ad523e39-1735-4f45-8802-1b24dff731b1 · inbound
PortraitGen: Exemplar-Driven GRPO with Dual-Reward Guidance for Photorealistic Portrait Generation Reward Models in Deep Reinforcement Learning: A Survey
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.