Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T14:52:27.301719Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 2 inbound Pith citation observations for arXiv:2502.01616.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T14:52:27.301719Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-03T12:29:40.316336Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T12:38:07.209843Z
15 of 15 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 89384354-e872-455e-bf1e-cc4fd718b371 · outbound
Preference VLM: Leveraging VLMs for Scalable Preference-Based Reinforcement Learning PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbb1aac8-dc7b-4ed0-ac89-f2ea6a603a04 · outbound
Preference VLM: Leveraging VLMs for Scalable Preference-Based Reinforcement Learning Observations are captured from Camera 2 and rendered as 300 × 300 images
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 77236849-523e-4e1c-b1f1-6242a4f10ca2 · outbound
Preference VLM: Leveraging VLMs for Scalable Preference-Based Reinforcement Learning Self-Instruct: Aligning Language Models with Self-Generated Instructions
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af46d44e-5259-47e5-8105-601d6c61ea28 · outbound
Preference VLM: Leveraging VLMs for Scalable Preference-Based Reinforcement Learning Language to Rewards for Robotic Skill Synthesis
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a9940bf-a6dd-433e-b83f-9670b9018ae3 · outbound
Preference VLM: Leveraging VLMs for Scalable Preference-Based Reinforcement Learning RIME: Robust Preference-based Reinforcement Learning with Noisy Preferences
Reference 1952
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bda703ef-bb9d-483f-997b-2915c27a2829 · outbound
Preference VLM: Leveraging VLMs for Scalable Preference-Based Reinforcement Learning Language Reward Modulation for Pretraining Reinforcement Learning
Reference 2004
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52287843-3473-4c86-9fc4-9eb2057dcd50 · outbound
Preference VLM: Leveraging VLMs for Scalable Preference-Based Reinforcement Learning AVID: Learning Multi-Stage Tasks via Pixel-Level Translation of Human Videos
Reference 2009
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b38c1319-5ace-4761-a5d0-bd20207763c1 · outbound
Preference VLM: Leveraging VLMs for Scalable Preference-Based Reinforcement Learning FuRL: Visual-Language Models as Fuzzy Rewards for Reinforcement Learning
Reference 2016
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0ed0ae6f-3252-4204-8da8-5239a319dd4d · outbound
Preference VLM: Leveraging VLMs for Scalable Preference-Based Reinforcement Learning Unresolved cited work
Reference 2018
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a09a5944-7778-4f91-a051-5a4b14dfcfad · outbound
Preference VLM: Leveraging VLMs for Scalable Preference-Based Reinforcement Learning Eureka: Human-Level Reward Design via Coding Large Language Models
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5df212af-6864-4c05-bcde-c381e175389c · outbound
Preference VLM: Leveraging VLMs for Scalable Preference-Based Reinforcement Learning Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd6314a9-eb3e-4343-9c59-039a6a44a791 · outbound
Preference VLM: Leveraging VLMs for Scalable Preference-Based Reinforcement Learning SURF: Semi-supervised Reward Learning with Data Augmentation for Feedback-efficient Preference-based Reinforcement Learning
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d1278a0-fa7f-4c21-85a8-724fb267ccb1 · outbound
Preference VLM: Leveraging VLMs for Scalable Preference-Based Reinforcement Learning Adam: A Method for Stochastic Optimization
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6333525d-a069-470e-a3c0-ad5c063c3b37 · outbound
Preference VLM: Leveraging VLMs for Scalable Preference-Based Reinforcement Learning L., Faust, A., Fiser, M., and Francis, A
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8d897723-7d68-4c35-bf67-316560bffee0 · outbound
Preference VLM: Leveraging VLMs for Scalable Preference-Based Reinforcement Learning The reward model is trained with a learning rate of 0.0003, a batch size of 128, and 200 update steps per iteration
Reference 3000
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 221239e0-a543-4473-9f65-15c87a0442a3 · inbound
Beyond Pixels: Learning Invariant Rewards for Real-World Robotics From a Few Demonstrations Preference VLM: Leveraging VLMs for Scalable Preference-Based Reinforcement Learning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation da783f53-7c16-476c-8fa3-c88d44d7d78e · inbound
CoRe: Combined Rewards with Vision-Language Model Feedback for Preference-Aligned Reinforcement Learning Preference VLM: Leveraging VLMs for Scalable Preference-Based Reinforcement Learning
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.