Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T10:46:28.386276Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 0 inbound Pith citation observations for arXiv:2608.03872.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T10:46:28.386276Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
40 of 40 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c6a5e125-819a-4c7c-81a0-a85a03ff8c17 · outbound
EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Precise and Dexterous Robotic Manipulation via Human-in-the-Loop Reinforcement Learning,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3dbb61d6-3dbf-49e8-a7ff-f65676dc94be · outbound
EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning SERL: A Software Suite for Sample-Efficient Robotic Reinforcement Learning,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 07bbcc99-6cde-46cd-8fe9-77c30a925837 · outbound
EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Efficient Online Reinforcement Learning with Offline Data,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7970e09c-8226-4a0b-84c3-cc097a6ca283 · outbound
EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Reinforcement Learning in Robotics: A Survey,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ee003f70-77fd-41ec-9553-583f223dc2f9 · outbound
EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b3096895-cdc3-49b6-a3f4-ac879ef16933 · outbound
EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning How to Train Your Robot with Deep Reinforcement Learning: Lessons We Have Learned,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 00b39d8e-ce23-45e3-bc7e-87952e9e97d3 · outbound
EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Addressing Function Approximation Error in Actor-Critic Methods,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 40e94431-dc38-4d66-b60b-a0f90c99a8b7 · outbound
EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Offline Reinforcement Learning with Implicit Q-Learning,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2a819dc0-b74a-4df3-8ca1-d924d20837c4 · outbound
EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 34b44d0c-441e-4c62-8145-4a128c3a3356 · outbound
EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning HG-DAgger: Interactive Imitation Learning with Human Experts,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f17c478f-2b73-4db6-9063-41c732a28843 · outbound
EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Deep Reinforcement Learning from Human Preferences,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4a128a76-9719-4ecd-acc5-a861f86a7e73 · outbound
EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Variational Inverse Control with Events: A General Framework for Data-Driven Reward Definition,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0483ecda-299b-411f-be2a-7c5315a3f4ac · outbound
EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Positive-Unlabeled Reward Learning,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0cb43490-8acb-4247-9fd0-dd6f41a666e9 · outbound
EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Human-Guided Online Reward Adaptation for Real-Robot Arm Manipulation,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 69bc05e6-d327-4c3b-9aed-b739e9514e80 · outbound
EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning On Calibration of Modern Neural Networks,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 723355f4-b2c4-4eb5-b93a-4753b1fbb54b · outbound
EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Learning from Imbalanced Data,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2ac28726-310c-4896-acc0-01e4ed9f700b · outbound
EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning A Survey on Concept Drift Adaptation,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 59d41c21-fca4-407a-857b-bdf9634e7584 · outbound
EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Can You Trust Your Model’s Uncertainty? Evaluating Predictive Uncertainty Under Dataset Shift,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d5cb31de-eda3-46a4-bd7d-c5bb94bd9ddd · outbound
EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Diffusion Policy: Visuomotor Policy Learning via Action Diffusion,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a22c2192-035a-45e7-90ae-1e414fee8020 · outbound
EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Implicit Behavioral Cloning,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 32dbd220-c69c-442c-be16-d6c59c2bd1f1 · outbound
EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Flow Matching for Generative Modeling,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9a1dcbfa-d030-4544-b0af-da1932723aba · outbound
EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation dbb56229-7bfe-4981-8084-ad4297a29541 · outbound
EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6d19b647-08af-4fa0-a909-304ec2ad1165 · outbound
EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Reinforcement Learning with Action Chunking
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83299f6f-d6a1-4a29-8389-38c1d2a70f88 · outbound
EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Diffusion Policy Policy Optimization,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d2f9b537-f368-497f-969b-1ce7105477e0 · outbound
EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning ReinFlow: Fine-Tuning Flow Matching Policy with Online Reinforcement Learning,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6923b95e-7b69-446d-afe5-258403ac8ab1 · outbound
EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Flow Matching Policy Gradients,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6cd7f075-b8fa-4009-92b8-47037a1cc47d · outbound
EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning SAC Flow: Sample-Efficient Reinforcement Learning of Flow-Based Policies via Velocity-Reparameterized Sequential Model- ing,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2517fc5d-0c76-4dc7-a093-946cd622f38c · outbound
EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Reinforcement Learning with Augmented Data,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e7cf9319-7bf0-4076-bc40-bdd9ec4bd6f4 · outbound
EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from Pixels,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ccc354a0-1743-4129-8397-473e10f00ace · outbound
EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bc970c47-367d-4ff0-a7e2-d1f5bbaa39a3 · outbound
EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning AugMix: A Simple Data Processing Method to Improve Robustness and Uncertainty,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f97f4f2d-835f-4c2c-a10a-38183256577b · outbound
EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Cosmos-Transfer1: Conditional World Generation with Adaptive Multimodal Control,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation dc7d9fb0-6292-4c23-88a6-5756ed7bfc09 · outbound
EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Experience Replay for Continual Learning,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 474f25e8-3476-42bd-b255-702e7abdd32d · outbound
EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Regularizing Action Policies for Smooth Control with Reinforcement Learning,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e85a1045-5005-4159-8b4c-1f52b905c0ff · outbound
EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7c8b414a-ce74-4b4e-82e7-1ad220966588 · outbound
EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Policy Invariance Under Reward Transformations: Theory and Application to Reward Shaping,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a2cf5f71-6aba-4db1-acfc-d02d664a6107 · outbound
EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Imitation Bootstrapped Rein- forcement Learning,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8d219ba2-7854-4e8b-bbaf-96a92aaa9b08 · outbound
EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Deep Residual Learning for Image Recognition,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e6f0ca36-6b27-4b5b-9b06-5f9778184faf · outbound
EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning UniRelight: Learning Joint Decomposition and Synthesis for Video Relighting
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.