Pith. sign in

Paper Citation Record · LEDGER

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning

As of 8 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 0 inbound Pith citation observations for arXiv:2608.03872.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.03872 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T10:46:28.386276Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

40 of 40 outbound references displayed

  • verified exact0
  • verified fuzzy38
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c6a5e125-819a-4c7c-81a0-a85a03ff8c17 · outbound

This paper cites Precise and Dexterous Robotic Manipulation via Human-in-the-Loop Reinforcement Learning,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Precise and Dexterous Robotic Manipulation via Human-in-the-Loop Reinforcement Learning,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:32.464484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:46:24.994176Z digest=sha256:f02e6f6ee3f3a77748c8690d3b1a0fa1e81d0209ebf6fcfb897b284bcadb466b

Observation 3dbb61d6-3dbf-49e8-a7ff-f65676dc94be · outbound

This paper cites SERL: A Software Suite for Sample-Efficient Robotic Reinforcement Learning,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning SERL: A Software Suite for Sample-Efficient Robotic Reinforcement Learning,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:32.455706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:46:25.057055Z digest=sha256:a4fa319e767450c2f2605a904e9d703be70649204af5758e38654e9e3b30538b

Observation 07bbcc99-6cde-46cd-8fe9-77c30a925837 · outbound

This paper cites Efficient Online Reinforcement Learning with Offline Data,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Efficient Online Reinforcement Learning with Offline Data,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:32.447346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:46:25.109193Z digest=sha256:9bd6242f579bd94e1fd41721bbcc51162b4af633b19b550cfc30cb24c1e0f1b5

Observation 7970e09c-8226-4a0b-84c3-cc097a6ca283 · outbound

This paper cites Reinforcement Learning in Robotics: A Survey,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Reinforcement Learning in Robotics: A Survey,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:32.438802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:46:25.182950Z digest=sha256:3007449cecf68849bb3c4abbebb112891296051bdecfa0b44c9fd63abc6bf582

Observation ee003f70-77fd-41ec-9553-583f223dc2f9 · outbound

This paper cites QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:32.430466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:46:25.255645Z digest=sha256:9f8ee3c328f9e28f0710b76c3d264ca1bcbcb0c29b5bf06ee484b5130f665c1d

Observation b3096895-cdc3-49b6-a3f4-ac879ef16933 · outbound

This paper cites How to Train Your Robot with Deep Reinforcement Learning: Lessons We Have Learned,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning How to Train Your Robot with Deep Reinforcement Learning: Lessons We Have Learned,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:32.421970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:46:25.360263Z digest=sha256:96eaad558b8746edf0991b96b019ff595556db3c20827028b6dd6e3a21158718

Observation 00b39d8e-ce23-45e3-bc7e-87952e9e97d3 · outbound

This paper cites Addressing Function Approximation Error in Actor-Critic Methods,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Addressing Function Approximation Error in Actor-Critic Methods,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:32.413050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:46:25.452881Z digest=sha256:1ab691b34084236b00571273012dee85258af0f681f4348ef2ac1157f1f1dccc

Observation 40e94431-dc38-4d66-b60b-a0f90c99a8b7 · outbound

This paper cites Offline Reinforcement Learning with Implicit Q-Learning,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Offline Reinforcement Learning with Implicit Q-Learning,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:32.404067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:46:25.495996Z digest=sha256:c11f4d3e97f85d9f44677a98de20bdc9299aa4dadbbbebf059658811e491f109

Observation 2a819dc0-b74a-4df3-8ca1-d924d20837c4 · outbound

This paper cites A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:32.395749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:46:25.561058Z digest=sha256:bdd6ab3541818775fd84504eb5652e55253df0c565eee816e6115735a6f0d780

Observation 34b44d0c-441e-4c62-8145-4a128c3a3356 · outbound

This paper cites HG-DAgger: Interactive Imitation Learning with Human Experts,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning HG-DAgger: Interactive Imitation Learning with Human Experts,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:32.386727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:46:25.622269Z digest=sha256:3bb4fcb2762ba66dce6cb77cc6c41df204616c7fe5c71629cdfe67ae680b73d0

Observation f17c478f-2b73-4db6-9063-41c732a28843 · outbound

This paper cites Deep Reinforcement Learning from Human Preferences,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Deep Reinforcement Learning from Human Preferences,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:32.378305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:46:25.690310Z digest=sha256:a20dfe9f3e564dfabd1863f6e864303994c950d7d3cc6c1c79b482893cef3859

Observation 4a128a76-9719-4ecd-acc5-a861f86a7e73 · outbound

This paper cites Variational Inverse Control with Events: A General Framework for Data-Driven Reward Definition,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Variational Inverse Control with Events: A General Framework for Data-Driven Reward Definition,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:32.369472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:46:25.781019Z digest=sha256:2e386ecbf17e0083c9b61b3f35bebad4a604691b254e2782705b601efc005fed

Observation 0483ecda-299b-411f-be2a-7c5315a3f4ac · outbound

This paper cites Positive-Unlabeled Reward Learning,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Positive-Unlabeled Reward Learning,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:32.360738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:46:25.865527Z digest=sha256:80bc462d5a5c8b7fa6abbc24e91e7bf8c73d2ed739f2e537a31629a2b19dfc6f

Observation 0cb43490-8acb-4247-9fd0-dd6f41a666e9 · outbound

This paper cites Human-Guided Online Reward Adaptation for Real-Robot Arm Manipulation,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Human-Guided Online Reward Adaptation for Real-Robot Arm Manipulation,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:32.350941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:46:25.947347Z digest=sha256:18d5b01e46d41b4a5d082f86b24ccc6767c2a22efcc6148dbc8add727bf535c9

Observation 69bc05e6-d327-4c3b-9aed-b739e9514e80 · outbound

This paper cites On Calibration of Modern Neural Networks,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning On Calibration of Modern Neural Networks,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:32.194875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:46:26.034319Z digest=sha256:1f78b2f3fce9a4c758cb07e70dda46c63ae5eeb06e6d63b3f4e53c19460490c6

Observation 723355f4-b2c4-4eb5-b93a-4753b1fbb54b · outbound

This paper cites Learning from Imbalanced Data,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Learning from Imbalanced Data,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:31.989060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:46:26.110725Z digest=sha256:f6fe84f97d6abf26f939bdbb166fbcb30d4e36dcd39bd43971ffe61b636646c9

Observation 2ac28726-310c-4896-acc0-01e4ed9f700b · outbound

This paper cites A Survey on Concept Drift Adaptation,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning A Survey on Concept Drift Adaptation,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:31.758765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:46:26.162134Z digest=sha256:7767d9ff72c9a4609bc0fb812f9084e1d7ee82bb00d074d33038c78a3e13d5a5

Observation 59d41c21-fca4-407a-857b-bdf9634e7584 · outbound

This paper cites Can You Trust Your Model’s Uncertainty? Evaluating Predictive Uncertainty Under Dataset Shift,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Can You Trust Your Model’s Uncertainty? Evaluating Predictive Uncertainty Under Dataset Shift,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:31.454474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:46:26.229088Z digest=sha256:f9e83a8b6f70c0dff7abde502684a7650adec8a77d7f20207bad4c3c9e3de1c4

Observation d5cb31de-eda3-46a4-bd7d-c5bb94bd9ddd · outbound

This paper cites Diffusion Policy: Visuomotor Policy Learning via Action Diffusion,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Diffusion Policy: Visuomotor Policy Learning via Action Diffusion,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:31.134034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:46:26.255218Z digest=sha256:40428fbe17927f2a7c313c587b0f28cbb871e959ff8bde79ff2fdabcca154ed7

Observation a22c2192-035a-45e7-90ae-1e414fee8020 · outbound

This paper cites Implicit Behavioral Cloning,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Implicit Behavioral Cloning,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:31.028914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:46:26.342791Z digest=sha256:1af5321b7a379e75b060b95c9d160f7c501205c59d2a5b0d9cd176e9817771cf

Observation 32dbd220-c69c-442c-be16-d6c59c2bd1f1 · outbound

This paper cites Flow Matching for Generative Modeling,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Flow Matching for Generative Modeling,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:30.835593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:46:26.396151Z digest=sha256:bc4dc4743d77a7756ba038af33ab17522c9f49728ece601214f6d93fdb592fda

Observation 9a1dcbfa-d030-4544-b0af-da1932723aba · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:30.713808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:46:26.458504Z digest=sha256:3a5cafb5825455d3bd20fc5e6bb0e4a988788ff51a6d4af461c38d8f14821f3e

Observation dbb56229-7bfe-4981-8084-ad4297a29541 · outbound

This paper cites Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:30.580894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:46:26.561614Z digest=sha256:63f28ce080c024ec7b8b3084b104436e183dae23fcb573872496bcd7b0c65400

Observation 6d19b647-08af-4fa0-a909-304ec2ad1165 · outbound

This paper cites Reinforcement Learning with Action Chunking.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Reinforcement Learning with Action Chunking

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T10:46:26.613179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:46:26.613179Z digest=sha256:8bbcb387513f5c1aee0fd4205d3f81df65d1cd7344cbe2de57812408dcb73e02

Observation 83299f6f-d6a1-4a29-8389-38c1d2a70f88 · outbound

This paper cites Diffusion Policy Policy Optimization,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Diffusion Policy Policy Optimization,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:30.513155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:46:26.667042Z digest=sha256:b5b4a4d5a43b216cf3d63280cf6928425e4995adfac907da7a80e1a364782202

Observation d2f9b537-f368-497f-969b-1ce7105477e0 · outbound

This paper cites ReinFlow: Fine-Tuning Flow Matching Policy with Online Reinforcement Learning,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning ReinFlow: Fine-Tuning Flow Matching Policy with Online Reinforcement Learning,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:30.409585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:46:26.770861Z digest=sha256:7207f41f652a9b9b9ac21b9a2f518611dd9c27a8b14f438397f2e3850efa9db7

Observation 6923b95e-7b69-446d-afe5-258403ac8ab1 · outbound

This paper cites Flow Matching Policy Gradients,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Flow Matching Policy Gradients,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:30.323687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:46:26.927731Z digest=sha256:23744c923c999882f0b1fd8c658c4a087be5ce98097dad18f62a61c99edcd8b5

Observation 6cd7f075-b8fa-4009-92b8-47037a1cc47d · outbound

This paper cites SAC Flow: Sample-Efficient Reinforcement Learning of Flow-Based Policies via Velocity-Reparameterized Sequential Model- ing,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning SAC Flow: Sample-Efficient Reinforcement Learning of Flow-Based Policies via Velocity-Reparameterized Sequential Model- ing,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:30.197362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:46:27.068200Z digest=sha256:c07308cbd82da0763efda2535bc7b490200167749589aa8971f067677e1a2312

Observation 2517fc5d-0c76-4dc7-a093-946cd622f38c · outbound

This paper cites Reinforcement Learning with Augmented Data,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Reinforcement Learning with Augmented Data,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:29.994478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:46:27.111855Z digest=sha256:f32b1e5b9cd1bfd2f4d64b5811ea9baa45d7034aeaec74818edcf66dd9fc6cc9

Observation e7cf9319-7bf0-4076-bc40-bdd9ec4bd6f4 · outbound

This paper cites Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from Pixels,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from Pixels,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:29.825849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:46:27.230961Z digest=sha256:6fe4f2580b0349baab839c6a375445c2d93d0c1be70f382efc106be357d2cac9

Observation ccc354a0-1743-4129-8397-473e10f00ace · outbound

This paper cites Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:29.674788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:46:27.392105Z digest=sha256:22c4a8b91758a6b84dfc3050cd712f91b3271a5ed316249ecebada2710f72538

Observation bc970c47-367d-4ff0-a7e2-d1f5bbaa39a3 · outbound

This paper cites AugMix: A Simple Data Processing Method to Improve Robustness and Uncertainty,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning AugMix: A Simple Data Processing Method to Improve Robustness and Uncertainty,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:29.554995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:46:27.549821Z digest=sha256:dd88ee91b51d6166c70e87e87ab08a231f2db288e3b6079613d1588be5b0250c

Observation f97f4f2d-835f-4c2c-a10a-38183256577b · outbound

This paper cites Cosmos-Transfer1: Conditional World Generation with Adaptive Multimodal Control,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Cosmos-Transfer1: Conditional World Generation with Adaptive Multimodal Control,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:29.420052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:46:27.685656Z digest=sha256:b2b7e6ec313a3d8f0a2897da5d2cd7fc8a51e9cbddd317f2882d373e5e82faf4

Observation dc7d9fb0-6292-4c23-88a6-5756ed7bfc09 · outbound

This paper cites Experience Replay for Continual Learning,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Experience Replay for Continual Learning,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:29.280319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:46:27.797275Z digest=sha256:5f88403cc7407de43ce704cf50bc81fac9d34a8594c4e3e2e801372db6ac98aa

Observation 474f25e8-3476-42bd-b255-702e7abdd32d · outbound

This paper cites Regularizing Action Policies for Smooth Control with Reinforcement Learning,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Regularizing Action Policies for Smooth Control with Reinforcement Learning,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:29.166038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:46:27.907550Z digest=sha256:cfcce95be017fd8a859edd8f6b890b6ff686409c4bffad31e5bf109656ff2e2a

Observation e85a1045-5005-4159-8b4c-1f52b905c0ff · outbound

This paper cites Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:29.009607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:46:28.019240Z digest=sha256:85b703c3116206a33dce04d6a70ce66ad122569c29ed2a2fd79603fe7c56303c

Observation 7c8b414a-ce74-4b4e-82e7-1ad220966588 · outbound

This paper cites Policy Invariance Under Reward Transformations: Theory and Application to Reward Shaping,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Policy Invariance Under Reward Transformations: Theory and Application to Reward Shaping,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:28.840976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:46:28.152810Z digest=sha256:b4e4862a1c082a1757c0f6e3a42e074c6acfca49164dbcd8f43c7434c694649e

Observation a2cf5f71-6aba-4db1-acfc-d02d664a6107 · outbound

This paper cites Imitation Bootstrapped Rein- forcement Learning,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Imitation Bootstrapped Rein- forcement Learning,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:28.686166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:46:28.219528Z digest=sha256:931d9b7d4fb4a94e30f4731532a3aaa43fcca1609e8e49cdcf07b6be82275f46

Observation 8d219ba2-7854-4e8b-bbaf-96a92aaa9b08 · outbound

This paper cites Deep Residual Learning for Image Recognition,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Deep Residual Learning for Image Recognition,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:28.536036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:46:28.301849Z digest=sha256:d630efae9d5ee16ecf101e3268fcb6610e0f00f7b413df66ff00fea5c1cd6b33

Observation e6f0ca36-6b27-4b5b-9b06-5f9778184faf · outbound

This paper cites UniRelight: Learning Joint Decomposition and Synthesis for Video Relighting.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning UniRelight: Learning Joint Decomposition and Synthesis for Video Relighting

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T10:46:28.386276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:46:28.386276Z digest=sha256:568a54b4b22831cc6d5726cebd779c2c03f367fdc565cadc572ffa28eef2b87a

Pith citing papers

No inbound Pith citation observations are available.