Pith. sign in

Paper Citation Record · LEDGER

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning

As of 15 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 0 inbound Pith citation observations for arXiv:2608.03872.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.03872 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T10:46:28.386276Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

40 of 40 outbound references displayed

  • verified exact0
  • verified fuzzy38
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c6a5e125-819a-4c7c-81a0-a85a03ff8c17 · outbound

This paper cites Precise and Dexterous Robotic Manipulation via Human-in-the-Loop Reinforcement Learning,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Precise and Dexterous Robotic Manipulation via Human-in-the-Loop Reinforcement Learning,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:32.464484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T10:46:24.994176Z digest=sha256:9c9df695a94c19f1c138e4c68e499b4d380334d238d62ea009465594ff5dd1df

Observation 3dbb61d6-3dbf-49e8-a7ff-f65676dc94be · outbound

This paper cites SERL: A Software Suite for Sample-Efficient Robotic Reinforcement Learning,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning SERL: A Software Suite for Sample-Efficient Robotic Reinforcement Learning,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:32.455706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T10:46:25.057055Z digest=sha256:d14c8e48f6579176ec688df3d444673f781760050efedd4428665be0197e9ec3

Observation 07bbcc99-6cde-46cd-8fe9-77c30a925837 · outbound

This paper cites Efficient Online Reinforcement Learning with Offline Data,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Efficient Online Reinforcement Learning with Offline Data,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:32.447346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T10:46:25.109193Z digest=sha256:2b2d24d5d88fb2ab514e45b64197ac9a90f71669955a1edb4935cc7eda9abebe

Observation 7970e09c-8226-4a0b-84c3-cc097a6ca283 · outbound

This paper cites Reinforcement Learning in Robotics: A Survey,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Reinforcement Learning in Robotics: A Survey,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:32.438802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T10:46:25.182950Z digest=sha256:3c81b66c033d6f0e8dce0c2a27d12e910946bdf0a316ace32206c54483032245

Observation ee003f70-77fd-41ec-9553-583f223dc2f9 · outbound

This paper cites QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:32.430466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T10:46:25.255645Z digest=sha256:74c60c55d1f5beba148dd488e70988ccea432dd112b5bcdd9474100aa4015765

Observation b3096895-cdc3-49b6-a3f4-ac879ef16933 · outbound

This paper cites How to Train Your Robot with Deep Reinforcement Learning: Lessons We Have Learned,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning How to Train Your Robot with Deep Reinforcement Learning: Lessons We Have Learned,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:32.421970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T10:46:25.360263Z digest=sha256:8cf5c104286bd707484dafb02578f307a09923c20507bc2ddc6e5cf5d5f8b843

Observation 00b39d8e-ce23-45e3-bc7e-87952e9e97d3 · outbound

This paper cites Addressing Function Approximation Error in Actor-Critic Methods,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Addressing Function Approximation Error in Actor-Critic Methods,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:32.413050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T10:46:25.452881Z digest=sha256:9bfc876aac333ac7fcfc4f7540150ecbf2360c72d62d6323383b05a4a2512273

Observation 40e94431-dc38-4d66-b60b-a0f90c99a8b7 · outbound

This paper cites Offline Reinforcement Learning with Implicit Q-Learning,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Offline Reinforcement Learning with Implicit Q-Learning,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:32.404067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T10:46:25.495996Z digest=sha256:f73ba8fadb5d1afdf53b507df70ceeb912a83033a4658492920c8792a0ab9273

Observation 2a819dc0-b74a-4df3-8ca1-d924d20837c4 · outbound

This paper cites A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:32.395749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T10:46:25.561058Z digest=sha256:c4f65908c9db06d95087ccd59a0386d698fe4f92b7ed3e418cf99a7bcdc98906

Observation 34b44d0c-441e-4c62-8145-4a128c3a3356 · outbound

This paper cites HG-DAgger: Interactive Imitation Learning with Human Experts,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning HG-DAgger: Interactive Imitation Learning with Human Experts,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:32.386727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T10:46:25.622269Z digest=sha256:4d99d6e9a116617a17e7a3c5a9698eec4b6c127eee7f5c1e1424f7b5840ad446

Observation f17c478f-2b73-4db6-9063-41c732a28843 · outbound

This paper cites Deep Reinforcement Learning from Human Preferences,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Deep Reinforcement Learning from Human Preferences,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:32.378305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T10:46:25.690310Z digest=sha256:3b8abb1c5676130c247c7bc7647b70d0240f1675e61777d3f51a465c73379587

Observation 4a128a76-9719-4ecd-acc5-a861f86a7e73 · outbound

This paper cites Variational Inverse Control with Events: A General Framework for Data-Driven Reward Definition,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Variational Inverse Control with Events: A General Framework for Data-Driven Reward Definition,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:32.369472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T10:46:25.781019Z digest=sha256:faf96e9ee92ebad39eb66e47a24d041ce63816c34f762a700d083fdef2c43df0

Observation 0483ecda-299b-411f-be2a-7c5315a3f4ac · outbound

This paper cites Positive-Unlabeled Reward Learning,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Positive-Unlabeled Reward Learning,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:32.360738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T10:46:25.865527Z digest=sha256:98d73d8fb63c6ebd1de18c88196fe86a2dca9254b0f39491330482dfaf5d0ee6

Observation 0cb43490-8acb-4247-9fd0-dd6f41a666e9 · outbound

This paper cites Human-Guided Online Reward Adaptation for Real-Robot Arm Manipulation,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Human-Guided Online Reward Adaptation for Real-Robot Arm Manipulation,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:32.350941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T10:46:25.947347Z digest=sha256:c27b20e5c3e4f413674276e03d50206575387f082669f4acfdc86f831a2a881f

Observation 69bc05e6-d327-4c3b-9aed-b739e9514e80 · outbound

This paper cites On Calibration of Modern Neural Networks,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning On Calibration of Modern Neural Networks,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:32.194875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T10:46:26.034319Z digest=sha256:cbedbca9178f55794b9d44cecd8c187827a9aa50403b79d6cb45df5c079c1729

Observation 723355f4-b2c4-4eb5-b93a-4753b1fbb54b · outbound

This paper cites Learning from Imbalanced Data,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Learning from Imbalanced Data,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:31.989060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T10:46:26.110725Z digest=sha256:29ac5a9d5b1488628912fad6fcde938090cb7b6e20c3d977e2fc4d8c3ee055b2

Observation 2ac28726-310c-4896-acc0-01e4ed9f700b · outbound

This paper cites A Survey on Concept Drift Adaptation,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning A Survey on Concept Drift Adaptation,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:31.758765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T10:46:26.162134Z digest=sha256:f03c0baf28acd11ae84ffaa6d7c5585e2ffdd925cfe434cf95895f8fcbebfef1

Observation 59d41c21-fca4-407a-857b-bdf9634e7584 · outbound

This paper cites Can You Trust Your Model’s Uncertainty? Evaluating Predictive Uncertainty Under Dataset Shift,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Can You Trust Your Model’s Uncertainty? Evaluating Predictive Uncertainty Under Dataset Shift,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:31.454474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T10:46:26.229088Z digest=sha256:8542992e8cb36374f2e4e97614b8852776ea8634f20578d304ac149c1f51ceae

Observation d5cb31de-eda3-46a4-bd7d-c5bb94bd9ddd · outbound

This paper cites Diffusion Policy: Visuomotor Policy Learning via Action Diffusion,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Diffusion Policy: Visuomotor Policy Learning via Action Diffusion,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:31.134034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T10:46:26.255218Z digest=sha256:c85a808094085595adf0b6d26c944dd1183c5a3bdcd44f8e25ac29d040934681

Observation a22c2192-035a-45e7-90ae-1e414fee8020 · outbound

This paper cites Implicit Behavioral Cloning,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Implicit Behavioral Cloning,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:31.028914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T10:46:26.342791Z digest=sha256:4c49d2e9526923e80a51bfbc6db6823fefe8c38d6df0b91a5ae7c8e5104695df

Observation 32dbd220-c69c-442c-be16-d6c59c2bd1f1 · outbound

This paper cites Flow Matching for Generative Modeling,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Flow Matching for Generative Modeling,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:30.835593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T10:46:26.396151Z digest=sha256:510e88d1d99fb5ea00ba4db9ef416b9c2ddebdf6776d84cddeb79221fd71156b

Observation 9a1dcbfa-d030-4544-b0af-da1932723aba · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:30.713808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T10:46:26.458504Z digest=sha256:cda1698006d02746ab7a9f12ac8f10dcc7abb553ab936eced1e9fb0633068b4e

Observation dbb56229-7bfe-4981-8084-ad4297a29541 · outbound

This paper cites Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:30.580894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T10:46:26.561614Z digest=sha256:78d56f1abe4e95e6b7c85f76fcc53f8a22d0fe4b0af8265145c9dcd08ace3fb3

Observation 6d19b647-08af-4fa0-a909-304ec2ad1165 · outbound

This paper cites Reinforcement Learning with Action Chunking.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Reinforcement Learning with Action Chunking

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T10:46:26.613179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:46:26.613179Z digest=sha256:eea8b4ef04075d83784f190557fe5d0752a85eb08c8ffdc8e0623a5266bc1fc0

Observation 83299f6f-d6a1-4a29-8389-38c1d2a70f88 · outbound

This paper cites Diffusion Policy Policy Optimization,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Diffusion Policy Policy Optimization,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:30.513155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T10:46:26.667042Z digest=sha256:55b2a606d41c7d439901f60e5e9846a90fed862aa4c81d03966786e6c908b5e1

Observation d2f9b537-f368-497f-969b-1ce7105477e0 · outbound

This paper cites ReinFlow: Fine-Tuning Flow Matching Policy with Online Reinforcement Learning,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning ReinFlow: Fine-Tuning Flow Matching Policy with Online Reinforcement Learning,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:30.409585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T10:46:26.770861Z digest=sha256:1f7501033cf56baa1db0112612e9f9290cf5cf7b086b2f33638e4763a4704ea8

Observation 6923b95e-7b69-446d-afe5-258403ac8ab1 · outbound

This paper cites Flow Matching Policy Gradients,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Flow Matching Policy Gradients,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:30.323687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T10:46:26.927731Z digest=sha256:9351308e852dccdb885d8488dbf00a323f836a86de093b727c8ffc74599ee07d

Observation 6cd7f075-b8fa-4009-92b8-47037a1cc47d · outbound

This paper cites SAC Flow: Sample-Efficient Reinforcement Learning of Flow-Based Policies via Velocity-Reparameterized Sequential Model- ing,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning SAC Flow: Sample-Efficient Reinforcement Learning of Flow-Based Policies via Velocity-Reparameterized Sequential Model- ing,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:30.197362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T10:46:27.068200Z digest=sha256:ab7763ec05c617bce16da41aa26154dfc902f8e2ea860ee842ff76d0e519b051

Observation 2517fc5d-0c76-4dc7-a093-946cd622f38c · outbound

This paper cites Reinforcement Learning with Augmented Data,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Reinforcement Learning with Augmented Data,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:29.994478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T10:46:27.111855Z digest=sha256:0e0b7cd46e78eabfc3c6eec1b843da3b03f634c8bdc5376c0994cd5afa266707

Observation e7cf9319-7bf0-4076-bc40-bdd9ec4bd6f4 · outbound

This paper cites Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from Pixels,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from Pixels,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:29.825849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T10:46:27.230961Z digest=sha256:a21ae128957f881b45bcf8baa6acac6f08f5ea901ecc939e2fda105e43fec66b

Observation ccc354a0-1743-4129-8397-473e10f00ace · outbound

This paper cites Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:29.674788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T10:46:27.392105Z digest=sha256:bc9a1e1351a3676e6279f5d9f6d87e2d9b5c815d2b46511c22c0a8ec9684dfac

Observation bc970c47-367d-4ff0-a7e2-d1f5bbaa39a3 · outbound

This paper cites AugMix: A Simple Data Processing Method to Improve Robustness and Uncertainty,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning AugMix: A Simple Data Processing Method to Improve Robustness and Uncertainty,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:29.554995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T10:46:27.549821Z digest=sha256:4d0864fb18ac4000dceb517c1c881d744d16c52c05ffa9f1ffbf0acd84dd1c01

Observation f97f4f2d-835f-4c2c-a10a-38183256577b · outbound

This paper cites Cosmos-Transfer1: Conditional World Generation with Adaptive Multimodal Control,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Cosmos-Transfer1: Conditional World Generation with Adaptive Multimodal Control,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:29.420052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T10:46:27.685656Z digest=sha256:984d595acb542d91ef1ee242c3943ebb0f67130da4f7289f73b7407e13100d3e

Observation dc7d9fb0-6292-4c23-88a6-5756ed7bfc09 · outbound

This paper cites Experience Replay for Continual Learning,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Experience Replay for Continual Learning,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:29.280319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T10:46:27.797275Z digest=sha256:f3971f83a907f64f75d19a6ceade0319873363e819dc9ceafb3975c841959fff

Observation 474f25e8-3476-42bd-b255-702e7abdd32d · outbound

This paper cites Regularizing Action Policies for Smooth Control with Reinforcement Learning,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Regularizing Action Policies for Smooth Control with Reinforcement Learning,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:29.166038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T10:46:27.907550Z digest=sha256:9c011de6f17013f0dc683a0f07c1059c043fe3ad96fb7fdc6a72c4df48a76434

Observation e85a1045-5005-4159-8b4c-1f52b905c0ff · outbound

This paper cites Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:29.009607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T10:46:28.019240Z digest=sha256:bd13266fde76bc07405b11090618d202a5d073a086021496223f630b35a85db8

Observation 7c8b414a-ce74-4b4e-82e7-1ad220966588 · outbound

This paper cites Policy Invariance Under Reward Transformations: Theory and Application to Reward Shaping,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Policy Invariance Under Reward Transformations: Theory and Application to Reward Shaping,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:28.840976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T10:46:28.152810Z digest=sha256:23c1cb566c9bed19d9f71a8fe6195cb8b1f0edb82d561eab9f012d12aeabddfe

Observation a2cf5f71-6aba-4db1-acfc-d02d664a6107 · outbound

This paper cites Imitation Bootstrapped Rein- forcement Learning,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Imitation Bootstrapped Rein- forcement Learning,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:28.686166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T10:46:28.219528Z digest=sha256:943c231324442f6b66a94f9d881595d1f6131cd296577c9ae01a9248fe07d67b

Observation 8d219ba2-7854-4e8b-bbaf-96a92aaa9b08 · outbound

This paper cites Deep Residual Learning for Image Recognition,.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Deep Residual Learning for Image Recognition,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:46:28.536036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T10:46:28.301849Z digest=sha256:435136b59b36d3a96955a79647a720555c9bfad0c59a0da8d124110afdec188a

Observation e6f0ca36-6b27-4b5b-9b06-5f9778184faf · outbound

This paper cites UniRelight: Learning Joint Decomposition and Synthesis for Video Relighting.

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning UniRelight: Learning Joint Decomposition and Synthesis for Video Relighting

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T10:46:28.386276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:46:28.386276Z digest=sha256:dcf2b1644cf0496006eddb8451c7699d01bb8f54aa8de66593707b88f58fbec5

Pith citing papers

No inbound Pith citation observations are available.