Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:45:23.355646Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 0 inbound Pith citation observations for arXiv:2505.21182.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:45:23.355646Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
44 of 44 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 2c4628ce-86d1-4cff-8765-84ec7d88f045 · outbound
Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Learning from negative feedback, or positive feedback or both
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 64791b30-10ef-4627-999a-145bda6aa78a · outbound
Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Ls-iq: Implicit reward regularization for inverse reinforcement learning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 205dd793-60b9-4425-8f25-4bd24ed16200 · outbound
Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Non-Adversarial Imitation Learning and its Connections to Adversarial Methods
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1912d9ba-2283-44e0-a095-7e5dfd3e11fa · outbound
Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Extrapolating beyond sub- optimal demonstrations via inverse reinforcement learning from observations
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 2ff03470-5320-4fbd-8bb1-bfeb0b9a8d81 · outbound
Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Diffusion policy: Visuomotor policy learning via action diffusion
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f2ea788-3d62-438b-88ee-8c8cc41ed391 · outbound
Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations D4rl: Datasets for deep data-driven reinforcement learning, 2020
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f188ec89-2133-4a54-a83c-56d677b9d520 · outbound
Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Learning robust rewards with adverserial inverse reinforcement learning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86f6ff67-01e4-47cf-b0ed-66feb0e9666e · outbound
Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Iq-learn: Inverse soft-q learning for imitation.Advances in Neural Information Processing Systems, 34:4028–4039, 2021
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d41efe19-f51f-4820-8ae0-df59e29f62bb · outbound
Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Extreme q-learning: Maxent rl without entropy
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1348ed5c-b626-47c2-ab7a-4e0dd31a737d · outbound
Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Offline safe reinforcement learning using trajectory classification
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e858749-43b9-4743-87e9-9156840f9d29 · outbound
Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Generative adversarial nets.Advances in neural information processing systems, 27, 2014
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac492e48-001b-49a3-84c5-adc1112e90ac · outbound
Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Soft Actor-Critic Algorithms and Applications
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f529ed5b-29a9-4cd9-8198-801404ef02b1 · outbound
Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Inverse preference learning: Preference-based rl without a reward function.Advances in Neural Information Processing Systems, 36, 2024
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 52247b9e-19c1-4021-b86a-6b26c97c4c89 · outbound
Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Generative adversarial imitation learning.Advances in neural information processing systems, 29, 2016
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d50e597-f61d-4d14-9411-548562ffd556 · outbound
Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Imitate the good and avoid the bad: An incremental approach to safe reinforcement learning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 18031182-88ee-4a67-8ca0-f491a21621ec · outbound
Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations SPRINQL: Sub-optimal demonstrations driven offline imitation learning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 901b9955-899b-4905-bd1f-374821406b14 · outbound
Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Safedice: offline safe imitation learning with non-preferred demonstra- tions.Advances in Neural Information Processing Systems, 36, 2024
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ebccd9e9-1a93-44c6-9eab-f3f86727b428 · outbound
Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Beyond reward: Offline preference-guided policy optimization
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7555e766-7186-488c-8900-13cc48bdb504 · outbound
Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Preference transformer: Modeling human preferences using transformers for rl
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8c6f081a-5732-4def-a972-acfccd6b7c0e · outbound
Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Lobs- dice: Offline learning from observation via stationary distribution correction estimation.Ad- vances in Neural Information Processing Systems, 35:8252–8264, 2022
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 929132a5-0aa2-4164-98bd-73051ca1f061 · outbound
Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Demodice: Offline imitation learning with supplementary imperfect demonstrations
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation bfa6dfa9-835d-4cee-bdb3-be81a1ea7a78 · outbound
Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Imitation learning via off-policy distribu- tion matching
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d7cd0d06-b351-4edb-9629-fd5a4efa2c3e · outbound
Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Offline Reinforcement Learning with Implicit Q-Learning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c410392f-0d3c-49c1-8657-0c1196d72813 · outbound
Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Optidice: Offline policy optimization via stationary distribution correction estimation
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04ffb083-fa8c-4686-9284-074cbef611d2 · outbound
Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Imitation learning from imperfection: Theoretical justifications and algorithms
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 4300174b-ee62-44f9-b5f0-eda6a886c6ce · outbound
Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Semantic loss guided data efficient supervised fine tuning for safe responses in LLMs
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 424ddd85-30e7-4b83-a606-3b869189a8f4 · outbound
Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Versatile offline imitation from observations and examples via regularized state-occupancy matching
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation fc3ca09a-e016-4aee-864b-8f1883a7923c · outbound
Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations ODICE: Revealing the mystery of distribution correction estimation via orthogonal-gradient update
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 30e0c133-e747-4e70-b5df-da0660b6d98f · outbound
Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Human-level control through deep reinforcement learning.nature, 518(7540):529–533, 2015
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cee96929-d64e-44c2-92d4-d7c3dda9849d · outbound
Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Learning multimodal rewards from rankings
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8c0e79a4-3bab-4f99-8b9d-9ac36a579a56 · outbound
Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations AlgaeDICE: Policy Gradient from Arbitrary Experience
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff04da8a-4818-467b-9f2a-93a3d7b4b752 · outbound
Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations John Wiley & Sons, 2014
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adfc0d2e-f24c-41f9-a220-8b6384514e8b · outbound
Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations SQIL: Imitation Learning via Reinforcement Learning with Sparse Rewards
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c709bc1-4fb8-4153-815e-febcac13d146 · outbound
Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations A reduction of imitation learning and structured prediction to no-regret online learning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1baaf247-61f6-47a5-ae5f-0ac77832350a · outbound
Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Dual rl: Unification and new methods for reinforcement and imitation learning
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 60918259-124d-4b73-b803-98d1a04efbdd · outbound
Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Value-Decomposition Networks For Cooperative Multi-Agent Learning
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c1093c0-bad7-441b-9fa6-b7006eeb0969 · outbound
Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Sutton and Andrew G
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6e1ccb1-e0e2-4803-9c4d-9ca5b907b114 · outbound
Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Behavioral Cloning from Observation
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a076efa-b485-4645-923a-c19e040ddb10 · outbound
Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Imitation learning from imperfect demonstration
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation df5684c9-6a3e-42b4-8da1-9a4b96a64f1c · outbound
Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Discriminator-weighted offline imitation learning from suboptimal demonstrations
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 4b9daa03-61a6-495c-9682-1b5c4454e25c · outbound
Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations How to leverage diverse demonstrations in offline imitation learning
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e736db86-5661-4b4c-b36d-6f2ba6a1e9fb · outbound
Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Confidence-aware imitation learning from demonstrations with varying optimality.Advances in Neural Information Pro- cessing Systems, 34:12340–12350, 2021
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 79f27f97-9f66-46a0-95b4-3fc3ab98fc96 · outbound
Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Learning fine-grained bimanual manipulation with low-cost hardware.Robotics: Science and Systems XIX, 2023
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bc48ea3-2bd1-4766-9042-4b38461a3eaf · outbound
Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations The ingredients of real world robotic reinforcement learning
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
No inbound Pith citation observations are available.