Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T18:20:25.838910Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 0 inbound Pith citation observations for arXiv:2507.08707.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T18:20:25.838910Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
38 of 38 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 3c3e263e-9bb5-4523-80cf-5f2812f2552e · outbound
SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Human-level control through deep reinforcement learning,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d76da8ca-b402-47e8-a52d-4225da23a541 · outbound
SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Unresolved cited work
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9659668-a760-456c-b3fd-a2c4c05e2c84 · outbound
SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Prioritized Experience Replay
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 232e92b0-299c-48a1-b84e-f9bd9d3527ec · outbound
SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Deep reinforcement learning with double q-learning,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8029730-e6e3-4ef3-9f85-0dc3f74ce68a · outbound
SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Dueling network architectures for deep reinforcement learning,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4627158-c1ed-486a-a136-c297b7d3b994 · outbound
SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Rainbow: Combining improvements in deep reinforcement learning,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f7fffa32-e1b7-4a6c-b266-2e628bf3c1fc · outbound
SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Agent57: Outperforming the atari human benchmark,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cd0789fd-9b5a-461b-aaad-6271c7a2b24b · outbound
SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Human-level Atari 200x faster
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d660e7e5-ed9c-48b2-9f6a-de95d503967d · outbound
SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Hind- sight experience replay,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca24fd49-4464-4192-9cba-c1e09a061b69 · outbound
SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Policy invariance under reward transformations: Theory and application to reward shaping,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f12da40c-1198-45e0-b5d2-6b873d5fc62c · outbound
SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Concrete Problems in AI Safety
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb9405a3-f9f2-4628-b6b6-fc5fd8a4bb8f · outbound
SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Algorithms for inverse reinforcement learning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 99d5a334-74b2-43f4-9d6f-0c0398f2c446 · outbound
SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations A survey of inverse reinforcement learning: Challenges, methods and progress,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf9d4eba-acb4-40ea-91d0-954e62a9e65d · outbound
SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Maximum entropy inverse reinforcement learning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1ec7ecbd-e008-40f9-bb35-1315b5ff8b45 · outbound
SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Relative entropy inverse reinforcement learning,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f91dc819-13e8-4b2b-b8e7-8bda180ffd72 · outbound
SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Guided cost learning: Deep inverse optimal control via policy optimization,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7c849bc3-582a-45e8-87a3-a85ab7d30ad1 · outbound
SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Preference-learning based inverse reinforcement learning for dialog control,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d7964256-0650-499d-bdeb-f811b26c75f1 · outbound
SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Model-free preference- based reinforcement learning,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cd855081-3c73-4431-8099-6ef536c9c79d · outbound
SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Reward learning from human preferences and demonstrations in atari,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bee54f88-2d1a-4a0e-8b97-7260168fc507 · outbound
SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Extrapolating beyond suboptimal demonstrations via inverse reinforcement learning from observations,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8ee9bfe2-6d15-484f-bfd3-47cde9b0ec2c · outbound
SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Learning Reward Functions by Integrating Human Demonstrations and Preferences
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4e3a1e0-6a47-45dd-93b3-b46eed409d7d · outbound
SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Better-than-demonstrator imitation learning via automatically-ranked demonstrations,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 326b094d-4a6c-4a6c-8014-ead0214a095c · outbound
SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Learning from suboptimal demonstration via self-supervised reward regression,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d43b979c-628b-4b0f-9e37-ea1c51da1bf6 · outbound
SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed52cbac-4d3b-4e44-81ce-a7deef39d4f6 · outbound
SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Pyquaticus capture the flag gymnasium,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e0446422-77f7-45bf-a6be-caaccd8f4c4e · outbound
SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Nested autonomy for unmanned marine vehicles with moos-ivp,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7429e3e1-cd43-4450-bc6e-024f70d379bc · outbound
SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Efficient training of artificial neural networks for autonomous navigation,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a11709ba-bd7c-4f6a-b7a1-b50978201d51 · outbound
SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Behavioral cloning from observation,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a55694fd-142c-4519-80b2-6d469ca7ed00 · outbound
SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Generative adversarial nets,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd703e8b-e143-4e15-97ea-ab212b85b628 · outbound
SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Generative adversarial imitation learning,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ea2561d4-701f-4003-ac5a-b8d56cddd217 · outbound
SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Learning Robust Rewards with Adversarial Inverse Reinforcement Learning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcdadd5c-ed15-4e7a-8265-1406a9c589bf · outbound
SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Inverse reinforcement learning for video games
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7ffeea1a-ab22-4eb1-879e-b2a9728b90f5 · outbound
SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations A survey of preference-based reinforcement learning methods,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a14d0fad-69fa-42c5-89df-5d23f111d458 · outbound
SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Optiongan: Learning joint reward-policy options using generative adversarial inverse reinforcement learning,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 441a28b5-767e-48b6-ab07-f752e206dd26 · outbound
SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Hierarchical relative entropy policy search,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 325fe346-047f-40c1-b2bd-063739624956 · outbound
SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Deep reinforcement learning from human preferences,
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e122e81-04f9-4e42-9508-a37889ed5663 · outbound
SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Inverse reinforcement learning from failure,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 395cd817-782a-463d-85a0-500b086c9701 · outbound
SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Available: https://doi.org/10.24963/ijcai.2018/687
Reference 4957
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.