Pith. sign in

Paper Citation Record · LEDGER

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations

As of 11 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 0 inbound Pith citation observations for arXiv:2507.08707.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.08707 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:20:25.838910Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

38 of 38 outbound references displayed

  • verified exact1
  • verified fuzzy20
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3c3e263e-9bb5-4523-80cf-5f2812f2552e · outbound

This paper cites Human-level control through deep reinforcement learning,.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Human-level control through deep reinforcement learning,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T18:20:25.710016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:20:25.710016Z digest=sha256:cc8017548621de9b13b96971a197961a4f58091eec23255736b3204100079fde

Observation d76da8ca-b402-47e8-a52d-4225da23a541 · outbound

This paper cites an unresolved cited work.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T18:20:25.714530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:20:25.714530Z digest=sha256:ff72cae2a488bcb3a7c9437ff919033018ada440a1fe7320cbd4f437d0eac65f

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:20:25.717886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:20:25.717886Z digest=sha256:6123f2800939691f35c4f7f46398ce75832d1d49eebc2fff5b7446940ebcac40

Observation 232e92b0-299c-48a1-b84e-f9bd9d3527ec · outbound

This paper cites Deep reinforcement learning with double q-learning,.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Deep reinforcement learning with double q-learning,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:20:25.721431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:20:25.721431Z digest=sha256:53e659d7d79c77d1046a32552eba088ba0b326e23459c5fa9c8d6c55f12003ea

Observation a8029730-e6e3-4ef3-9f85-0dc3f74ce68a · outbound

This paper cites Dueling network architectures for deep reinforcement learning,.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Dueling network architectures for deep reinforcement learning,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:20:25.724923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:20:25.724923Z digest=sha256:f5535114489371f4888b7523ab8d22f205adc9d950fe2f0db051b877773dd5f2

Observation d4627158-c1ed-486a-a136-c297b7d3b994 · outbound

This paper cites Rainbow: Combining improvements in deep reinforcement learning,.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Rainbow: Combining improvements in deep reinforcement learning,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:20:26.171698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T18:20:25.728465Z digest=sha256:cc4aa5a6749797f384cdc450eda5745a7551a1515e57ca07ea60336289fe930d

Observation f7fffa32-e1b7-4a6c-b266-2e628bf3c1fc · outbound

This paper cites Agent57: Outperforming the atari human benchmark,.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Agent57: Outperforming the atari human benchmark,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:20:26.161510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T18:20:25.731984Z digest=sha256:4a52977c590fc394ef30d571dccdfc061a6e25111ffb3d321f17bab8237b858e

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:20:25.735117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:20:25.735117Z digest=sha256:63d9b8c2b492c550ff71d3e0aecbe7efb2a60bc7676f8eff8dd69a60cca75301

Observation d660e7e5-ed9c-48b2-9f6a-de95d503967d · outbound

This paper cites Hind- sight experience replay,.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Hind- sight experience replay,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T18:20:25.738520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:20:25.738520Z digest=sha256:ce9a77cdcf3eea4f831756362ffdc55b2a1f940e069cdabaa588ba57b377a2f4

Observation ca24fd49-4464-4192-9cba-c1e09a061b69 · outbound

This paper cites Policy invariance under reward transformations: Theory and application to reward shaping,.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Policy invariance under reward transformations: Theory and application to reward shaping,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T18:20:25.741654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:20:25.741654Z digest=sha256:fd3f97504f255be2eacc5a808c09c61ff83936ed97dad0db69ae08042a79d3fb

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T18:20:25.744740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:20:25.744740Z digest=sha256:8ef303970536c969162065b302dbcde7b254b3ea1e600bdfa54f1d0edc3d68a6

Observation fb9405a3-f9f2-4628-b6b6-fc5fd8a4bb8f · outbound

This paper cites Algorithms for inverse reinforcement learning.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Algorithms for inverse reinforcement learning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:20:26.139975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T18:20:25.748757Z digest=sha256:f2a48871ab65b722bc43791aed718feda091d14174becd16d7ec6c3dcfc8b900

Observation 99d5a334-74b2-43f4-9d6f-0c0398f2c446 · outbound

This paper cites A survey of inverse reinforcement learning: Challenges, methods and progress,.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations A survey of inverse reinforcement learning: Challenges, methods and progress,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:20:25.751966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:20:25.751966Z digest=sha256:ddc8028fcf4c946dca51aa88ac7dd0581899ecdf5d89d21746e817d90834759e

Observation bf9d4eba-acb4-40ea-91d0-954e62a9e65d · outbound

This paper cites Maximum entropy inverse reinforcement learning.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Maximum entropy inverse reinforcement learning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:20:26.124150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T18:20:25.754940Z digest=sha256:8d196760fa172f0fd044da5078f7cd5d7cb0bb1e5766b77299da5bb3393220fc

Observation 1ec7ecbd-e008-40f9-bb35-1315b5ff8b45 · outbound

This paper cites Relative entropy inverse reinforcement learning,.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Relative entropy inverse reinforcement learning,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:20:26.114120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T18:20:25.757931Z digest=sha256:0cd5287c7ccb6b928e217a94164ab485811afb2ffa28f3487fefbb61e05e6489

Observation f91dc819-13e8-4b2b-b8e7-8bda180ffd72 · outbound

This paper cites Guided cost learning: Deep inverse optimal control via policy optimization,.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Guided cost learning: Deep inverse optimal control via policy optimization,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:20:26.104633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T18:20:25.760986Z digest=sha256:935197f50a1d425eee48c4480485d84be71a936c46c2eb25431fffc70419f38f

Observation 7c849bc3-582a-45e8-87a3-a85ab7d30ad1 · outbound

This paper cites Preference-learning based inverse reinforcement learning for dialog control,.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Preference-learning based inverse reinforcement learning for dialog control,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:20:26.092792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T18:20:25.764010Z digest=sha256:67bb15c4ba7c08d739b5bae397c2dbb2eb659909ff784948a1b2fa8bab476670

Observation d7964256-0650-499d-bdeb-f811b26c75f1 · outbound

This paper cites Model-free preference- based reinforcement learning,.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Model-free preference- based reinforcement learning,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:20:26.081885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T18:20:25.766940Z digest=sha256:cb53e855225aad8ab4d99b6d02bcc33f5ddc0642522c8719c7bce3ea175695d4

Observation cd855081-3c73-4431-8099-6ef536c9c79d · outbound

This paper cites Reward learning from human preferences and demonstrations in atari,.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Reward learning from human preferences and demonstrations in atari,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T18:20:25.770793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:20:25.770793Z digest=sha256:cb47afdaa2515e2e3ab984b26e7050e00316d98a060e5112fde07bafa964029d

Observation bee54f88-2d1a-4a0e-8b97-7260168fc507 · outbound

This paper cites Extrapolating beyond suboptimal demonstrations via inverse reinforcement learning from observations,.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Extrapolating beyond suboptimal demonstrations via inverse reinforcement learning from observations,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:20:26.066841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T18:20:25.774286Z digest=sha256:a7320f40fff26adf3edfde003b16a5c8be2e845965b97cf90081690e1ad570e8

Observation 8ee9bfe2-6d15-484f-bfd3-47cde9b0ec2c · outbound

This paper cites Learning Reward Functions by Integrating Human Demonstrations and Preferences.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Learning Reward Functions by Integrating Human Demonstrations and Preferences

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T18:20:25.778489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:20:25.778489Z digest=sha256:15bf815a3ffa8dc20f46ba9ca51c07b97d7b69317c6421abad166454736ee08c

Observation f4e3a1e0-6a47-45dd-93b3-b46eed409d7d · outbound

This paper cites Better-than-demonstrator imitation learning via automatically-ranked demonstrations,.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Better-than-demonstrator imitation learning via automatically-ranked demonstrations,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:20:26.056843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T18:20:25.782420Z digest=sha256:fc39d19a5ace5699c71c763cc720d610f81e26975cbc509baa6e0914e69d9265

Observation 326b094d-4a6c-4a6c-8014-ead0214a095c · outbound

This paper cites Learning from suboptimal demonstration via self-supervised reward regression,.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Learning from suboptimal demonstration via self-supervised reward regression,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:20:26.046640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T18:20:25.786139Z digest=sha256:1c0e23383654dfb96c15401fc6f0219c8461c09fd92fd25e006a32f568d9dceb

Observation d43b979c-628b-4b0f-9e37-ea1c51da1bf6 · outbound

This paper cites Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning,.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T18:20:25.789375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:20:25.789375Z digest=sha256:5e1e0fa851abddca692d3ebb42b39eb73b44c3e11fea6e6cb7b06597776db48e

Observation ed52cbac-4d3b-4e44-81ce-a7deef39d4f6 · outbound

This paper cites Pyquaticus capture the flag gymnasium,.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Pyquaticus capture the flag gymnasium,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:20:26.029779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T18:20:25.792772Z digest=sha256:01583df8854192a866e4858393355838337211cd37a7d83806a7c805747af1e2

Observation e0446422-77f7-45bf-a6be-caaccd8f4c4e · outbound

This paper cites Nested autonomy for unmanned marine vehicles with moos-ivp,.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Nested autonomy for unmanned marine vehicles with moos-ivp,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:20:26.019423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T18:20:25.796031Z digest=sha256:474e00b4837f91f004fe712131c500e41b2fc6ff8a460bac9194d103ddb38cb6

Observation 7429e3e1-cd43-4450-bc6e-024f70d379bc · outbound

This paper cites Efficient training of artificial neural networks for autonomous navigation,.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Efficient training of artificial neural networks for autonomous navigation,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:20:26.008756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T18:20:25.799195Z digest=sha256:114ecd60d7db7da21f320ca3867b4c3ed9e6a310dcee4ae5edccb5a209705010

Observation a11709ba-bd7c-4f6a-b7a1-b50978201d51 · outbound

This paper cites Behavioral cloning from observation,.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Behavioral cloning from observation,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:20:25.997906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T18:20:25.802431Z digest=sha256:50869190518b0e13826bdb98db162efef4cf8b82b34d40a922500b58dfded7f6

Observation a55694fd-142c-4519-80b2-6d469ca7ed00 · outbound

This paper cites Generative adversarial nets,.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Generative adversarial nets,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T18:20:25.809841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:20:25.809841Z digest=sha256:041317bb750763c398f7d6208db47ec0d5859b55d7c65d861f31866155b79b91

Observation cd703e8b-e143-4e15-97ea-ab212b85b628 · outbound

This paper cites Generative adversarial imitation learning,.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Generative adversarial imitation learning,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:20:25.980808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T18:20:25.813120Z digest=sha256:2f8110c0a63c4abfd3f9e2ec6c576618e52e7b51650252338f8ab76695e1d657

Observation ea2561d4-701f-4003-ac5a-b8d56cddd217 · outbound

This paper cites Learning Robust Rewards with Adversarial Inverse Reinforcement Learning.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Learning Robust Rewards with Adversarial Inverse Reinforcement Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T18:20:25.816962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:20:25.816962Z digest=sha256:2d24e0417d7b82fab4fd2e48a415eabdbac11f4d056f8691157f7bd5dc512ae4

Observation fcdadd5c-ed15-4e7a-8265-1406a9c589bf · outbound

This paper cites Inverse reinforcement learning for video games.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Inverse reinforcement learning for video games

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:20:25.879792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T18:20:25.820875Z digest=sha256:fd03a71a49a9f4b13ab7e2ba42e275b12d312c181f7212f594bfeabe67f22d2c

Observation 7ffeea1a-ab22-4eb1-879e-b2a9728b90f5 · outbound

This paper cites A survey of preference-based reinforcement learning methods,.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations A survey of preference-based reinforcement learning methods,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:20:25.970626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T18:20:25.824428Z digest=sha256:c608c94ed7cad956d3db4254c3cc911e3efdd186433c3e65cce0071e84e404f0

Observation a14d0fad-69fa-42c5-89df-5d23f111d458 · outbound

This paper cites Optiongan: Learning joint reward-policy options using generative adversarial inverse reinforcement learning,.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Optiongan: Learning joint reward-policy options using generative adversarial inverse reinforcement learning,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:20:25.959978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T18:20:25.827400Z digest=sha256:b6d0d60a1579ee93ae063fb0903f1a612f818650cb015e0b77d0e27322f82f72

Observation 441a28b5-767e-48b6-ab07-f752e206dd26 · outbound

This paper cites Hierarchical relative entropy policy search,.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Hierarchical relative entropy policy search,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:20:25.949493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T18:20:25.831645Z digest=sha256:1e4f9146663ed6f608feb5be5785ddd65ff0100b6ce3f86ad30d518b57a8244d

Observation 325fe346-047f-40c1-b2bd-063739624956 · outbound

This paper cites Deep reinforcement learning from human preferences,.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Deep reinforcement learning from human preferences,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T18:20:25.835801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:20:25.835801Z digest=sha256:d05a2e238222561745f80f272daa4cfb98fbb2d32ca1a55187b27cbdfa5025d4

Observation 5e122e81-04f9-4e42-9508-a37889ed5663 · outbound

This paper cites Inverse reinforcement learning from failure,.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Inverse reinforcement learning from failure,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:20:25.934336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T18:20:25.838910Z digest=sha256:84ddadec2691a2c25d428d9f827d4bc5f68463f364250fbf73f78843506fbffd

Observation 395cd817-782a-463d-85a0-500b086c9701 · outbound

This paper cites Available: https://doi.org/10.24963/ijcai.2018/687.

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations Available: https://doi.org/10.24963/ijcai.2018/687

Reference 4957

Resolution
unresolved
no resolver link, observed 2026-08-06T18:20:25.805843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:20:25.805843Z digest=sha256:08b92e3378e754238918e763283cd2d90f2cb0dd4dc42624b2c437a32f3be7ee

Pith citing papers

No inbound Pith citation observations are available.