Pith. sign in

Paper Citation Record · LEDGER

Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations

As of 15 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 0 inbound Pith citation observations for arXiv:2505.21182.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.21182 v1

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:45:23.355646Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

44 of 44 outbound references displayed

  • verified exact0
  • verified fuzzy23
  • unresolved20
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2c4628ce-86d1-4cff-8765-84ec7d88f045 · outbound

This paper cites Learning from negative feedback, or positive feedback or both.

Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Learning from negative feedback, or positive feedback or both

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:27.082515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:45:20.474848Z digest=sha256:43e3fe4ace2b84717d2d1dc696fd05bb9d2b38c51042d2a6bc131ce250ab0eed

Observation 64791b30-10ef-4627-999a-145bda6aa78a · outbound

This paper cites Ls-iq: Implicit reward regularization for inverse reinforcement learning.

Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Ls-iq: Implicit reward regularization for inverse reinforcement learning

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:26.925023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:45:20.546632Z digest=sha256:6c20a85d51215890dd0f790bbb4eba0be7a79bb290b8dfbd142dc6550201cccb

Observation 205dd793-60b9-4425-8f25-4bd24ed16200 · outbound

This paper cites Non-Adversarial Imitation Learning and its Connections to Adversarial Methods.

Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Non-Adversarial Imitation Learning and its Connections to Adversarial Methods

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:20.621562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:45:20.621562Z digest=sha256:96bb9c97e444c33a3072383bc2f669b15a1c2e9a275c492e1f1b55436fc568ff

Observation 1912d9ba-2283-44e0-a095-7e5dfd3e11fa · outbound

This paper cites Extrapolating beyond sub- optimal demonstrations via inverse reinforcement learning from observations.

Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Extrapolating beyond sub- optimal demonstrations via inverse reinforcement learning from observations

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:26.765071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:45:20.703436Z digest=sha256:5b2e2b774c0dc2ddc4fdea420992eb2acb606acdde0242d3161fb17e8122713f

Observation 2ff03470-5320-4fbd-8bb1-bfeb0b9a8d81 · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.

Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Diffusion policy: Visuomotor policy learning via action diffusion

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:20.779872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:45:20.779872Z digest=sha256:8f0c883ea29898358779301af6c41d6f6c57bb3a65a0390c955d07db6f5b4ea3

Observation 9f2ea788-3d62-438b-88ee-8c8cc41ed391 · outbound

This paper cites D4rl: Datasets for deep data-driven reinforcement learning, 2020.

Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations D4rl: Datasets for deep data-driven reinforcement learning, 2020

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:20.857897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:45:20.857897Z digest=sha256:139efbb34ebd22d521c7925dd1d06f1a8cb69c8a34400838011285ad4317ad34

Observation f188ec89-2133-4a54-a83c-56d677b9d520 · outbound

This paper cites Learning robust rewards with adverserial inverse reinforcement learning.

Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Learning robust rewards with adverserial inverse reinforcement learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:20.920843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:45:20.920843Z digest=sha256:4796770906c436688188a626a711e851073823bc4c2d8d35e6fe2f710d1a87a7

Observation 86f6ff67-01e4-47cf-b0ed-66feb0e9666e · outbound

This paper cites Iq-learn: Inverse soft-q learning for imitation.Advances in Neural Information Processing Systems, 34:4028–4039, 2021.

Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Iq-learn: Inverse soft-q learning for imitation.Advances in Neural Information Processing Systems, 34:4028–4039, 2021

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:26.595388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:45:20.987020Z digest=sha256:55dfb8debcdc256cacea6a5a235267da915d326171390240ca8dca6c1276db1e

Observation d41efe19-f51f-4820-8ae0-df59e29f62bb · outbound

This paper cites Extreme q-learning: Maxent rl without entropy.

Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Extreme q-learning: Maxent rl without entropy

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:21.055628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:45:21.055628Z digest=sha256:11550d7f0ece045149afdaeaae4d81101a88f91fb65c70cea84fba32a3c7eb38

Observation 1348ed5c-b626-47c2-ab7a-4e0dd31a737d · outbound

This paper cites Offline safe reinforcement learning using trajectory classification.

Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Offline safe reinforcement learning using trajectory classification

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:21.109441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:45:21.109441Z digest=sha256:7d3fb68c5643d6697424a06bc96028ba466fc260549671413e45989893cd81b0

Observation 6e858749-43b9-4743-87e9-9156840f9d29 · outbound

This paper cites Generative adversarial nets.Advances in neural information processing systems, 27, 2014.

Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Generative adversarial nets.Advances in neural information processing systems, 27, 2014

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:21.208925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:45:21.208925Z digest=sha256:5e6f6664ca31522cf2df1f9cdd67114d552e0ab029cac8e763abd46aff98cb56

Observation ac492e48-001b-49a3-84c5-adc1112e90ac · outbound

This paper cites Soft Actor-Critic Algorithms and Applications.

Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Soft Actor-Critic Algorithms and Applications

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:21.257006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:45:21.257006Z digest=sha256:a9e701f968bf4026135d33c78c65e7dfed6d85a3060892e0bac36b7453012af0

Observation f529ed5b-29a9-4cd9-8198-801404ef02b1 · outbound

This paper cites Inverse preference learning: Preference-based rl without a reward function.Advances in Neural Information Processing Systems, 36, 2024.

Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Inverse preference learning: Preference-based rl without a reward function.Advances in Neural Information Processing Systems, 36, 2024

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:26.457263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:45:21.308512Z digest=sha256:9165bad3ad57bb240e02028479c79a101312017f63d05b2ce08e041e52e05581

Observation 52247b9e-19c1-4021-b86a-6b26c97c4c89 · outbound

This paper cites Generative adversarial imitation learning.Advances in neural information processing systems, 29, 2016.

Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Generative adversarial imitation learning.Advances in neural information processing systems, 29, 2016

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:21.394551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:45:21.394551Z digest=sha256:2c30dfcf5affc3437d348af8794ea1841e0303cf094f029d0887e855ad6acb76

Observation 4d50e597-f61d-4d14-9411-548562ffd556 · outbound

This paper cites Imitate the good and avoid the bad: An incremental approach to safe reinforcement learning.

Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Imitate the good and avoid the bad: An incremental approach to safe reinforcement learning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:26.322307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:45:21.474304Z digest=sha256:ba57feb9549eaec51c017e6cea719c8fda1db25791a46aa4ae68a511bed86f06

Observation 18031182-88ee-4a67-8ca0-f491a21621ec · outbound

This paper cites SPRINQL: Sub-optimal demonstrations driven offline imitation learning.

Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations SPRINQL: Sub-optimal demonstrations driven offline imitation learning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:26.199634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:45:21.552971Z digest=sha256:c39d14979e63708111ecbdf3b7299b207d8abfae2afae9a5921dd4fd0b373071

Observation 901b9955-899b-4905-bd1f-374821406b14 · outbound

This paper cites Safedice: offline safe imitation learning with non-preferred demonstra- tions.Advances in Neural Information Processing Systems, 36, 2024.

Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Safedice: offline safe imitation learning with non-preferred demonstra- tions.Advances in Neural Information Processing Systems, 36, 2024

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:26.034700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:45:21.625195Z digest=sha256:3923d63c423b0f0e138a6f79dbb170f35c4e6336ef27cc558bebac55f4c0fb42

Observation ebccd9e9-1a93-44c6-9eab-f3f86727b428 · outbound

This paper cites Beyond reward: Offline preference-guided policy optimization.

Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Beyond reward: Offline preference-guided policy optimization

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:25.922860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:45:21.725329Z digest=sha256:c2723ca82b0acd2bf5d07a688647ff722b60d0fe459dd535f11899fee1f6f1f4

Observation 7555e766-7186-488c-8900-13cc48bdb504 · outbound

This paper cites Preference transformer: Modeling human preferences using transformers for rl.

Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Preference transformer: Modeling human preferences using transformers for rl

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:25.802081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:45:21.776080Z digest=sha256:f1b8894c39708d10d937d74fefd45498136a50760e304ce980ceb6365d2da22f

Observation 8c6f081a-5732-4def-a972-acfccd6b7c0e · outbound

This paper cites Lobs- dice: Offline learning from observation via stationary distribution correction estimation.Ad- vances in Neural Information Processing Systems, 35:8252–8264, 2022.

Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Lobs- dice: Offline learning from observation via stationary distribution correction estimation.Ad- vances in Neural Information Processing Systems, 35:8252–8264, 2022

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:25.682139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:45:21.841038Z digest=sha256:e22939bab8b5d88db665c128e4e4ab208fe3c7ab0d89fd4a4650cac9321b9bdd

Observation 929132a5-0aa2-4164-98bd-73051ca1f061 · outbound

This paper cites Demodice: Offline imitation learning with supplementary imperfect demonstrations.

Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Demodice: Offline imitation learning with supplementary imperfect demonstrations

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:25.533048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:45:21.906581Z digest=sha256:9f944ad1ff7c2510e03302a002e0cd429f370f601e74373ce743b89bfda6c342

Observation bfa6dfa9-835d-4cee-bdb3-be81a1ea7a78 · outbound

This paper cites Imitation learning via off-policy distribu- tion matching.

Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Imitation learning via off-policy distribu- tion matching

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:25.377199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:45:21.998421Z digest=sha256:6c5bc4949bba449b5a34bcc224d13737d34df980312e0e4b3773e767506afe03

Observation d7cd0d06-b351-4edb-9629-fd5a4efa2c3e · outbound

This paper cites Offline Reinforcement Learning with Implicit Q-Learning.

Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Offline Reinforcement Learning with Implicit Q-Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:22.055691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:45:22.055691Z digest=sha256:5ad4bd31c8b002cbcc170d694d8652aaa79135ca1be36a33ce1b76c6a511963d

Observation c410392f-0d3c-49c1-8657-0c1196d72813 · outbound

This paper cites Optidice: Offline policy optimization via stationary distribution correction estimation.

Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Optidice: Offline policy optimization via stationary distribution correction estimation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:22.126583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:45:22.126583Z digest=sha256:7ee8fc5841ffac44f721fa90aea8a97d52ac8ee37f707d09d00dcac07c8310d0

Observation 04ffb083-fa8c-4686-9284-074cbef611d2 · outbound

This paper cites Imitation learning from imperfection: Theoretical justifications and algorithms.

Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Imitation learning from imperfection: Theoretical justifications and algorithms

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:25.226569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:45:22.191612Z digest=sha256:f0433401b13c37b5bf72798ab12de21005978f3b452b77cee25e8571eaf3b7fb

Observation 4300174b-ee62-44f9-b5f0-eda6a886c6ce · outbound

This paper cites Semantic loss guided data efficient supervised fine tuning for safe responses in LLMs.

Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Semantic loss guided data efficient supervised fine tuning for safe responses in LLMs

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:25.070431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:45:22.244655Z digest=sha256:323e77c4fb35b69013174622086fe9820086ce5e29ccedf79a2080a621a32a00

Observation 424ddd85-30e7-4b83-a606-3b869189a8f4 · outbound

This paper cites Versatile offline imitation from observations and examples via regularized state-occupancy matching.

Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Versatile offline imitation from observations and examples via regularized state-occupancy matching

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:24.903313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:45:22.299475Z digest=sha256:a32465ca7073bf6c9c115110735a6d3bd27998ebfa05b83f6d3806d06ba9f561

Observation fc3ca09a-e016-4aee-864b-8f1883a7923c · outbound

This paper cites ODICE: Revealing the mystery of distribution correction estimation via orthogonal-gradient update.

Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations ODICE: Revealing the mystery of distribution correction estimation via orthogonal-gradient update

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:24.765969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:45:22.357370Z digest=sha256:a7015f0cbf44946edcd64508b0e8b185f08b8660c6616c50c9845bbdca78dfd8

Observation 30e0c133-e747-4e70-b5df-da0660b6d98f · outbound

This paper cites Human-level control through deep reinforcement learning.nature, 518(7540):529–533, 2015.

Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Human-level control through deep reinforcement learning.nature, 518(7540):529–533, 2015

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:22.428158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:45:22.428158Z digest=sha256:6d314a2483a338ac1f59c4568a4eff7bfa21ff4ac4c387173020c63b0a989436

Observation cee96929-d64e-44c2-92d4-d7c3dda9849d · outbound

This paper cites Learning multimodal rewards from rankings.

Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Learning multimodal rewards from rankings

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:24.635468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:45:22.483683Z digest=sha256:8115ddd10df6c1584e3e8b57423dff29802f996106d990aa105dd5a6797029e5

Observation 8c0e79a4-3bab-4f99-8b9d-9ac36a579a56 · outbound

This paper cites AlgaeDICE: Policy Gradient from Arbitrary Experience.

Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations AlgaeDICE: Policy Gradient from Arbitrary Experience

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:22.530387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:45:22.530387Z digest=sha256:173bd764414204247ee0a384fe0790260bb64eabf5a2d1382210d40e0e06fffc

Observation ff04da8a-4818-467b-9f2a-93a3d7b4b752 · outbound

This paper cites John Wiley & Sons, 2014.

Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations John Wiley & Sons, 2014

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:22.584986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:45:22.584986Z digest=sha256:e4371cee6f3471b430de631a62b9b8221bd156c32071f871fbe7894b3b9f5dab

Observation adfc0d2e-f24c-41f9-a220-8b6384514e8b · outbound

This paper cites SQIL: Imitation Learning via Reinforcement Learning with Sparse Rewards.

Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations SQIL: Imitation Learning via Reinforcement Learning with Sparse Rewards

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:22.644916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:45:22.644916Z digest=sha256:15eece45cd812ffbff2727ad30b1158b3ccb41e94d7d69171295dd765f024e51

Observation 2c709bc1-4fb8-4153-815e-febcac13d146 · outbound

This paper cites A reduction of imitation learning and structured prediction to no-regret online learning.

Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations A reduction of imitation learning and structured prediction to no-regret online learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:22.713611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:45:22.713611Z digest=sha256:7cfe10b3673e060876c3f80382cc2b9eef8b9e03a15e20f2cf79f78a1ffd192a

Observation 1baaf247-61f6-47a5-ae5f-0ac77832350a · outbound

This paper cites Dual rl: Unification and new methods for reinforcement and imitation learning.

Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Dual rl: Unification and new methods for reinforcement and imitation learning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:24.463802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:45:22.811006Z digest=sha256:332234bc6fd2426ecd882fd83f132de1788c04161e52c184c96ffe4108881e34

Observation 60918259-124d-4b73-b803-98d1a04efbdd · outbound

This paper cites Value-Decomposition Networks For Cooperative Multi-Agent Learning.

Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Value-Decomposition Networks For Cooperative Multi-Agent Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:22.895040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:45:22.895040Z digest=sha256:511b081af0f527ab19c242e8bb83af54887250d2a99e9568c4e77813d30121cf

Observation 5c1093c0-bad7-441b-9fa6-b7006eeb0969 · outbound

This paper cites Sutton and Andrew G.

Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Sutton and Andrew G

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:22.948033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:45:22.948033Z digest=sha256:1eb0de0a48e018641b777e563dc7271111cf80762335624c6984b1063a239dde

Observation e6e1ccb1-e0e2-4803-9c4d-9ca5b907b114 · outbound

This paper cites Behavioral Cloning from Observation.

Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Behavioral Cloning from Observation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:23.005235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:45:23.005235Z digest=sha256:52da4330a60bcb875b733b3362223f6ee06f3e60dfddab948aee2eec6bd20b7b

Observation 9a076efa-b485-4645-923a-c19e040ddb10 · outbound

This paper cites Imitation learning from imperfect demonstration.

Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Imitation learning from imperfect demonstration

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:24.346386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:45:23.063705Z digest=sha256:912d0c1e789102c1f22b8d4143a69a01ca1d4df178f7397234ed23a7d43a7329

Observation df5684c9-6a3e-42b4-8da1-9a4b96a64f1c · outbound

This paper cites Discriminator-weighted offline imitation learning from suboptimal demonstrations.

Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Discriminator-weighted offline imitation learning from suboptimal demonstrations

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:24.184828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:45:23.118107Z digest=sha256:585a2f7e3c6464af3e382831dbc491c3bbd1b02b7c6f1538dcfea5989df0a083

Observation 4b9daa03-61a6-495c-9682-1b5c4454e25c · outbound

This paper cites How to leverage diverse demonstrations in offline imitation learning.

Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations How to leverage diverse demonstrations in offline imitation learning

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:24.054467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:45:23.181502Z digest=sha256:6d96fe415d1a451f7ffee395ccd53f5c629bc4ffa0c50bedd275722599020af5

Observation e736db86-5661-4b4c-b36d-6f2ba6a1e9fb · outbound

This paper cites Confidence-aware imitation learning from demonstrations with varying optimality.Advances in Neural Information Pro- cessing Systems, 34:12340–12350, 2021.

Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Confidence-aware imitation learning from demonstrations with varying optimality.Advances in Neural Information Pro- cessing Systems, 34:12340–12350, 2021

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:23.927034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:45:23.235768Z digest=sha256:6f0d58174ab5022bf0a9f84a86d269bc8cc3c3e2437de608d3aca899f8dc84c1

Observation 79f27f97-9f66-46a0-95b4-3fc3ab98fc96 · outbound

This paper cites Learning fine-grained bimanual manipulation with low-cost hardware.Robotics: Science and Systems XIX, 2023.

Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations Learning fine-grained bimanual manipulation with low-cost hardware.Robotics: Science and Systems XIX, 2023

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:23.301455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:45:23.301455Z digest=sha256:6ede7332a4558b82ac2390b568ae2c27d8306d8ebf432678a557a98c95d36f8f

Observation 2bc48ea3-2bd1-4766-9042-4b38461a3eaf · outbound

This paper cites The ingredients of real world robotic reinforcement learning.

Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations The ingredients of real world robotic reinforcement learning

Reference 44

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T13:45:23.733063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:45:23.355646Z digest=sha256:f2db2b073fa98ad7b3b9eef0a999527b47cc427847993d622b953537430bc38d

Pith citing papers

No inbound Pith citation observations are available.