Pith. sign in

Paper Citation Record · LEDGER

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration

As of 12 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 0 inbound Pith citation observations for arXiv:2506.20307.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.20307 v1

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:01:05.049084Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

60 of 60 outbound references displayed

  • verified exact3
  • verified fuzzy49
  • unresolved6
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 29b64a58-902b-42e7-b8c8-7c39dbb83145 · outbound

This paper cites VO Q L: Towards Optimal Regret in Model-free RL with Nonlinear Function Approximation.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration VO Q L: Towards Optimal Regret in Model-free RL with Nonlinear Function Approximation

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:06.164194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:01:04.751681Z digest=sha256:ea94bd55464f243a074f6689afd4d51b6d38b85eee7b83f25a919b643b98ddef

Observation c2935928-7172-4033-b4de-e0536c1748d6 · outbound

This paper cites Hence the right-hand side of the above inequality is a martingale difference sequence.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Hence the right-hand side of the above inequality is a martingale difference sequence

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.385296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:01:05.038586Z digest=sha256:b1acd05a6d12ed9da2f6c759f133c37d0a7c219827fd8c622c4d8e1469866626

Observation d1430f1a-554f-4981-8768-5fe5da62a566 · outbound

This paper cites Mitigating Covariate Shift in Imitation Learning via Offline Data Without Great Coverage.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Mitigating Covariate Shift in Imitation Learning via Offline Data Without Great Coverage

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:04.784230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:04.784230Z digest=sha256:c147296797aae70e3a9becda3406304a89ccab138047cb84649f0dd9ad0f9a8d

Observation df7ddbfb-ad7c-403d-8beb-40225c342e4c · outbound

This paper cites Deep reinforcement learning from human preferences.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Deep reinforcement learning from human preferences

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:06.103150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:01:04.794263Z digest=sha256:57a238557f6218a271ba32635747bccfb0590163f8ebaf716745523cf15a2b44

Observation 71e633ff-c691-4c30-90cf-0455709637f0 · outbound

This paper cites Guided cost learning: Deep inverse optimal control via policy optimization.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Guided cost learning: Deep inverse optimal control via policy optimization

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:06.072586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:01:04.804541Z digest=sha256:aea93b464163a84c76afc47da3e03e4d5bd19ce23ea7197ff42c89163308c032

Observation 7a4acd9a-846d-455d-91cc-1b8b51483044 · outbound

This paper cites Efficient Bias-Span-Constrained Exploration-Exploitation in Reinforcement Learning.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Efficient Bias-Span-Constrained Exploration-Exploitation in Reinforcement Learning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:06.057428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:01:04.809354Z digest=sha256:a8f998a27f4f520e31d30cd38c674dea40da063cac464680a39e64f99040b93a

Observation 45e1ea72-f27c-4164-9f2b-1e274f29568d · outbound

This paper cites Ess- InfoGAIL: Semi-supervised Imitation Learning from Imbalanced Demonstrations.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Ess- InfoGAIL: Semi-supervised Imitation Learning from Imbalanced Demonstrations

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:06.041045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:01:04.813739Z digest=sha256:dace2745ec31f70b232e1618d8e8d99574d3e746bdc85bffd3767b92add6dee0

Observation cd84756d-95f2-4c82-a4d3-743b27f40256 · outbound

This paper cites Deep Q-learning from Demonstrations.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Deep Q-learning from Demonstrations

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:06.010041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:01:04.827906Z digest=sha256:fd76410cb138f30ebe80ce5942facf45146258757d71291ed521b35beb622bcd

Observation b133a3c7-4d75-4bfa-9866-a868c4e96ac3 · outbound

This paper cites Generative adversarial imitation learning.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Generative adversarial imitation learning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.993258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:01:04.832292Z digest=sha256:1a5d0a3cebf34e3a625e2d742474de47db8c73543de899b99b9af5518c3a1ae5

Observation 91b69499-6cc8-42b8-ac7c-3b9f59206ec2 · outbound

This paper cites 3D Perception based Imitation Learning under Limited Demonstration for Laparoscope Control in Robotic Surgery.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration 3D Perception based Imitation Learning under Limited Demonstration for Laparoscope Control in Robotic Surgery

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.959186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:01:04.841127Z digest=sha256:364e6ad7658e0a46e06a98ab771b64c79e795d7df8308ef40fdab39db3c94930

Observation c8fc85d8-5ad0-41f3-ad6c-8e10520e6407 · outbound

This paper cites Visual imitation learning with patch rewards.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Visual imitation learning with patch rewards

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.928351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:01:04.849975Z digest=sha256:daeb9e0c6a08db93f2274994846cd90d0d8c511c7805516c7764e8467b0b82cb

Observation 742c03b6-3d2f-41c0-bdf4-0f047d30e529 · outbound

This paper cites Online Learning: A Modern Introduction Using Convex Optimization.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Online Learning: A Modern Introduction Using Convex Optimization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:04.854645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:04.854645Z digest=sha256:f08a359205e87c826e19353594fcb283f07f3fb56154bb018e0252d352391b26

Observation 3ff3c01e-d0ab-4d8d-9825-ff9e557cf0af · outbound

This paper cites Online Trajectory Planning in Dynamic Environments for Surgical Task Automation.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Online Trajectory Planning in Dynamic Environments for Surgical Task Automation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.911845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:01:04.859792Z digest=sha256:0d2daeca0da987e2b84d0969971a9100fb9a365a8bddf4c3769bca6301667a69

Observation da17bb64-0cca-4218-9a21-124f22abb2bf · outbound

This paper cites Variational discriminator bottleneck: Improving imitation learning, inverse RL, and GANs by constraining information flow.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Variational discriminator bottleneck: Improving imitation learning, inverse RL, and GANs by constraining information flow

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.877453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:01:04.868617Z digest=sha256:25423480dc689b2be869530cafea30a119584943e1edf7c5680ebe0632dc7fd7

Observation f879d5e7-62b6-4b22-ae25-c9e88a38667d · outbound

This paper cites Alvinn: An autonomous land vehicle in a neural network.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Alvinn: An autonomous land vehicle in a neural network

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.860689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:01:04.873024Z digest=sha256:ec22ba25d4b198c83651d74a4bae9dee567f9fcf2a28753902309269d7654ef7

Observation 3c0ea454-4000-404a-bf68-3febb61d8b92 · outbound

This paper cites Eluder dimension and the sample complexity of optimistic ex- ploration.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Eluder dimension and the sample complexity of optimistic ex- ploration

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.829117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:01:04.887573Z digest=sha256:55bd2b817a4e5ed76f94f1ac0d14db3e4ad327ed8c36e86cb60f2f3560a0f238

Observation 41ed20c2-c928-4329-931b-d583343495c6 · outbound

This paper cites State entropy maximization with random encoders for efficient exploration.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration State entropy maximization with random encoders for efficient exploration

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.812856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:01:04.900895Z digest=sha256:437fe61caf32f737711edc5b6299eb2633cd547bf95b5b9d80162547cdd1d1ad

Observation 8062e568-92f6-4ff9-93b5-cb188b226393 · outbound

This paper cites Optimistic policy optimization with bandit feedback.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Optimistic policy optimization with bandit feedback

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.797507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:01:04.905816Z digest=sha256:3f2edb49edfeae5f841cf4301eb85907765983799cc35f5ce062c2581a1d0d8a

Observation 1b72b1f6-3135-4fe2-9106-8d1b85e3984f · outbound

This paper cites Online apprenticeship learning.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Online apprenticeship learning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.781367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:01:04.910538Z digest=sha256:6cd1475ddfcdef4ca84c5f632f6898e42d3ea7e58df299e132d31e767c5e963c

Observation 5525c199-3a0c-463a-bfe5-386352dddaef · outbound

This paper cites Error bounds of imitating policies and environments.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Error bounds of imitating policies and environments

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.715637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:01:04.930935Z digest=sha256:2fdd76e80442fbeed6d7adaef682b09644afdd6c8e11d229f8e62f30ef31ce31

Observation f4a1949e-5356-4d75-8e55-016851457701 · outbound

This paper cites Planning for sample efficient imitation learning.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Planning for sample efficient imitation learning

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.700048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:01:04.935986Z digest=sha256:3533f7f1e35c2b831246737a9389197813bfbd52fe188be5881d968c78935ef4

Observation 9301e7f9-9829-4f86-b4f6-69119740d485 · outbound

This paper cites Intrinsic reward driven imitation learning via generative model.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Intrinsic reward driven imitation learning via generative model

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.684029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:01:04.940955Z digest=sha256:b718de35163e683888c5b67d6947a87f17a10b95cb8c492cc5cd172c2a39e12e

Observation 41454812-50d6-4cb2-ac40-31fe30f4c44a · outbound

This paper cites Confidence-Aware Imitation Learning from Demonstrations with Varying Optimality.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Confidence-Aware Imitation Learning from Demonstrations with Varying Optimality

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.668551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:01:04.946103Z digest=sha256:71378c81410dfdd685fb442d5fc55e7d519ffb22c472bd1ff737e95f051d3eb6

Observation 2e8b9210-f32b-4cad-8c97-6cdc6ee02110 · outbound

This paper cites Deep imitation learning for complex manipulation tasks from virtual reality teleoperation.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Deep imitation learning for complex manipulation tasks from virtual reality teleoperation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.652735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:01:04.952083Z digest=sha256:4688930057258786d7b3e373e48b9331dc90c65392e89ea254359dc3b9554a31

Observation 56e791c1-fa1b-4b82-9e48-a56c187039e8 · outbound

This paper cites Generative adversarial imitation learning with neural network parameterization: Global optimality and convergence rate.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Generative adversarial imitation learning with neural network parameterization: Global optimality and convergence rate

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.636759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:01:04.957872Z digest=sha256:c788790caecc472b0401f3bf181d771a476646a9c6a0581f873514c2e00cb9d0

Observation 0a00e1f4-25c7-4b9f-9038-864e591e6e5e · outbound

This paper cites A Nearly Optimal and Low-Switching Algorithm for Reinforcement Learning with General Function Approximation.arXiv preprint arXiv:2311.15238,.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration A Nearly Optimal and Low-Switching Algorithm for Reinforcement Learning with General Function Approximation.arXiv preprint arXiv:2311.15238,

Reference 43

Resolution
verified exact
raw_fallback, observed 2026-08-06T23:01:05.182636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:01:04.962898Z digest=sha256:a37b4253339865f05948493f6d04d1f6f93c66f2aa1f2263cd5a842520ee862e

Observation 460a07f9-20da-4e1a-bd98-1d8dd9cbda12 · outbound

This paper cites Self-adaptive imitation learning: Learning tasks with delayed rewards from sub-optimal demonstrations.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Self-adaptive imitation learning: Learning tasks with delayed rewards from sub-optimal demonstrations

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.620954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:01:04.968464Z digest=sha256:c2c541fa90b9453de2e350c38b50d1deddc7ccbaad4b974757f25ea2b6715499

Observation 29734fda-d086-4df7-8427-4b4a258fd90f · outbound

This paper cites Maximum entropy inverse reinforcement learning.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Maximum entropy inverse reinforcement learning

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.604679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:01:04.973262Z digest=sha256:0b45f634561f2c3a213ca51e1ee56d7a81c4c5b8a22647c302d5068004a1ecd5

Observation 048c802e-77c7-494e-9aee-26f4ad98270c · outbound

This paper cites Table 3: Comparison of three reward components with three attributes.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Table 3: Comparison of three reward components with three attributes

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.569475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:01:04.983256Z digest=sha256:eff92b1b16df2c64f9261a7ff0d19d3b66fc2b1dc65574c3c6068fbe86a3a433

Observation 6b90e6a3-99bb-4dcd-b744-7602181ce3b6 · outbound

This paper cites 14 Published as a conference paper at ICLR 2025 Table 4: Demonstration lengths in the Atari environment.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration 14 Published as a conference paper at ICLR 2025 Table 4: Demonstration lengths in the Atari environment

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.551681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:01:04.987773Z digest=sha256:9bf4277763d15b80c735de01b0f3f8e8db2f70fe4c23a5c8292824d3442e1b07

Observation 24521a19-8289-447d-955c-3a17c41e9cac · outbound

This paper cites VDB constrains the information flow in the discriminator using an information bottleneck.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration VDB constrains the information flow in the discriminator using an information bottleneck

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.535824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:01:04.992820Z digest=sha256:1595e7ae04e59d319bce5565496ddbe1bc0bf6ab7cd799c8ea8ee4dde9726a7c

Observation e1069b43-fe6c-4bd7-93d9-ebe3814cbe2b · outbound

This paper cites The module is composed of several neural networks, including recognition network qϕ(z|st, st+1), a generative network pθ(st+1|z, st), and prior network pθ(z|st).

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration The module is composed of several neural networks, including recognition network qϕ(z|st, st+1), a generative network pθ(st+1|z, st), and prior network pθ(z|st)

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.519977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:01:04.997706Z digest=sha256:073740f524b7badc1f88e94b677e2810225b0970677cf51afe7947c444a2f480

Observation 8ce758f5-e18f-48de-b356-14e6c64c46bd · outbound

This paper cites State entropy estimate as bonus.Following Seo et al.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration State entropy estimate as bonus.Following Seo et al

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.504115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:01:05.002708Z digest=sha256:3305329ddb7f57b4316f4052646e9e8ef018db8d56fb9bf87729250f5c29cb1b

Observation 626a3192-7096-406e-b367-eccaccf50ef2 · outbound

This paper cites We use the same hyperparameter setting for the different experiments within a domain (Atari vs MuJoCo), apart from a multiplicative constant based on the range of the curiosity.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration We use the same hyperparameter setting for the different experiments within a domain (Atari vs MuJoCo), apart from a multiplicative constant based on the range of the curiosity

Reference 53

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T23:01:05.470126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:01:05.012779Z digest=sha256:e4a2583e55fb0e6ca0170a6c485cf94b431ef510dc4e1f25389fce82b7cc4e29

Observation b5b7e079-11c4-4f60-8379-ce67da842a0c · outbound

This paper cites Improve vs Expert.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Improve vs Expert

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.453937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:01:05.018130Z digest=sha256:6336b0eaad6c3b5529aeadfe4e0c8f3d0321776d57136b1d509b99f2a1d70285

Observation de71d438-3aa6-43c2-953c-46ad03772d60 · outbound

This paper cites Improve vs GIRIL.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Improve vs GIRIL

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.437406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:01:05.023291Z digest=sha256:20d34ac86340ccee07edf02aed09dd01efe4eaaaf4ee0dd80889e74565c9d37a

Observation 95168769-48c7-4292-8631-a800bebb5d95 · outbound

This paper cites (2023), and (3) the analysis leads to a sublinear batch-regret, which is stronger than a sample complexity bound.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration (2023), and (3) the analysis leads to a sublinear batch-regret, which is stronger than a sample complexity bound

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.403114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:01:05.033529Z digest=sha256:978a0593dd76cd74a466580eb9212ddff5748ad6acffdedbf8976b45cd9b99a0

Observation b7b8e7e6-2f8c-446e-af65-50d250a2a372 · outbound

This paper cites Then with probability at least1−2δ, KX k=1 J(π ∗,erk)−J(π,er k) ≤Hlog|A|/η+ηH 3K/2 + 2HK· p 8H 2 log(H· NF (ϵF )/δ) + 4ϵF N+γ·O s H N · H 2 +γ γ dimN (Fh) log(1/δ) !.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Then with probability at least1−2δ, KX k=1 J(π ∗,erk)−J(π,er k) ≤Hlog|A|/η+ηH 3K/2 + 2HK· p 8H 2 log(H· NF (ϵF )/δ) + 4ϵF N+γ·O s H N · H 2 +γ γ dimN (Fh) log(1/δ) !

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.368525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:01:05.043470Z digest=sha256:4b5e247493bbf77dc4d8ce71db88be16b961b9eb471c3b18b81addd1f4bfddb0

Observation c65daf9e-c0af-46ef-aa7d-b44bcf0bcbee · outbound

This paper cites Lemma D.2(Self-normalized bound for scalar-valued martingales).Consider random variables (vn|n∈N) adapted to the filtration (Hn :n= 0,1, ...).

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Lemma D.2(Self-normalized bound for scalar-valued martingales).Consider random variables (vn|n∈N) adapted to the filtration (Hn :n= 0,1, ...)

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.352347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:01:05.049084Z digest=sha256:a11ba38f1dfed3be186f529d8f59adc69dba911c37e074a81b3443ffbe1de0f9

Observation 15ae85a1-92f2-4ddd-a953-0fecbb8f89a3 · outbound

This paper cites Table 16 shows the MLP architectures, i.e., GIRIL’s encoder and decoder and V AIL’s discriminator.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Table 16 shows the MLP architectures, i.e., GIRIL’s encoder and decoder and V AIL’s discriminator

Reference 100

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T23:01:05.419654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:01:05.028464Z digest=sha256:39066d3e613759c0a7aafedd2ce5cc32e34b52b7f27aacf8d04326d1852ea3b0

Observation 3dca2e0c-e864-40b9-ab17-07e8e35015c4 · outbound

This paper cites Toward the fundamental limits of imitation learning.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Toward the fundamental limits of imitation learning

Reference 1988

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.844783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:01:04.877390Z digest=sha256:7372ee21aa01fd66f937c76731d421ccec7bedc70cd69c8384830b560ed2f714

Observation ed4f5999-4c14-45f7-a129-809e339dc0d1 · outbound

This paper cites Relative entropy inverse reinforcement learning.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Relative entropy inverse reinforcement learning

Reference 2002

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:06.149046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:01:04.757294Z digest=sha256:f06534b3f835725b4eaa61c93f4c670245fb8449e24e628111b95697d620d2f4

Observation 5d5afa12-317f-4884-ad1f-fee39e977ae9 · outbound

This paper cites Learning structured output representation using deep conditional generative models.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Learning structured output representation using deep conditional generative models

Reference 2003

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.765510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:01:04.915792Z digest=sha256:800df2203864cece0f87b53f11bc3e42f2d07808f1a0b1dc2c574dd096bd1c0f

Observation b5aefac0-2d17-4260-a7a9-046d7ec9f654 · outbound

This paper cites and Bagnell, J.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration and Bagnell, J

Reference 2008

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.586670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:01:04.978061Z digest=sha256:7dfe18da82d9a282dc5f662487a26aa8bf1d267304b164090f4c0ddf45ba1501

Observation 49dc83bb-418a-48f0-a380-a4f038286589 · outbound

This paper cites Provably efficient reinforcement learning with linear function approximation.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Provably efficient reinforcement learning with linear function approximation

Reference 2010

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.975912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:01:04.836637Z digest=sha256:f43290bf4eb1c22fc0d26bef4cbecf504e04271f3a7c3e7bb5080ab2eb91fdc7

Observation d20c8eb2-84bc-45df-a5cb-92e43fa43df2 · outbound

This paper cites OpenAI Gym.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration OpenAI Gym

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:04.762311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:04.762311Z digest=sha256:fce526e56977023876f9515f91f05548142fdabbd6e92eaf680d1c134d7d21bc

Observation ffb10c27-9f1e-4529-a1b7-1c24fc875d59 · outbound

This paper cites Proximal point imitation learning.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Proximal point imitation learning

Reference 2012

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.731530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:01:04.925993Z digest=sha256:a69967f8601e40a9aacb570382ddd763fc1ecad5c81d9bbe7ec68e4eb4a37ed3

Observation 35da6538-f151-47cf-8194-bc640d1ed319 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Proximal Policy Optimization Algorithms

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:04.893867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:04.893867Z digest=sha256:ad1e8a098637dfb458af65f9fd353a345386bc6dfa9251753f0c54f342a6eb40

Observation 9b272b49-a52a-4f0a-9472-499bc4ac3db2 · outbound

This paper cites Curiosity-driven exploration by self-supervised prediction.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Curiosity-driven exploration by self-supervised prediction

Reference 2014

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.895423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:01:04.864194Z digest=sha256:237e161cc65f6763e779bbf87fa551571dacdaa16e2520957a3119412c7463d1

Observation e4ddb88d-b381-4354-b3c4-c767fd83b6c0 · outbound

This paper cites Mujoco: A physics engine for model-based control.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Mujoco: A physics engine for model-based control

Reference 2015

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.748282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:01:04.921365Z digest=sha256:510ff058938844e07a160eaf5b2909106b5a4100be2c1b06f01a9416289df52f

Observation 22552bf1-f381-4fa2-b776-3e639ded67e3 · outbound

This paper cites Extrapolating beyond sub- optimal demonstrations via inverse reinforcement learning from observations.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Extrapolating beyond sub- optimal demonstrations via inverse reinforcement learning from observations

Reference 2016

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:06.133533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:01:04.768492Z digest=sha256:98874704017204f3c731cb0207c61b1663f36d3850249e2b33e54fa7d654c491

Observation 43d0ada8-eb40-4889-bc2d-e2d7d8a241d2 · outbound

This paper cites End-to- end driving via conditional imitation learning.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration End-to- end driving via conditional imitation learning

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:06.088173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:01:04.800160Z digest=sha256:87e779da8afdfe7baa94260b726612f4cc5dbe34b060984920661f0272c9ba6d

Observation 8fc4b90e-23d8-42d3-8a74-998af7883bec · outbound

This paper cites On the Global Convergence of Imitation Learning: A Case for Linear Quadratic Regulator.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration On the Global Convergence of Imitation Learning: A Case for Linear Quadratic Regulator

Reference 2018

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:01:05.300489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:01:04.778839Z digest=sha256:a19addff0a660556bce085faaffdc1483ec09622e921c5ef6330a46835ba03eb

Observation d5d1a19d-f8db-45e8-8aa7-21be4d6cd5cb · outbound

This paper cites Large-Scale Study of Curiosity-Driven Learning.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Large-Scale Study of Curiosity-Driven Learning

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:04.773534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:04.773534Z digest=sha256:3b0aa6f858aa03e4a5c588fa36280f6ac3586f79a5fc013fe284220ab0b54bfb

Observation e83f2d93-b30d-44a7-afce-4c234c4597cf · outbound

This paper cites Hybrid Inverse Reinforcement Learning.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Hybrid Inverse Reinforcement Learning

Reference 2020

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:01:05.226854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:01:04.882451Z digest=sha256:3470d1e5fc53741ff0ea88372849fac2e0508a289cba3c5fac92e0254e266e3f

Observation 82a4e234-d90e-4be9-9f52-028324fb64e7 · outbound

This paper cites On Computation and Generalization of Generative Adversarial Imitation Learning.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration On Computation and Generalization of Generative Adversarial Imitation Learning

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:06.117953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:01:04.789475Z digest=sha256:13746d2f42405437995d005eab3a9e03906a6cfbe26c2458e2454888a108d561

Observation f1698227-9077-41f4-b4ab-59ce9bea3dab · outbound

This paper cites Imitation Learning from Imperfection: Theoretical Justifications and Algorithms.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Imitation Learning from Imperfection: Theoretical Justifications and Algorithms

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.943710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:01:04.845606Z digest=sha256:710cc0e82ff416994d03998ed82104f2d02f179dd74ff61dad4316a970a66446

Observation 47b4d814-94b0-4d36-972e-c16de17b9058 · outbound

This paper cites IQ-Learn: Inverse soft-Q Learning for Imitation.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration IQ-Learn: Inverse soft-Q Learning for Imitation

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:06.025396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:01:04.818299Z digest=sha256:c954dceddf822d1e7cc5f81ad2935179cc4c13dd0ba0465228ff9f3b1eea9ea1

Observation 8276749f-f8d1-469e-8c3e-e8d28e216c89 · outbound

This paper cites Exploration via Elliptical Episodic Bonuses.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Exploration via Elliptical Episodic Bonuses

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:04.822802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:04.822802Z digest=sha256:7db331b60dcb593815f2e591ed0de1639d3796249c0da37174c534743d9ac987

Observation 009e47ab-f081-4445-8d6d-f6bf83f30407 · outbound

This paper cites For a fair comparison, we used an identical policy network for all methods.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration For a fair comparison, we used an identical policy network for all methods

Reference 5184

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.488424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:01:05.007817Z digest=sha256:c4146fe3d616d40c128cccece5d78c15c753dc5fce74da3168e426056d5aa900

Pith citing papers

No inbound Pith citation observations are available.