Pith. sign in

Paper Citation Record · LEDGER

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration

As of 17 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 0 inbound Pith citation observations for arXiv:2506.20307.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.20307 v1

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:01:05.049084Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

60 of 60 outbound references displayed

  • verified exact3
  • verified fuzzy49
  • unresolved6
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 29b64a58-902b-42e7-b8c8-7c39dbb83145 · outbound

This paper cites VO Q L: Towards Optimal Regret in Model-free RL with Nonlinear Function Approximation.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration VO Q L: Towards Optimal Regret in Model-free RL with Nonlinear Function Approximation

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:06.164194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:01:04.751681Z digest=sha256:6df35761f7b406eccd07869ab61d17a1166f080cb5d593033cc5a05caf1ae26d

Observation c2935928-7172-4033-b4de-e0536c1748d6 · outbound

This paper cites Hence the right-hand side of the above inequality is a martingale difference sequence.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Hence the right-hand side of the above inequality is a martingale difference sequence

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.385296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:01:05.038586Z digest=sha256:3e929bd6709f4e1a316332ddf433672d7ff0fee767c8b46d65e561ffb56a928e

Observation d1430f1a-554f-4981-8768-5fe5da62a566 · outbound

This paper cites Mitigating Covariate Shift in Imitation Learning via Offline Data Without Great Coverage.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Mitigating Covariate Shift in Imitation Learning via Offline Data Without Great Coverage

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:04.784230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:04.784230Z digest=sha256:cf7f06f0e574a109b1d827b3a8451e97cb217a4bf3f1d296a99fab0e85db236a

Observation df7ddbfb-ad7c-403d-8beb-40225c342e4c · outbound

This paper cites Deep reinforcement learning from human preferences.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Deep reinforcement learning from human preferences

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:06.103150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:01:04.794263Z digest=sha256:9c4369927a0e3846248ddbb36dffcfff8e6ff6a9fbaf4667ada63d0beb53deb8

Observation 71e633ff-c691-4c30-90cf-0455709637f0 · outbound

This paper cites Guided cost learning: Deep inverse optimal control via policy optimization.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Guided cost learning: Deep inverse optimal control via policy optimization

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:06.072586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:01:04.804541Z digest=sha256:bef8cfc40c9b5aef7602ab2254d0dcb687a8c36c6bb20ba30950de8f075bd06c

Observation 7a4acd9a-846d-455d-91cc-1b8b51483044 · outbound

This paper cites Efficient Bias-Span-Constrained Exploration-Exploitation in Reinforcement Learning.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Efficient Bias-Span-Constrained Exploration-Exploitation in Reinforcement Learning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:06.057428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:01:04.809354Z digest=sha256:df28755bab2c8b27fcc01892a47cba6a6b917a565ecf7c938e4bac13348bc514

Observation 45e1ea72-f27c-4164-9f2b-1e274f29568d · outbound

This paper cites Ess- InfoGAIL: Semi-supervised Imitation Learning from Imbalanced Demonstrations.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Ess- InfoGAIL: Semi-supervised Imitation Learning from Imbalanced Demonstrations

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:06.041045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:01:04.813739Z digest=sha256:a170360fba1d8d5b4b3d95d6ebcbfa28a01006bda119263d430ead7dd677c7df

Observation cd84756d-95f2-4c82-a4d3-743b27f40256 · outbound

This paper cites Deep Q-learning from Demonstrations.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Deep Q-learning from Demonstrations

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:06.010041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:01:04.827906Z digest=sha256:057682324448369c007d886975cf80d65aabed9e919a662fa7690795eaaead63

Observation b133a3c7-4d75-4bfa-9866-a868c4e96ac3 · outbound

This paper cites Generative adversarial imitation learning.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Generative adversarial imitation learning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.993258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:01:04.832292Z digest=sha256:3f066ed34f4b7415dda2e0c5b4ae361a9268feb38fa81cebcfe22ec2ca15c967

Observation 91b69499-6cc8-42b8-ac7c-3b9f59206ec2 · outbound

This paper cites 3D Perception based Imitation Learning under Limited Demonstration for Laparoscope Control in Robotic Surgery.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration 3D Perception based Imitation Learning under Limited Demonstration for Laparoscope Control in Robotic Surgery

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.959186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:01:04.841127Z digest=sha256:5bf05fcd3a31b897c6ec5bef294735497725ec79f6cbbc55e0bba3030fa43965

Observation c8fc85d8-5ad0-41f3-ad6c-8e10520e6407 · outbound

This paper cites Visual imitation learning with patch rewards.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Visual imitation learning with patch rewards

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.928351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:01:04.849975Z digest=sha256:345029e8f308a575106a362893651f7e835e9b96fa47b9484ba9e8931296ddaa

Observation 742c03b6-3d2f-41c0-bdf4-0f047d30e529 · outbound

This paper cites Online Learning: A Modern Introduction Using Convex Optimization.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Online Learning: A Modern Introduction Using Convex Optimization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:04.854645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:04.854645Z digest=sha256:2fa8e2b0af63c03abf1bb911d9c7f4e7176fc54fc3116fdbab654a049aca5bcc

Observation 3ff3c01e-d0ab-4d8d-9825-ff9e557cf0af · outbound

This paper cites Online Trajectory Planning in Dynamic Environments for Surgical Task Automation.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Online Trajectory Planning in Dynamic Environments for Surgical Task Automation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.911845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:01:04.859792Z digest=sha256:f1f95db30bece482b72608e98324af302a43f28f11fdf7a6125ff79927a3a7a3

Observation da17bb64-0cca-4218-9a21-124f22abb2bf · outbound

This paper cites Variational discriminator bottleneck: Improving imitation learning, inverse RL, and GANs by constraining information flow.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Variational discriminator bottleneck: Improving imitation learning, inverse RL, and GANs by constraining information flow

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.877453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:01:04.868617Z digest=sha256:faf4776bc6577fef8c42249256888f55dd9f1ae42006c7197291687fcd5943e6

Observation f879d5e7-62b6-4b22-ae25-c9e88a38667d · outbound

This paper cites Alvinn: An autonomous land vehicle in a neural network.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Alvinn: An autonomous land vehicle in a neural network

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.860689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:01:04.873024Z digest=sha256:59b2cbb82382e73a949cded19a9e2ec88b7aad73645924bbd97172c3a2df96aa

Observation 3c0ea454-4000-404a-bf68-3febb61d8b92 · outbound

This paper cites Eluder dimension and the sample complexity of optimistic ex- ploration.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Eluder dimension and the sample complexity of optimistic ex- ploration

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.829117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:01:04.887573Z digest=sha256:03f9a3f14649f4c4e8044ce8fdfe174abc82f4a1a056fad47d454b246025ffea

Observation 41ed20c2-c928-4329-931b-d583343495c6 · outbound

This paper cites State entropy maximization with random encoders for efficient exploration.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration State entropy maximization with random encoders for efficient exploration

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.812856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:01:04.900895Z digest=sha256:2014567d8676f664c01f1278cb85fa3a1eb44a5a90bf81cb3aaae866fc31fadd

Observation 8062e568-92f6-4ff9-93b5-cb188b226393 · outbound

This paper cites Optimistic policy optimization with bandit feedback.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Optimistic policy optimization with bandit feedback

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.797507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:01:04.905816Z digest=sha256:1a22de5f91520e0ac9bcdd72581a1d18adc5b6d9a8180d627985c30f2fb2f78e

Observation 1b72b1f6-3135-4fe2-9106-8d1b85e3984f · outbound

This paper cites Online apprenticeship learning.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Online apprenticeship learning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.781367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:01:04.910538Z digest=sha256:ac382a9d26bc972ee4576a997a25fb41e4f721e91b8746ca00dfa15c8ca221bb

Observation 5525c199-3a0c-463a-bfe5-386352dddaef · outbound

This paper cites Error bounds of imitating policies and environments.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Error bounds of imitating policies and environments

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.715637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:01:04.930935Z digest=sha256:703ce01ef437ee87111d266fc008cd6aff482ae26f647f798a5b7b0cee08fcb3

Observation f4a1949e-5356-4d75-8e55-016851457701 · outbound

This paper cites Planning for sample efficient imitation learning.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Planning for sample efficient imitation learning

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.700048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:01:04.935986Z digest=sha256:dafcc4b45385623933f1b2e3c650f3bb738b19ba3902e05b5ee1c3132ce27730

Observation 9301e7f9-9829-4f86-b4f6-69119740d485 · outbound

This paper cites Intrinsic reward driven imitation learning via generative model.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Intrinsic reward driven imitation learning via generative model

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.684029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:01:04.940955Z digest=sha256:88f4edaf3d123195e2566020783b8ebfc8c034ccd3d02c7ee00a26b1db5e0301

Observation 41454812-50d6-4cb2-ac40-31fe30f4c44a · outbound

This paper cites Confidence-Aware Imitation Learning from Demonstrations with Varying Optimality.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Confidence-Aware Imitation Learning from Demonstrations with Varying Optimality

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.668551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:01:04.946103Z digest=sha256:535e5e916b4cc0e615baec574e76fe29e35ed521d8a1f643439070166423f080

Observation 2e8b9210-f32b-4cad-8c97-6cdc6ee02110 · outbound

This paper cites Deep imitation learning for complex manipulation tasks from virtual reality teleoperation.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Deep imitation learning for complex manipulation tasks from virtual reality teleoperation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.652735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:01:04.952083Z digest=sha256:19c93a609c5db9eb8812dba71c730b399e7b98eb85397d8a320f52a5781af0f3

Observation 56e791c1-fa1b-4b82-9e48-a56c187039e8 · outbound

This paper cites Generative adversarial imitation learning with neural network parameterization: Global optimality and convergence rate.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Generative adversarial imitation learning with neural network parameterization: Global optimality and convergence rate

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.636759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:01:04.957872Z digest=sha256:a0b96a532b04f9e6a8319e0c47cfafdedfa52d1437a5fed54ee01fb57c8f53f2

Observation 0a00e1f4-25c7-4b9f-9038-864e591e6e5e · outbound

This paper cites A Nearly Optimal and Low-Switching Algorithm for Reinforcement Learning with General Function Approximation.arXiv preprint arXiv:2311.15238,.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration A Nearly Optimal and Low-Switching Algorithm for Reinforcement Learning with General Function Approximation.arXiv preprint arXiv:2311.15238,

Reference 43

Resolution
verified exact
raw_fallback, observed 2026-08-06T23:01:05.182636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:01:04.962898Z digest=sha256:c55c2f7187e989d6e6ccff87ee7e1a55cffc5976babec848f92ae9b188966aeb

Observation 460a07f9-20da-4e1a-bd98-1d8dd9cbda12 · outbound

This paper cites Self-adaptive imitation learning: Learning tasks with delayed rewards from sub-optimal demonstrations.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Self-adaptive imitation learning: Learning tasks with delayed rewards from sub-optimal demonstrations

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.620954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:01:04.968464Z digest=sha256:c909faeac9f3a0a84dc5765babe53647dcf50650d2303f64a73e8411c0583bc4

Observation 29734fda-d086-4df7-8427-4b4a258fd90f · outbound

This paper cites Maximum entropy inverse reinforcement learning.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Maximum entropy inverse reinforcement learning

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.604679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:01:04.973262Z digest=sha256:aac2a743b9c608748bc83bf165f16be6e12d84bc36645d5e83af63cade7f1478

Observation 048c802e-77c7-494e-9aee-26f4ad98270c · outbound

This paper cites Table 3: Comparison of three reward components with three attributes.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Table 3: Comparison of three reward components with three attributes

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.569475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:01:04.983256Z digest=sha256:cc7ac280c1ab1712dcfae2b9c0558b1b8a1698976cdc1686279d020cd0176fea

Observation 6b90e6a3-99bb-4dcd-b744-7602181ce3b6 · outbound

This paper cites 14 Published as a conference paper at ICLR 2025 Table 4: Demonstration lengths in the Atari environment.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration 14 Published as a conference paper at ICLR 2025 Table 4: Demonstration lengths in the Atari environment

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.551681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:01:04.987773Z digest=sha256:1fa2d4448fb0b47d83ef8aa698a2e75a679055dade1ecdc55db324b96878fb16

Observation 24521a19-8289-447d-955c-3a17c41e9cac · outbound

This paper cites VDB constrains the information flow in the discriminator using an information bottleneck.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration VDB constrains the information flow in the discriminator using an information bottleneck

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.535824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:01:04.992820Z digest=sha256:7dfc4353490de7e490f56964b794c5e86b0d8e629e11bdb9c8aefff04640b619

Observation e1069b43-fe6c-4bd7-93d9-ebe3814cbe2b · outbound

This paper cites The module is composed of several neural networks, including recognition network qϕ(z|st, st+1), a generative network pθ(st+1|z, st), and prior network pθ(z|st).

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration The module is composed of several neural networks, including recognition network qϕ(z|st, st+1), a generative network pθ(st+1|z, st), and prior network pθ(z|st)

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.519977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:01:04.997706Z digest=sha256:db180011d88aeb90cc361bd9cbaa42214651b9b6748559fb379a8540642d7ef3

Observation 8ce758f5-e18f-48de-b356-14e6c64c46bd · outbound

This paper cites State entropy estimate as bonus.Following Seo et al.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration State entropy estimate as bonus.Following Seo et al

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.504115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:01:05.002708Z digest=sha256:23bc236d2d57d46d4f130746571fa659a862e8f629c4d121632af135eca374e3

Observation 626a3192-7096-406e-b367-eccaccf50ef2 · outbound

This paper cites We use the same hyperparameter setting for the different experiments within a domain (Atari vs MuJoCo), apart from a multiplicative constant based on the range of the curiosity.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration We use the same hyperparameter setting for the different experiments within a domain (Atari vs MuJoCo), apart from a multiplicative constant based on the range of the curiosity

Reference 53

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T23:01:05.470126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:01:05.012779Z digest=sha256:eef7c2932d307703926a14ace52837292d340c788b42b31216c52dc8cb693ce1

Observation b5b7e079-11c4-4f60-8379-ce67da842a0c · outbound

This paper cites Improve vs Expert.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Improve vs Expert

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.453937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:01:05.018130Z digest=sha256:47a5b966128ab84426d9406255be102dab860ff219d39a14314a581d69d523d7

Observation de71d438-3aa6-43c2-953c-46ad03772d60 · outbound

This paper cites Improve vs GIRIL.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Improve vs GIRIL

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.437406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:01:05.023291Z digest=sha256:6e04f9a18808947057b9d4bb89ecfafc973723a9073066509bb6cec91637fba8

Observation 95168769-48c7-4292-8631-a800bebb5d95 · outbound

This paper cites (2023), and (3) the analysis leads to a sublinear batch-regret, which is stronger than a sample complexity bound.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration (2023), and (3) the analysis leads to a sublinear batch-regret, which is stronger than a sample complexity bound

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.403114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:01:05.033529Z digest=sha256:8945c09eac6144732f9f37fdb7b52579282a5aeb39d54f4d6f80709dc7b4ee50

Observation b7b8e7e6-2f8c-446e-af65-50d250a2a372 · outbound

This paper cites Then with probability at least1−2δ, KX k=1 J(π ∗,erk)−J(π,er k) ≤Hlog|A|/η+ηH 3K/2 + 2HK· p 8H 2 log(H· NF (ϵF )/δ) + 4ϵF N+γ·O s H N · H 2 +γ γ dimN (Fh) log(1/δ) !.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Then with probability at least1−2δ, KX k=1 J(π ∗,erk)−J(π,er k) ≤Hlog|A|/η+ηH 3K/2 + 2HK· p 8H 2 log(H· NF (ϵF )/δ) + 4ϵF N+γ·O s H N · H 2 +γ γ dimN (Fh) log(1/δ) !

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.368525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:01:05.043470Z digest=sha256:c53c68b8595016ed211d0df6f6f16efd31bb1d0dbd48bd470906c9e7c0c8aee9

Observation c65daf9e-c0af-46ef-aa7d-b44bcf0bcbee · outbound

This paper cites Lemma D.2(Self-normalized bound for scalar-valued martingales).Consider random variables (vn|n∈N) adapted to the filtration (Hn :n= 0,1, ...).

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Lemma D.2(Self-normalized bound for scalar-valued martingales).Consider random variables (vn|n∈N) adapted to the filtration (Hn :n= 0,1, ...)

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.352347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:01:05.049084Z digest=sha256:5856cc3d5d87eac3503f1398a49f859a4c9488e294631c191b188ede30b66104

Observation 15ae85a1-92f2-4ddd-a953-0fecbb8f89a3 · outbound

This paper cites Table 16 shows the MLP architectures, i.e., GIRIL’s encoder and decoder and V AIL’s discriminator.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Table 16 shows the MLP architectures, i.e., GIRIL’s encoder and decoder and V AIL’s discriminator

Reference 100

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T23:01:05.419654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:01:05.028464Z digest=sha256:1130d4d5eeed2eb7bcafbd358bdbd71b0d59e41a46fb0c2ff224ef5c43ade3a4

Observation 3dca2e0c-e864-40b9-ab17-07e8e35015c4 · outbound

This paper cites Toward the fundamental limits of imitation learning.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Toward the fundamental limits of imitation learning

Reference 1988

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.844783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:01:04.877390Z digest=sha256:55b0991b4b053f9f1e4a7abbcc0424e02c94253bfc3fc066c402c7b05ddb930d

Observation ed4f5999-4c14-45f7-a129-809e339dc0d1 · outbound

This paper cites Relative entropy inverse reinforcement learning.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Relative entropy inverse reinforcement learning

Reference 2002

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:06.149046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:01:04.757294Z digest=sha256:6674573e5d2ac51cb746c87713681a43c8d40549ee28666edf106183dd4393da

Observation 5d5afa12-317f-4884-ad1f-fee39e977ae9 · outbound

This paper cites Learning structured output representation using deep conditional generative models.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Learning structured output representation using deep conditional generative models

Reference 2003

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.765510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:01:04.915792Z digest=sha256:56ffc7f35c3f9e4408d2a67757892a84c8d97d695974f63c3677584e688e3672

Observation b5aefac0-2d17-4260-a7a9-046d7ec9f654 · outbound

This paper cites and Bagnell, J.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration and Bagnell, J

Reference 2008

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.586670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:01:04.978061Z digest=sha256:50f589dcc196401370aef1d1aa92968d03bd6cb0d7e8b8db0cff367ec31c55f5

Observation 49dc83bb-418a-48f0-a380-a4f038286589 · outbound

This paper cites Provably efficient reinforcement learning with linear function approximation.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Provably efficient reinforcement learning with linear function approximation

Reference 2010

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.975912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:01:04.836637Z digest=sha256:0711d68352769dfccd61173241b83c5ae4d52d95cf7e20a1c6ee8c210dbe62c4

Observation d20c8eb2-84bc-45df-a5cb-92e43fa43df2 · outbound

This paper cites OpenAI Gym.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration OpenAI Gym

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:04.762311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:04.762311Z digest=sha256:ef2c578770391dfaffffe400841f6cba6d39a0e52a7e1230f9c3e4ced7922e4c

Observation ffb10c27-9f1e-4529-a1b7-1c24fc875d59 · outbound

This paper cites Proximal point imitation learning.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Proximal point imitation learning

Reference 2012

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.731530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:01:04.925993Z digest=sha256:8a639c376424d2bc6c643e6ca808f1013a9c74fa9becc09e58b38d44b2c0c882

Observation 35da6538-f151-47cf-8194-bc640d1ed319 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Proximal Policy Optimization Algorithms

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:04.893867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:04.893867Z digest=sha256:0a164c0a1fa9e0aa11da561ea445ee380120c80b9d4bfeb3b493e09b964228bb

Observation 9b272b49-a52a-4f0a-9472-499bc4ac3db2 · outbound

This paper cites Curiosity-driven exploration by self-supervised prediction.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Curiosity-driven exploration by self-supervised prediction

Reference 2014

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.895423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:01:04.864194Z digest=sha256:3f3c00cb22078b28ebda3572447bfdcf2af6f5b7ff0d0d9777895e11e7cd1283

Observation e4ddb88d-b381-4354-b3c4-c767fd83b6c0 · outbound

This paper cites Mujoco: A physics engine for model-based control.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Mujoco: A physics engine for model-based control

Reference 2015

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.748282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:01:04.921365Z digest=sha256:bdc48f7549274ab01e58dd2fd4ee51da1b51c7a4605a75c71214240964b27d19

Observation 22552bf1-f381-4fa2-b776-3e639ded67e3 · outbound

This paper cites Extrapolating beyond sub- optimal demonstrations via inverse reinforcement learning from observations.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Extrapolating beyond sub- optimal demonstrations via inverse reinforcement learning from observations

Reference 2016

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:06.133533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:01:04.768492Z digest=sha256:ca2979b5eb41abecf41fc496b25c189884b10d6f16e5b8efb46f03e3e68ea0c6

Observation 43d0ada8-eb40-4889-bc2d-e2d7d8a241d2 · outbound

This paper cites End-to- end driving via conditional imitation learning.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration End-to- end driving via conditional imitation learning

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:06.088173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:01:04.800160Z digest=sha256:d868614d43150b412fd77ee600e3812a426cb2811cdd0d0b992f315a37b7b82c

Observation 8fc4b90e-23d8-42d3-8a74-998af7883bec · outbound

This paper cites On the Global Convergence of Imitation Learning: A Case for Linear Quadratic Regulator.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration On the Global Convergence of Imitation Learning: A Case for Linear Quadratic Regulator

Reference 2018

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:01:05.300489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:01:04.778839Z digest=sha256:11bd47c10a7cfe50c8617ef77216e9f74d2d9c6f3c2cc29e01986789c2dbd899

Observation d5d1a19d-f8db-45e8-8aa7-21be4d6cd5cb · outbound

This paper cites Large-Scale Study of Curiosity-Driven Learning.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Large-Scale Study of Curiosity-Driven Learning

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:04.773534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:04.773534Z digest=sha256:fe3ef702f11cfda9616b87b9aa0aaaba1e7971003640737cfe92e3acc575e8d4

Observation e83f2d93-b30d-44a7-afce-4c234c4597cf · outbound

This paper cites Hybrid Inverse Reinforcement Learning.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Hybrid Inverse Reinforcement Learning

Reference 2020

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:01:05.226854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:01:04.882451Z digest=sha256:6b0c75a6606a0a80a5b5453f3b7621599a12be425d44d24fa15d54555cce53ce

Observation 82a4e234-d90e-4be9-9f52-028324fb64e7 · outbound

This paper cites On Computation and Generalization of Generative Adversarial Imitation Learning.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration On Computation and Generalization of Generative Adversarial Imitation Learning

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:06.117953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:01:04.789475Z digest=sha256:7125fb9ebeb919fe5960b51044519d37d9eeb3bb982efaef0901edf5de7b43de

Observation f1698227-9077-41f4-b4ab-59ce9bea3dab · outbound

This paper cites Imitation Learning from Imperfection: Theoretical Justifications and Algorithms.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Imitation Learning from Imperfection: Theoretical Justifications and Algorithms

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.943710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:01:04.845606Z digest=sha256:5687f9fa52da1aa683a61c123d517039ae9e16f2dad4bb6c5dc38d95c117bfd5

Observation 47b4d814-94b0-4d36-972e-c16de17b9058 · outbound

This paper cites IQ-Learn: Inverse soft-Q Learning for Imitation.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration IQ-Learn: Inverse soft-Q Learning for Imitation

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:06.025396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:01:04.818299Z digest=sha256:7a6b22bdc7f5baefe8a1a3cf80109cca58cfd132c615c798e94fdc5ec8926da0

Observation 8276749f-f8d1-469e-8c3e-e8d28e216c89 · outbound

This paper cites Exploration via Elliptical Episodic Bonuses.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Exploration via Elliptical Episodic Bonuses

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:04.822802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:04.822802Z digest=sha256:110b689c8e54c02d07438f3dccf38d0628b4632616bb847c4e691e56e57039bb

Observation 009e47ab-f081-4445-8d6d-f6bf83f30407 · outbound

This paper cites For a fair comparison, we used an identical policy network for all methods.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration For a fair comparison, we used an identical policy network for all methods

Reference 5184

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.488424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:01:05.007817Z digest=sha256:e6051ff12c28512992e3714bdc2a268997ef3a3e9513f48ba81b322a75b72047

Pith citing papers

No inbound Pith citation observations are available.