Pith. sign in

Paper Citation Record · LEDGER

Reinforcement Learning via Implicit Imitation Guidance

As of 9 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 8 inbound Pith citation observations for arXiv:2506.07505.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.07505 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:37:46.402936Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T04:27:43.858412Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T05:12:05.250025Z

Reference resolution

23 of 23 outbound references displayed

  • verified exact3
  • verified fuzzy4
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 54098eb7-f739-4b53-b9d7-90eaa0d274cc · outbound

This paper cites Efficient Online Reinforcement Learning with Offline Data.

Reinforcement Learning via Implicit Imitation Guidance Efficient Online Reinforcement Learning with Offline Data

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:46.315134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:46.315134Z digest=sha256:fd25f0699d5a989b40c914a5b6f540f10426334638dad7d35c991a8bf919037a

Observation c0c7ca42-7fe9-4e31-b0ba-49f239dff920 · outbound

This paper cites epochs per update.

Reinforcement Learning via Implicit Imitation Guidance epochs per update

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:37:46.651015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:37:46.399321Z digest=sha256:2de509b24a859b63529b691b42de5ed5ddab6f4cba5658b071add0b32248e395

Observation fe3a2452-6743-4417-a163-f222b9a72661 · outbound

This paper cites worse", which are successful demonstrations collected by inexperienced operators to incorporate additional diversity. Note that even though it is labeled.

Reinforcement Learning via Implicit Imitation Guidance worse", which are successful demonstrations collected by inexperienced operators to incorporate additional diversity. Note that even though it is labeled

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:37:46.640381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:37:46.402936Z digest=sha256:d1bde9382e5a02ff4e90a9385a8a83a6000fc17feac5ab06774ce1994f1d9a45

Observation 532c19c2-6b0f-4037-997e-8780c6c955fc · outbound

This paper cites Imitation Bootstrapped Reinforcement Learning.

Reinforcement Learning via Implicit Imitation Guidance Imitation Bootstrapped Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:46.332349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:46.332349Z digest=sha256:4ad4f1c9909d4e014ffe679dd2f988ffbcca72eb7311bda46f5c08386ead0d3a

Observation 06c7343a-0aec-4818-8248-c1f59cbebcfd · outbound

This paper cites Offline Reinforcement Learning with Implicit Q-Learning.

Reinforcement Learning via Implicit Imitation Guidance Offline Reinforcement Learning with Implicit Q-Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:46.336559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:46.336559Z digest=sha256:794d7a19f0ccc831669399759879c6dcbe69ad00654ecb998d061215fbff1b41

Observation b2329927-ac49-4b9c-b574-274bd0fd61ed · outbound

This paper cites Offline Retraining for Online RL: Decoupled Policy Learning to Mitigate Exploration Bias.

Reinforcement Learning via Implicit Imitation Guidance Offline Retraining for Online RL: Decoupled Policy Learning to Mitigate Exploration Bias

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:46.344696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:46.344696Z digest=sha256:cfdd87250949f4b1d1f7540027b78e652a7adb7c7593a084f56a056e93aee39f

Observation a13b6330-ea3a-4e6e-8793-87762ccd29da · outbound

This paper cites Over- coming exploration in reinforcement learning with demonstrations.

Reinforcement Learning via Implicit Imitation Guidance Over- coming exploration in reinforcement learning with demonstrations

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:37:46.673242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:37:46.349077Z digest=sha256:122619d4b7581ce15644e22f4920c1e4b2a12321f9aa892aab76033268dbc271

Observation 2d76a449-f55e-45b6-9a25-4fe50f74881b · outbound

This paper cites Computational Theories of Curiosity-Driven Learning.

Reinforcement Learning via Implicit Imitation Guidance Computational Theories of Curiosity-Driven Learning

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:37:46.541101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:37:46.356660Z digest=sha256:9bc693ca81f16377c431f0773ce67e9ad58e4d56f18933890b8a89ffc3134365

Observation aeaa12dd-e2cd-41c4-8703-41c978fb7710 · outbound

This paper cites Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations.

Reinforcement Learning via Implicit Imitation Guidance Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:46.368469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:46.368469Z digest=sha256:1c72dcb367f01732c3c16964e7268bb81ddfc93f9efb2acb097f9824b28de9b0

Observation 8a60e0dd-1666-43bd-8c84-0e5cc8d313af · outbound

This paper cites Schmidhuber.

Reinforcement Learning via Implicit Imitation Guidance Schmidhuber

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:37:46.661903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:37:46.376056Z digest=sha256:9e295e253332bc5d0fa353f34c292e24d0a477e660cd97b9066725c094e6f3e5

Observation 13036d68-8079-4468-a573-ac508a627241 · outbound

This paper cites Parrot: Data-Driven Behavioral Priors for Reinforcement Learning.

Reinforcement Learning via Implicit Imitation Guidance Parrot: Data-Driven Behavioral Priors for Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:46.379573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:46.379573Z digest=sha256:e6b844b22d18ba6cdc504e90eefae3a75acd0facd42a3306725a015b7b838797

Observation 0b408ba0-19b1-410c-bd43-e0e121d4bbbf · outbound

This paper cites Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient.

Reinforcement Learning via Implicit Imitation Guidance Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:46.383535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:46.383535Z digest=sha256:6a46e792c7a72972de9118c7033e3b26ea3b6a3cb49ba0d8128ed42eeb17c1b6

Observation 4474229d-0083-4b59-87ee-9d01368b8b76 · outbound

This paper cites Leveraging Demonstrations for Deep Reinforcement Learning on Robotics Problems with Sparse Rewards.

Reinforcement Learning via Implicit Imitation Guidance Leveraging Demonstrations for Deep Reinforcement Learning on Robotics Problems with Sparse Rewards

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:46.387296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:46.387296Z digest=sha256:1255c06d333ec5067ed1fb4d23ee0819d1909c126996bb8cd0490afe4acd3236

Observation 0066d2d0-370f-462f-b277-9688d2779557 · outbound

This paper cites Learning latent state representation for speeding up exploration.

Reinforcement Learning via Implicit Imitation Guidance Learning latent state representation for speeding up exploration

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:37:46.450428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:37:46.391312Z digest=sha256:4808451882a0da6135d2dfc25acce07a274f7738c4d84a738249e16d57cfe43e

Observation 5da8b341-7bea-4905-9769-d2ff68f8c0dd · outbound

This paper cites Policy Expansion for Bridging Offline-to-Online Reinforcement Learning.

Reinforcement Learning via Implicit Imitation Guidance Policy Expansion for Bridging Offline-to-Online Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:46.395221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:46.395221Z digest=sha256:fc78e3ebcf133634666eff034c167ae386a5e2f2081ee5e58da405b8c71e459d

Observation 14ab7d25-fb58-442a-b0b2-d46ed3aa6d56 · outbound

This paper cites Exploration by Random Network Distillation.

Reinforcement Learning via Implicit Imitation Guidance Exploration by Random Network Distillation

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:46.319807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:46.319807Z digest=sha256:138b6762051af54ea3758e46d46ee4458fe1a164ecb7455f7ca30f0ac6eaa5a7

Observation fd8c01a0-028f-4001-ae53-4883703f0cc9 · outbound

This paper cites Learning by Playing - Solving Sparse Reward Tasks from Scratch.

Reinforcement Learning via Implicit Imitation Guidance Learning by Playing - Solving Sparse Reward Tasks from Scratch

Reference 2017

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:37:46.496733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:37:46.372311Z digest=sha256:bf76823dac8375e8ab76ededc47948923e1d04b3ba641fcac3a57829e95836df

Observation 4197caaf-9d9d-41ea-b2dc-43b242038ebc · outbound

This paper cites AWAC: Accelerating Online Reinforcement Learning with Offline Datasets.

Reinforcement Learning via Implicit Imitation Guidance AWAC: Accelerating Online Reinforcement Learning with Offline Datasets

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:46.352606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:46.352606Z digest=sha256:7b7ceee48e315f546860e09155e1da0488b9de329ee4ba9ea26b49a51523dd67

Observation 084fe3ea-9951-45c1-9f97-44b59aa7afea · outbound

This paper cites Self-Supervised Exploration via Disagreement.

Reinforcement Learning via Implicit Imitation Guidance Self-Supervised Exploration via Disagreement

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:46.364579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:46.364579Z digest=sha256:538c8344a3f92ccad168eab1e1a556f1ba911f2a37ee9d8397791dc9db510a53

Observation 18e26509-53b6-4ac0-a579-ec7add5f2be0 · outbound

This paper cites Go-Explore: a New Approach for Hard-Exploration Problems.

Reinforcement Learning via Implicit Imitation Guidance Go-Explore: a New Approach for Hard-Exploration Problems

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:46.323819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:46.323819Z digest=sha256:be85b1f87d4723eb2b0c4c50529b275989ca577168859c1a35b38558decb69a9

Observation 59835236-da50-444b-87cb-09f63dc28f21 · outbound

This paper cites Efficient Exploration via State Marginal Matching.

Reinforcement Learning via Implicit Imitation Guidance Efficient Exploration via State Marginal Matching

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:46.340669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:46.340669Z digest=sha256:275cf7ac230ee2794018b37d23aaa6b4e39965b777c245189b7d13a05facd8f5

Observation 88953db2-4e1a-4aa0-8f4b-bd80c719fd01 · outbound

This paper cites Making Efficient Use of Demonstrations to Solve Hard Exploration Problems.

Reinforcement Learning via Implicit Imitation Guidance Making Efficient Use of Demonstrations to Solve Hard Exploration Problems

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:46.360602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:46.360602Z digest=sha256:97ab647d7ece9b969823af7061eb8da6843769e251f2d3c6c73b2e8f61088f65

Observation 88e13b0c-c32f-4da1-9331-d149b31a9721 · outbound

This paper cites MoDem: Accelerating Visual Model-Based Reinforcement Learning with Demonstrations.

Reinforcement Learning via Implicit Imitation Guidance MoDem: Accelerating Visual Model-Based Reinforcement Learning with Demonstrations

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:46.328115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:46.328115Z digest=sha256:092532a3f00a360fe083c736c9e9f0953ee8d92bcfec28a42af1ce051a159aa6

Pith citing papers

Observation f7f7acb0-042e-4fab-b930-ab01711ff06d · inbound

EXPO: Stable Reinforcement Learning with Expressive Policies cites this paper.

EXPO: Stable Reinforcement Learning with Expressive Policies Reinforcement Learning via Implicit Imitation Guidance

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:12:05.252079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T05:09:02.111308Z digest=sha256:1f4bd53147e37d377adeefcab66d3de5e1cb59722fb8f319e10ab5e0f5b82964

Observation 912f030f-4650-4d47-b137-548db7d98738 · inbound

Value Flows cites this paper.

Value Flows Reinforcement Learning via Implicit Imitation Guidance

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T11:01:27.929388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:01:27.929388Z digest=sha256:18e301950be1888c97e6b1d9b69a6667408453040bb422c38ac4618919e896b9

Observation 73bb73d3-8468-4d6c-b222-9bd2e736b647 · inbound

Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving cites this paper.

Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving Reinforcement Learning via Implicit Imitation Guidance

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T11:59:59.505114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T11:56:40.836234Z digest=sha256:2e352825ba1bcd6daa5aee3738380c7e30fbcca672aee49b812948753e80e570

Observation 740ec016-c10b-4d8c-b869-42282c0b9244 · inbound

Incremental Residual Reinforcement Learning Toward Real-World Learning for Social Navigation cites this paper.

Incremental Residual Reinforcement Learning Toward Real-World Learning for Social Navigation Reinforcement Learning via Implicit Imitation Guidance

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:05:57.992073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T17:48:36.372298Z digest=sha256:05ad2f8ef9ca8af9d08a4468ffa82a7cd461c9f2cebb1e3198caae60c2f4a9bf

Observation 510bd843-417c-4782-93d9-602cb0adf3fb · inbound

Incremental Residual Reinforcement Learning Toward Real-World Learning for Social Navigation cites this paper.

Incremental Residual Reinforcement Learning Toward Real-World Learning for Social Navigation Reinforcement Learning via Implicit Imitation Guidance

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-13T00:08:53.253285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:08:53.253285Z digest=sha256:76ab51f129eff02080f9c3825640e4d91f1cf1dfd22646163bfe7699ecde9aee

Observation 2eee268b-bfe8-4404-b5b0-805b788a900a · inbound

FASTER: Value-Guided Sampling for Fast RL cites this paper.

FASTER: Value-Guided Sampling for Fast RL Reinforcement Learning via Implicit Imitation Guidance

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-10T02:48:27.234174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T02:47:36.475845Z digest=sha256:0b7d90f28f87c217422f6733b5fcc4bbf0e1e97f04c527ac4857d13aaa7eac47

Observation 6dc67857-44d7-4bbf-85f2-b9e352933bc6 · inbound

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? cites this paper.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Reinforcement Learning via Implicit Imitation Guidance

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-30T11:06:22.032376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T11:06:22.032376Z digest=sha256:b17507c1c256cbf90ca8bfa7b3a270c01ed8cf611b379768160a5cdf1d89b4b1

Observation 49b28c9e-8205-4ea0-b167-2c60e716a353 · inbound

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? cites this paper.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Reinforcement Learning via Implicit Imitation Guidance

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.858412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.858412Z digest=sha256:d9a51da8739631e96fad8b7eba2f06c3f47fffd289ab7d151f0694d9d17c9474