Pith. sign in

Paper Citation Record · LEDGER

Relay Policy Learning: Solving Long-Horizon Tasks via Imitation and Reinforcement Learning

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 30 inbound Pith citation observations for arXiv:1910.11956.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1910.11956 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 30 of 30 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:26:30.940459Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:39:51.078263Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b0592225-ea90-47e1-bf98-8704da4577f9 · inbound

D4RL: Datasets for Deep Data-Driven Reinforcement Learning cites this paper.

D4RL: Datasets for Deep Data-Driven Reinforcement Learning Relay Policy Learning: Solving Long-Horizon Tasks via Imitation and Reinforcement Learning

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:19:17.414978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T23:19:17.322890Z digest=sha256:120dc7be7036c18c94905d62cda0a1d62b3e26e2940c63c617ad57cdd3e131d6

Observation 198e1864-ad6b-4f8a-b3c7-0b888f799c45 · inbound

VIP: Towards Universal Visual Reward and Representation via Value-Implicit Pre-Training cites this paper.

VIP: Towards Universal Visual Reward and Representation via Value-Implicit Pre-Training Relay Policy Learning: Solving Long-Horizon Tasks via Imitation and Reinforcement Learning

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-15T04:42:52.711721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T04:42:52.627166Z digest=sha256:04d87baa5699b36d314317f5783a16f19e51cfb569c5907002728ec21fedc754

Observation 4defba8b-2764-4aeb-aa63-2183159555b1 · inbound

Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation cites this paper.

Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation Relay Policy Learning: Solving Long-Horizon Tasks via Imitation and Reinforcement Learning

Reference 120

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T16:32:05.604253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-13T16:32:05.538507Z digest=sha256:038fbdc90a9eb0baa6c313498ac0920a96dc904f7ff3fc93915f2bbeb1e248d2

Observation de2072f5-0ece-47fb-96a7-13ab1a87fafa · inbound

Diffusion Policy Policy Optimization cites this paper.

Diffusion Policy Policy Optimization Relay Policy Learning: Solving Long-Horizon Tasks via Imitation and Reinforcement Learning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:48:14.875918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:b2de18847dbe7a26f6baa23990b03cb3b4e258d45b67828534792bb6cfbe123e

Observation 5a0b22ca-5241-4357-9b91-ab026ad64f38 · inbound

Direct Preference Optimization for Primitive-Enabled Hierarchical RL: A Bilevel Approach cites this paper.

Direct Preference Optimization for Primitive-Enabled Hierarchical RL: A Bilevel Approach Relay Policy Learning: Solving Long-Horizon Tasks via Imitation and Reinforcement Learning

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-23T18:33:19.665531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T18:30:38.818711Z digest=sha256:67eb4569de296531a30f212e42e6e9539def7026480c08b45cf4e27bd16031e8

Observation f361a955-fa05-4f92-93bc-d4217624d2d4 · inbound

Block-wise Adaptive Caching for Accelerating Diffusion Policy cites this paper.

Block-wise Adaptive Caching for Accelerating Diffusion Policy Relay Policy Learning: Solving Long-Horizon Tasks via Imitation and Reinforcement Learning

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:37:14.015757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T09:36:09.790248Z digest=sha256:d1618b8320210bed7d7abcc5769545372a0a51f112979410ce16ce0c1ba16a8e

Observation df18f18a-39af-47af-a11c-a273e36f4d61 · inbound

Steering Robots with Inference-Time Interactions cites this paper.

Steering Robots with Inference-Time Interactions Relay Policy Learning: Solving Long-Horizon Tasks via Imitation and Reinforcement Learning

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:30.940459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:30.940459Z digest=sha256:7d28ae316a5e0a6e3a083e1046682e7db5e9bb616b2e5c75cc7066fa86c77570

Observation 85db7f8c-2df1-49c7-836d-3b681e7bcffb · inbound

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers cites this paper.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers Relay Policy Learning: Solving Long-Horizon Tasks via Imitation and Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:23.545915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:23.545915Z digest=sha256:6a185e5d0498ad3f7d3c4bce7a736b53db3886196b94d1fb06de4ccd5857e64a

Observation 70a90a27-ee19-4c55-90ab-040b1ee96e71 · inbound

FOUNDER: Grounding Foundation Models in World Models for Open-Ended Embodied Decision Making cites this paper.

FOUNDER: Grounding Foundation Models in World Models for Open-Ended Embodied Decision Making Relay Policy Learning: Solving Long-Horizon Tasks via Imitation and Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T17:12:27.957201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:12:27.957201Z digest=sha256:b90f7f0e36fdcb9f23a8d8f8d4340b87b0e50991548a428df20357f2d4fd830d

Observation 59e65321-a80a-46f8-8a06-056ab68d8390 · inbound

Learning Temporal Abstractions via Variational Homomorphisms in Option-Induced Abstract MDPs cites this paper.

Learning Temporal Abstractions via Variational Homomorphisms in Option-Induced Abstract MDPs Relay Policy Learning: Solving Long-Horizon Tasks via Imitation and Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T15:17:34.911775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:17:34.911775Z digest=sha256:e833ea86c951e17a0611ece73f2126da2fb7766a99f98a05412019bef80e0e9c

Observation e1c9652e-a846-42bb-b6cc-536c9be01cd3 · inbound

D3P: Dynamic Denoising Diffusion Policy via Reinforcement Learning cites this paper.

D3P: Dynamic Denoising Diffusion Policy via Reinforcement Learning Relay Policy Learning: Solving Long-Horizon Tasks via Imitation and Reinforcement Learning

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-05T22:37:38.862346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:37:38.862346Z digest=sha256:991837d01cfa50ceb5df7ebb92cbaba91db58008f1375fd13d60150c14306e3e

Observation 17b2e3e5-a7c8-4e3a-ae6b-d8f26c2d3b74 · inbound

Learning Upper Lower Value Envelopes to Shape Online RL: A Principled Approach cites this paper.

Learning Upper Lower Value Envelopes to Shape Online RL: A Principled Approach Relay Policy Learning: Solving Long-Horizon Tasks via Imitation and Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T08:45:48.991607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:45:48.991607Z digest=sha256:e9f90e7b8854f4aba379f33e144f6be7056c05693dace9405be5b323e2a26c81

Observation e5a808f4-74b4-44be-b620-5437f98b78f3 · inbound

DASIP: Dynamic Test-Time Compute Scaling for Robot Control with Stochastic Interpolant Policies cites this paper.

DASIP: Dynamic Test-Time Compute Scaling for Robot Control with Stochastic Interpolant Policies Relay Policy Learning: Solving Long-Horizon Tasks via Imitation and Reinforcement Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T20:13:27.607487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:13:27.607487Z digest=sha256:694ab47169f6d0170930ccecf986ff535701278847fe6f572bdda3589ca7d1a8

Observation 5bea7f2a-15cc-46e8-a837-adb104d549b1 · inbound

Sparse ActionGen: Accelerating Diffusion Policy with Real-time Pruning cites this paper.

Sparse ActionGen: Accelerating Diffusion Policy with Real-time Pruning Relay Policy Learning: Solving Long-Horizon Tasks via Imitation and Reinforcement Learning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-21T16:34:16.328023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T16:33:41.439450Z digest=sha256:9c2d71f9858f27710754b0723fce83c023c59e8c6661560c8c81a542908e39c5

Observation 3c387d72-f425-4ff5-bcf3-e6344341bcfd · inbound

Efficient Multi-Objective Planning with Weighted Maximization Using Large Neighbourhood Search cites this paper.

Efficient Multi-Objective Planning with Weighted Maximization Using Large Neighbourhood Search Relay Policy Learning: Solving Long-Horizon Tasks via Imitation and Reinforcement Learning

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:20:50.996667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T19:13:45.010530Z digest=sha256:1f074897ce8b79e8fce77045311485532a614a9dcb8984e1a7c5a8aca828e681

Observation b7d900f5-c284-46da-a984-bf67f9623a63 · inbound

Learning 3D Representations for Spatial Intelligence from Unposed Multi-View Images cites this paper.

Learning 3D Representations for Spatial Intelligence from Unposed Multi-View Images Relay Policy Learning: Solving Long-Horizon Tasks via Imitation and Reinforcement Learning

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:56:03.625672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T15:44:39.322417Z digest=sha256:1511880a7ae172825ed6c9dd6402044fb27089004a88bc648b4ca5bf982bc6b1

Observation 0950a9e9-687e-47f0-a34d-48ea79c64616 · inbound

ScoRe-Flow: Complete Distributional Control via Score-Based Reinforcement Learning for Flow Matching cites this paper.

ScoRe-Flow: Complete Distributional Control via Score-Based Reinforcement Learning for Flow Matching Relay Policy Learning: Solving Long-Horizon Tasks via Imitation and Reinforcement Learning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:26:00.645341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T16:02:36.261596Z digest=sha256:7e93372463321625ab51b4abf4cb696c29ea0a6596af8c3080f16cb0a789a85d

Observation 915a78e5-f88f-4878-a8fc-5cca01a41ca9 · inbound

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies cites this paper.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Relay Policy Learning: Solving Long-Horizon Tasks via Imitation and Reinforcement Learning

Reference 147

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:10:42.470603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-08T18:48:56.075160Z digest=sha256:4a650904e38d3e178623a8c0d66948f5fb6a71e04e09dd0defab540f0ca162e3

Observation 242da48d-3abb-4790-885d-79a84a51b9fd · inbound

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies cites this paper.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Relay Policy Learning: Solving Long-Horizon Tasks via Imitation and Reinforcement Learning

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-01T00:05:09.684233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:7357dd1ee44efc032c711de3375e3d7b7fdce9bdce8e8d6a95aeb8f9eaf661e2

Observation 88d10d64-f776-49cd-babb-95b915500970 · inbound

Learning Reactive Dexterous Grasping via Hierarchical Task-Space RL Planning and Joint-Space QP Control cites this paper.

Learning Reactive Dexterous Grasping via Hierarchical Task-Space RL Planning and Joint-Space QP Control Relay Policy Learning: Solving Long-Horizon Tasks via Imitation and Reinforcement Learning

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:51:29.433062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-07T15:57:58.686493Z digest=sha256:9f7dc735553941ae30ec8d365a4cc45c38d2c72c5d07e64463995adf9f74ab75

Observation 6e350794-77b1-49ad-9769-58a7a6d14469 · inbound

World Action Models: The Next Frontier in Embodied AI cites this paper.

World Action Models: The Next Frontier in Embodied AI Relay Policy Learning: Solving Long-Horizon Tasks via Imitation and Reinforcement Learning

Reference 224

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:18.000999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T05:01:16.802019Z digest=sha256:4b641ccc24c9aa07d064c7bfb5f025efc458ffad5107ffaa3bf5757ef77f7120

Observation 1dfa8652-7124-4f99-a039-7d29d0201766 · inbound

DiLA: Disentangled Latent Action World Models cites this paper.

DiLA: Disentangled Latent Action World Models Relay Policy Learning: Solving Long-Horizon Tasks via Imitation and Reinforcement Learning

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:38:56.392763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T19:35:37.527479Z digest=sha256:a08fb2f7af8db499a94f8a502109882a54aaa776f3bf3a698d35f32cfa900f5e

Observation a0bcfbde-a451-4380-94f0-5510b9f0adff · inbound

Drift Flow Matching cites this paper.

Drift Flow Matching Relay Policy Learning: Solving Long-Horizon Tasks via Imitation and Reinforcement Learning

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:33:21.456573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T14:31:05.979325Z digest=sha256:a0f1c846228fb2ae432c195d00a270aca6a58bdb1e7ae3c50ca6ed9ad3209c93

Observation f0d5cfea-ba06-45d2-831a-0b89a5816f8f · inbound

Understanding Multimodal Failure in Action-Chunking Behavioral Cloning cites this paper.

Understanding Multimodal Failure in Action-Chunking Behavioral Cloning Relay Policy Learning: Solving Long-Horizon Tasks via Imitation and Reinforcement Learning

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:14:45.504390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T08:14:28.115848Z digest=sha256:0f646385f9d94668e6e1cb912e4d77ea53e25fe6a9403fef91df21074a0ceec3

Observation 619cd848-645e-45ef-9a86-fcc2a7408db6 · inbound

Automating Potential-based Reward Shaping with Vision Language Model Guidance cites this paper.

Automating Potential-based Reward Shaping with Vision Language Model Guidance Relay Policy Learning: Solving Long-Horizon Tasks via Imitation and Reinforcement Learning

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:39:51.080320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T05:00:24.832807Z digest=sha256:81cd41da843dce535160cc4ca0a73ed0b5f0fb2d03a79b39dc3872dc57ec2141

Observation 5d93897c-4d3d-4532-b647-8fb3651a43ca · inbound

Learning Task-Sufficient World Models by Synergizing Agentic Exploration and Structured Modeling cites this paper.

Learning Task-Sufficient World Models by Synergizing Agentic Exploration and Structured Modeling Relay Policy Learning: Solving Long-Horizon Tasks via Imitation and Reinforcement Learning

Reference 135

Resolution
unresolved
no resolver link, observed 2026-07-11T19:24:48.899301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T19:24:48.899301Z digest=sha256:25c010bba6595332d331b7f5ae621415f1f1fddf6447aa09924be9fa501bd7e5

Observation fd3fb91f-ada2-4ef3-9ee7-e3907b692c60 · inbound

Dual-Process Atomic Skill Learning: Decoupling Semantic Reasoning and Real-Time Control cites this paper.

Dual-Process Atomic Skill Learning: Decoupling Semantic Reasoning and Real-Time Control Relay Policy Learning: Solving Long-Horizon Tasks via Imitation and Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-14T10:23:25.158961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:23:25.158961Z digest=sha256:e9ac9f188572b72e559a89ef3d054fe17110e09a2808456adf81bc3e45b35ced

Observation 42ade7d4-095b-4500-ba52-3515edd58d5f · inbound

Relative Value Learning cites this paper.

Relative Value Learning Relay Policy Learning: Solving Long-Horizon Tasks via Imitation and Reinforcement Learning

Reference 115

Resolution
unresolved
no resolver link, observed 2026-08-01T08:32:00.377284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T08:32:00.377284Z digest=sha256:b0668031bce019e1651eff4879676ff6f7efe33f5ce00b3e6f9a374a0dcb941b

Observation 86aeffef-3846-4676-acbd-1de15e085da4 · inbound

Data Pyramid for Embodied Manipulation cites this paper.

Data Pyramid for Embodied Manipulation Relay Policy Learning: Solving Long-Horizon Tasks via Imitation and Reinforcement Learning

Reference 132

Resolution
unresolved
no resolver link, observed 2026-07-31T06:18:55.529553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:18:55.529553Z digest=sha256:6d7dac85c96adc548d64c18b9ab8f1e2f5cdc1ca59a6f8dece8c591dbd7077b6

Observation a92d4533-c64d-43b5-9267-6193ab9d44c5 · inbound

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills cites this paper.

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills Relay Policy Learning: Solving Long-Horizon Tasks via Imitation and Reinforcement Learning

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-04T19:45:31.346266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:45:31.346266Z digest=sha256:508a68edff5d9a01f0e98a7f8a22139b9a62240cd7c34507cedb0328f0b30984