Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T07:56:36.541257Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 0 inbound Pith citation observations for arXiv:2607.21302.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T07:56:36.541257Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
65 of 65 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 341f92dd-3769-4d31-8a59-b7bf920ffaab · outbound
Expert Behavior Prior Reinforcement Learning Diverse imitation learning via self-organizing generative models,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6a8681b-336e-43ca-80ee-7433cefb0314 · outbound
Expert Behavior Prior Reinforcement Learning Markov balance satisfaction improves performance in strictly batch offline imitation learning,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbcf8448-4dcf-4eda-8814-79ce693b9ed4 · outbound
Expert Behavior Prior Reinforcement Learning X-IL: Exploring the Design Space of Imitation Learning Policies
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0c9f0f7-1bf0-4375-86bb-5fa0567a7e39 · outbound
Expert Behavior Prior Reinforcement Learning Augmenting decision with hypothesis in reinforcement learning,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83e7b3f0-9c47-41bc-8384-a76252bf6064 · outbound
Expert Behavior Prior Reinforcement Learning Why so pessimistic? estimating uncertainties for offline rl through ensembles, and why their independence matters,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cee8cfde-a681-4a64-b018-10afd8a8aa72 · outbound
Expert Behavior Prior Reinforcement Learning Epistemic bellman operators,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb72262a-016a-474f-a607-9083bebed730 · outbound
Expert Behavior Prior Reinforcement Learning Behavior priors for efficient reinforcement learning,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b766335-3a59-4364-ac77-73b6d6037151 · outbound
Expert Behavior Prior Reinforcement Learning Pre-training goal-based models for sample-efficient reinforcement learning,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f28a8c9-f018-48b8-ab9b-19bfaa08277a · outbound
Expert Behavior Prior Reinforcement Learning BLEND: Behavior-guided Neural Population Dynamics Modeling via Privileged Knowledge Distillation
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b915a079-3c54-4e29-9cef-2a701c37ffc7 · outbound
Expert Behavior Prior Reinforcement Learning Jump-start reinforcement learning,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6081f832-9100-423c-bd19-aca97dfae6e3 · outbound
Expert Behavior Prior Reinforcement Learning Scaling proprioceptive-visual learning with heterogeneous pre-trained transformers,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 346620d3-e6f3-4c03-b532-fb9a0a79556e · outbound
Expert Behavior Prior Reinforcement Learning Learn to supervise: Deep reinforcement learning-based prototype refinement for few-shot motor fault diagnosis,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94b59805-fd35-4a10-bf2b-3d3929510011 · outbound
Expert Behavior Prior Reinforcement Learning Efficient online reinforcement learning with offline data,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6d2af2e-c156-48bf-8659-1371ea1dd623 · outbound
Expert Behavior Prior Reinforcement Learning Leveraging offline data in online reinforcement learning,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afb0555d-6cf0-413f-8b5c-7a1f828849d6 · outbound
Expert Behavior Prior Reinforcement Learning Enhancing Reinforcement Learning Agents with Local Guides
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8506d4f5-30cb-49ef-9af4-cfcda1e306c8 · outbound
Expert Behavior Prior Reinforcement Learning Leveraging demonstrations to improve online learning: Quality matters,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32933b84-abbe-4abc-a2d2-22c5d7e74678 · outbound
Expert Behavior Prior Reinforcement Learning Iterative regularized policy optimization with imperfect demonstrations,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91d4652f-f4bf-4688-9b00-7c6f2ee3783f · outbound
Expert Behavior Prior Reinforcement Learning Constraint- adaptive policy switching for offline safe reinforcement learning,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e5e2604-fdd2-4c8f-90cb-bba73b3c3757 · outbound
Expert Behavior Prior Reinforcement Learning Residual skill policies: Learning an adaptable skill-based action space for rein- forcement learning for robotics,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8653a43-5ca3-403c-9606-da4b949fe954 · outbound
Expert Behavior Prior Reinforcement Learning Leveraging Skills from Unlabeled Prior Data for Efficient Online Exploration
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e6a71d8-ecec-40a5-83f5-548cde8187e0 · outbound
Expert Behavior Prior Reinforcement Learning Policy regularization with dataset constraint for offline reinforcement learning,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81e955d9-8414-4ee5-a57c-db274c0f9a2d · outbound
Expert Behavior Prior Reinforcement Learning Accelerating exploration with unlabeled prior data,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d2a733e-aa25-4c28-b5b3-f7949b75cba8 · outbound
Expert Behavior Prior Reinforcement Learning Cross-domain offline policy adaptation with optimal transport and dataset constraint,
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6dc0356f-c43f-41cb-87fc-0962c26c76f9 · outbound
Expert Behavior Prior Reinforcement Learning Policy gradient for rectangular robust markov decision processes,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ff1a6f6-c2e5-4349-98f8-215a34f9e7fb · outbound
Expert Behavior Prior Reinforcement Learning Reinforcement learning: An introduction,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b62b83be-8784-46e8-a09a-9f9566b69423 · outbound
Expert Behavior Prior Reinforcement Learning Is q-learning provably efficient?
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea73e4e2-ee0a-4120-ac1e-bbf5940092a4 · outbound
Expert Behavior Prior Reinforcement Learning Actor-critic alignment for offline-to-online re- inforcement learning,
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5feb362e-7a47-4d7a-991e-4671991e57bf · outbound
Expert Behavior Prior Reinforcement Learning Adaptive policy learning for offline-to-online reinforcement learning,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21f7a3da-bb6e-4f61-94fc-e2a21133bdbc · outbound
Expert Behavior Prior Reinforcement Learning Optimistic critic reconstruction and constrained fine-tuning for general offline-to-online rl,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23ab0009-a461-439e-b9bf-0698c090ced3 · outbound
Expert Behavior Prior Reinforcement Learning Tree-based batch mode rein- forcement learning,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fa71875-f4e4-4297-a91c-5c32be4a3538 · outbound
Expert Behavior Prior Reinforcement Learning Mildly conservative q-learning for offline reinforcement learning,
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1447621-db8d-4328-ba4e-95b364d4eae0 · outbound
Expert Behavior Prior Reinforcement Learning Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7f0e50a-4191-49fe-be71-a2cb55e7e547 · outbound
Expert Behavior Prior Reinforcement Learning De-pessimism offline reinforcement learning via value compensation,
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09c13e59-946a-4bba-947d-eaffdbddc620 · outbound
Expert Behavior Prior Reinforcement Learning Cal-ql: Calibrated offline rl pre-training for efficient online fine-tuning,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 822d88c3-11b9-4f6d-954b-86655039d45c · outbound
Expert Behavior Prior Reinforcement Learning AWAC: Accelerating Online Reinforcement Learning with Offline Datasets
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f20a930-d0d0-4395-8ffa-a72e6d36787b · outbound
Expert Behavior Prior Reinforcement Learning Don’t start from scratch: Leveraging prior data to automate robotic reinforcement learning,
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5072f68-74b8-47aa-8502-26451274812d · outbound
Expert Behavior Prior Reinforcement Learning Behavior prior representation learning for offline reinforce- ment learning,
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03b83487-1f98-4c71-9d77-7bbfcd9df768 · outbound
Expert Behavior Prior Reinforcement Learning One ACT Play: Single Demonstration Behavior Cloning with Action Chunking Transformers
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 163ac0ef-2bef-4e25-9473-990c566512c1 · outbound
Expert Behavior Prior Reinforcement Learning Policy optimization with demonstrations,
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dba72cb6-62df-4dd0-96f0-98af582b032c · outbound
Expert Behavior Prior Reinforcement Learning Reinforcement Learning with Sparse Rewards using Guidance from Offline Demonstration
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5893c168-8d3b-444e-9354-fbbfb894a289 · outbound
Expert Behavior Prior Reinforcement Learning Goal-conditioned on-policy reinforcement learning,
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fa04518-1d3f-411f-92cd-b92d001e8d69 · outbound
Expert Behavior Prior Reinforcement Learning Recurrent experience replay in distributed reinforcement learning,
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdd59388-1c54-4fdd-aff1-630c621322fb · outbound
Expert Behavior Prior Reinforcement Learning Learning Sparse Control Tasks from Pixels by Latent Nearest-Neighbor-Guided Explorations
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a641db0-4e66-40b2-b67b-68049ba7f9b3 · outbound
Expert Behavior Prior Reinforcement Learning Theoretically principled deep rl acceleration via nearest neighbor function approximation,
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62b46839-b198-4750-b1c5-8e89c68393cc · outbound
Expert Behavior Prior Reinforcement Learning A review of recurrent neural net- works: Lstm cells and network architectures,
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 504f9b9c-e716-4a8d-a92f-ab32a0141259 · outbound
Expert Behavior Prior Reinforcement Learning Frustratingly easy regularization on representation can boost deep reinforcement learning,
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f4f9eb3-2fd6-44e7-8dd6-80ec3b3fde8e · outbound
Expert Behavior Prior Reinforcement Learning Q-learning with nearest neighbors,
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44512d60-f245-4f66-b622-b0dcfd475e60 · outbound
Expert Behavior Prior Reinforcement Learning Improving policy exploitation in online reinforcement learning with instant retrospect action,
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa4d02f3-c43d-474d-a816-8729a6e9bd6c · outbound
Expert Behavior Prior Reinforcement Learning Seizing serendipity: exploiting the value of past success in off-policy actor-critic,
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8476f44-e366-4a04-8c2b-2f01e5b160b9 · outbound
Expert Behavior Prior Reinforcement Learning Offline-boosted actor-critic: Adaptively blending optimal historical behaviors in deep off-policy rl,
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ea73369-3cd0-4d66-9a5c-adc939e45cd0 · outbound
Expert Behavior Prior Reinforcement Learning Off-policy deep reinforcement learning without exploration,
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3fa0b96-1174-4cda-a03e-5c0168346b18 · outbound
Expert Behavior Prior Reinforcement Learning Auto-Encoding Variational Bayes
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b910d697-74b6-4a3b-b22e-3b59658c6b1d · outbound
Expert Behavior Prior Reinforcement Learning Reparameterized policy learning for multimodal trajectory optimization,
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce2529a7-64a0-4bbf-9ea1-37190d9bb340 · outbound
Expert Behavior Prior Reinforcement Learning DIDI: Diffusion-Guided Diversity for Offline Behavioral Generation
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee6eb6eb-f9d9-4fcd-8a93-026744b08b85 · outbound
Expert Behavior Prior Reinforcement Learning Addressing function approxima- tion error in actor-critic methods,
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a77a6591-4287-4227-93d6-34a572ae9783 · outbound
Expert Behavior Prior Reinforcement Learning Unresolved cited work
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80fa20c8-a962-4b8a-a3e5-a9eca32fc7bc · outbound
Expert Behavior Prior Reinforcement Learning Continuous control with deep reinforcement learning
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9301c955-d440-4711-82d7-a5e3f9ba95c9 · outbound
Expert Behavior Prior Reinforcement Learning Soft actor-critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor,
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c158b7c2-32ec-4edf-9475-07c3e811f996 · outbound
Expert Behavior Prior Reinforcement Learning Unresolved cited work
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7e0ee36-8abc-45c2-a41f-3faedef10911 · outbound
Expert Behavior Prior Reinforcement Learning Reinforcement learning with stochastic reward machines,
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef020d41-844a-4952-b14b-0a9e874f033b · outbound
Expert Behavior Prior Reinforcement Learning Distributional soft actor-critic: Off-policy reinforcement learning for addressing value estimation errors
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1fd98d0-4f57-49cb-b69a-1f3c4f1e8399 · outbound
Expert Behavior Prior Reinforcement Learning Softmax deep double deterministic policy gradients,
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b60f18e-79ec-4e30-9aa6-98c606b6ff08 · outbound
Expert Behavior Prior Reinforcement Learning Logit standardization in knowledge distillation,
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa6be7b5-bd2b-4c1e-8ea0-157f77c92cb7 · outbound
Expert Behavior Prior Reinforcement Learning Distilling the Knowledge in a Neural Network
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88022178-65f0-4374-bf61-c4e6ad1621bd · outbound
Expert Behavior Prior Reinforcement Learning Deep reinforcement learning at the edge of the statistical precipice,
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.