Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-26T14:38:43.803126Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 92 of 92 outbound references and 0 inbound Pith citation observations for arXiv:2606.21136.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-26T14:38:43.803126Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
92 of 92 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4e1e084f-d42c-4d15-8c27-efee6ae807ed · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Learning to predict by the methods of temporal differences.Machine learning, 3(1):9–44, 1988
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff74459d-d860-40c7-9619-804708ea7e88 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Q-learning.Machine learning, 8(3):279–292, 1992
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba73f511-03e2-4d80-8765-caf891bd3712 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Human-level control through deep reinforcement learning.nature, 518(7540):529–533, 2015
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e382c61-6075-45ae-aa07-420ead382bff · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Mastering the game of go with deep neural networks and tree search.nature, 529 (7587):484–489, 2016
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f55a9e2f-b1f2-42c0-81dd-473285e312d4 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Mastering the game of go without human knowledge.nature, 550(7676):354–359, 2017
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be62d276-2bb8-4f7c-af7e-585c5ab08cb3 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Reinforcement learning: An introduction
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae82ef4b-7692-4cc2-934c-ce59fde0d832 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching A survey of temporal credit assignment in deep reinforcement learning.Transac- tions on Machine Learning Research, 2024
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b00a4a8-39fe-4285-9d8c-1f59e5150319 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Optimizing agent behavior over long time scales by transporting value.Nature communications, 10(1):5223, 2019
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07f7f2ef-5880-4bae-bcdd-52adab8a9046 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Is value learning really the main bottleneck in offline rl? InThe Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7801efc-da11-4091-95a0-75c045ca8e20 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Convergence of stochastic iterative dynamic programming algorithms.Advances in neural information processing systems, 6, 1993
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6dc7829-2c1c-4320-8ace-02411a00062b · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Analysis of temporal-diffference learning with function approximation.Advances in neural information processing systems, 9, 1996
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e598c191-f7e7-4545-8bb3-f826596fb6d8 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Multi-step rein- forcement learning: A unifying algorithm
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2142cf8-e62d-401c-ae6a-8026fa99b18a · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Horizon reduction makes rl scalable
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f4c06a0-09a3-48b7-b229-9b7bf11dbe76 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Td_gamma: Re-evaluating complex backups in temporal difference learning.Advances in Neural Information Processing Systems, 24, 2011
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 708e183f-ed42-454a-b096-38193cc25133 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Coarse-to-fine q-network with action sequence for data- efficient reinforcement learning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d64887ef-8370-4b33-81f0-fd90dc2d4a36 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Chunking the Critic: A Transformer-based Soft Actor-Critic with N-Step Returns
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b314f8f2-367f-454b-8054-e97eab2948b9 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Reinforcement learning with action chunking
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5cf95b01-4642-4459-960c-ea4c00a90007 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Decoupled q-chunking, 2025
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 304cfa38-a896-471a-b1d1-0ea007335653 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching A distributional perspective on reinforce- ment learning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e5bde1c-3f2e-40d2-b3ba-3c60aff19b1a · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching floq: Training critics via flow-matching for scaling compute in value-based RL
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31bba480-bef7-4b32-a0ec-592890d8afaf · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Value flows
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90b7df33-985d-4b4c-9c57-0b75a609992c · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Temporal abstrac- tion in reinforcement learning with the successor representation.Journal of machine learning research, 24(80):1–69, 2023
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc610f11-2da9-4a5a-8baf-a3bf1363ba8c · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Sutton, Doina Precup, and Satinder Singh
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18711e13-d8de-4b1d-a722-fefa0124b2b5 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching University of Massachusetts Amherst, 2000
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c84c3a0-0fb7-453f-999c-4844beb5be74 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Learning options in reinforcement learning
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f68be6ab-c195-4e38-b399-fc036e222120 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching The option-critic architecture
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16a766cd-d281-43e8-9a9e-e3491c41970d · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Learning abstract options
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5476a698-bdff-4e43-a81f-d6e87c7d1dd3 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching A policy-guided imitation approach for offline reinforcement learning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation edc20e7e-0272-4e3a-9485-7a371328241f · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Hiql: Offline goal- conditioned rl with latent states as actions, 2023
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c4d5da7-c5d8-462a-a0eb-69993c53d6d7 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Data-efficient hierarchical reinforcement learning.Advances in neural information processing systems, 31, 2018
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9444543-ae92-498b-bc53-25a75908d319 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Ogbench: Bench- marking offline goal-conditioned rl
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15df0640-1466-44d1-b213-cd980375b363 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching John Wiley & Sons, 2007
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e3b6525-ab96-4955-b1f3-e4cf6f3b5291 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Bridging the gap be- tween value and policy based reinforcement learning
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32410af2-2484-493c-ac67-0d8f4fb1f62f · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Conservative q-learning for offline reinforcement learning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f49ce69-4a9b-4f36-8801-bc0886a6ec61 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Offline reinforcement learning with implicit q-learning
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad3787dc-9707-42f6-91be-64a5e6061c1d · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Offline rl with no ood actions: In-sample learning via implicit value regularization
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38891176-8fe2-4cc5-b1e0-7324e646a8f7 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Learning from delayed rewards, 1989
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fdfc6f1-b556-4700-9766-e69ae6162343 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Incremental multi-step q-learning
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4ca57ab-451c-41c1-a73d-75a949f403ad · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Policy evalu- ation using theω-return.Advances in Neural Information Processing Systems, 28, 2015
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9be2d08a-025b-4f4d-89a8-aadaf58ff3d4 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Learning fine-grained bimanual manipulation with low-cost hardware.Robotics: Science and Systems, 2023
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 098e19ef-733f-49ff-8979-0b591cb32a84 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching John Wiley & Sons, 2014
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c65e21b-56fe-4008-b99e-c9d1dd2951b0 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Feudal networks for hierarchical reinforcement learn- ing
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae9b5635-871e-457a-93cb-9fe23948a0d1 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 881da501-79f0-4fe3-a9b2-76559a3fee98 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Extreme q-learning: Maxent rl without entropy
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e937a338-424f-481c-a77c-bdb858e017fd · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Safe offline reinforcement learning with feasibility-guided diffusion model
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b26c65ea-5dab-42ad-a761-5e4a2d6130ac · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Dichoto- mous diffusion policy optimization
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd856d21-b402-41df-99f3-f030f5c4b6fd · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Deep unsuper- vised learning using nonequilibrium thermodynamics
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 588ff4bf-ef48-450c-a38c-58afd39d4eb2 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Denoising diffusion probabilistic models.Ad- vances in neural information processing systems, 33:6840–6851, 2020
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8080cfe6-8dba-455b-b8e4-d4061ed368ce · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Unresolved cited work
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71eead67-ea6e-40ff-bcbb-78a974e02e04 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Towards robust zero-shot reinforcement learning
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72199c30-b291-4b1f-9523-1f37db045b96 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Efficient online reinforcement learning for diffusion policy
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31be72bd-bcb8-4bb3-9973-cd3629d14b17 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Flow q-learning
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5746736b-c25a-4f6f-88d5-7e86dee960e4 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Diffusion-dice: In-sample diffusion guidance for offline reinforcement learning
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e0498c7-4982-48ce-ab94-d84277a20b40 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 06ad89fe-348f-4091-ae54-899e989a3a24 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Off-policy deep reinforcement learning without exploration
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0173a308-994c-4d60-b43a-620b7b7f4b88 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Stabilizing off- policy q-learning via bootstrapping error reduction
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b65b53b8-fe2a-4a6a-8a6e-af728e135e53 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching A minimalist approach to offline reinforcement learn- ing.Advances in neural information processing systems, 34:20132–20145, 2021
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a1fd1ef-a8d6-4731-a10b-3b872ebc7660 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Proto: Iterative policy regularized offline-to-online reinforcement learning, 2023
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b70cf63-911a-41bf-af07-cfc5c7aa73a1 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching When data geometry meets deep function: Generalizing offline reinforcement learning
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4e0ea1e-f596-4343-aeee-7ad9ed6d2a58 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Look beneath the surface: Exploiting fundamental symmetry for sample-efficient offline rl
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7865b4a-80aa-4a77-b2b0-6b87e55f6e83 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Offline reinforcement learn- ing with fisher divergence critic regularization
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f14b3940-3c9a-4446-9803-3b91f3cec9b2 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching When to trust your simulator: Dynamics-aware hybrid offline-and-online reinforcement learning
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3adfc898-376f-458e-812a-bad7d63bf04c · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Rorl: Robust offline reinforcement learning via conservative smoothing
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c16bef7-267b-4161-bf1f-c396698d5c4c · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Constraints penalized q-learning for safe offline reinforcement learning
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16df5df6-ea72-47e1-b3cc-ab6fbdea108d · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Offline multi-agent rein- forcement learning with implicit global-to-local value regularization
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bcea7810-e5fa-464a-8ea8-52a2f0726b30 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Weiss, Niru Maheswaranathan, and Surya Ganguli
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02edb6e5-09fc-4ad7-81ce-4d9904f0f004 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Denoising diffusion probabilistic models
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf71de4e-6427-4726-b52b-3167ca623e42 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Score-based generative modeling through stochastic evolution equations in hilbert spaces
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 799d3ff8-c920-46ac-b4f8-4c9176c15309 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Unresolved cited work
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5da92b9-19de-43f3-882b-7b698606081c · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Hunt, and Mingyuan Zhou
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59bc57ae-4a49-462e-873e-54bd34b449bb · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Idql: Implicit q-learning as an actor-critic method with diffusion policies, 2023
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d35d20e2-c7b4-43ee-b794-fb299c5f7583 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Offline reinforcement learn- ing via high-fidelity generative behavior modeling
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21c33a3f-4228-4782-9fe3-864a6f59f690 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Contrastive energy prediction for exact energy-guided diffusion sampling in offline reinforcement learning
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81df0d6c-a69b-4cbb-a1c1-11176ca99791 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Diffusion guidance is a con- trollable policy improvement operator, 2025
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cb012a8-55af-4b86-aa0a-7ebd057a9adb · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Scaling offline rl via efficient and expressive shortcut models
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c90a53ab-ab74-4496-8243-7fbe4cbb3d5b · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Q-learning with Adjoint Matching
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 06bc1572-333a-4214-ac87-2181deef392a · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Unresolved cited work
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0846e62-9c6c-4efd-9975-7901c3a05b69 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Unleashing the potential of diffusion models for end-to-end autonomous driving
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39f2e062-feb3-4415-a3e6-3d051d790040 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Stop regressing: Training value functions via classification for scalable deep RL
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e827b47e-0581-4be6-ada2-c585ff96e94a · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Dsac: Distributional soft actor-critic for risk-sensitive reinforcement learning.Journal of Artificial Intelligence Research, 83, 2025
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e04b4a0a-bdf1-4d31-8e8f-2cd1177d323a · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Q-transformer: Scalable offline reinforce- ment learning via autoregressive q-functions
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a5d6b76-7af4-46c9-92b8-89d097e7321d · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Offline actor-critic reinforcement learn- ing scales to large models
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4052f0cd-cda5-41bb-a93e-29ee9eb05508 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Mixtures of experts unlock parameter scaling for deep RL
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a4196ac-39d2-4c34-bca3-3a22c13e5172 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Feudal reinforcement learning.Advances in neural information processing systems, 5, 1992
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30dba910-41d5-402c-863a-bdc1178884e7 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Hierarchical reinforcement learning with the maxq value function de- composition.Journal of artificial intelligence research, 13:227–303, 2000
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5d3140c-8e9e-4821-830e-dfa71a4c4511 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Strategic attentive writer for learning macro-actions.Advances in neural information processing systems, 29, 2016
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 630a80df-79c5-4d5a-a630-b65b3a758232 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Hierarchical reinforce- ment learning: A comprehensive survey.ACM Computing Surveys (CSUR), 54(5):1–35, 2021
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5105a454-9f35-41bf-b6b5-7b96e6cdf342 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation.Ad- vances in neural information processing systems, 29, 2016
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6f4b1e2-e1bc-4ea2-a2b9-7eb0ccbc1183 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching MuJoCo Menagerie: A collection of high-quality simulation models for MuJoCo, 2022
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74a291ea-9b75-44df-9e83-d3fcb5ffe1e1 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Meta-world: A benchmark and evaluation for multi-task and meta reinforce- ment learning
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bb7b916-bd7b-45b0-a94d-d9306db0fd03 · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching Mujoco: A physics engine for model-based control
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63ce8e93-0fc9-4d48-b2b4-d8a1a5d2acca · outbound
Horizon Adaptive Offline Policy Learning via Value Stitching i−1X t=0 γtrt + k−1X t=i γtrt s0 =s,(s k, k) # (21) =E i∼Unif{1,...,k−1},s i∼π Eπ
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.