Pith. sign in

Paper Citation Record · LEDGER

Horizon Adaptive Offline Policy Learning via Value Stitching

As of 17 August 2026, this Paper Citation Record lists 92 of 92 outbound references and 0 inbound Pith citation observations for arXiv:2606.21136.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.21136 v1

Coverage vector

measured 92 of 92 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-26T14:38:43.803126Z

measured 92 of 92 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

92 of 92 outbound references displayed

  • verified exact3
  • verified fuzzy0
  • unresolved88
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4e1e084f-d42c-4d15-8c27-efee6ae807ed · outbound

This paper cites Learning to predict by the methods of temporal differences.Machine learning, 3(1):9–44, 1988.

Horizon Adaptive Offline Policy Learning via Value Stitching Learning to predict by the methods of temporal differences.Machine learning, 3(1):9–44, 1988

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:5309941cb344c1d8c6672ff3425a3b35ddf31be48613f6113a0e170d7aeb9bed

Observation ff74459d-d860-40c7-9619-804708ea7e88 · outbound

This paper cites Q-learning.Machine learning, 8(3):279–292, 1992.

Horizon Adaptive Offline Policy Learning via Value Stitching Q-learning.Machine learning, 8(3):279–292, 1992

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:3e0b0b84c568d86546fd896fa1912f22ebdeb71a527251a3cef58b6bf3fbafbb

Observation ba73f511-03e2-4d80-8765-caf891bd3712 · outbound

This paper cites Human-level control through deep reinforcement learning.nature, 518(7540):529–533, 2015.

Horizon Adaptive Offline Policy Learning via Value Stitching Human-level control through deep reinforcement learning.nature, 518(7540):529–533, 2015

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:753d47a12214219219ff4d5a473ec3e3d3f4a1597454c169944cf93292d6b9b7

Observation 8e382c61-6075-45ae-aa07-420ead382bff · outbound

This paper cites Mastering the game of go with deep neural networks and tree search.nature, 529 (7587):484–489, 2016.

Horizon Adaptive Offline Policy Learning via Value Stitching Mastering the game of go with deep neural networks and tree search.nature, 529 (7587):484–489, 2016

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:05d7759f180df9701c2bdc0fc9f0aa791f411e23c2159dc616cef25ccad8aefa

Observation f55a9e2f-b1f2-42c0-81dd-473285e312d4 · outbound

This paper cites Mastering the game of go without human knowledge.nature, 550(7676):354–359, 2017.

Horizon Adaptive Offline Policy Learning via Value Stitching Mastering the game of go without human knowledge.nature, 550(7676):354–359, 2017

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:165ddb161fa5727aae6639b49e302428a4aae78920067289c2e304d6d7c547df

Observation be62d276-2bb8-4f7c-af7e-585c5ab08cb3 · outbound

This paper cites Reinforcement learning: An introduction.

Horizon Adaptive Offline Policy Learning via Value Stitching Reinforcement learning: An introduction

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:094796466b03cab7167ad598478172038b0824ac6aad9f3ea1c0c0560dc7b06c

Observation ae82ef4b-7692-4cc2-934c-ce59fde0d832 · outbound

This paper cites A survey of temporal credit assignment in deep reinforcement learning.Transac- tions on Machine Learning Research, 2024.

Horizon Adaptive Offline Policy Learning via Value Stitching A survey of temporal credit assignment in deep reinforcement learning.Transac- tions on Machine Learning Research, 2024

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:b36e09bbf470a4ff44a1fa3752631d2354bc9c53526c5ceaa504068f9b787c09

Observation 8b00a4a8-39fe-4285-9d8c-1f59e5150319 · outbound

This paper cites Optimizing agent behavior over long time scales by transporting value.Nature communications, 10(1):5223, 2019.

Horizon Adaptive Offline Policy Learning via Value Stitching Optimizing agent behavior over long time scales by transporting value.Nature communications, 10(1):5223, 2019

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:5e4ce22980ea4ba5f60c6f12a6666f61317d99bb30af3fcdc2b44c5a62f0b7d8

Observation 07f7f2ef-5880-4bae-bcdd-52adab8a9046 · outbound

This paper cites Is value learning really the main bottleneck in offline rl? InThe Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024.

Horizon Adaptive Offline Policy Learning via Value Stitching Is value learning really the main bottleneck in offline rl? InThe Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:56118677676bc4c0fd7cb076dfd32e3009daa8c8a7bb66e071b4e7264da41012

Observation f7801efc-da11-4091-95a0-75c045ca8e20 · outbound

This paper cites Convergence of stochastic iterative dynamic programming algorithms.Advances in neural information processing systems, 6, 1993.

Horizon Adaptive Offline Policy Learning via Value Stitching Convergence of stochastic iterative dynamic programming algorithms.Advances in neural information processing systems, 6, 1993

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:2504acbbd9eb0c6ea121449378cedad7854e03d0e03badd916d1693759228673

Observation d6dc7829-2c1c-4320-8ace-02411a00062b · outbound

This paper cites Analysis of temporal-diffference learning with function approximation.Advances in neural information processing systems, 9, 1996.

Horizon Adaptive Offline Policy Learning via Value Stitching Analysis of temporal-diffference learning with function approximation.Advances in neural information processing systems, 9, 1996

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:bd93eacabfebe48b7a8fd579e71903de655fc1fab7c3c7d7b21098d692b2fa5f

Observation e598c191-f7e7-4545-8bb3-f826596fb6d8 · outbound

This paper cites Multi-step rein- forcement learning: A unifying algorithm.

Horizon Adaptive Offline Policy Learning via Value Stitching Multi-step rein- forcement learning: A unifying algorithm

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:05565adf3833b303aa1a21569893b5f7fc8197c3d4eb3bac467e7afff2eb19f2

Observation a2142cf8-e62d-401c-ae6a-8026fa99b18a · outbound

This paper cites Horizon reduction makes rl scalable.

Horizon Adaptive Offline Policy Learning via Value Stitching Horizon reduction makes rl scalable

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:5790d9e3409779f8a3f121a5b7c64c23970c0b8ec1318370b94dc048269dca4b

Observation 2f4c06a0-09a3-48b7-b229-9b7bf11dbe76 · outbound

This paper cites Td_gamma: Re-evaluating complex backups in temporal difference learning.Advances in Neural Information Processing Systems, 24, 2011.

Horizon Adaptive Offline Policy Learning via Value Stitching Td_gamma: Re-evaluating complex backups in temporal difference learning.Advances in Neural Information Processing Systems, 24, 2011

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:f9edf21acf4912ac0fe010eae31f252c2faf892af43402cbba9ab37c325d4b3a

Observation 708e183f-ed42-454a-b096-38193cc25133 · outbound

This paper cites Coarse-to-fine q-network with action sequence for data- efficient reinforcement learning.

Horizon Adaptive Offline Policy Learning via Value Stitching Coarse-to-fine q-network with action sequence for data- efficient reinforcement learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:dc000ec9d28e43eb494a6928b8945e534690ed07c4d7a15c076637ef3517f6bb

Observation d64887ef-8370-4b33-81f0-fd90dc2d4a36 · outbound

This paper cites Chunking the Critic: A Transformer-based Soft Actor-Critic with N-Step Returns.

Horizon Adaptive Offline Policy Learning via Value Stitching Chunking the Critic: A Transformer-based Soft Actor-Critic with N-Step Returns

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-04T06:19:37.921595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:d4316d06e3fc1d878a0e4717bb9332100f3d67c91e9bd5a5a7a37afb5d0ab640

Observation b314f8f2-367f-454b-8054-e97eab2948b9 · outbound

This paper cites Reinforcement learning with action chunking.

Horizon Adaptive Offline Policy Learning via Value Stitching Reinforcement learning with action chunking

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:9b13d61df2c8e73e38e875b0ce9e785c7fd02d51836d5f09ee492ad9f27d7560

Observation 5cf95b01-4642-4459-960c-ea4c00a90007 · outbound

This paper cites Decoupled q-chunking, 2025.

Horizon Adaptive Offline Policy Learning via Value Stitching Decoupled q-chunking, 2025

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:052aa29451e5c606ca34a1683d77768f42b38c595e325af595f784216ad25940

Observation 304cfa38-a896-471a-b1d1-0ea007335653 · outbound

This paper cites A distributional perspective on reinforce- ment learning.

Horizon Adaptive Offline Policy Learning via Value Stitching A distributional perspective on reinforce- ment learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:e3840598be072cfc5bff25dd83c6dd07999d9712c8530bd13514096569611c78

Observation 9e5bde1c-3f2e-40d2-b3ba-3c60aff19b1a · outbound

This paper cites floq: Training critics via flow-matching for scaling compute in value-based RL.

Horizon Adaptive Offline Policy Learning via Value Stitching floq: Training critics via flow-matching for scaling compute in value-based RL

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:185aafef8038ebf70a4073852e58603ccc6f0546bec198ab0813b73f4fd028f9

Observation 31bba480-bef7-4b32-a0ec-592890d8afaf · outbound

This paper cites Value flows.

Horizon Adaptive Offline Policy Learning via Value Stitching Value flows

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:1df81408405751254af823760ebc1a58969e31a5c65b229f00d551368e481c39

Observation 90b7df33-985d-4b4c-9c57-0b75a609992c · outbound

This paper cites Temporal abstrac- tion in reinforcement learning with the successor representation.Journal of machine learning research, 24(80):1–69, 2023.

Horizon Adaptive Offline Policy Learning via Value Stitching Temporal abstrac- tion in reinforcement learning with the successor representation.Journal of machine learning research, 24(80):1–69, 2023

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:446f118d078c67bc6838d90feea788e46b9335020fb5dde79a547861ef19a784

Observation fc610f11-2da9-4a5a-8baf-a3bf1363ba8c · outbound

This paper cites Sutton, Doina Precup, and Satinder Singh.

Horizon Adaptive Offline Policy Learning via Value Stitching Sutton, Doina Precup, and Satinder Singh

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:50e0b5fc027896faa6987080ef4619f2aeb626eef78f248a09ab4b9e75afa14b

Observation 18711e13-d8de-4b1d-a722-fefa0124b2b5 · outbound

This paper cites University of Massachusetts Amherst, 2000.

Horizon Adaptive Offline Policy Learning via Value Stitching University of Massachusetts Amherst, 2000

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:6f0ad108d519f5cedbf94401693952bc8897beb36c502f8142b7acced718e4de

Observation 2c84c3a0-0fb7-453f-999c-4844beb5be74 · outbound

This paper cites Learning options in reinforcement learning.

Horizon Adaptive Offline Policy Learning via Value Stitching Learning options in reinforcement learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:d8ddacf660f620a5fdccfb3706c94be58ad36fd876fb4c0f37c3f0a41b6689b9

Observation f68be6ab-c195-4e38-b399-fc036e222120 · outbound

This paper cites The option-critic architecture.

Horizon Adaptive Offline Policy Learning via Value Stitching The option-critic architecture

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:539955d6b508da519a5ce79ed392bdeea215932d02431f66187154bff29b5980

Observation 16a766cd-d281-43e8-9a9e-e3491c41970d · outbound

This paper cites Learning abstract options.

Horizon Adaptive Offline Policy Learning via Value Stitching Learning abstract options

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:c479826c40cf3106535316c3a8ec07904679bbfb23ea9c6701b28d22c36a8a76

Observation 5476a698-bdff-4e43-a81f-d6e87c7d1dd3 · outbound

This paper cites A policy-guided imitation approach for offline reinforcement learning.

Horizon Adaptive Offline Policy Learning via Value Stitching A policy-guided imitation approach for offline reinforcement learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:64b607d0a1e7d8df77f3a673223bc1efa0ebc291e5566426700e908fc7d15765

Observation edc20e7e-0272-4e3a-9485-7a371328241f · outbound

This paper cites Hiql: Offline goal- conditioned rl with latent states as actions, 2023.

Horizon Adaptive Offline Policy Learning via Value Stitching Hiql: Offline goal- conditioned rl with latent states as actions, 2023

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:001c7f37cd21c63b4c3a6cb3dc93ba5def563b97341ee45f0ebd92e49f917d9d

Observation 9c4d5da7-c5d8-462a-a0eb-69993c53d6d7 · outbound

This paper cites Data-efficient hierarchical reinforcement learning.Advances in neural information processing systems, 31, 2018.

Horizon Adaptive Offline Policy Learning via Value Stitching Data-efficient hierarchical reinforcement learning.Advances in neural information processing systems, 31, 2018

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:43ae3d415ee5d566301f9575fb52d5f3a7f90547c8e06f3dfea8d564368f7428

Observation b9444543-ae92-498b-bc53-25a75908d319 · outbound

This paper cites Ogbench: Bench- marking offline goal-conditioned rl.

Horizon Adaptive Offline Policy Learning via Value Stitching Ogbench: Bench- marking offline goal-conditioned rl

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:09c92f2be409d098d3dd66b41fbd816bff7a35c5d902458c2e5a04fb5c5819bb

Observation 15df0640-1466-44d1-b213-cd980375b363 · outbound

This paper cites John Wiley & Sons, 2007.

Horizon Adaptive Offline Policy Learning via Value Stitching John Wiley & Sons, 2007

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:3e75379eb4750b59bcf555c546dee6397e9c423e6f4a18fb256bda47cf305acd

Observation 9e3b6525-ab96-4955-b1f3-e4cf6f3b5291 · outbound

This paper cites Bridging the gap be- tween value and policy based reinforcement learning.

Horizon Adaptive Offline Policy Learning via Value Stitching Bridging the gap be- tween value and policy based reinforcement learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:ac8c466bae6f0577cdfb37bf070fb6dbaab54408cd72a2ad9cf76befe94c8bce

Observation 32410af2-2484-493c-ac67-0d8f4fb1f62f · outbound

This paper cites Conservative q-learning for offline reinforcement learning.

Horizon Adaptive Offline Policy Learning via Value Stitching Conservative q-learning for offline reinforcement learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:a0cf9fdd29c5c91535a960c3758711de9a750f42f97ae1c70b9ac4a3690a244d

Observation 7f49ce69-4a9b-4f36-8801-bc0886a6ec61 · outbound

This paper cites Offline reinforcement learning with implicit q-learning.

Horizon Adaptive Offline Policy Learning via Value Stitching Offline reinforcement learning with implicit q-learning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:58cbf764188266733ac1f33ba8d7aaadfbdcf17fa99e0fe8cbc868c25cf274e2

Observation ad3787dc-9707-42f6-91be-64a5e6061c1d · outbound

This paper cites Offline rl with no ood actions: In-sample learning via implicit value regularization.

Horizon Adaptive Offline Policy Learning via Value Stitching Offline rl with no ood actions: In-sample learning via implicit value regularization

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:128ced3b2046dec1353d1b932185b7015973178e06c0f75ef7b58dba09cab1e1

Observation 38891176-8fe2-4cc5-b1e0-7324e646a8f7 · outbound

This paper cites Learning from delayed rewards, 1989.

Horizon Adaptive Offline Policy Learning via Value Stitching Learning from delayed rewards, 1989

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:6117951dc01c8b76738455adb12d9029902e3956053b8cbe8160c20e17ce2227

Observation 1fdfc6f1-b556-4700-9766-e69ae6162343 · outbound

This paper cites Incremental multi-step q-learning.

Horizon Adaptive Offline Policy Learning via Value Stitching Incremental multi-step q-learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:01470e8c7764a17a07533af67602ec0cdf2443e57ac0d3c66091c6a76ae0ecca

Observation f4ca57ab-451c-41c1-a73d-75a949f403ad · outbound

This paper cites Policy evalu- ation using theω-return.Advances in Neural Information Processing Systems, 28, 2015.

Horizon Adaptive Offline Policy Learning via Value Stitching Policy evalu- ation using theω-return.Advances in Neural Information Processing Systems, 28, 2015

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:e36d4775fdf6a1ea8fb5154fbff7224726d167c5099d081d20edd4e4e3d3eb6a

Observation 9be2d08a-025b-4f4d-89a8-aadaf58ff3d4 · outbound

This paper cites Learning fine-grained bimanual manipulation with low-cost hardware.Robotics: Science and Systems, 2023.

Horizon Adaptive Offline Policy Learning via Value Stitching Learning fine-grained bimanual manipulation with low-cost hardware.Robotics: Science and Systems, 2023

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:b13fd6a0807733218e54896b745e16bd3c928ba7ebe461013a57acc896412329

Observation 098e19ef-733f-49ff-8979-0b591cb32a84 · outbound

This paper cites John Wiley & Sons, 2014.

Horizon Adaptive Offline Policy Learning via Value Stitching John Wiley & Sons, 2014

Reference 41

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:cd0caa7851b689be708d3bd771b8cd2704214fff4b2fbd57ece3ef1526f1d35e

Observation 1c65e21b-56fe-4008-b99e-c9d1dd2951b0 · outbound

This paper cites Feudal networks for hierarchical reinforcement learn- ing.

Horizon Adaptive Offline Policy Learning via Value Stitching Feudal networks for hierarchical reinforcement learn- ing

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:e9cdfa5f6c1a5e89f0e75e43ef2a0db5657868414a9bf8c1e93b56d75f4e6ab6

Observation ae9b5635-871e-457a-93cb-9fe23948a0d1 · outbound

This paper cites Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning.

Horizon Adaptive Offline Policy Learning via Value Stitching Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-07-04T06:19:37.924825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:e83b2efba4f706314eb7bdd3479642e251444bb40e70b23d28360f6f123693af

Observation 881da501-79f0-4fe3-a9b2-76559a3fee98 · outbound

This paper cites Extreme q-learning: Maxent rl without entropy.

Horizon Adaptive Offline Policy Learning via Value Stitching Extreme q-learning: Maxent rl without entropy

Reference 44

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:63031a42b9e6dac23ebde9269a47ed2e60ec1190cab1e140772bcea0bbb0dac3

Observation e937a338-424f-481c-a77c-bdb858e017fd · outbound

This paper cites Safe offline reinforcement learning with feasibility-guided diffusion model.

Horizon Adaptive Offline Policy Learning via Value Stitching Safe offline reinforcement learning with feasibility-guided diffusion model

Reference 45

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:d05fa4519aca627aa77e4780dfb20fa10b515f29c4db22e3096ca5e90ea07058

Observation b26c65ea-5dab-42ad-a761-5e4a2d6130ac · outbound

This paper cites Dichoto- mous diffusion policy optimization.

Horizon Adaptive Offline Policy Learning via Value Stitching Dichoto- mous diffusion policy optimization

Reference 46

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:2a5293c800098ceb42c925d9838384f483f59fb5909e59dfdde1f35ccbfbddc2

Observation cd856d21-b402-41df-99f3-f030f5c4b6fd · outbound

This paper cites Deep unsuper- vised learning using nonequilibrium thermodynamics.

Horizon Adaptive Offline Policy Learning via Value Stitching Deep unsuper- vised learning using nonequilibrium thermodynamics

Reference 47

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:a4a888211ec09f0a9829d302e3a17b2bc83feb616b759a58b9053c4ef67384b8

Observation 588ff4bf-ef48-450c-a38c-58afd39d4eb2 · outbound

This paper cites Denoising diffusion probabilistic models.Ad- vances in neural information processing systems, 33:6840–6851, 2020.

Horizon Adaptive Offline Policy Learning via Value Stitching Denoising diffusion probabilistic models.Ad- vances in neural information processing systems, 33:6840–6851, 2020

Reference 48

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:66af588c8a0d0807dff9ed4699f2457e7fecbacf759800fb9c2170cd82141e16

Observation 8080cfe6-8dba-455b-b8e4-d4061ed368ce · outbound

This paper cites an unresolved cited work.

Horizon Adaptive Offline Policy Learning via Value Stitching Unresolved cited work

Reference 49

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:8aac8649609084824f54fddf3393910abdc9858a7e2723bdcb58640014c1dd49

Observation 71eead67-ea6e-40ff-bcbb-78a974e02e04 · outbound

This paper cites Towards robust zero-shot reinforcement learning.

Horizon Adaptive Offline Policy Learning via Value Stitching Towards robust zero-shot reinforcement learning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:262902595cccff1cfbdc94cbdd5f961f6c0f87965bd7f407793ab8f64d96b4ed

Observation 72199c30-b291-4b1f-9523-1f37db045b96 · outbound

This paper cites Efficient online reinforcement learning for diffusion policy.

Horizon Adaptive Offline Policy Learning via Value Stitching Efficient online reinforcement learning for diffusion policy

Reference 51

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:4be46ed23899c80e8e0b0086bda0e7fc22bdceb72d0ca99d5760a9613575de44

Observation 31be72bd-bcb8-4bb3-9973-cd3629d14b17 · outbound

This paper cites Flow q-learning.

Horizon Adaptive Offline Policy Learning via Value Stitching Flow q-learning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:8c8fe6998314c5619dba421294b241c4541422dd070c710fdf125010a1bd6e9f

Observation 5746736b-c25a-4f6f-88d5-7e86dee960e4 · outbound

This paper cites Diffusion-dice: In-sample diffusion guidance for offline reinforcement learning.

Horizon Adaptive Offline Policy Learning via Value Stitching Diffusion-dice: In-sample diffusion guidance for offline reinforcement learning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:4deeb2a235e680b1ec4471c1b2dbfe5e0634d63e387f56894f8b343ce519239a

Observation 7e0498c7-4982-48ce-ab94-d84277a20b40 · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

Horizon Adaptive Offline Policy Learning via Value Stitching Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 54

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T06:19:37.931325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:53d2e0a45cb8d24fd71f31a1b6bed7b6db72a30170fd50b496f504a0061d9e03

Observation 06ad89fe-348f-4091-ae54-899e989a3a24 · outbound

This paper cites Off-policy deep reinforcement learning without exploration.

Horizon Adaptive Offline Policy Learning via Value Stitching Off-policy deep reinforcement learning without exploration

Reference 55

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:9c2c40aa590d425f3a934bfdcf9299d22b41cbf1bbccfe7681193a36c7365c26

Observation 0173a308-994c-4d60-b43a-620b7b7f4b88 · outbound

This paper cites Stabilizing off- policy q-learning via bootstrapping error reduction.

Horizon Adaptive Offline Policy Learning via Value Stitching Stabilizing off- policy q-learning via bootstrapping error reduction

Reference 56

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:0eae2febe286c53144ab7560b2e5c8770c7f8fba38495538f0f2f9a8df9ad93e

Observation b65b53b8-fe2a-4a6a-8a6e-af728e135e53 · outbound

This paper cites A minimalist approach to offline reinforcement learn- ing.Advances in neural information processing systems, 34:20132–20145, 2021.

Horizon Adaptive Offline Policy Learning via Value Stitching A minimalist approach to offline reinforcement learn- ing.Advances in neural information processing systems, 34:20132–20145, 2021

Reference 57

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:3645ed8a14de907002fe96655fb73a0c4a707e826ea00aafff9153a7825a207d

Observation 2a1fd1ef-a8d6-4731-a10b-3b872ebc7660 · outbound

This paper cites Proto: Iterative policy regularized offline-to-online reinforcement learning, 2023.

Horizon Adaptive Offline Policy Learning via Value Stitching Proto: Iterative policy regularized offline-to-online reinforcement learning, 2023

Reference 58

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:4937743339a794edc13334f338bc8aee86fdfad1b1ba4a396c22bf4f850ebbd1

Observation 2b70cf63-911a-41bf-af07-cfc5c7aa73a1 · outbound

This paper cites When data geometry meets deep function: Generalizing offline reinforcement learning.

Horizon Adaptive Offline Policy Learning via Value Stitching When data geometry meets deep function: Generalizing offline reinforcement learning

Reference 59

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:d05f1cac1e1ba9a05b54ade6b168ebd7e184f6372eda20abc0bbb03a70a0087d

Observation f4e0ea1e-f596-4343-aeee-7ad9ed6d2a58 · outbound

This paper cites Look beneath the surface: Exploiting fundamental symmetry for sample-efficient offline rl.

Horizon Adaptive Offline Policy Learning via Value Stitching Look beneath the surface: Exploiting fundamental symmetry for sample-efficient offline rl

Reference 60

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:2e7eb82885abaaa63476f97feaa4737ce143250b851b0f8124b4cd44d040d31c

Observation a7865b4a-80aa-4a77-b2b0-6b87e55f6e83 · outbound

This paper cites Offline reinforcement learn- ing with fisher divergence critic regularization.

Horizon Adaptive Offline Policy Learning via Value Stitching Offline reinforcement learn- ing with fisher divergence critic regularization

Reference 61

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:68987ff70ce2e8387298fb2aa085ad6be80d026ab8b85aeef44777d7cd92f2bd

Observation f14b3940-3c9a-4446-9803-3b91f3cec9b2 · outbound

This paper cites When to trust your simulator: Dynamics-aware hybrid offline-and-online reinforcement learning.

Horizon Adaptive Offline Policy Learning via Value Stitching When to trust your simulator: Dynamics-aware hybrid offline-and-online reinforcement learning

Reference 62

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:a5a8006e347de79c78807936f8f0b1fac391fecc4ea71e0dad3d809f01604910

Observation 3adfc898-376f-458e-812a-bad7d63bf04c · outbound

This paper cites Rorl: Robust offline reinforcement learning via conservative smoothing.

Horizon Adaptive Offline Policy Learning via Value Stitching Rorl: Robust offline reinforcement learning via conservative smoothing

Reference 63

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:364ea7ddc329eb7bd19d4c8cb59bf782c18f46391bd822fbf2f33b607d8461ec

Observation 9c16bef7-267b-4161-bf1f-c396698d5c4c · outbound

This paper cites Constraints penalized q-learning for safe offline reinforcement learning.

Horizon Adaptive Offline Policy Learning via Value Stitching Constraints penalized q-learning for safe offline reinforcement learning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:bbd86235ba94c447589e3236900d9241a147807476bc922c45feecfa33aef141

Observation 16df5df6-ea72-47e1-b3cc-ab6fbdea108d · outbound

This paper cites Offline multi-agent rein- forcement learning with implicit global-to-local value regularization.

Horizon Adaptive Offline Policy Learning via Value Stitching Offline multi-agent rein- forcement learning with implicit global-to-local value regularization

Reference 65

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:c9a734ba6bd840adaeb7e742007a85635620d18745fbfca3f7329764b4b5a710

Observation bcea7810-e5fa-464a-8ea8-52a2f0726b30 · outbound

This paper cites Weiss, Niru Maheswaranathan, and Surya Ganguli.

Horizon Adaptive Offline Policy Learning via Value Stitching Weiss, Niru Maheswaranathan, and Surya Ganguli

Reference 66

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:dde4a44d24c9f32a98009d50eb1021c9ea568c9eabb9341e7fa49016ae4d78af

Observation 02edb6e5-09fc-4ad7-81ce-4d9904f0f004 · outbound

This paper cites Denoising diffusion probabilistic models.

Horizon Adaptive Offline Policy Learning via Value Stitching Denoising diffusion probabilistic models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:16d2f189624b61669a8e85d00bf5368173941e1f99abbaf823c1da2a56bc7bfd

Observation bf71de4e-6427-4726-b52b-3167ca623e42 · outbound

This paper cites Score-based generative modeling through stochastic evolution equations in hilbert spaces.

Horizon Adaptive Offline Policy Learning via Value Stitching Score-based generative modeling through stochastic evolution equations in hilbert spaces

Reference 68

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:0d2801c0ee4a276d3a7e3deb0269aa884b12664fa3f6ec55db8618de27a8a6b0

Observation 799d3ff8-c920-46ac-b4f8-4c9176c15309 · outbound

This paper cites an unresolved cited work.

Horizon Adaptive Offline Policy Learning via Value Stitching Unresolved cited work

Reference 69

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:fc3edf2128993fe7c597f3534daabd1450b71fec64434cd09d78e6fe7a784aab

Observation a5da92b9-19de-43f3-882b-7b698606081c · outbound

This paper cites Hunt, and Mingyuan Zhou.

Horizon Adaptive Offline Policy Learning via Value Stitching Hunt, and Mingyuan Zhou

Reference 70

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:f5a292c1051767fe642366d8d786c57edda6b5b21f18a7369ebafbcfcb14e9c7

Observation 59bc57ae-4a49-462e-873e-54bd34b449bb · outbound

This paper cites Idql: Implicit q-learning as an actor-critic method with diffusion policies, 2023.

Horizon Adaptive Offline Policy Learning via Value Stitching Idql: Implicit q-learning as an actor-critic method with diffusion policies, 2023

Reference 71

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:9ade301018931ef4da66f9ae1243a6f4e65f066caa5d7e6aafae8462b057bcd1

Observation d35d20e2-c7b4-43ee-b794-fb299c5f7583 · outbound

This paper cites Offline reinforcement learn- ing via high-fidelity generative behavior modeling.

Horizon Adaptive Offline Policy Learning via Value Stitching Offline reinforcement learn- ing via high-fidelity generative behavior modeling

Reference 72

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:3ca271e3aff37edb620ddd6d26f0ae68015c01825dc5137181f91d158ed00e13

Observation 21c33a3f-4228-4782-9fe3-864a6f59f690 · outbound

This paper cites Contrastive energy prediction for exact energy-guided diffusion sampling in offline reinforcement learning.

Horizon Adaptive Offline Policy Learning via Value Stitching Contrastive energy prediction for exact energy-guided diffusion sampling in offline reinforcement learning

Reference 73

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:210da115d0dd04ca368624350ba160ffb5e2f3d1dcee0df12df237484c7ae7ea

Observation 81df0d6c-a69b-4cbb-a1c1-11176ca99791 · outbound

This paper cites Diffusion guidance is a con- trollable policy improvement operator, 2025.

Horizon Adaptive Offline Policy Learning via Value Stitching Diffusion guidance is a con- trollable policy improvement operator, 2025

Reference 74

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:a57ec1d93a8f7a7be55a3c567907f7361b60eb17cd628b99f84722262ea621cb

Observation 8cb012a8-55af-4b86-aa0a-7ebd057a9adb · outbound

This paper cites Scaling offline rl via efficient and expressive shortcut models.

Horizon Adaptive Offline Policy Learning via Value Stitching Scaling offline rl via efficient and expressive shortcut models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:399b6c96b7fed8a9337d315da711f12398b97ea9be79cd608ce7bd6dbaaa8baf

Observation c90a53ab-ab74-4496-8243-7fbe4cbb3d5b · outbound

This paper cites Q-learning with Adjoint Matching.

Horizon Adaptive Offline Policy Learning via Value Stitching Q-learning with Adjoint Matching

Reference 76

Resolution
verified exact
local_arxiv, observed 2026-07-04T06:19:37.928185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:541d0316064adfe7f21ff5d504cf770050a958ca46a20128a95dbe9322da0719

Observation 06bc1572-333a-4214-ac87-2181deef392a · outbound

This paper cites an unresolved cited work.

Horizon Adaptive Offline Policy Learning via Value Stitching Unresolved cited work

Reference 77

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:fe04e1ed79c1b8307827fe5ce2c21d8d87bc21df3f835f4fe79eea0bdccf9175

Observation e0846e62-9c6c-4efd-9975-7901c3a05b69 · outbound

This paper cites Unleashing the potential of diffusion models for end-to-end autonomous driving.

Horizon Adaptive Offline Policy Learning via Value Stitching Unleashing the potential of diffusion models for end-to-end autonomous driving

Reference 78

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:60223445c00d9dd1fcb61cfd361019964b15b246f5170e84e1fc6c9676fd0df9

Observation 39f2e062-feb3-4415-a3e6-3d051d790040 · outbound

This paper cites Stop regressing: Training value functions via classification for scalable deep RL.

Horizon Adaptive Offline Policy Learning via Value Stitching Stop regressing: Training value functions via classification for scalable deep RL

Reference 79

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:10bdfdc31e9b8366354d1c2bc07e9e2e3944a8cfd7ae3c06673e4e070c2782ad

Observation e827b47e-0581-4be6-ada2-c585ff96e94a · outbound

This paper cites Dsac: Distributional soft actor-critic for risk-sensitive reinforcement learning.Journal of Artificial Intelligence Research, 83, 2025.

Horizon Adaptive Offline Policy Learning via Value Stitching Dsac: Distributional soft actor-critic for risk-sensitive reinforcement learning.Journal of Artificial Intelligence Research, 83, 2025

Reference 80

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:1b7d75d048d6a85cf6f8502b82fae50034708e4679f803190ccf3ef8366d0065

Observation e04b4a0a-bdf1-4d31-8e8f-2cd1177d323a · outbound

This paper cites Q-transformer: Scalable offline reinforce- ment learning via autoregressive q-functions.

Horizon Adaptive Offline Policy Learning via Value Stitching Q-transformer: Scalable offline reinforce- ment learning via autoregressive q-functions

Reference 81

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:ce3f6af58fdbffce54ee22956beec97428998d55644c57264871212b8badcd19

Observation 6a5d6b76-7af4-46c9-92b8-89d097e7321d · outbound

This paper cites Offline actor-critic reinforcement learn- ing scales to large models.

Horizon Adaptive Offline Policy Learning via Value Stitching Offline actor-critic reinforcement learn- ing scales to large models

Reference 82

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:2850318e6a235762ce9c4d1c5472afe05e3edc1994a6d57426cfe9faacdf4ab7

Observation 4052f0cd-cda5-41bb-a93e-29ee9eb05508 · outbound

This paper cites Mixtures of experts unlock parameter scaling for deep RL.

Horizon Adaptive Offline Policy Learning via Value Stitching Mixtures of experts unlock parameter scaling for deep RL

Reference 83

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:256d22471d3deb38f8d5d891052c92160cf56c2a6e8691c441656686dbd29100

Observation 5a4196ac-39d2-4c34-bca3-3a22c13e5172 · outbound

This paper cites Feudal reinforcement learning.Advances in neural information processing systems, 5, 1992.

Horizon Adaptive Offline Policy Learning via Value Stitching Feudal reinforcement learning.Advances in neural information processing systems, 5, 1992

Reference 84

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:8914bc41254ef78682d4430454b46a93e7133ec7627116d1542e6d59478efe91

Observation 30dba910-41d5-402c-863a-bdc1178884e7 · outbound

This paper cites Hierarchical reinforcement learning with the maxq value function de- composition.Journal of artificial intelligence research, 13:227–303, 2000.

Horizon Adaptive Offline Policy Learning via Value Stitching Hierarchical reinforcement learning with the maxq value function de- composition.Journal of artificial intelligence research, 13:227–303, 2000

Reference 85

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:32f05eb3e123303a14408e5bb8127c9bb7e38759e544fe4b556a9b8d03b8ee31

Observation e5d3140c-8e9e-4821-830e-dfa71a4c4511 · outbound

This paper cites Strategic attentive writer for learning macro-actions.Advances in neural information processing systems, 29, 2016.

Horizon Adaptive Offline Policy Learning via Value Stitching Strategic attentive writer for learning macro-actions.Advances in neural information processing systems, 29, 2016

Reference 86

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:a21c358221c483232e294b5bf405f0d22bb7d455572ddfd1d77745ead38c2d13

Observation 630a80df-79c5-4d5a-a630-b65b3a758232 · outbound

This paper cites Hierarchical reinforce- ment learning: A comprehensive survey.ACM Computing Surveys (CSUR), 54(5):1–35, 2021.

Horizon Adaptive Offline Policy Learning via Value Stitching Hierarchical reinforce- ment learning: A comprehensive survey.ACM Computing Surveys (CSUR), 54(5):1–35, 2021

Reference 87

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:5340a6a74dc55c378999c8802179904e0aaec2ed7ad1a288e5f9af5bab39cd34

Observation 5105a454-9f35-41bf-b6b5-7b96e6cdf342 · outbound

This paper cites Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation.Ad- vances in neural information processing systems, 29, 2016.

Horizon Adaptive Offline Policy Learning via Value Stitching Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation.Ad- vances in neural information processing systems, 29, 2016

Reference 88

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:554cbeee770b279636ecf587ed2aa2f750d1307c9b77e5242c41dcc3fae5ded6

Observation c6f4b1e2-e1bc-4ea2-a2b9-7eb0ccbc1183 · outbound

This paper cites MuJoCo Menagerie: A collection of high-quality simulation models for MuJoCo, 2022.

Horizon Adaptive Offline Policy Learning via Value Stitching MuJoCo Menagerie: A collection of high-quality simulation models for MuJoCo, 2022

Reference 89

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:c1bfcbc4593af212f983399660ad47cdb4924a7e1baa3709baeedb1a23e75b55

Observation 74a291ea-9b75-44df-9e83-d3fcb5ffe1e1 · outbound

This paper cites Meta-world: A benchmark and evaluation for multi-task and meta reinforce- ment learning.

Horizon Adaptive Offline Policy Learning via Value Stitching Meta-world: A benchmark and evaluation for multi-task and meta reinforce- ment learning

Reference 90

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:18c7d68d34caad63850d6001956bed4c158889dc0dddc38ea7c3bb1fc8640522

Observation 4bb7b916-bd7b-45b0-a94d-d9306db0fd03 · outbound

This paper cites Mujoco: A physics engine for model-based control.

Horizon Adaptive Offline Policy Learning via Value Stitching Mujoco: A physics engine for model-based control

Reference 91

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:8dc61a82672e4752521346719c21c4285aed5e76413cc1140d845eb37ec50e42

Observation 63ce8e93-0fc9-4d48-b2b4-d8a1a5d2acca · outbound

This paper cites i−1X t=0 γtrt + k−1X t=i γtrt s0 =s,(s k, k) # (21) =E i∼Unif{1,...,k−1},s i∼π Eπ.

Horizon Adaptive Offline Policy Learning via Value Stitching i−1X t=0 γtrt + k−1X t=i γtrt s0 =s,(s k, k) # (21) =E i∼Unif{1,...,k−1},s i∼π Eπ

Reference 92

Resolution
unresolved
no resolver link, observed 2026-06-26T14:38:43.803126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:38:43.803126Z digest=sha256:3e43092deb2fa45b42c4da689f6fee61f177e281ebbd1cf8e7b04ca698f0020e

Pith citing papers

No inbound Pith citation observations are available.