Pith. sign in

Paper Citation Record · LEDGER

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization

As of 14 August 2026, this Paper Citation Record lists 83 of 83 outbound references and 0 inbound Pith citation observations for arXiv:2506.00795.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00795 v3

Coverage vector

measured 83 of 83 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:05:26.101533Z

measured 83 of 83 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

83 of 83 outbound references displayed

  • verified exact1
  • verified fuzzy41
  • unresolved41
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ac9d45cc-7a9f-41d9-b013-0c0be567dee7 · outbound

This paper cites Deep reinforcement learning at the edge of the statistical precipice.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Deep reinforcement learning at the edge of the statistical precipice

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.717787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.717787Z digest=sha256:86d82ed7bcfb1de0a86350496b2c4e4ed78ed5e8cb6547fc42a276e09d8ab842

Observation 0c629fa2-b676-4ebf-a403-7d3f06ebde26 · outbound

This paper cites On the estimation of production frontiers: maximum likelihood estimation of the parameters of a discontinuous density function.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization On the estimation of production frontiers: maximum likelihood estimation of the parameters of a discontinuous density function

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.723026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.723026Z digest=sha256:462ac4b81ef6cf12598991ac6a7e09a8b4ab4467957aefb2e6faea1c2f1a19a8

Observation a6022746-c35d-493e-ab30-e9d86c9fab73 · outbound

This paper cites Hindsight experience replay.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Hindsight experience replay

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.728070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.728070Z digest=sha256:0e3c8d41adea8011cb6e3b8831434df191aa113570a9fc0d640b8a8016caae61

Observation 0eec7250-8f73-4841-8163-0584315efc57 · outbound

This paper cites Learning Successor States and Goal-Dependent Values: A Mathematical Viewpoint.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Learning Successor States and Goal-Dependent Values: A Mathematical Viewpoint

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.733026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.733026Z digest=sha256:1bae6cb4549d35a97dbd67da3f0319744db91b23d374b607b1d2f637dbae4694

Observation d0775145-8232-416d-ad3c-985501eafa1c · outbound

This paper cites Accelerating goal-conditioned reinforcement learning algorithms and research.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Accelerating goal-conditioned reinforcement learning algorithms and research

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:27.123474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.737574Z digest=sha256:06a9db8b311617c44cacc845101b26ba9ba4ee5364b59eb93bbfa7a9f8030ae9

Observation fb9f4169-e126-4d26-8760-662962c2a265 · outbound

This paper cites Offline rl without off-policy evaluation.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Offline rl without off-policy evaluation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.742340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.742340Z digest=sha256:5fb4370ffd99df97c24d758e9b6f6b084a5f181c594e3bc18af65cc0a4c304e3

Observation dc72a2a9-7eda-4695-8aeb-59ff29e0fce6 · outbound

This paper cites When does return-conditioned supervised learning work for offline reinforcement learning? Advances in Neural Information Processing Systems, 35: 0 1542--1553, 2022.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization When does return-conditioned supervised learning work for offline reinforcement learning? Advances in Neural Information Processing Systems, 35: 0 1542--1553, 2022

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:27.100466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.747851Z digest=sha256:e75f833dfb0fde4e34e1853ea14585ccd615208777ae2dff6efa5718981058cc

Observation 1900dc52-3d95-4889-bea5-aa5ba811b9a9 · outbound

This paper cites Mamba as Decision Maker: Exploring Multi-scale Sequence Modeling in Offline Reinforcement Learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Mamba as Decision Maker: Exploring Multi-scale Sequence Modeling in Offline Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.752674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.752674Z digest=sha256:7e98a9fa044aef80db21ca03aed7db4b1e2fcf17248d30460682b2e19ba1dd56

Observation da09bc93-76e4-4615-9e8e-9c8fe5094a62 · outbound

This paper cites Goal-conditioned reinforcement learning with imagined subgoals.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Goal-conditioned reinforcement learning with imagined subgoals

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:27.088998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.757238Z digest=sha256:689108509e2e533d865df2718e9a7c3ac9f668cd4ecac09f60c548d1081b9de8

Observation d46ca468-7388-4700-9cb4-81b704157246 · outbound

This paper cites On the statistical benefits of temporal difference learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization On the statistical benefits of temporal difference learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:27.073974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.761572Z digest=sha256:c4ebe1af2fb30e6585a30ec68266346d7de1be07831217014c1342769cbb85be

Observation 05f055f6-a886-4d96-9ed4-26e1783a705a · outbound

This paper cites Decision transformer: Reinforcement learning via sequence modeling.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Decision transformer: Reinforcement learning via sequence modeling

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.766013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.766013Z digest=sha256:90f7d6386db5e843973f109db2654e4b349b04bd9454688ecb03ea07901cf3f4

Observation 40ee4e50-4a6d-48ab-a17f-d624539b2950 · outbound

This paper cites Goal-conditioned imitation learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Goal-conditioned imitation learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.770153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.770153Z digest=sha256:cf898daa2eb609ddfe28131fb92a742b5223a5f1c88582ac7397faa0d9f86d60

Observation b26c610a-459b-44af-9acb-69a25dc55670 · outbound

This paper cites RvS: What is Essential for Offline RL via Supervised Learning?.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization RvS: What is Essential for Offline RL via Supervised Learning?

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.775369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.775369Z digest=sha256:1df5855c431164dcaddbb8042440590a0dcf9eb1e93062a75a0c9c55313a2a68

Observation bbad15d8-706f-4444-9661-dcec1aa5946d · outbound

This paper cites C-Learning: Learning to Achieve Goals via Recursive Classification.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization C-Learning: Learning to Achieve Goals via Recursive Classification

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.785189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.785189Z digest=sha256:40985ce7647b4e702886065beab991cf744c2ba4e4b4ff6ff406ace59320abda

Observation 6ffaf888-2f16-4a8a-8ffd-9d2ed56273b1 · outbound

This paper cites Imitating past successes can be very suboptimal.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Imitating past successes can be very suboptimal

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:27.041974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.797884Z digest=sha256:a1c27385442320eb900881287b65467d68543b8836af63c76d4439beaa512ee9

Observation 1b94d514-e345-483e-ab0f-72493bc2707a · outbound

This paper cites Contrastive learning as goal-conditioned reinforcement learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Contrastive learning as goal-conditioned reinforcement learning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:27.026589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.806254Z digest=sha256:03c99841749393bfe4fb42b1e141c2dbd6fea0b155db16f8536a24f4a1cbb388

Observation 8408698a-ec22-445f-8d72-8d4e4f62cfc2 · outbound

This paper cites Inference via interpolation: Contrastive representations provably enable planning and inference.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Inference via interpolation: Contrastive representations provably enable planning and inference

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:27.009713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.812950Z digest=sha256:531bec51686e2507ff66a9aec2263c4f22d3e472bc3de095fc230aaec03bfdf1

Observation 4e9db73b-ece7-43c6-9a15-e5a3a1b6028b · outbound

This paper cites D4RL: Datasets for Deep Data-Driven Reinforcement Learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization D4RL: Datasets for Deep Data-Driven Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.820597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.820597Z digest=sha256:468fab3e8ac88f4da08e88d95ddb46d08088d986d9eee3cef08d79dc02da00c7

Observation 6833a941-aa87-4b87-b910-42a1c54cfb00 · outbound

This paper cites A minimalist approach to offline reinforcement learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization A minimalist approach to offline reinforcement learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.828824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.828824Z digest=sha256:3bbe1d661183fb2b29f09fadaceb4a02312b2cdcfff7e99065754be6385612b6

Observation 68403ab2-45bf-4103-8b10-8bdc7ef7f3bc · outbound

This paper cites Learning to reach goals via iterated supervised learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Learning to reach goals via iterated supervised learning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.986664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.836914Z digest=sha256:6a13e061c3950dda25af0cda1a4a7b70ae78251158539a3cb4d56e811d41e491

Observation 2e8b731c-3160-440d-9f66-c4cc18ed1276 · outbound

This paper cites Closing the gap between TD learning and supervised learning - a generalisation point of view.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Closing the gap between TD learning and supervised learning - a generalisation point of view

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.973047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.845241Z digest=sha256:76f2f1d6eab9753eb638acbb5f61698065c8cb61f2c6872bf7829b21384693fa

Observation 36e41fa6-edff-4f22-af06-65760ef38fe7 · outbound

This paper cites Distance Weighted Supervised Learning for Offline Interaction Data.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Distance Weighted Supervised Learning for Offline Interaction Data

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.851471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.851471Z digest=sha256:04fa2bf6afe081477e094ae5ed75a56f6c3b2d558a3dd7d309e29cff6d1b71d2

Observation f7e355dd-8db1-498a-bbed-21c4b1f7bf24 · outbound

This paper cites Diffused task-agnostic milestone planner.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Diffused task-agnostic milestone planner

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.960312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.856909Z digest=sha256:d0bf57c59328d97a1633a395e2e7a1c3ffe7a8e62d1e1f06114041fc5158770f

Observation 77b97c48-a143-4bf1-abcd-26486fc6b117 · outbound

This paper cites Decision Mamba: Reinforcement Learning via Hybrid Selective Sequence Modeling.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Decision Mamba: Reinforcement Learning via Hybrid Selective Sequence Modeling

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:05:26.394932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.860929Z digest=sha256:a7d901ef5014c06713d755092acdaaf8fac60d95ee3cb6ea07024f8e615c4068

Observation d2b47e03-40b7-4d65-91e9-1887d2c8aa79 · outbound

This paper cites Learning to reach goals via diffusion.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Learning to reach goals via diffusion

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.949447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.865082Z digest=sha256:222300c4dde3314fc70d0eabbd6d47baca6099763847a4f7e45d8c28c020a795

Observation 95e9d632-ee99-466d-abfc-0287ee48a294 · outbound

This paper cites Efficient planning in a compact latent action space.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Efficient planning in a compact latent action space

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.938587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.869691Z digest=sha256:af3f0d2f3ae227fe5f5b3d1bed2fccdf78c11b7e42340a451bc7ad6e53d331dd

Observation 333012e5-a28a-43fa-9fbd-f51f78732e31 · outbound

This paper cites Adaptive q -aid for conditional supervised learning in offline reinforcement learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Adaptive q -aid for conditional supervised learning in offline reinforcement learning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.927879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.874197Z digest=sha256:0424e104d410601b80f55dca46849a7725bee15fdefadabd8df90611e00ce622

Observation 7cb756c9-6bfd-4cd5-993b-d1ddc942d258 · outbound

This paper cites Imitating Graph-Based Planning with Goal-Conditioned Policies.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Imitating Graph-Based Planning with Goal-Conditioned Policies

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.878623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.878623Z digest=sha256:d0b5d0a6f794b0e0246ec5d58742e1a2134c9e19cc93e5926d67a8dbacb4ea5c

Observation 991722db-63c3-41eb-83da-9416adc5569c · outbound

This paper cites Adam: A method for stochastic optimization.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Adam: A method for stochastic optimization

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.916192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.883186Z digest=sha256:4003e5038ec1e0db6ec988c4e2459a0e41dfe814d3532faae2e359995a9e8284

Observation 3f22aaa2-4fed-44ab-a087-56d7c8e1c94a · outbound

This paper cites An Introduction to Variational Autoencoders.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization An Introduction to Variational Autoencoders

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.887182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.887182Z digest=sha256:9eafa4e2965709e5bb04e58a1d6e57b39880ec75617025387a1f00f58bd4e955

Observation 0d1a09a9-e890-4949-a7ad-ce801bfbbe15 · outbound

This paper cites Offline Reinforcement Learning with Implicit Q-Learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Offline Reinforcement Learning with Implicit Q-Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.890730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.890730Z digest=sha256:52b5ad9fa3cfc7e91fa6c6ce59f014b9a08d98ef70d38fff136a01c4ca238d84

Observation 6cac748e-1609-4fa2-8267-547de17c0893 · outbound

This paper cites Stabilizing off-policy q-learning via bootstrapping error reduction.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Stabilizing off-policy q-learning via bootstrapping error reduction

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.894463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.894463Z digest=sha256:d729063fdf865a020ed548e47a4c9c96ed2400a798a4654624979837f1b5f82e

Observation e129b63b-db0d-4810-b0ba-659543a77553 · outbound

This paper cites Conservative q-learning for offline reinforcement learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Conservative q-learning for offline reinforcement learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.898096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.898096Z digest=sha256:6dcb5193543c268d299bd50a97d8ccbd949aa06e6337f389e6444c27d712f7bc

Observation 298c4dee-8b0d-474e-afad-c57e481b4223 · outbound

This paper cites Multi-game decision transformers.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Multi-game decision transformers

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.890075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.901393Z digest=sha256:663895ad1918f583f299badce79da19534d195fa99d6aa353842c93caacebf1d

Observation 4da2b345-5e0d-4d95-a039-aed07cae0242 · outbound

This paper cites Dhrl: a graph-based approach for long-horizon and sparse hierarchical reinforcement learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Dhrl: a graph-based approach for long-horizon and sparse hierarchical reinforcement learning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.877877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.904989Z digest=sha256:d09b910cabd7bb70658232c8c18f67eade454117800f8acb736a896f4142f50e

Observation d0173034-cc60-47ef-acb7-cf86d44b445c · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.909047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.909047Z digest=sha256:959ca39ebbb5b04fda58ad227b2fbe5b9f8f50a4ffb901d795bb261f414e9929

Observation 338ffa32-a5ea-4dcf-8a0d-aae0f181c06a · outbound

This paper cites Hierarchical planning through goal-conditioned offline reinforcement learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Hierarchical planning through goal-conditioned offline reinforcement learning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.865153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.912606Z digest=sha256:40d5d74db26531a2ec7aa301890e3619f9b0f325a253d6661f3abecfeea65897

Observation 2b3f252c-1f37-403b-ba3f-b2186e6ef988 · outbound

This paper cites DIDI: Diffusion-Guided Diversity for Offline Behavioral Generation.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization DIDI: Diffusion-Guided Diversity for Offline Behavioral Generation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.916444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.916444Z digest=sha256:b0a41caeeef1798c03d277955197f826df8345f684f552a8c307d35e232be513

Observation ab0ce511-34ba-4e69-b886-0908976a8819 · outbound

This paper cites Beyond ood state actions: Supported cross-domain offline reinforcement learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Beyond ood state actions: Supported cross-domain offline reinforcement learning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.850298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.920305Z digest=sha256:b4085fa664172163c86de4b38263717df328239a562788e450cc31b9e333186e

Observation d1625f51-008d-47f1-a75b-80e9ec423a6f · outbound

This paper cites Least squares quantization in pcm.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Least squares quantization in pcm

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.924152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.924152Z digest=sha256:f0937a119f64cd027c6eba9938ae023fd60302b0dcb68b95d326474afc5b8111

Observation 05f6f010-fa26-43cc-9226-12825383213b · outbound

This paper cites Decoupled Weight Decay Regularization.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Decoupled Weight Decay Regularization

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.928161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.928161Z digest=sha256:e3391f7fe453f70d24b3f70b1eab487e0833d5c2a5d924a3dcc09d2ae4977bd6

Observation 2a9bff04-ecfd-42a1-be87-47952d4c0d07 · outbound

This paper cites Generative Trajectory Stitching through Diffusion Composition.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Generative Trajectory Stitching through Diffusion Composition

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.931670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.931670Z digest=sha256:fae7d49905f0c527332cc7ee0dbed7d9fde645c24124b4cbe2d89d4a83d331f0

Observation 0ee3c592-d739-4461-ae0c-81d25979e657 · outbound

This paper cites Decision mamba: A multi-grained state space model with self-evolution regularization for offline rl.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Decision mamba: A multi-grained state space model with self-evolution regularization for offline rl

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.823899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.936092Z digest=sha256:bdc3474e515b36cc9f0146c4f567221e0235a75c7739a9f531b553e41c62313d

Observation 9161b203-157a-43c2-a75a-e6c58b3bf523 · outbound

This paper cites Learning latent plans from play.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Learning latent plans from play

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.807672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.939663Z digest=sha256:1ab226cae8c10bf6f70dd26e1dd8ecc88bb502df37536a1c02d0e6f0b283c696

Observation 7760b84f-ef4a-4c07-b158-dd9ed90e7346 · outbound

This paper cites How Far I'll Go: Offline Goal-Conditioned Reinforcement Learning via $f$-Advantage Regression.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization How Far I'll Go: Offline Goal-Conditioned Reinforcement Learning via $f$-Advantage Regression

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.943342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.943342Z digest=sha256:3902d1ca60ad8953ae97ac7be89c0391aee659ebb87547953932296869c69e42

Observation c121fe31-7789-4cdd-bd85-6da077e82bd2 · outbound

This paper cites VIP : Towards universal visual reward and representation via value-implicit pre-training.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization VIP : Towards universal visual reward and representation via value-implicit pre-training

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.787378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.947029Z digest=sha256:bd5d4ee78555de5ed9f8f7f5d8a8b8830838192c9a650cf6c8d11b7bf3aa43e7

Observation 81535569-6821-4cf1-b589-a68adfbf95cf · outbound

This paper cites Human-level control through deep reinforcement learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Human-level control through deep reinforcement learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.950640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.950640Z digest=sha256:013a198e6221c3e85c0bf442e6842b2c3d1f0111372090900bc7d5da5de6bacd

Observation b836f8df-5d17-418a-9393-f6a1a8ccaa13 · outbound

This paper cites Learning Temporal Distances: Contrastive Successor Features Can Provide a Metric Structure for Decision-Making.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Learning Temporal Distances: Contrastive Successor Features Can Provide a Metric Structure for Decision-Making

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.954383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.954383Z digest=sha256:82e6b9d6d4351037f1feef1d1aad36f2e89c9c5025431356d746ba5bbb21c097

Observation d110249c-705f-4d70-a3d4-ceec22004d39 · outbound

This paper cites Asymmetric least squares estimation and testing.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Asymmetric least squares estimation and testing

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.763153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.958453Z digest=sha256:ffbc79724b7e191960f5fbede09a86e48f6ffc8f206aefacc34b0047abf06ab5

Observation 6fff31ef-0216-4c08-a905-3cc793992213 · outbound

This paper cites Decision Mamba: Reinforcement Learning via Sequence Modeling with Selective State Spaces.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Decision Mamba: Reinforcement Learning via Sequence Modeling with Selective State Spaces

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.961759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.961759Z digest=sha256:c4438ce32e909a25bd31ce08df62b4317efceb6965b9f11fbfb740beb0b92983

Observation 49179c0b-5541-4a55-b2de-51de3d7635b3 · outbound

This paper cites Hiql: Offline goal-conditioned rl with latent states as actions.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Hiql: Offline goal-conditioned rl with latent states as actions

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.966571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.966571Z digest=sha256:15e571019e39874791ec7ce8d18a804e0756fdc0b540ee70f9b4fc6d89d90aea

Observation f5f5575a-0a9d-432b-9163-2d9147e6f3df · outbound

This paper cites Foundation Policies with Hilbert Representations.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Foundation Policies with Hilbert Representations

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.970149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.970149Z digest=sha256:df0fe62a730e5a5fe445d1347ee8d8b935c318f4931a285af93d01306309bf99

Observation 0b6d3909-af36-4f36-b8d2-07deef469bd8 · outbound

This paper cites Ogbench: Benchmarking offline goal-conditioned rl.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Ogbench: Benchmarking offline goal-conditioned rl

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.739722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.975726Z digest=sha256:fdef3c706590432ff2498748fa8eb0de927e137aab19158f59215004764add93

Observation 03cde593-9e7a-4b38-b8c0-d2b09ffe2fe9 · outbound

This paper cites A survey on offline reinforcement learning: Taxonomy, review, and open problems.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization A survey on offline reinforcement learning: Taxonomy, review, and open problems

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.727985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.980393Z digest=sha256:b36ba53d02da34dcf87694322300d671bc0a615046fc88cdf714c1a4a17fe8fe

Observation 738e246b-af48-488e-bd3d-3f9c6928ce3d · outbound

This paper cites Goal-Conditioned Imitation Learning using Score-based Diffusion Policies.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Goal-Conditioned Imitation Learning using Score-based Diffusion Policies

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.984273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.984273Z digest=sha256:14321daa55ce90c70c66b8bfe98d2174c9be168ecfed75608ab0fe4960b247d4

Observation 0047489f-1d2b-4ace-91ba-3d824f4a0da4 · outbound

This paper cites Stochastic backpropagation and approximate inference in deep generative models.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Stochastic backpropagation and approximate inference in deep generative models

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.716004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.988205Z digest=sha256:694a01c1ec89e377fc184e4a778b82b3c4a3b036d107501c11620ba68acfcdd6

Observation f43cd76c-e99b-47ef-a3bc-6bd55c4b6680 · outbound

This paper cites Outcome-driven reinforcement learning via variational inference.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Outcome-driven reinforcement learning via variational inference

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.704287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.994354Z digest=sha256:fa65dedfa16ce8e97be1a5ca1b75ca1ce574d3a5e6ade9741e03f113c0e15847

Observation 5d75398f-09e4-484f-a1a0-d8b10e8421f6 · outbound

This paper cites Reinforcement learning upside down: Don't predict rewards -- just map them to actions, 2020.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Reinforcement learning upside down: Don't predict rewards -- just map them to actions, 2020

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.692774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.998229Z digest=sha256:0568176da6cf89a74f16d7ca93977ca59b3f990241aa37de33c6a98a96601786

Observation a0ad97e9-02cb-4edd-8211-cb21fa4dab4d · outbound

This paper cites Rapid exploration for open-world navigation with latent goal models.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Rapid exploration for open-world navigation with latent goal models

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.681724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:05:26.002289Z digest=sha256:430732b605a7bd845b40174e94a8c11fb08aa97fdb55bbab2369afc755dca9ef

Observation dfc8bf9e-530b-400e-a12c-6bd0d5fc4e6f · outbound

This paper cites Score models for offline goal-conditioned reinforcement learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Score models for offline goal-conditioned reinforcement learning

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.670794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:05:26.006634Z digest=sha256:b8720a92c396bba96f35dc7d68f605c9daa008ce4bf8dcd82d72514ca6ed8dc0

Observation a9e70b2b-72cf-44f6-996f-cfd0632dad84 · outbound

This paper cites Geoadditive expectile regression.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Geoadditive expectile regression

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.659954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:05:26.010660Z digest=sha256:bbeac998d5bcfb34ed696983336aedeb619dbeda9382d9222cc2f147cea65f5b

Observation d9c00165-7d4f-4aeb-9178-49005ea0b73e · outbound

This paper cites Learning structured output representation using deep conditional generative models.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Learning structured output representation using deep conditional generative models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:26.014377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:26.014377Z digest=sha256:71fb1e733ed2ebb50f0109502e1555f986391785c5e73ea34af3b397bbb1c474

Observation 31cf407e-cc0a-4685-ae51-73d4d982f234 · outbound

This paper cites Gymnasium (mar 2023), 2023.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Gymnasium (mar 2023), 2023

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.639864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:05:26.018409Z digest=sha256:8045d15b576cb6715d7ace342cee832964f19591a54c5fa41b7e42cb65caa89a

Observation 15c15ca4-5ecf-46a0-abb1-078650721ce1 · outbound

This paper cites Deep Reinforcement Learning and the Deadly Triad.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Deep Reinforcement Learning and the Deadly Triad

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:26.022196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:26.022196Z digest=sha256:14316b30e3f8b184713f1e7a06a9e8cf5460bedafe4fb8ef74b86744289b1475

Observation 7e47b5a1-f2c6-4aec-b59c-e20c712aadcc · outbound

This paper cites Attention is all you need.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Attention is all you need

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:26.026210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:26.026210Z digest=sha256:7a37f09a8de21b499c34cbb564902cebcd4f78b9a237a27ba9d40e7285a01ece

Observation c20b3ea7-d6cd-4f80-a058-c02026c8f25b · outbound

This paper cites GOP lan: Goal-conditioned offline reinforcement learning by planning with learned models.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization GOP lan: Goal-conditioned offline reinforcement learning by planning with learned models

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.619349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:05:26.029722Z digest=sha256:299fa1e28b0e5c33ac77a0f7b8f2f77dc53e46d17b1f8fefc1db871b25be6d6f

Observation 34fa6ff0-f0f9-4b0f-bac5-f5c5f9efce2b · outbound

This paper cites Optimal goal-reaching reinforcement learning via quasimetric learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Optimal goal-reaching reinforcement learning via quasimetric learning

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.607776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:05:26.033672Z digest=sha256:13897b7a3d1a7737eac5619f9aa48090820ca41567c3e624cdc270095b4c465e

Observation 4064caf9-0460-452a-a414-d3ea921924ee · outbound

This paper cites Critic-guided decision transformer for offline reinforcement learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Critic-guided decision transformer for offline reinforcement learning

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.596625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:05:26.038097Z digest=sha256:62c7c53a9025dc9571a234a62a6596be694d4603725fb8064f9332664e4a6066

Observation 1f240bd1-3f82-44a7-900f-bb0dcb6eb6ec · outbound

This paper cites Supported policy optimization for offline reinforcement learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Supported policy optimization for offline reinforcement learning

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.584880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:05:26.041485Z digest=sha256:61679ab9751a0845213adbc1876fc9924e5ba9c6c260f8fc5aa83de1ef3781cb

Observation e8e42e3d-1cdb-42c8-9228-c14463d124a2 · outbound

This paper cites Elastic Decision Transformer.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Elastic Decision Transformer

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:26.045615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:26.045615Z digest=sha256:b9b0e2615f631bce04791bf3756b7192ed82eaee67571543d19bf5b2e03c6954

Observation ca8c62b3-2b08-4809-9704-568102ecae9d · outbound

This paper cites Q-learning decision transformer: Leveraging dynamic programming for conditional sequence modelling in offline rl.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Q-learning decision transformer: Leveraging dynamic programming for conditional sequence modelling in offline rl

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.573852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:05:26.050121Z digest=sha256:24817c3320f760de16457ab75103a44f1ab7b7ee61daed1b96c96ce0725b489a

Observation 158ff2de-8365-47d0-809c-dfad1cc405e2 · outbound

This paper cites Rethinking Goal-conditioned Supervised Learning and Its Connection to Offline RL.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Rethinking Goal-conditioned Supervised Learning and Its Connection to Offline RL

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:26.054551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:26.054551Z digest=sha256:c9cab7298687c70e43ea84aa136389127f680265a0988447e36e87631307befc

Observation db0fa5c6-5581-4b25-87fb-013fcef04677 · outbound

This paper cites What is essential for unseen goal generalization of offline goal-conditioned rl? In International Conference on Machine Learning, pages 39543--39571.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization What is essential for unseen goal generalization of offline goal-conditioned rl? In International Conference on Machine Learning, pages 39543--39571

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.561418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:05:26.058489Z digest=sha256:ac02cc59f25c3d037b1b004bde0640cb2431e4a335b9b000c734679a62f63b3f

Observation 77db9b91-9fca-46a6-9296-56fca0933486 · outbound

This paper cites Swapped goal-conditioned offline reinforcement learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Swapped goal-conditioned offline reinforcement learning

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:26.062222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:26.062222Z digest=sha256:708dde6154386ce49fef973afb7d11433b73da787d6ecff40cc19a98cb687181

Observation 2f8dd9a5-e9c1-4d76-af4a-b181000ab2fb · outbound

This paper cites Breadth-first exploration on adaptive grid for reinforcement learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Breadth-first exploration on adaptive grid for reinforcement learning

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:26.066117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:26.066117Z digest=sha256:4af63ea8f4544687df13a507774cfd8f381fed6e8a8d794fe74dae6c4632ce30

Observation 7fbcd7cd-bb20-4774-be4e-c2841fc8cbee · outbound

This paper cites Goal-conditioned predictive coding for offline reinforcement learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Goal-conditioned predictive coding for offline reinforcement learning

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.539901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:05:26.069683Z digest=sha256:14cfb1dbba06f90d3ce8f98ee7874969c93b1d8c1e531dc2168f2975a78ec553

Observation b155832e-295f-432b-bfcd-1de1323f54b8 · outbound

This paper cites Stabilizing Contrastive RL: Techniques for Robotic Goal Reaching from Offline Data.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Stabilizing Contrastive RL: Techniques for Robotic Goal Reaching from Offline Data

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:26.072975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:26.072975Z digest=sha256:dc05dbd7873448c27cc1a5b53fa46ec0da26970040fbf3089235e8992b1a1bac

Observation 1b355547-cf41-4fd4-860b-354833c8961b · outbound

This paper cites Contrastive difference predictive coding.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Contrastive difference predictive coding

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.526281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:05:26.076893Z digest=sha256:5b656e01415f2174e6f0b00ecc1dddba5ebf3064c70ae30d723f8597579d7c6c

Observation 1007d93e-a759-4670-a9ae-ce0f34545b99 · outbound

This paper cites Online decision transformer.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Online decision transformer

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.514001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:05:26.081476Z digest=sha256:254f49655968046c899c646e22075584227117130229a3482ff3ee9019e1dbcc

Observation 2de92e6a-ec70-4159-9a24-cf51af5f8e47 · outbound

This paper cites Behavior Proximal Policy Optimization.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Behavior Proximal Policy Optimization

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:26.085023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:26.085023Z digest=sha256:5a0bf55e9e41e3e960e7e7de9aa4d9b50e728b47902072c347a1ff931fb7fc2e

Observation 1267e842-1074-4441-b705-b7277e3483cd · outbound

This paper cites Reinformer: Max-Return Sequence Modeling for Offline RL.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Reinformer: Max-Return Sequence Modeling for Offline RL

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:26.091332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:26.091332Z digest=sha256:75277beb1f5c7a21cad6d1d58e1deca03935f5ccc030656e1164d3684e313500

Observation ad6054d5-8274-4dd4-b327-6f158fb0db69 · outbound

This paper cites Revisiting the design choices in max-return sequence modeling, 2025.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Revisiting the design choices in max-return sequence modeling, 2025

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.500024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:05:26.095894Z digest=sha256:88c75abc0b0cb584ea80eb3813cc1f6047868649215a72046cb066a519f85b46

Observation fc578017-cce3-4fd5-bb09-46cadeecab2d · outbound

This paper cites Maximum entropy inverse reinforcement learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Maximum entropy inverse reinforcement learning

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:26.101533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:26.101533Z digest=sha256:039a0bdb788066b6e46e6f9d386e2c2d871c64b463124df43e37f75779164afe

Pith citing papers

No inbound Pith citation observations are available.