Pith. sign in

Paper Citation Record · LEDGER

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization

As of 10 August 2026, this Paper Citation Record lists 83 of 83 outbound references and 0 inbound Pith citation observations for arXiv:2506.00795.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00795 v3

Coverage vector

measured 83 of 83 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:05:26.101533Z

measured 83 of 83 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

83 of 83 outbound references displayed

  • verified exact1
  • verified fuzzy41
  • unresolved41
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ac9d45cc-7a9f-41d9-b013-0c0be567dee7 · outbound

This paper cites Deep reinforcement learning at the edge of the statistical precipice.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Deep reinforcement learning at the edge of the statistical precipice

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.717787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.717787Z digest=sha256:86d82ed7bcfb1de0a86350496b2c4e4ed78ed5e8cb6547fc42a276e09d8ab842

Observation 0c629fa2-b676-4ebf-a403-7d3f06ebde26 · outbound

This paper cites On the estimation of production frontiers: maximum likelihood estimation of the parameters of a discontinuous density function.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization On the estimation of production frontiers: maximum likelihood estimation of the parameters of a discontinuous density function

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.723026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.723026Z digest=sha256:462ac4b81ef6cf12598991ac6a7e09a8b4ab4467957aefb2e6faea1c2f1a19a8

Observation a6022746-c35d-493e-ab30-e9d86c9fab73 · outbound

This paper cites Hindsight experience replay.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Hindsight experience replay

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.728070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.728070Z digest=sha256:0e3c8d41adea8011cb6e3b8831434df191aa113570a9fc0d640b8a8016caae61

Observation 0eec7250-8f73-4841-8163-0584315efc57 · outbound

This paper cites Learning Successor States and Goal-Dependent Values: A Mathematical Viewpoint.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Learning Successor States and Goal-Dependent Values: A Mathematical Viewpoint

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.733026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.733026Z digest=sha256:1bae6cb4549d35a97dbd67da3f0319744db91b23d374b607b1d2f637dbae4694

Observation d0775145-8232-416d-ad3c-985501eafa1c · outbound

This paper cites Accelerating goal-conditioned reinforcement learning algorithms and research.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Accelerating goal-conditioned reinforcement learning algorithms and research

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:27.123474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.737574Z digest=sha256:2502f80cf032916a5d5e293f4a56c3e2cde767272c03f688a29d4f14541becf7

Observation fb9f4169-e126-4d26-8760-662962c2a265 · outbound

This paper cites Offline rl without off-policy evaluation.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Offline rl without off-policy evaluation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.742340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.742340Z digest=sha256:5fb4370ffd99df97c24d758e9b6f6b084a5f181c594e3bc18af65cc0a4c304e3

Observation dc72a2a9-7eda-4695-8aeb-59ff29e0fce6 · outbound

This paper cites When does return-conditioned supervised learning work for offline reinforcement learning? Advances in Neural Information Processing Systems, 35: 0 1542--1553, 2022.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization When does return-conditioned supervised learning work for offline reinforcement learning? Advances in Neural Information Processing Systems, 35: 0 1542--1553, 2022

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:27.100466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.747851Z digest=sha256:068981f60652840ca2ee7ed4ffd40bb1c5af3fe059e07948c62e1fe55413fea2

Observation 1900dc52-3d95-4889-bea5-aa5ba811b9a9 · outbound

This paper cites Mamba as Decision Maker: Exploring Multi-scale Sequence Modeling in Offline Reinforcement Learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Mamba as Decision Maker: Exploring Multi-scale Sequence Modeling in Offline Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.752674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.752674Z digest=sha256:5eadd1b0df76f7a804bdd4fef1523740cfc1b796a6173afd03788cc6c681dbb5

Observation da09bc93-76e4-4615-9e8e-9c8fe5094a62 · outbound

This paper cites Goal-conditioned reinforcement learning with imagined subgoals.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Goal-conditioned reinforcement learning with imagined subgoals

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:27.088998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.757238Z digest=sha256:5e0f78cadd150d9f52f9c3061dba40fb776e48b83a6ee59dd8bd35a2ede0767f

Observation d46ca468-7388-4700-9cb4-81b704157246 · outbound

This paper cites On the statistical benefits of temporal difference learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization On the statistical benefits of temporal difference learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:27.073974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.761572Z digest=sha256:57ed5d9f79675bde58d7883e60017ca6e09f5dc5e5717f47be26bf9e1b78fc37

Observation 05f055f6-a886-4d96-9ed4-26e1783a705a · outbound

This paper cites Decision transformer: Reinforcement learning via sequence modeling.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Decision transformer: Reinforcement learning via sequence modeling

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.766013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.766013Z digest=sha256:90f7d6386db5e843973f109db2654e4b349b04bd9454688ecb03ea07901cf3f4

Observation 40ee4e50-4a6d-48ab-a17f-d624539b2950 · outbound

This paper cites Goal-conditioned imitation learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Goal-conditioned imitation learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.770153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.770153Z digest=sha256:cf898daa2eb609ddfe28131fb92a742b5223a5f1c88582ac7397faa0d9f86d60

Observation b26c610a-459b-44af-9acb-69a25dc55670 · outbound

This paper cites RvS: What is Essential for Offline RL via Supervised Learning?.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization RvS: What is Essential for Offline RL via Supervised Learning?

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.775369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.775369Z digest=sha256:1236d69cc387aa7ef640915d85ef2aca64ab15a1a4951f22974a3e0dda03854d

Observation bbad15d8-706f-4444-9661-dcec1aa5946d · outbound

This paper cites C-Learning: Learning to Achieve Goals via Recursive Classification.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization C-Learning: Learning to Achieve Goals via Recursive Classification

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.785189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.785189Z digest=sha256:3b8d4dea5cd937cb3c0de35f85aedc4234cce51997c6f79ea97a85842e98d513

Observation 6ffaf888-2f16-4a8a-8ffd-9d2ed56273b1 · outbound

This paper cites Imitating past successes can be very suboptimal.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Imitating past successes can be very suboptimal

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:27.041974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.797884Z digest=sha256:ce073162c4b9052fdeb617d288277e96e6e68334595e7bde79db72c6d17c096b

Observation 1b94d514-e345-483e-ab0f-72493bc2707a · outbound

This paper cites Contrastive learning as goal-conditioned reinforcement learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Contrastive learning as goal-conditioned reinforcement learning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:27.026589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.806254Z digest=sha256:e48a52befcff97dfcf1a0c8bb4a4d5d9b1781575ac6c81c8d1d4c44373d851b1

Observation 8408698a-ec22-445f-8d72-8d4e4f62cfc2 · outbound

This paper cites Inference via interpolation: Contrastive representations provably enable planning and inference.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Inference via interpolation: Contrastive representations provably enable planning and inference

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:27.009713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.812950Z digest=sha256:e42acc4cb86044104e8308b6c66272d73f8930ef3ef6e0de75151b691bc9d526

Observation 4e9db73b-ece7-43c6-9a15-e5a3a1b6028b · outbound

This paper cites D4RL: Datasets for Deep Data-Driven Reinforcement Learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization D4RL: Datasets for Deep Data-Driven Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.820597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.820597Z digest=sha256:86f35efe837ef7b2b02514d6677b2d4a3e3b59b6e67a35f4b34b9afa52e7d6f4

Observation 6833a941-aa87-4b87-b910-42a1c54cfb00 · outbound

This paper cites A minimalist approach to offline reinforcement learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization A minimalist approach to offline reinforcement learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.828824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.828824Z digest=sha256:3bbe1d661183fb2b29f09fadaceb4a02312b2cdcfff7e99065754be6385612b6

Observation 68403ab2-45bf-4103-8b10-8bdc7ef7f3bc · outbound

This paper cites Learning to reach goals via iterated supervised learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Learning to reach goals via iterated supervised learning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.986664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.836914Z digest=sha256:2b2e61ea42a24f2e083370bb3b9c1716fc40601828001635f4592adbe09dcf49

Observation 2e8b731c-3160-440d-9f66-c4cc18ed1276 · outbound

This paper cites Closing the gap between TD learning and supervised learning - a generalisation point of view.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Closing the gap between TD learning and supervised learning - a generalisation point of view

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.973047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.845241Z digest=sha256:768da47ae208233dd6f545b4dd649d1c14cafb546dae5c3e758f9881b478c38d

Observation 36e41fa6-edff-4f22-af06-65760ef38fe7 · outbound

This paper cites Distance Weighted Supervised Learning for Offline Interaction Data.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Distance Weighted Supervised Learning for Offline Interaction Data

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.851471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.851471Z digest=sha256:442947647e5848d48e1345ee95f31013d781f55c9cd404f8561180e47e43939c

Observation f7e355dd-8db1-498a-bbed-21c4b1f7bf24 · outbound

This paper cites Diffused task-agnostic milestone planner.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Diffused task-agnostic milestone planner

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.960312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.856909Z digest=sha256:0181c6ed5f46352cc8b2b7a63ce144c3201be181f8c065b15b66b7b3319b963a

Observation 77b97c48-a143-4bf1-abcd-26486fc6b117 · outbound

This paper cites Decision Mamba: Reinforcement Learning via Hybrid Selective Sequence Modeling.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Decision Mamba: Reinforcement Learning via Hybrid Selective Sequence Modeling

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:05:26.394932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.860929Z digest=sha256:4f1dd4b57512d35e96b469ac99cea7382ff68329955e2e0cb51aee90ade0ca9a

Observation d2b47e03-40b7-4d65-91e9-1887d2c8aa79 · outbound

This paper cites Learning to reach goals via diffusion.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Learning to reach goals via diffusion

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.949447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.865082Z digest=sha256:17a2b8d4fc123794603397f9c6f54fb023d6697f78bdcdce5e77abaa905706a9

Observation 95e9d632-ee99-466d-abfc-0287ee48a294 · outbound

This paper cites Efficient planning in a compact latent action space.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Efficient planning in a compact latent action space

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.938587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.869691Z digest=sha256:fbfe66cdef5ab078fae771892969000c8ebd0cde1c313c856bdc99c2d9398919

Observation 333012e5-a28a-43fa-9fbd-f51f78732e31 · outbound

This paper cites Adaptive q -aid for conditional supervised learning in offline reinforcement learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Adaptive q -aid for conditional supervised learning in offline reinforcement learning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.927879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.874197Z digest=sha256:a4b0db325fa38b2fba1f3a86e10c55dda8b212eebbef37516155b40e6e1051b2

Observation 7cb756c9-6bfd-4cd5-993b-d1ddc942d258 · outbound

This paper cites Imitating Graph-Based Planning with Goal-Conditioned Policies.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Imitating Graph-Based Planning with Goal-Conditioned Policies

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.878623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.878623Z digest=sha256:141d905eb082f2f416ae8760922b2a90fc5ec02fb8f5473b9fab590163053f69

Observation 991722db-63c3-41eb-83da-9416adc5569c · outbound

This paper cites Adam: A method for stochastic optimization.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Adam: A method for stochastic optimization

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.916192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.883186Z digest=sha256:3e7d40bee1d17cf3d869c6b95b16d7b4741e9487883b2f14dc502056514ada69

Observation 3f22aaa2-4fed-44ab-a087-56d7c8e1c94a · outbound

This paper cites An Introduction to Variational Autoencoders.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization An Introduction to Variational Autoencoders

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.887182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.887182Z digest=sha256:5bb8bc3f529b7ee07572dbfef7a6c5112846ab6edfa2fa48f8bc76fe0036c753

Observation 0d1a09a9-e890-4949-a7ad-ce801bfbbe15 · outbound

This paper cites Offline Reinforcement Learning with Implicit Q-Learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Offline Reinforcement Learning with Implicit Q-Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.890730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.890730Z digest=sha256:64ff3a2f7350f0f8e9300a3c21a35977ca910c00c9eab5469b77ae38dc19b747

Observation 6cac748e-1609-4fa2-8267-547de17c0893 · outbound

This paper cites Stabilizing off-policy q-learning via bootstrapping error reduction.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Stabilizing off-policy q-learning via bootstrapping error reduction

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.894463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.894463Z digest=sha256:d729063fdf865a020ed548e47a4c9c96ed2400a798a4654624979837f1b5f82e

Observation e129b63b-db0d-4810-b0ba-659543a77553 · outbound

This paper cites Conservative q-learning for offline reinforcement learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Conservative q-learning for offline reinforcement learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.898096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.898096Z digest=sha256:6dcb5193543c268d299bd50a97d8ccbd949aa06e6337f389e6444c27d712f7bc

Observation 298c4dee-8b0d-474e-afad-c57e481b4223 · outbound

This paper cites Multi-game decision transformers.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Multi-game decision transformers

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.890075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.901393Z digest=sha256:0d6f5f845a9384f77441e3cee03608548e3a5d39bf9022fc4d4739c7e49a76f1

Observation 4da2b345-5e0d-4d95-a039-aed07cae0242 · outbound

This paper cites Dhrl: a graph-based approach for long-horizon and sparse hierarchical reinforcement learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Dhrl: a graph-based approach for long-horizon and sparse hierarchical reinforcement learning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.877877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.904989Z digest=sha256:27478b18a62b98e78e41dfb27d9aaadb975f1b30aad8c679b0106992167627a0

Observation d0173034-cc60-47ef-acb7-cf86d44b445c · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.909047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.909047Z digest=sha256:959ca39ebbb5b04fda58ad227b2fbe5b9f8f50a4ffb901d795bb261f414e9929

Observation 338ffa32-a5ea-4dcf-8a0d-aae0f181c06a · outbound

This paper cites Hierarchical planning through goal-conditioned offline reinforcement learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Hierarchical planning through goal-conditioned offline reinforcement learning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.865153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.912606Z digest=sha256:48bec0b1721b2949ef06dfaa6450f977d3e4fd43f975c4c7d4802d082831fb11

Observation 2b3f252c-1f37-403b-ba3f-b2186e6ef988 · outbound

This paper cites DIDI: Diffusion-Guided Diversity for Offline Behavioral Generation.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization DIDI: Diffusion-Guided Diversity for Offline Behavioral Generation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.916444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.916444Z digest=sha256:915275f2920b1b736081ab3734fb392cbf783f2c52038cb72ab3daafd2490571

Observation ab0ce511-34ba-4e69-b886-0908976a8819 · outbound

This paper cites Beyond ood state actions: Supported cross-domain offline reinforcement learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Beyond ood state actions: Supported cross-domain offline reinforcement learning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.850298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.920305Z digest=sha256:9511a072074af5fa4f0e58141750a4f43c3bd8e7dc27126ace469673acb88bd2

Observation d1625f51-008d-47f1-a75b-80e9ec423a6f · outbound

This paper cites Least squares quantization in pcm.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Least squares quantization in pcm

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.924152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.924152Z digest=sha256:f0937a119f64cd027c6eba9938ae023fd60302b0dcb68b95d326474afc5b8111

Observation 05f6f010-fa26-43cc-9226-12825383213b · outbound

This paper cites Decoupled Weight Decay Regularization.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Decoupled Weight Decay Regularization

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.928161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.928161Z digest=sha256:e3391f7fe453f70d24b3f70b1eab487e0833d5c2a5d924a3dcc09d2ae4977bd6

Observation 2a9bff04-ecfd-42a1-be87-47952d4c0d07 · outbound

This paper cites Generative Trajectory Stitching through Diffusion Composition.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Generative Trajectory Stitching through Diffusion Composition

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.931670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.931670Z digest=sha256:fae7d49905f0c527332cc7ee0dbed7d9fde645c24124b4cbe2d89d4a83d331f0

Observation 0ee3c592-d739-4461-ae0c-81d25979e657 · outbound

This paper cites Decision mamba: A multi-grained state space model with self-evolution regularization for offline rl.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Decision mamba: A multi-grained state space model with self-evolution regularization for offline rl

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.823899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.936092Z digest=sha256:a06b7c8de24484d3d232b5478b77148deb383b4ab1279ce85ebb4513d1ba4adc

Observation 9161b203-157a-43c2-a75a-e6c58b3bf523 · outbound

This paper cites Learning latent plans from play.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Learning latent plans from play

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.807672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.939663Z digest=sha256:2c634570b16fefed929e4b4898b4b33968a6c0aac865918dafde1d3ac45f0e54

Observation 7760b84f-ef4a-4c07-b158-dd9ed90e7346 · outbound

This paper cites How Far I'll Go: Offline Goal-Conditioned Reinforcement Learning via $f$-Advantage Regression.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization How Far I'll Go: Offline Goal-Conditioned Reinforcement Learning via $f$-Advantage Regression

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.943342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.943342Z digest=sha256:0c1e1762d81b4b1b58d22b0b651c4c0bdb0877fb54977a053365313077cc1f5d

Observation c121fe31-7789-4cdd-bd85-6da077e82bd2 · outbound

This paper cites VIP : Towards universal visual reward and representation via value-implicit pre-training.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization VIP : Towards universal visual reward and representation via value-implicit pre-training

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.787378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.947029Z digest=sha256:ac5cc4276e15da186e8f69a5b38f7e823fe7fd8476c374b1ca7b557d9be88c5b

Observation 81535569-6821-4cf1-b589-a68adfbf95cf · outbound

This paper cites Human-level control through deep reinforcement learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Human-level control through deep reinforcement learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.950640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.950640Z digest=sha256:013a198e6221c3e85c0bf442e6842b2c3d1f0111372090900bc7d5da5de6bacd

Observation b836f8df-5d17-418a-9393-f6a1a8ccaa13 · outbound

This paper cites Learning Temporal Distances: Contrastive Successor Features Can Provide a Metric Structure for Decision-Making.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Learning Temporal Distances: Contrastive Successor Features Can Provide a Metric Structure for Decision-Making

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.954383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.954383Z digest=sha256:0e0c55d4a68f0fd26dfd7ad69aac9337c92a53335b2a7713df1602930cc91a28

Observation d110249c-705f-4d70-a3d4-ceec22004d39 · outbound

This paper cites Asymmetric least squares estimation and testing.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Asymmetric least squares estimation and testing

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.763153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.958453Z digest=sha256:1defa2ba44928448cf221a2c9f1e71aa942e6cb0f0cc346e99e6b5acd789f49f

Observation 6fff31ef-0216-4c08-a905-3cc793992213 · outbound

This paper cites Decision Mamba: Reinforcement Learning via Sequence Modeling with Selective State Spaces.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Decision Mamba: Reinforcement Learning via Sequence Modeling with Selective State Spaces

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.961759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.961759Z digest=sha256:4326178dba5198129cc8012d425116625ac77fca1d27e8a51ab0bde123c51a66

Observation 49179c0b-5541-4a55-b2de-51de3d7635b3 · outbound

This paper cites Hiql: Offline goal-conditioned rl with latent states as actions.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Hiql: Offline goal-conditioned rl with latent states as actions

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.966571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.966571Z digest=sha256:15e571019e39874791ec7ce8d18a804e0756fdc0b540ee70f9b4fc6d89d90aea

Observation f5f5575a-0a9d-432b-9163-2d9147e6f3df · outbound

This paper cites Foundation Policies with Hilbert Representations.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Foundation Policies with Hilbert Representations

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.970149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.970149Z digest=sha256:db547df0d02795151a22a2827bd5665ba49234065a9a4a974576ffa6f9f30278

Observation 0b6d3909-af36-4f36-b8d2-07deef469bd8 · outbound

This paper cites Ogbench: Benchmarking offline goal-conditioned rl.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Ogbench: Benchmarking offline goal-conditioned rl

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.739722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.975726Z digest=sha256:d00e35b33f010feab74776517f6522785f49ed2fbac0aa2c3a0535472950a8ca

Observation 03cde593-9e7a-4b38-b8c0-d2b09ffe2fe9 · outbound

This paper cites A survey on offline reinforcement learning: Taxonomy, review, and open problems.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization A survey on offline reinforcement learning: Taxonomy, review, and open problems

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.727985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.980393Z digest=sha256:4ae971e645d9d7cf2f63c41ab8b58bbac9a80e540c1fb552137c15f0568ade24

Observation 738e246b-af48-488e-bd3d-3f9c6928ce3d · outbound

This paper cites Goal-Conditioned Imitation Learning using Score-based Diffusion Policies.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Goal-Conditioned Imitation Learning using Score-based Diffusion Policies

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.984273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.984273Z digest=sha256:a2f96df92e22655b83833aef80dc2967d678efac7bc87df894961f6a6f65c77b

Observation 0047489f-1d2b-4ace-91ba-3d824f4a0da4 · outbound

This paper cites Stochastic backpropagation and approximate inference in deep generative models.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Stochastic backpropagation and approximate inference in deep generative models

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.716004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.988205Z digest=sha256:228b3d472af7ea6e5457782c378689ccb7558b66bcfa98c2807e509b90b7d570

Observation f43cd76c-e99b-47ef-a3bc-6bd55c4b6680 · outbound

This paper cites Outcome-driven reinforcement learning via variational inference.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Outcome-driven reinforcement learning via variational inference

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.704287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.994354Z digest=sha256:d70b39eba1cee3c2c9297f9f4e3299161ca24777eba0187c54b6cc8916500060

Observation 5d75398f-09e4-484f-a1a0-d8b10e8421f6 · outbound

This paper cites Reinforcement learning upside down: Don't predict rewards -- just map them to actions, 2020.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Reinforcement learning upside down: Don't predict rewards -- just map them to actions, 2020

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.692774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.998229Z digest=sha256:717414f01260282b443a7cfc83f6745663f7d2fcac1370c566d1500c51c1b588

Observation a0ad97e9-02cb-4edd-8211-cb21fa4dab4d · outbound

This paper cites Rapid exploration for open-world navigation with latent goal models.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Rapid exploration for open-world navigation with latent goal models

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.681724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:05:26.002289Z digest=sha256:4e3658a76b1a8a24f2005d4d8e6d3beb613e79c3e3faa6a75cb6c24d757dd937

Observation dfc8bf9e-530b-400e-a12c-6bd0d5fc4e6f · outbound

This paper cites Score models for offline goal-conditioned reinforcement learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Score models for offline goal-conditioned reinforcement learning

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.670794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:05:26.006634Z digest=sha256:ac6cba887abb5a13132f5eaeb52f64acd4d1791f0a9b7ab9090b1fabefad23a8

Observation a9e70b2b-72cf-44f6-996f-cfd0632dad84 · outbound

This paper cites Geoadditive expectile regression.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Geoadditive expectile regression

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.659954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:05:26.010660Z digest=sha256:3fddb9b316980a8d38f3f03e8f2cba4da6dd572d2d53a9c8e16a231873c266ac

Observation d9c00165-7d4f-4aeb-9178-49005ea0b73e · outbound

This paper cites Learning structured output representation using deep conditional generative models.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Learning structured output representation using deep conditional generative models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:26.014377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:26.014377Z digest=sha256:71fb1e733ed2ebb50f0109502e1555f986391785c5e73ea34af3b397bbb1c474

Observation 31cf407e-cc0a-4685-ae51-73d4d982f234 · outbound

This paper cites Gymnasium (mar 2023), 2023.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Gymnasium (mar 2023), 2023

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.639864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:05:26.018409Z digest=sha256:63b2a728cf41481de0e010e338bd08c67adefad7ec8ce8bda4c006b45e08f241

Observation 15c15ca4-5ecf-46a0-abb1-078650721ce1 · outbound

This paper cites Deep Reinforcement Learning and the Deadly Triad.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Deep Reinforcement Learning and the Deadly Triad

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:26.022196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:26.022196Z digest=sha256:8354f57f7aefdeb745f7745d7e8a8b8f2f0e882ee57abc9d6b3cf3078fb60294

Observation 7e47b5a1-f2c6-4aec-b59c-e20c712aadcc · outbound

This paper cites Attention is all you need.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Attention is all you need

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:26.026210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:26.026210Z digest=sha256:7a37f09a8de21b499c34cbb564902cebcd4f78b9a237a27ba9d40e7285a01ece

Observation c20b3ea7-d6cd-4f80-a058-c02026c8f25b · outbound

This paper cites GOP lan: Goal-conditioned offline reinforcement learning by planning with learned models.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization GOP lan: Goal-conditioned offline reinforcement learning by planning with learned models

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.619349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:05:26.029722Z digest=sha256:1dd1b26113e3c969a800dea806a5f207d1eaac5eec9fac542483736f2d3488bf

Observation 34fa6ff0-f0f9-4b0f-bac5-f5c5f9efce2b · outbound

This paper cites Optimal goal-reaching reinforcement learning via quasimetric learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Optimal goal-reaching reinforcement learning via quasimetric learning

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.607776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:05:26.033672Z digest=sha256:9e006c34b856445660f52079d20ee40b50aa7c80d176ff477b532655f55b6e61

Observation 4064caf9-0460-452a-a414-d3ea921924ee · outbound

This paper cites Critic-guided decision transformer for offline reinforcement learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Critic-guided decision transformer for offline reinforcement learning

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.596625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:05:26.038097Z digest=sha256:001777a92803841e7770f712bf9b8a4f645cd0cdafd57c51280c7c7802951b43

Observation 1f240bd1-3f82-44a7-900f-bb0dcb6eb6ec · outbound

This paper cites Supported policy optimization for offline reinforcement learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Supported policy optimization for offline reinforcement learning

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.584880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:05:26.041485Z digest=sha256:1f4f5e7dc655b222094f06cf16067966b02d0ed2faf501e26d37f73a41a9cfe8

Observation e8e42e3d-1cdb-42c8-9228-c14463d124a2 · outbound

This paper cites Elastic Decision Transformer.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Elastic Decision Transformer

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:26.045615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:26.045615Z digest=sha256:52a2aa8df670e6fbde8cece8b58ae2c1efcb339a969e808279fdcefac331ed52

Observation ca8c62b3-2b08-4809-9704-568102ecae9d · outbound

This paper cites Q-learning decision transformer: Leveraging dynamic programming for conditional sequence modelling in offline rl.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Q-learning decision transformer: Leveraging dynamic programming for conditional sequence modelling in offline rl

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.573852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:05:26.050121Z digest=sha256:eb7961863209aa2f74681e4528997a285ca9eeb5d2b1ea16211fd971b853faf4

Observation 158ff2de-8365-47d0-809c-dfad1cc405e2 · outbound

This paper cites Rethinking Goal-conditioned Supervised Learning and Its Connection to Offline RL.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Rethinking Goal-conditioned Supervised Learning and Its Connection to Offline RL

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:26.054551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:26.054551Z digest=sha256:04572f55e0449ae0b9b99687e78770db4c0966931b5297cb458ceb766edf89e2

Observation db0fa5c6-5581-4b25-87fb-013fcef04677 · outbound

This paper cites What is essential for unseen goal generalization of offline goal-conditioned rl? In International Conference on Machine Learning, pages 39543--39571.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization What is essential for unseen goal generalization of offline goal-conditioned rl? In International Conference on Machine Learning, pages 39543--39571

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.561418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:05:26.058489Z digest=sha256:3301d8e203982315876bb74e8b1eaca8fe3032294f097732c95ac9297d40b33d

Observation 77db9b91-9fca-46a6-9296-56fca0933486 · outbound

This paper cites Swapped goal-conditioned offline reinforcement learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Swapped goal-conditioned offline reinforcement learning

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:26.062222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:26.062222Z digest=sha256:bb6fe448aa0401e4f5d5362b8c897a1ec7ab59391a2ea5ab053085778a53b470

Observation 2f8dd9a5-e9c1-4d76-af4a-b181000ab2fb · outbound

This paper cites Breadth-first exploration on adaptive grid for reinforcement learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Breadth-first exploration on adaptive grid for reinforcement learning

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:26.066117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:26.066117Z digest=sha256:4af63ea8f4544687df13a507774cfd8f381fed6e8a8d794fe74dae6c4632ce30

Observation 7fbcd7cd-bb20-4774-be4e-c2841fc8cbee · outbound

This paper cites Goal-conditioned predictive coding for offline reinforcement learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Goal-conditioned predictive coding for offline reinforcement learning

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.539901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:05:26.069683Z digest=sha256:8f78534f93e9014de9d0bf5451167c749ea970ea31dca8c633e2788574ba9cf6

Observation b155832e-295f-432b-bfcd-1de1323f54b8 · outbound

This paper cites Stabilizing Contrastive RL: Techniques for Robotic Goal Reaching from Offline Data.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Stabilizing Contrastive RL: Techniques for Robotic Goal Reaching from Offline Data

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:26.072975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:26.072975Z digest=sha256:a6ed17abd7d82913cf5bcd87c97461029a91d28220ffc48ead7df2218521fda2

Observation 1b355547-cf41-4fd4-860b-354833c8961b · outbound

This paper cites Contrastive difference predictive coding.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Contrastive difference predictive coding

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.526281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:05:26.076893Z digest=sha256:f5a05294b04c867b8a0a10207b3f7a8252c78cf57f46309610ba269d873ff5c4

Observation 1007d93e-a759-4670-a9ae-ce0f34545b99 · outbound

This paper cites Online decision transformer.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Online decision transformer

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.514001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:05:26.081476Z digest=sha256:64b9c59ac21a242daaf1f61a8885b670566b7f342b4014d6d8455c5c668c0b5f

Observation 2de92e6a-ec70-4159-9a24-cf51af5f8e47 · outbound

This paper cites Behavior Proximal Policy Optimization.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Behavior Proximal Policy Optimization

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:26.085023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:26.085023Z digest=sha256:2d147d51e4dec3265015894073ec2e694dfb505908317a99f2aef42cbf11d128

Observation 1267e842-1074-4441-b705-b7277e3483cd · outbound

This paper cites Reinformer: Max-Return Sequence Modeling for Offline RL.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Reinformer: Max-Return Sequence Modeling for Offline RL

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:26.091332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:26.091332Z digest=sha256:7334725c61a98096844155477fa73a3ddb88fc733d2207579a3406aa5bbc8308

Observation ad6054d5-8274-4dd4-b327-6f158fb0db69 · outbound

This paper cites Revisiting the design choices in max-return sequence modeling, 2025.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Revisiting the design choices in max-return sequence modeling, 2025

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.500024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:05:26.095894Z digest=sha256:0126151d05f644d722789dcacab0963395b2a43b3192991e008ec5c6d7d570da

Observation fc578017-cce3-4fd5-bb09-46cadeecab2d · outbound

This paper cites Maximum entropy inverse reinforcement learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Maximum entropy inverse reinforcement learning

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:26.101533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:26.101533Z digest=sha256:039a0bdb788066b6e46e6f9d386e2c2d871c64b463124df43e37f75779164afe

Pith citing papers

No inbound Pith citation observations are available.