Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:19:10.275156Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 0 inbound Pith citation observations for arXiv:2506.08463.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:19:10.275156Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
43 of 43 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation fcfc7757-4214-4f91-b109-8f106280f373 · outbound
How to Provably Improve Return Conditioned Supervised Learning? When does return-conditioned supervised learning work for offline reinforcement learning? Advances in Neural Information Processing Systems, 35:1542–1553, 2022
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5b9ca6b3-c313-4672-b65b-880e92d71eaa · outbound
How to Provably Improve Return Conditioned Supervised Learning? Decision transformer: Reinforcement learning 1For two distributionsPandQ,P≪QmeansPis absolutely continuous w.r.t
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 632f3e8c-a2d3-4cd7-99cb-80ca75d930d9 · outbound
How to Provably Improve Return Conditioned Supervised Learning? Rvs: What is essential for offline rl via supervised learning? InInternational Conference on Learning Representations,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8059f672-d218-4f2a-bb9e-0686a58617fe · outbound
How to Provably Improve Return Conditioned Supervised Learning? Imitating past successes can be very suboptimal.Advances in Neural Information Processing Systems, 35:6047–6059, 2022
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2ecf1c2e-1149-4524-a53e-55a7f1a8a4b7 · outbound
How to Provably Improve Return Conditioned Supervised Learning? D4RL: Datasets for Deep Data-Driven Reinforcement Learning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cbb6611-c082-4463-9026-07bfd459cf24 · outbound
How to Provably Improve Return Conditioned Supervised Learning? A minimalist approach to offline reinforcement learning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5655c842-202c-4de9-a6f0-f83ed89e05dd · outbound
How to Provably Improve Return Conditioned Supervised Learning? Generalized decision transformer for of- fline hindsight information matching
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e2aa48c5-3aa6-4a9d-95c2-bd9a5f79d2c2 · outbound
How to Provably Improve Return Conditioned Supervised Learning? Act: empowering decision transformer with dynamic programming via advantage conditioning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3269d05b-354c-48d7-afa8-da6aaa6f9c35 · outbound
How to Provably Improve Return Conditioned Supervised Learning? DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c368ab03-1013-4091-9e99-89f1c5170047 · outbound
How to Provably Improve Return Conditioned Supervised Learning? Q-value regularized transformer for offline reinforcement learning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 01b99dd9-f6c6-4d7e-8696-ab0812b89bc3 · outbound
How to Provably Improve Return Conditioned Supervised Learning? Offline reinforcement learning as one big sequence modeling problem.Advances in neural information processing systems, 34:1273– 1286, 2021
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a250da26-9d7a-4fe9-8485-5e895f9eb332 · outbound
How to Provably Improve Return Conditioned Supervised Learning? Is pessimism provably efficient for offline rl? In International Conference on Machine Learning, pages 5084–5096
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7f6df767-4bcd-4739-8a4c-fab9b5d99d93 · outbound
How to Provably Improve Return Conditioned Supervised Learning? Reinforcement learning in robotics: A survey
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cd59261a-7881-4ac4-bc9b-a754d3a0f442 · outbound
How to Provably Improve Return Conditioned Supervised Learning? Quantile regression.Journal of economic perspectives, 15(4):143–156, 2001
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f822402f-8bc0-4912-83f8-edbe38a56fe4 · outbound
How to Provably Improve Return Conditioned Supervised Learning? Offline reinforcement learning with implicit q-learning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6c5875f2-64b1-4600-ae59-09958a242422 · outbound
How to Provably Improve Return Conditioned Supervised Learning? Reward-Conditioned Policies
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8b33b93-0360-41c1-ae74-3d9cf32ad1c5 · outbound
How to Provably Improve Return Conditioned Supervised Learning? Conservative q-learning for offline reinforcement learning.Advances in Neural Information Processing Systems, 33: 1179–1191, 2020
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 767bf042-42de-4d6b-abd8-08006375deda · outbound
How to Provably Improve Return Conditioned Supervised Learning? Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fe8ce3c-613d-4c92-9751-86e2c1fe269f · outbound
How to Provably Improve Return Conditioned Supervised Learning? Deep rein- forcement learning for dynamic treatment regimes on medical registry data
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8858455d-b390-4bb0-acae-6186137f7915 · outbound
How to Provably Improve Return Conditioned Supervised Learning? Deep spatial q- learning for infectious disease control.Journal of Agricultural, Biological and Environmental Statistics, 28(4):749–773, 2023
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9b0d61b0-f99f-49c0-a347-8858d116fba9 · outbound
How to Provably Improve Return Conditioned Supervised Learning? Human-level control through deep reinforcement learning.nature, 518(7540):529–533, 2015
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation be7d3c4b-a6ba-4257-915f-69937ccbad93 · outbound
How to Provably Improve Return Conditioned Supervised Learning? Asymmetric least squares estimation and testing
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 90b11ce6-50d2-4a37-b5d2-e924eca9d161 · outbound
How to Provably Improve Return Conditioned Supervised Learning? You can’t count on luck: Why decision transformers and rvs fail in stochastic environments.Advances in neural information processing systems, 35:38966–38979, 2022
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5717991e-91ec-4bfa-ba3a-df3a467f88cf · outbound
How to Provably Improve Return Conditioned Supervised Learning? Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cca5d79-1339-4d7a-a4c9-664a40f50954 · outbound
How to Provably Improve Return Conditioned Supervised Learning? Reinforcement learning in robotic applications: a comprehensive survey.Artificial Intelligence Review, 55(2):945–990, 2022
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1e69c1d7-7926-44ab-b3ac-bb44ef679d6c · outbound
How to Provably Improve Return Conditioned Supervised Learning? Training Agents using Upside-Down Reinforcement Learning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2afd2b9d-7214-4a41-9df1-ea4084ea1283 · outbound
How to Provably Improve Return Conditioned Supervised Learning? MIT press,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 721c7b5e-3a48-45fe-9077-1e0d3ef24a7a · outbound
How to Provably Improve Return Conditioned Supervised Learning? Temporal difference learning and td-gammon.Communications of the ACM, 38(3):58–68, 1995
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 87aa220a-a2ed-477e-a513-63bba5759d4e · outbound
How to Provably Improve Return Conditioned Supervised Learning? Grandmaster level in starcraft ii using multi-agent reinforcement learning.nature, 575(7782):350–354, 2019
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6141da23-8251-45f2-a0d1-7b5dbb5fea7e · outbound
How to Provably Improve Return Conditioned Supervised Learning? Return augmented decision transformer for off-dynamics reinforcement learning.arXiv preprint arXiv:2410.23450, 2024
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff035164-e227-4e2b-a65f-8cbc31daef0d · outbound
How to Provably Improve Return Conditioned Supervised Learning? Q-learning.Machine learning, 8:279–292, 1992
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 53389b36-65f8-4b59-9e7f-1c542c47ee22 · outbound
How to Provably Improve Return Conditioned Supervised Learning? Elastic decision transformer.Advances in Neural Information Processing Systems, 36, 2024
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c631914c-ae8a-4a8d-9b28-261d8d59acbb · outbound
How to Provably Improve Return Conditioned Supervised Learning? A policy-guided imitation approach for offline reinforcement learning.Advances in neural information processing systems, 35: 4085–4098, 2022
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f7854296-a476-4262-841e-cfb35771bd49 · outbound
How to Provably Improve Return Conditioned Supervised Learning? Q-learning decision transformer: Leveraging dynamic programming for conditional sequence modelling in offline rl
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 43b19276-ba88-4e84-8dcb-7e19465c5716 · outbound
How to Provably Improve Return Conditioned Supervised Learning? Dichotomy of control: Sepa- rating what you can control from what you cannot
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a50769dc-a148-47b9-9a8a-774bd96bf1c9 · outbound
How to Provably Improve Return Conditioned Supervised Learning? Reinforcement learning in healthcare: A survey.ACM Computing Surveys (CSUR), 55(1):1–36, 2021
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0b35ec78-9d8d-40d5-bf81-1186ecea1ed2 · outbound
How to Provably Improve Return Conditioned Supervised Learning? Online decision transformer
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 544670a7-2e87-4095-8ce6-248819a8584e · outbound
How to Provably Improve Return Conditioned Supervised Learning? How does goal relabeling improve sample efficiency? InProceedings of the 41st International Conference on Machine Learning, pages 61246–61266
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 77d681dd-db1b-4cca-95f7-8e794631c989 · outbound
How to Provably Improve Return Conditioned Supervised Learning? Unresolved cited work
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 496aea22-1320-433e-a1cb-28b8a0d6fb68 · outbound
How to Provably Improve Return Conditioned Supervised Learning? Reinformer: Max- return sequence modeling for offline rl
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 55a97671-7fdc-4e24-9518-930d3c5a772d · outbound
How to Provably Improve Return Conditioned Supervised Learning? Starting from the second stage, π⋆ starts stitching the performance of different trajectories
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cd6d66ec-ab1a-41b8-a2fb-2b6f7c8dac22 · outbound
How to Provably Improve Return Conditioned Supervised Learning? We formulate the loss function as LDT−R 2CSL =E τ [−logπ θ(·|τ)−λH(π θ(·|τ))]
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation da474c82-2cbc-41c7-847d-50e88d987a07 · outbound
How to Provably Improve Return Conditioned Supervised Learning? Unresolved cited work
Reference 2022
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
No inbound Pith citation observations are available.