Pith. sign in

Paper Citation Record · LEDGER

How to Provably Improve Return Conditioned Supervised Learning?

As of 11 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 0 inbound Pith citation observations for arXiv:2506.08463.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.08463 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:19:10.275156Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

43 of 43 outbound references displayed

  • verified exact0
  • verified fuzzy34
  • unresolved8
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fcfc7757-4214-4f91-b109-8f106280f373 · outbound

This paper cites When does return-conditioned supervised learning work for offline reinforcement learning? Advances in Neural Information Processing Systems, 35:1542–1553, 2022.

How to Provably Improve Return Conditioned Supervised Learning? When does return-conditioned supervised learning work for offline reinforcement learning? Advances in Neural Information Processing Systems, 35:1542–1553, 2022

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:19.174686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:19:04.103676Z digest=sha256:34b1d0910aeda03ab0a1333035b161b843e28fe60ec28d0dc60e92dba07c154c

Observation 5b9ca6b3-c313-4672-b65b-880e92d71eaa · outbound

This paper cites Decision transformer: Reinforcement learning 1For two distributionsPandQ,P≪QmeansPis absolutely continuous w.r.t.

How to Provably Improve Return Conditioned Supervised Learning? Decision transformer: Reinforcement learning 1For two distributionsPandQ,P≪QmeansPis absolutely continuous w.r.t

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:18.989128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:19:04.192132Z digest=sha256:0a150c671fa18abe74c824b38e445ef8fa71e2a6757899335ffddaa19e8a1627

Observation 632f3e8c-a2d3-4cd7-99cb-80ca75d930d9 · outbound

This paper cites Rvs: What is essential for offline rl via supervised learning? InInternational Conference on Learning Representations,.

How to Provably Improve Return Conditioned Supervised Learning? Rvs: What is essential for offline rl via supervised learning? InInternational Conference on Learning Representations,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:18.778572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:19:04.330667Z digest=sha256:9b700c919b839883025b70087bcd1921c426e68a9e84db2960c1adf7e8fd0a3b

Observation 8059f672-d218-4f2a-bb9e-0686a58617fe · outbound

This paper cites Imitating past successes can be very suboptimal.Advances in Neural Information Processing Systems, 35:6047–6059, 2022.

How to Provably Improve Return Conditioned Supervised Learning? Imitating past successes can be very suboptimal.Advances in Neural Information Processing Systems, 35:6047–6059, 2022

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:18.350689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:19:04.670336Z digest=sha256:141cfca84502d0b6347386a7af0457b12e57fd71f8d9c6e1e76ceeb03689e042

Observation 2ecf1c2e-1149-4524-a53e-55a7f1a8a4b7 · outbound

This paper cites D4RL: Datasets for Deep Data-Driven Reinforcement Learning.

How to Provably Improve Return Conditioned Supervised Learning? D4RL: Datasets for Deep Data-Driven Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:19:04.868201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:19:04.868201Z digest=sha256:4644256d1006c17f48644cd543edf61d10172366ccd40a71d67cd3f3e949420c

Observation 6cbb6611-c082-4463-9026-07bfd459cf24 · outbound

This paper cites A minimalist approach to offline reinforcement learning.

How to Provably Improve Return Conditioned Supervised Learning? A minimalist approach to offline reinforcement learning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:18.154249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:19:05.017626Z digest=sha256:0e02e568bdd64745a0ce211776dd72a0ca3a362e1e28fd14321917e5af4a7bc8

Observation 5655c842-202c-4de9-a6f0-f83ed89e05dd · outbound

This paper cites Generalized decision transformer for of- fline hindsight information matching.

How to Provably Improve Return Conditioned Supervised Learning? Generalized decision transformer for of- fline hindsight information matching

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:17.981506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:19:05.154444Z digest=sha256:d52cf21cb17defe13834cca0495ad32221369a26cb32168b8d7bfdf8ff5ad0de

Observation e2aa48c5-3aa6-4a9d-95c2-bd9a5f79d2c2 · outbound

This paper cites Act: empowering decision transformer with dynamic programming via advantage conditioning.

How to Provably Improve Return Conditioned Supervised Learning? Act: empowering decision transformer with dynamic programming via advantage conditioning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:17.798842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:19:05.328078Z digest=sha256:aa157d9daf9b6cfb5a084c312b6eaec168f6afad976ee6f8b879bcc7eb3d1969

Observation 3269d05b-354c-48d7-afa8-da6aaa6f9c35 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

How to Provably Improve Return Conditioned Supervised Learning? DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T05:19:05.443153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:19:05.443153Z digest=sha256:4af5dd7ee939cbd8e061a8394b59030069483a891e82a7cf4fe04db1a67c264c

Observation c368ab03-1013-4091-9e99-89f1c5170047 · outbound

This paper cites Q-value regularized transformer for offline reinforcement learning.

How to Provably Improve Return Conditioned Supervised Learning? Q-value regularized transformer for offline reinforcement learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:17.590710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:19:05.597675Z digest=sha256:ab1e4a49f20ae8601eb621876ccb58911c57fc22b35f6f8f36f7ef52b098800f

Observation 01b99dd9-f6c6-4d7e-8696-ab0812b89bc3 · outbound

This paper cites Offline reinforcement learning as one big sequence modeling problem.Advances in neural information processing systems, 34:1273– 1286, 2021.

How to Provably Improve Return Conditioned Supervised Learning? Offline reinforcement learning as one big sequence modeling problem.Advances in neural information processing systems, 34:1273– 1286, 2021

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:17.389293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:19:05.742356Z digest=sha256:5c07ad2ce5f65ce50bb646a7dfe4cc449fcc857c4431981592e49ec083332cb5

Observation a250da26-9d7a-4fe9-8485-5e895f9eb332 · outbound

This paper cites Is pessimism provably efficient for offline rl? In International Conference on Machine Learning, pages 5084–5096.

How to Provably Improve Return Conditioned Supervised Learning? Is pessimism provably efficient for offline rl? In International Conference on Machine Learning, pages 5084–5096

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:17.179796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:19:05.905336Z digest=sha256:ea6b9c4b7888689142fdecda95b4a012b78ddb5bab1f2fc24c718efbfedcfd15

Observation 7f6df767-4bcd-4739-8a4c-fab9b5d99d93 · outbound

This paper cites Reinforcement learning in robotics: A survey.

How to Provably Improve Return Conditioned Supervised Learning? Reinforcement learning in robotics: A survey

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:16.979615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:19:06.058393Z digest=sha256:6a4e5d8f0445bf0a788af6fcabb38a62f70ab9b7d491132875275cc619318061

Observation cd59261a-7881-4ac4-bc9b-a754d3a0f442 · outbound

This paper cites Quantile regression.Journal of economic perspectives, 15(4):143–156, 2001.

How to Provably Improve Return Conditioned Supervised Learning? Quantile regression.Journal of economic perspectives, 15(4):143–156, 2001

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:16.814473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:19:06.236473Z digest=sha256:2fd256eeb276f717e07921e5bce0a18620e14934f32c1cdf308e4421ece3e6e8

Observation f822402f-8bc0-4912-83f8-edbe38a56fe4 · outbound

This paper cites Offline reinforcement learning with implicit q-learning.

How to Provably Improve Return Conditioned Supervised Learning? Offline reinforcement learning with implicit q-learning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:16.614230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:19:06.387415Z digest=sha256:c20a1710475766046a4f934d96f96fbff2f28e90dc550f9005737f757fe39fd3

Observation 6c5875f2-64b1-4600-ae59-09958a242422 · outbound

This paper cites Reward-Conditioned Policies.

How to Provably Improve Return Conditioned Supervised Learning? Reward-Conditioned Policies

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:19:06.527500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:19:06.527500Z digest=sha256:3bdba05f4fc1c6d8fda3f1a28b1ec0f7d6d75a2af778128d79af59e6d1800aef

Observation b8b33b93-0360-41c1-ae74-3d9cf32ad1c5 · outbound

This paper cites Conservative q-learning for offline reinforcement learning.Advances in Neural Information Processing Systems, 33: 1179–1191, 2020.

How to Provably Improve Return Conditioned Supervised Learning? Conservative q-learning for offline reinforcement learning.Advances in Neural Information Processing Systems, 33: 1179–1191, 2020

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:16.417121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:19:06.681264Z digest=sha256:8be48f3a53e005330500ec853c441c1da4123e03fb397fe290830d37df3a2fc3

Observation 767bf042-42de-4d6b-abd8-08006375deda · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

How to Provably Improve Return Conditioned Supervised Learning? Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:19:06.860433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:19:06.860433Z digest=sha256:ee1a8fc0113a48117d7d5db50caf2318d2d971a365c21cf0b4c39bf42b50d058

Observation 2fe8ce3c-613d-4c92-9751-86e2c1fe269f · outbound

This paper cites Deep rein- forcement learning for dynamic treatment regimes on medical registry data.

How to Provably Improve Return Conditioned Supervised Learning? Deep rein- forcement learning for dynamic treatment regimes on medical registry data

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:16.248503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:19:07.000640Z digest=sha256:66022912b78be87f10ae66ac1e9f4c3c90dc627a90bf52cb6efed2084e302498

Observation 8858455d-b390-4bb0-acae-6186137f7915 · outbound

This paper cites Deep spatial q- learning for infectious disease control.Journal of Agricultural, Biological and Environmental Statistics, 28(4):749–773, 2023.

How to Provably Improve Return Conditioned Supervised Learning? Deep spatial q- learning for infectious disease control.Journal of Agricultural, Biological and Environmental Statistics, 28(4):749–773, 2023

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:16.058975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:19:07.146714Z digest=sha256:2cf29fa8711a20d94a1a935171fc7200553c015d97d6d69e1ed82965e54317dc

Observation 9b0d61b0-f99f-49c0-a347-8858d116fba9 · outbound

This paper cites Human-level control through deep reinforcement learning.nature, 518(7540):529–533, 2015.

How to Provably Improve Return Conditioned Supervised Learning? Human-level control through deep reinforcement learning.nature, 518(7540):529–533, 2015

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:15.881289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:19:07.322593Z digest=sha256:8d5a7a2eef145e337683d8b94293a3d873989dba7f550fa575d30cc95d7d1900

Observation be7d3c4b-a6ba-4257-915f-69937ccbad93 · outbound

This paper cites Asymmetric least squares estimation and testing.

How to Provably Improve Return Conditioned Supervised Learning? Asymmetric least squares estimation and testing

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:13.905451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:19:07.467721Z digest=sha256:ae812100da38566530ac961444ec250c9e28152e69eda8b725a0233637a657e3

Observation 90b11ce6-50d2-4a37-b5d2-e924eca9d161 · outbound

This paper cites You can’t count on luck: Why decision transformers and rvs fail in stochastic environments.Advances in neural information processing systems, 35:38966–38979, 2022.

How to Provably Improve Return Conditioned Supervised Learning? You can’t count on luck: Why decision transformers and rvs fail in stochastic environments.Advances in neural information processing systems, 35:38966–38979, 2022

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:13.732547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:19:07.618289Z digest=sha256:933fe316e320e604cf2dffd9047a2556a2b3b6f3acf987c2ab36787d84a2973e

Observation 5717991e-91ec-4bfa-ba3a-df3a467f88cf · outbound

This paper cites Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions.

How to Provably Improve Return Conditioned Supervised Learning? Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:19:07.756599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:19:07.756599Z digest=sha256:9a47edfa2977a800fa6bce39d0468183b10af7be638bc0127513d2c2bcb282ee

Observation 1cca5d79-1339-4d7a-a4c9-664a40f50954 · outbound

This paper cites Reinforcement learning in robotic applications: a comprehensive survey.Artificial Intelligence Review, 55(2):945–990, 2022.

How to Provably Improve Return Conditioned Supervised Learning? Reinforcement learning in robotic applications: a comprehensive survey.Artificial Intelligence Review, 55(2):945–990, 2022

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:13.536857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:19:07.874637Z digest=sha256:b3d81ced986aae99466b39a3ab68b9b9a85334d6ed06240283f3b6304f260973

Observation 1e69c1d7-7926-44ab-b3ac-bb44ef679d6c · outbound

This paper cites Training Agents using Upside-Down Reinforcement Learning.

How to Provably Improve Return Conditioned Supervised Learning? Training Agents using Upside-Down Reinforcement Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:19:08.014844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:19:08.014844Z digest=sha256:db6b6b3fcf14ca413c962851dae5f175b8ac9054efbda947cdacb2e53667cf63

Observation 2afd2b9d-7214-4a41-9df1-ea4084ea1283 · outbound

This paper cites MIT press,.

How to Provably Improve Return Conditioned Supervised Learning? MIT press,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:13.380306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:19:08.194854Z digest=sha256:6cd3a9960cfccdd1d1acc783b312f0259ee2c8d855e8249ba83ec6908cfb9e30

Observation 721c7b5e-3a48-45fe-9077-1e0d3ef24a7a · outbound

This paper cites Temporal difference learning and td-gammon.Communications of the ACM, 38(3):58–68, 1995.

How to Provably Improve Return Conditioned Supervised Learning? Temporal difference learning and td-gammon.Communications of the ACM, 38(3):58–68, 1995

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:13.180697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:19:08.280991Z digest=sha256:a77def70c77ddd68e9857b9eb85a791514ec1e2d5ee42869eb597b16247bce9d

Observation 87aa220a-a2ed-477e-a513-63bba5759d4e · outbound

This paper cites Grandmaster level in starcraft ii using multi-agent reinforcement learning.nature, 575(7782):350–354, 2019.

How to Provably Improve Return Conditioned Supervised Learning? Grandmaster level in starcraft ii using multi-agent reinforcement learning.nature, 575(7782):350–354, 2019

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:13.005163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:19:08.452464Z digest=sha256:e22e57cb6c8f180654b23c7117fca6672866883eadcbc4d57b172115bea1faeb

Observation 6141da23-8251-45f2-a0d1-7b5dbb5fea7e · outbound

This paper cites Return augmented decision transformer for off-dynamics reinforcement learning.arXiv preprint arXiv:2410.23450, 2024.

How to Provably Improve Return Conditioned Supervised Learning? Return augmented decision transformer for off-dynamics reinforcement learning.arXiv preprint arXiv:2410.23450, 2024

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:19:08.615517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:19:08.615517Z digest=sha256:7a8f1bc2f9dc0a34ae57d10db9840f5bae433dc2eab375e978ba316dc2691a7d

Observation ff035164-e227-4e2b-a65f-8cbc31daef0d · outbound

This paper cites Q-learning.Machine learning, 8:279–292, 1992.

How to Provably Improve Return Conditioned Supervised Learning? Q-learning.Machine learning, 8:279–292, 1992

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:12.809949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:19:08.701647Z digest=sha256:41a52e79e96dcb15886f37971b6cd6fdea1572cc30d2b7c2dc43503ef0685b20

Observation 53389b36-65f8-4b59-9e7f-1c542c47ee22 · outbound

This paper cites Elastic decision transformer.Advances in Neural Information Processing Systems, 36, 2024.

How to Provably Improve Return Conditioned Supervised Learning? Elastic decision transformer.Advances in Neural Information Processing Systems, 36, 2024

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:12.608655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:19:08.859217Z digest=sha256:e4577d664e739fea730179e3af52ad1f561e26f587867ff9166fcf91d925112e

Observation c631914c-ae8a-4a8d-9b28-261d8d59acbb · outbound

This paper cites A policy-guided imitation approach for offline reinforcement learning.Advances in neural information processing systems, 35: 4085–4098, 2022.

How to Provably Improve Return Conditioned Supervised Learning? A policy-guided imitation approach for offline reinforcement learning.Advances in neural information processing systems, 35: 4085–4098, 2022

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:12.409417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:19:09.061745Z digest=sha256:e5b7c1f2b25c312d7f1e72b23995a1b30892381a1e4069537842185fd27ef714

Observation f7854296-a476-4262-841e-cfb35771bd49 · outbound

This paper cites Q-learning decision transformer: Leveraging dynamic programming for conditional sequence modelling in offline rl.

How to Provably Improve Return Conditioned Supervised Learning? Q-learning decision transformer: Leveraging dynamic programming for conditional sequence modelling in offline rl

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:12.219928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:19:09.142012Z digest=sha256:477b706432c17afba7ffd194b5a532427a47d6fe2890e9358e3e8958379a55ae

Observation 43b19276-ba88-4e84-8dcb-7e19465c5716 · outbound

This paper cites Dichotomy of control: Sepa- rating what you can control from what you cannot.

How to Provably Improve Return Conditioned Supervised Learning? Dichotomy of control: Sepa- rating what you can control from what you cannot

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:12.062449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:19:09.297843Z digest=sha256:11b0951f2e07baefc5e1cc151d5044f51b2cc16460090518d0c47f70d499a9a3

Observation a50769dc-a148-47b9-9a8a-774bd96bf1c9 · outbound

This paper cites Reinforcement learning in healthcare: A survey.ACM Computing Surveys (CSUR), 55(1):1–36, 2021.

How to Provably Improve Return Conditioned Supervised Learning? Reinforcement learning in healthcare: A survey.ACM Computing Surveys (CSUR), 55(1):1–36, 2021

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:11.895118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:19:09.413750Z digest=sha256:e15828879ebd6cc2502ad6e5e5dbee1df2ab6ef30e0b1482c6f66ba1f9044832

Observation 0b35ec78-9d8d-40d5-bf81-1186ecea1ed2 · outbound

This paper cites Online decision transformer.

How to Provably Improve Return Conditioned Supervised Learning? Online decision transformer

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:11.717619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:19:09.528698Z digest=sha256:3af6aae0e980376d141ce64a35fe5c2f32f55b413189e3a798299a82ca22e2fc

Observation 544670a7-2e87-4095-8ce6-248819a8584e · outbound

This paper cites How does goal relabeling improve sample efficiency? InProceedings of the 41st International Conference on Machine Learning, pages 61246–61266.

How to Provably Improve Return Conditioned Supervised Learning? How does goal relabeling improve sample efficiency? InProceedings of the 41st International Conference on Machine Learning, pages 61246–61266

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:11.514778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:19:09.716896Z digest=sha256:c6c0b74a80fbd9493abac293377948edd65c527ebb5438e98e5f1e1a555c8ea7

Observation 77d681dd-db1b-4cca-95f7-8e794631c989 · outbound

This paper cites an unresolved cited work.

How to Provably Improve Return Conditioned Supervised Learning? Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:19:11.309789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:19:09.811315Z digest=sha256:371f412ff390053b6e0989acfeff47cfbb9bd8498bb447cfb66f2954d27ca197

Observation 496aea22-1320-433e-a1cb-28b8a0d6fb68 · outbound

This paper cites Reinformer: Max- return sequence modeling for offline rl.

How to Provably Improve Return Conditioned Supervised Learning? Reinformer: Max- return sequence modeling for offline rl

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:11.115470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:19:09.935960Z digest=sha256:5aad09754adc5393717f7dc0fa80b980dbd1a688329d8ddefe76de248f73d9b6

Observation 55a97671-7fdc-4e24-9518-930d3c5a772d · outbound

This paper cites Starting from the second stage, π⋆ starts stitching the performance of different trajectories.

How to Provably Improve Return Conditioned Supervised Learning? Starting from the second stage, π⋆ starts stitching the performance of different trajectories

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:10.911021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:19:10.095418Z digest=sha256:cf2f1771e2409e85cdf5f925e904ebae39e2c01c5e5fb7dcb30f803cf5eb6679

Observation cd6d66ec-ab1a-41b8-a2fb-2b6f7c8dac22 · outbound

This paper cites We formulate the loss function as LDT−R 2CSL =E τ [−logπ θ(·|τ)−λH(π θ(·|τ))].

How to Provably Improve Return Conditioned Supervised Learning? We formulate the loss function as LDT−R 2CSL =E τ [−logπ θ(·|τ)−λH(π θ(·|τ))]

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:10.716042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:19:10.275156Z digest=sha256:d5e81faed91423d2b948c3b1af4ec97c2bf07ee1888772bc8e8b44a16cec67f0

Observation da474c82-2cbc-41c7-847d-50e88d987a07 · outbound

This paper cites an unresolved cited work.

How to Provably Improve Return Conditioned Supervised Learning? Unresolved cited work

Reference 2022

Resolution
parse uncertain
raw_fallback, observed 2026-08-07T05:19:18.566541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:19:04.464264Z digest=sha256:09680a0754b3a429b854f8a0ee5068b0c5944eed757bba6feccddcb016e82441

Pith citing papers

No inbound Pith citation observations are available.