Pith. sign in

Paper Citation Record · LEDGER

How to Provably Improve Return Conditioned Supervised Learning?

As of 8 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 0 inbound Pith citation observations for arXiv:2506.08463.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.08463 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:19:10.275156Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

43 of 43 outbound references displayed

  • verified exact0
  • verified fuzzy34
  • unresolved8
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fcfc7757-4214-4f91-b109-8f106280f373 · outbound

This paper cites When does return-conditioned supervised learning work for offline reinforcement learning? Advances in Neural Information Processing Systems, 35:1542–1553, 2022.

How to Provably Improve Return Conditioned Supervised Learning? When does return-conditioned supervised learning work for offline reinforcement learning? Advances in Neural Information Processing Systems, 35:1542–1553, 2022

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:19.174686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:19:04.103676Z digest=sha256:4c30512b328428b0ea00ce16dd9e5c6da70ba31b6799a26f00671967a105e441

Observation 5b9ca6b3-c313-4672-b65b-880e92d71eaa · outbound

This paper cites Decision transformer: Reinforcement learning 1For two distributionsPandQ,P≪QmeansPis absolutely continuous w.r.t.

How to Provably Improve Return Conditioned Supervised Learning? Decision transformer: Reinforcement learning 1For two distributionsPandQ,P≪QmeansPis absolutely continuous w.r.t

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:18.989128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:19:04.192132Z digest=sha256:65d7461ff015e3b378cd4c17060713d83ceafc9c9fb3076531433c830371e3cb

Observation 632f3e8c-a2d3-4cd7-99cb-80ca75d930d9 · outbound

This paper cites Rvs: What is essential for offline rl via supervised learning? InInternational Conference on Learning Representations,.

How to Provably Improve Return Conditioned Supervised Learning? Rvs: What is essential for offline rl via supervised learning? InInternational Conference on Learning Representations,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:18.778572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:19:04.330667Z digest=sha256:fd25d350fe7d543927dbff0d9fa3b9cc62ffc08b03699f603e05b06504ec9e3d

Observation 8059f672-d218-4f2a-bb9e-0686a58617fe · outbound

This paper cites Imitating past successes can be very suboptimal.Advances in Neural Information Processing Systems, 35:6047–6059, 2022.

How to Provably Improve Return Conditioned Supervised Learning? Imitating past successes can be very suboptimal.Advances in Neural Information Processing Systems, 35:6047–6059, 2022

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:18.350689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:19:04.670336Z digest=sha256:5c1f5ef4c69733182ceccbac1eeaf753fb593e64db3497ffced13256e5cb2c36

Observation 2ecf1c2e-1149-4524-a53e-55a7f1a8a4b7 · outbound

This paper cites D4RL: Datasets for Deep Data-Driven Reinforcement Learning.

How to Provably Improve Return Conditioned Supervised Learning? D4RL: Datasets for Deep Data-Driven Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:19:04.868201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:19:04.868201Z digest=sha256:42fccd4bdab6b627ac01c3909e2a7c58f6722b6df3bbd5fedfb2452038011ee6

Observation 6cbb6611-c082-4463-9026-07bfd459cf24 · outbound

This paper cites A minimalist approach to offline reinforcement learning.

How to Provably Improve Return Conditioned Supervised Learning? A minimalist approach to offline reinforcement learning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:18.154249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:19:05.017626Z digest=sha256:42aa5ce74ef1099ea629952c61e0075bb2e1f533df30411664a5627d12fb1b2b

Observation 5655c842-202c-4de9-a6f0-f83ed89e05dd · outbound

This paper cites Generalized decision transformer for of- fline hindsight information matching.

How to Provably Improve Return Conditioned Supervised Learning? Generalized decision transformer for of- fline hindsight information matching

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:17.981506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:19:05.154444Z digest=sha256:5cece3859be280b74fe1caa9638dc0038eed130ba3ef0e94ac4f4d67fbac3525

Observation e2aa48c5-3aa6-4a9d-95c2-bd9a5f79d2c2 · outbound

This paper cites Act: empowering decision transformer with dynamic programming via advantage conditioning.

How to Provably Improve Return Conditioned Supervised Learning? Act: empowering decision transformer with dynamic programming via advantage conditioning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:17.798842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:19:05.328078Z digest=sha256:2bb8b35c4a2d29711c60050da2128092a43604414ac3fb6ab6ed362f2ff6e249

Observation 3269d05b-354c-48d7-afa8-da6aaa6f9c35 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

How to Provably Improve Return Conditioned Supervised Learning? DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T05:19:05.443153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:19:05.443153Z digest=sha256:f21d31c6767a9fcb14de6ada1daa8905cded2f804f92c8655acd39455230a6b4

Observation c368ab03-1013-4091-9e99-89f1c5170047 · outbound

This paper cites Q-value regularized transformer for offline reinforcement learning.

How to Provably Improve Return Conditioned Supervised Learning? Q-value regularized transformer for offline reinforcement learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:17.590710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:19:05.597675Z digest=sha256:93b7364bb2ff7334a69c30ec351ad7120d6d5f70061f4947e1de4e2667b6e915

Observation 01b99dd9-f6c6-4d7e-8696-ab0812b89bc3 · outbound

This paper cites Offline reinforcement learning as one big sequence modeling problem.Advances in neural information processing systems, 34:1273– 1286, 2021.

How to Provably Improve Return Conditioned Supervised Learning? Offline reinforcement learning as one big sequence modeling problem.Advances in neural information processing systems, 34:1273– 1286, 2021

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:17.389293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:19:05.742356Z digest=sha256:b562deedcc9e753b76fea1a05296774ed55d6e6d45d81011b356f5d5ec67a2c1

Observation a250da26-9d7a-4fe9-8485-5e895f9eb332 · outbound

This paper cites Is pessimism provably efficient for offline rl? In International Conference on Machine Learning, pages 5084–5096.

How to Provably Improve Return Conditioned Supervised Learning? Is pessimism provably efficient for offline rl? In International Conference on Machine Learning, pages 5084–5096

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:17.179796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:19:05.905336Z digest=sha256:0121ddc756b09cb618c32b4455dd0a61e71ad4847f4fbd36ff93f1e714256c4b

Observation 7f6df767-4bcd-4739-8a4c-fab9b5d99d93 · outbound

This paper cites Reinforcement learning in robotics: A survey.

How to Provably Improve Return Conditioned Supervised Learning? Reinforcement learning in robotics: A survey

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:16.979615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:19:06.058393Z digest=sha256:1d944017d54ddcbfa2e0458b8cb183a53270a2a702d9885e3e81749611464aba

Observation cd59261a-7881-4ac4-bc9b-a754d3a0f442 · outbound

This paper cites Quantile regression.Journal of economic perspectives, 15(4):143–156, 2001.

How to Provably Improve Return Conditioned Supervised Learning? Quantile regression.Journal of economic perspectives, 15(4):143–156, 2001

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:16.814473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:19:06.236473Z digest=sha256:d1fb601f7d7b6093e063962c301d4cddd81880226d17029ce73bf787f9487c32

Observation f822402f-8bc0-4912-83f8-edbe38a56fe4 · outbound

This paper cites Offline reinforcement learning with implicit q-learning.

How to Provably Improve Return Conditioned Supervised Learning? Offline reinforcement learning with implicit q-learning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:16.614230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:19:06.387415Z digest=sha256:8040c7294c14f9ad0e84cc2667645668eeb06b69c2c9f26d972735e7f2e7e82a

Observation 6c5875f2-64b1-4600-ae59-09958a242422 · outbound

This paper cites Reward-Conditioned Policies.

How to Provably Improve Return Conditioned Supervised Learning? Reward-Conditioned Policies

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:19:06.527500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:19:06.527500Z digest=sha256:250869a73a87e49d9a3c6aca9446f3c57e8f8694330f65dfe5a2e68fccbb9f27

Observation b8b33b93-0360-41c1-ae74-3d9cf32ad1c5 · outbound

This paper cites Conservative q-learning for offline reinforcement learning.Advances in Neural Information Processing Systems, 33: 1179–1191, 2020.

How to Provably Improve Return Conditioned Supervised Learning? Conservative q-learning for offline reinforcement learning.Advances in Neural Information Processing Systems, 33: 1179–1191, 2020

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:16.417121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:19:06.681264Z digest=sha256:bba0f15fde5832c6c55e333b5c9e14d8f023cbef63e80ac15ab60f98c823a037

Observation 767bf042-42de-4d6b-abd8-08006375deda · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

How to Provably Improve Return Conditioned Supervised Learning? Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:19:06.860433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:19:06.860433Z digest=sha256:277b6b051f96e6a30937d77db6a39ae4b8f2592a2dee272a0a2224e985e708ed

Observation 2fe8ce3c-613d-4c92-9751-86e2c1fe269f · outbound

This paper cites Deep rein- forcement learning for dynamic treatment regimes on medical registry data.

How to Provably Improve Return Conditioned Supervised Learning? Deep rein- forcement learning for dynamic treatment regimes on medical registry data

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:16.248503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:19:07.000640Z digest=sha256:4cee091b1e51e81fec8f36824a59c8ff5185c8ead05f06f515a16dbf73f4be13

Observation 8858455d-b390-4bb0-acae-6186137f7915 · outbound

This paper cites Deep spatial q- learning for infectious disease control.Journal of Agricultural, Biological and Environmental Statistics, 28(4):749–773, 2023.

How to Provably Improve Return Conditioned Supervised Learning? Deep spatial q- learning for infectious disease control.Journal of Agricultural, Biological and Environmental Statistics, 28(4):749–773, 2023

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:16.058975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:19:07.146714Z digest=sha256:9aa18ff0129b62905a41ff8ae2492e240ccf6b33863c540a2db496fb05903154

Observation 9b0d61b0-f99f-49c0-a347-8858d116fba9 · outbound

This paper cites Human-level control through deep reinforcement learning.nature, 518(7540):529–533, 2015.

How to Provably Improve Return Conditioned Supervised Learning? Human-level control through deep reinforcement learning.nature, 518(7540):529–533, 2015

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:15.881289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:19:07.322593Z digest=sha256:09f6037f6772fe4b034d7c6e7d810c39be181e3547ac22ee04cbeb081ac00836

Observation be7d3c4b-a6ba-4257-915f-69937ccbad93 · outbound

This paper cites Asymmetric least squares estimation and testing.

How to Provably Improve Return Conditioned Supervised Learning? Asymmetric least squares estimation and testing

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:13.905451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:19:07.467721Z digest=sha256:2f42e3f3f67b215a749b6ed20ec8042ef33dcca7c6f7559239a466c562937f9e

Observation 90b11ce6-50d2-4a37-b5d2-e924eca9d161 · outbound

This paper cites You can’t count on luck: Why decision transformers and rvs fail in stochastic environments.Advances in neural information processing systems, 35:38966–38979, 2022.

How to Provably Improve Return Conditioned Supervised Learning? You can’t count on luck: Why decision transformers and rvs fail in stochastic environments.Advances in neural information processing systems, 35:38966–38979, 2022

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:13.732547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:19:07.618289Z digest=sha256:4aeb21e195ce8c241982d8d379bed99379f2a1f374febc973543f9cf14f398a2

Observation 5717991e-91ec-4bfa-ba3a-df3a467f88cf · outbound

This paper cites Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions.

How to Provably Improve Return Conditioned Supervised Learning? Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:19:07.756599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:19:07.756599Z digest=sha256:0365370e9093a58ea3d484cdb2858eba527893e8a049564051cbacffaf777d75

Observation 1cca5d79-1339-4d7a-a4c9-664a40f50954 · outbound

This paper cites Reinforcement learning in robotic applications: a comprehensive survey.Artificial Intelligence Review, 55(2):945–990, 2022.

How to Provably Improve Return Conditioned Supervised Learning? Reinforcement learning in robotic applications: a comprehensive survey.Artificial Intelligence Review, 55(2):945–990, 2022

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:13.536857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:19:07.874637Z digest=sha256:5081d9c919b96692074a3dc1b1df9647136e499fa81be9e6f0fe517e784cf6bc

Observation 1e69c1d7-7926-44ab-b3ac-bb44ef679d6c · outbound

This paper cites Training Agents using Upside-Down Reinforcement Learning.

How to Provably Improve Return Conditioned Supervised Learning? Training Agents using Upside-Down Reinforcement Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:19:08.014844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:19:08.014844Z digest=sha256:5688dbfe763b32a19081cdfab913d29634d5995d47226210e2ce9cd41970e450

Observation 2afd2b9d-7214-4a41-9df1-ea4084ea1283 · outbound

This paper cites MIT press,.

How to Provably Improve Return Conditioned Supervised Learning? MIT press,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:13.380306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:19:08.194854Z digest=sha256:0515fa526466f2b61c7468db0ec8cb1e3661e8c27a63f4dd17ca7bf059ba391f

Observation 721c7b5e-3a48-45fe-9077-1e0d3ef24a7a · outbound

This paper cites Temporal difference learning and td-gammon.Communications of the ACM, 38(3):58–68, 1995.

How to Provably Improve Return Conditioned Supervised Learning? Temporal difference learning and td-gammon.Communications of the ACM, 38(3):58–68, 1995

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:13.180697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:19:08.280991Z digest=sha256:d91e79945a32f5160582e12b1fd6d87dd18bf2fe6b2b58f00960eaf9cb3b4d46

Observation 87aa220a-a2ed-477e-a513-63bba5759d4e · outbound

This paper cites Grandmaster level in starcraft ii using multi-agent reinforcement learning.nature, 575(7782):350–354, 2019.

How to Provably Improve Return Conditioned Supervised Learning? Grandmaster level in starcraft ii using multi-agent reinforcement learning.nature, 575(7782):350–354, 2019

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:13.005163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:19:08.452464Z digest=sha256:52703383c33267e64d89f93570f974f83255ee3fa9d8bc5c5eff6d13f6b09b8c

Observation 6141da23-8251-45f2-a0d1-7b5dbb5fea7e · outbound

This paper cites Return augmented decision transformer for off-dynamics reinforcement learning.arXiv preprint arXiv:2410.23450, 2024.

How to Provably Improve Return Conditioned Supervised Learning? Return augmented decision transformer for off-dynamics reinforcement learning.arXiv preprint arXiv:2410.23450, 2024

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:19:08.615517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:19:08.615517Z digest=sha256:8b9aadfabf8973504db9feb9a7c113b0b43b021a09d16ee7213e95e19a675e3f

Observation ff035164-e227-4e2b-a65f-8cbc31daef0d · outbound

This paper cites Q-learning.Machine learning, 8:279–292, 1992.

How to Provably Improve Return Conditioned Supervised Learning? Q-learning.Machine learning, 8:279–292, 1992

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:12.809949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:19:08.701647Z digest=sha256:c60a4901603e312088857c79195ffd582a32a593bdcb64380cb2455599e19080

Observation 53389b36-65f8-4b59-9e7f-1c542c47ee22 · outbound

This paper cites Elastic decision transformer.Advances in Neural Information Processing Systems, 36, 2024.

How to Provably Improve Return Conditioned Supervised Learning? Elastic decision transformer.Advances in Neural Information Processing Systems, 36, 2024

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:12.608655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:19:08.859217Z digest=sha256:c8df8f678edb6251467d97a7a6c5881efb08c0b876fa6360bf676a2e92718a5b

Observation c631914c-ae8a-4a8d-9b28-261d8d59acbb · outbound

This paper cites A policy-guided imitation approach for offline reinforcement learning.Advances in neural information processing systems, 35: 4085–4098, 2022.

How to Provably Improve Return Conditioned Supervised Learning? A policy-guided imitation approach for offline reinforcement learning.Advances in neural information processing systems, 35: 4085–4098, 2022

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:12.409417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:19:09.061745Z digest=sha256:0e61e88633b7839274cd856f3315cf2a94da904714bec89cdcd2228823fc7528

Observation f7854296-a476-4262-841e-cfb35771bd49 · outbound

This paper cites Q-learning decision transformer: Leveraging dynamic programming for conditional sequence modelling in offline rl.

How to Provably Improve Return Conditioned Supervised Learning? Q-learning decision transformer: Leveraging dynamic programming for conditional sequence modelling in offline rl

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:12.219928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:19:09.142012Z digest=sha256:2697edf430612cb93d5d77efe53aec3126272bbdaaa2a0a9f74d882ad1f1489c

Observation 43b19276-ba88-4e84-8dcb-7e19465c5716 · outbound

This paper cites Dichotomy of control: Sepa- rating what you can control from what you cannot.

How to Provably Improve Return Conditioned Supervised Learning? Dichotomy of control: Sepa- rating what you can control from what you cannot

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:12.062449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:19:09.297843Z digest=sha256:45861cb50aac67a4affe2fc7f38bb10922f2c0e823eb58cd09fbcf891bc2b52b

Observation a50769dc-a148-47b9-9a8a-774bd96bf1c9 · outbound

This paper cites Reinforcement learning in healthcare: A survey.ACM Computing Surveys (CSUR), 55(1):1–36, 2021.

How to Provably Improve Return Conditioned Supervised Learning? Reinforcement learning in healthcare: A survey.ACM Computing Surveys (CSUR), 55(1):1–36, 2021

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:11.895118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:19:09.413750Z digest=sha256:76db2a396641d8a3147dad46031d0940230191dd7d02acd32c8a73cfc5351d4e

Observation 0b35ec78-9d8d-40d5-bf81-1186ecea1ed2 · outbound

This paper cites Online decision transformer.

How to Provably Improve Return Conditioned Supervised Learning? Online decision transformer

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:11.717619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:19:09.528698Z digest=sha256:ade6991053467d54e8b3b6a6f5596218b9d23ce0d4af710939bab87de66eb75e

Observation 544670a7-2e87-4095-8ce6-248819a8584e · outbound

This paper cites How does goal relabeling improve sample efficiency? InProceedings of the 41st International Conference on Machine Learning, pages 61246–61266.

How to Provably Improve Return Conditioned Supervised Learning? How does goal relabeling improve sample efficiency? InProceedings of the 41st International Conference on Machine Learning, pages 61246–61266

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:11.514778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:19:09.716896Z digest=sha256:70540b7ab28eddb0a4e8cac387a5d14e22f057b36eeca12706facf0e077c61a0

Observation 77d681dd-db1b-4cca-95f7-8e794631c989 · outbound

This paper cites an unresolved cited work.

How to Provably Improve Return Conditioned Supervised Learning? Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:19:11.309789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:19:09.811315Z digest=sha256:6b9483556723620e3f76a9d3c2b7687604eb1e02ca31fafc552fd94e1f02889f

Observation 496aea22-1320-433e-a1cb-28b8a0d6fb68 · outbound

This paper cites Reinformer: Max- return sequence modeling for offline rl.

How to Provably Improve Return Conditioned Supervised Learning? Reinformer: Max- return sequence modeling for offline rl

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:11.115470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:19:09.935960Z digest=sha256:22be2bc662312a99ed100370303f3615a2402ee2ee13084a14cb494fa1257efa

Observation 55a97671-7fdc-4e24-9518-930d3c5a772d · outbound

This paper cites Starting from the second stage, π⋆ starts stitching the performance of different trajectories.

How to Provably Improve Return Conditioned Supervised Learning? Starting from the second stage, π⋆ starts stitching the performance of different trajectories

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:10.911021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:19:10.095418Z digest=sha256:8a9a30f78a04484062c389220a389e8843ed6bce745d5b84c4c69244a70289e8

Observation cd6d66ec-ab1a-41b8-a2fb-2b6f7c8dac22 · outbound

This paper cites We formulate the loss function as LDT−R 2CSL =E τ [−logπ θ(·|τ)−λH(π θ(·|τ))].

How to Provably Improve Return Conditioned Supervised Learning? We formulate the loss function as LDT−R 2CSL =E τ [−logπ θ(·|τ)−λH(π θ(·|τ))]

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:10.716042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:19:10.275156Z digest=sha256:47f93c084fc2e65428b7e4ba463bba1d0a850ddaad2c0b0ee4a9aac82a709f51

Observation da474c82-2cbc-41c7-847d-50e88d987a07 · outbound

This paper cites an unresolved cited work.

How to Provably Improve Return Conditioned Supervised Learning? Unresolved cited work

Reference 2022

Resolution
parse uncertain
raw_fallback, observed 2026-08-07T05:19:18.566541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:19:04.464264Z digest=sha256:252599fae7873cc089e395e65cbff488f2a7a35045750806b6ec6fa7ff9ccb17

Pith citing papers

No inbound Pith citation observations are available.