Pith. sign in

Paper Citation Record · LEDGER

Fully Offline Reinforcement Learning

As of 20 August 2026, this Paper Citation Record lists 83 of 83 outbound references and 2 inbound Pith citation observations for arXiv:2505.22442.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22442 v3

Coverage vector

measured 83 of 83 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:16:05.159203Z

measured 85 of 85 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T15:25:39.829506Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T01:48:51.121586Z

Reference resolution

83 of 83 outbound references displayed

  • verified exact11
  • verified fuzzy41
  • unresolved28
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7f4fa7dd-2576-4833-918f-082dad491144 · outbound

This paper cites Aitchison.

Fully Offline Reinforcement Learning Aitchison

Reference 1

Resolution
verified exact
raw_fallback, observed 2026-08-07T13:16:09.424246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:15:57.315944Z digest=sha256:2c135bd647f5ae8970dd164384f1b41f5ce2fa726efe02e9b3822e5367d9dd3f

Observation aa23fa33-78ea-43be-ad09-ec04a2ec318a · outbound

This paper cites Alaa and Mihaela van der Schaar.

Fully Offline Reinforcement Learning Alaa and Mihaela van der Schaar

Reference 2

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T13:16:09.081121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:15:57.356190Z digest=sha256:19aac08cce4c7a5ac5e3882fcec28ee18c07ed5f629f424755dedc402b9911f4

Observation faa3004a-74fb-466a-b59d-28e8b50b6332 · outbound

This paper cites Uncertainty-based offline reinforcement learning with diversified q-ensemble.

Fully Offline Reinforcement Learning Uncertainty-based offline reinforcement learning with diversified q-ensemble

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:15:57.413976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:15:57.413976Z digest=sha256:d16a36c422d6d08586c1d696ef3bff56b57ce0b78df31f388e0f3a792fe94d78

Observation 3ad69933-f655-4d69-8dee-5f498d7ff47c · outbound

This paper cites Asymptotically minimax bayes predictive densities.The Annals of Statistics, 34(6):2921–2938, 2006.

Fully Offline Reinforcement Learning Asymptotically minimax bayes predictive densities.The Annals of Statistics, 34(6):2921–2938, 2006

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:15:57.474090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:15:57.474090Z digest=sha256:bfa3eeef67545561722ea662df2c17b8c61d0b6b1a104b5c23ceb961e8dafe2c

Observation 02430707-3cef-41eb-acea-41eedfcd62cc · outbound

This paper cites Augmented world models facilitate zero-shot dynamics generalization from a single offline environment.

Fully Offline Reinforcement Learning Augmented world models facilitate zero-shot dynamics generalization from a single offline environment

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:15:57.536060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:15:57.536060Z digest=sha256:300d8eb2ff4b022e9144645fb48c346ed9b3e2fb11418b64959682ea2f95ad8b

Observation b48111ec-9df1-470f-b256-4122791e1c06 · outbound

This paper cites Information-theoretic characterization of bayes performance and the choice of priors in parametric and nonparametric problems.

Fully Offline Reinforcement Learning Information-theoretic characterization of bayes performance and the choice of priors in parametric and nonparametric problems

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:21.255216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:15:57.590263Z digest=sha256:b72bba4dff33c02dc5c8855389eab78ed2bc2b2a29705f1f3b970362f6011e55

Observation ab150b87-375f-4aca-b853-765cb9f73e2f · outbound

This paper cites Barron.The Exponential Convergence of Posterior Probabilities with Implications for Bayes Estimators of Density Functions.

Fully Offline Reinforcement Learning Barron.The Exponential Convergence of Posterior Probabilities with Implications for Bayes Estimators of Density Functions

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:20.989180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:15:57.757860Z digest=sha256:6287bc49416d623f03337ea18b0ed86586c227fb24a78bf0c63901e0e8c2d894

Observation d16ea67a-4b7d-4dcf-83cb-b64d4ddd1675 · outbound

This paper cites Bass.Real Analysis for Graduate Students, chapter 21.

Fully Offline Reinforcement Learning Bass.Real Analysis for Graduate Students, chapter 21

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:20.820274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:15:57.822068Z digest=sha256:5b134959e7905860dc1f932422fde99fedb7d167c0e93bdcaa0c6466f55687f0

Observation 14fff60b-933b-4eff-8483-c0f9dc237f61 · outbound

This paper cites A Tutorial on Meta-Reinforcement Learning.

Fully Offline Reinforcement Learning A Tutorial on Meta-Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:15:57.942175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:15:57.942175Z digest=sha256:ec8d26022b3dda9114e4ea39eae1e73fe854a719c3ded6c2827b6fcfd99b09d8

Observation c275b9b1-b9bd-4ad5-a6de-e2c89d65fe86 · outbound

This paper cites A problem in the sequential design of experiments.Sankhy ¯a: The Indian Journal of Statistics (1933-1960), 16(3/4):221–229, 1956.

Fully Offline Reinforcement Learning A problem in the sequential design of experiments.Sankhy ¯a: The Indian Journal of Statistics (1933-1960), 16(3/4):221–229, 1956

Reference 10

Resolution
verified exact
raw_fallback, observed 2026-08-07T13:16:08.501815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:15:58.001783Z digest=sha256:aaf1a19e74dd1ed9cc650a0dca422d47120ad08825d5736d35947350406514e3

Observation 53dbe3be-d34e-4b8f-b8c4-eeae6d108a02 · outbound

This paper cites Dynamic programming and stochastic control processes.Information and Control, 1(3):228–239, 1958.

Fully Offline Reinforcement Learning Dynamic programming and stochastic control processes.Information and Control, 1(3):228–239, 1958

Reference 11

Resolution
malformed identifier
no resolver link, observed 2026-08-07T13:15:58.052512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:15:58.052512Z digest=sha256:44bb884e86b8d548e858de477d37ef9a5d1d03c531ea10e59002fcf7266cb68c

Observation ad5273c6-cfb0-4297-8b5b-d36395eea30e · outbound

This paper cites Foster, and Daniel M.

Fully Offline Reinforcement Learning Foster, and Daniel M

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:20.427736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:15:58.116894Z digest=sha256:3fefc7370d915cc6f0497a0bae76777d0ee6cea7943d2953f906f70057a5cb65

Observation 3f3dade5-de7b-468d-8deb-ae49eabbfc43 · outbound

This paper cites On the foundations of statistical inference.Journal of the American Statistical Association, 57(298):269–306, 1962.

Fully Offline Reinforcement Learning On the foundations of statistical inference.Journal of the American Statistical Association, 57(298):269–306, 1962

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:15:58.171525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:15:58.171525Z digest=sha256:25a0ac8540cfcef167370e9b23dcf906ba5d2154d2293a0d3d9749ff3f06e196

Observation 50ccbea6-642b-4c41-abfb-6ccd7d75f62c · outbound

This paper cites an unresolved cited work.

Fully Offline Reinforcement Learning Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:16:20.279450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:15:58.236104Z digest=sha256:7cdb4c4fea2c665cedf2db8d15098fc37eec6f908252bf62711ef8574fe27dbf

Observation de0a4ee4-4b43-46b8-9c14-1c86432ce338 · outbound

This paper cites Bayes adaptive monte carlo tree search for of- fline model-based reinforcement learning, 2024.

Fully Offline Reinforcement Learning Bayes adaptive monte carlo tree search for of- fline model-based reinforcement learning, 2024

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:20.135149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:15:58.297640Z digest=sha256:78be87e73afdd990dbb84192ec8896a875cfe638667d808642e0b4203d37b09b

Observation cfe50663-3e99-4113-815f-551afb1c29ee · outbound

This paper cites Conser- vative uncertainty estimation by fitting prior networks.

Fully Offline Reinforcement Learning Conser- vative uncertainty estimation by fitting prior networks

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:19.983352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:15:58.372662Z digest=sha256:ce4ca795ec067dfb50735fd8fb155dc1d7e0d38db4bfb816d8908fd6d9cb107a

Observation 016295ea-6bda-4420-99f2-2c9d7ae85950 · outbound

This paper cites Clarke and A.R.

Fully Offline Reinforcement Learning Clarke and A.R

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:19.743980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:15:58.448968Z digest=sha256:bb55c06e8d91fe04aba328bf6c01fa52ac883edbb3d844cecca7500ffca5fe7c

Observation 4ba4c40f-40e6-47f5-a44a-23eb99b56a5f · outbound

This paper cites an unresolved cited work.

Fully Offline Reinforcement Learning Unresolved cited work

Reference 18

Resolution
verified exact
raw_fallback, observed 2026-08-07T13:16:07.959162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:15:58.500423Z digest=sha256:89ae485a45bf77b89ab80db79aa369d311dd1c47ea5454c1d2ae3738dcf50108

Observation 54158c56-5f57-431f-be73-7a763611224d · outbound

This paper cites Observation of a markov process through a noisy channel.PhD Thesis, 1962.

Fully Offline Reinforcement Learning Observation of a markov process through a noisy channel.PhD Thesis, 1962

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:19.551859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:15:58.563857Z digest=sha256:8325e0ae36e3e4e31ddb727da19a02683bb7bedcd329e25fab535f3d17e2f1a4

Observation d5475ef1-f8d4-4265-9802-3e75ba2d6176 · outbound

This paper cites Fast reinforcement learning via slow reinforcement learning.

Fully Offline Reinforcement Learning Fast reinforcement learning via slow reinforcement learning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:19.341391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:15:58.618468Z digest=sha256:ab494adc5ebc37f0761932294ef7bbe7e420d3367e129936651a09ce7c7dbfc6

Observation fd95ba64-4cef-4f53-a399-142fcde22945 · outbound

This paper cites PhD thesis, 2002.

Fully Offline Reinforcement Learning PhD thesis, 2002

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:19.101212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:15:58.684957Z digest=sha256:a834f3e625192401e37e303b0b3259be8d41ff44ec61e10f82f6bd2f51e7695e

Observation b34d3302-8720-46c0-88bf-8404861c1c57 · outbound

This paper cites Bayesian exploration networks.

Fully Offline Reinforcement Learning Bayesian exploration networks

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:18.931137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:15:58.801124Z digest=sha256:a77af2990c49c857c3e4ce75b78789a708d06927539a65d5620d0428a6d0da86

Observation af8f42f6-eebf-41e4-a8ed-29b0c9113c05 · outbound

This paper cites D4rl: Datasets for deep data-driven reinforcement learning, 2020.

Fully Offline Reinforcement Learning D4rl: Datasets for deep data-driven reinforcement learning, 2020

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:18.703635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:15:58.874936Z digest=sha256:e0cd39461394d864746bd56f9297784f43d77e53a683280f18881ae486012c75

Observation 37eca544-1166-4b74-b6d2-90890599c766 · outbound

This paper cites A minimalist approach to offline reinforcement learn- ing.

Fully Offline Reinforcement Learning A minimalist approach to offline reinforcement learn- ing

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:18.497264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:15:58.962879Z digest=sha256:eaf1f197ec3543a008a56b73bba3696f75b789b4d72b1f2d91504258fd74fd10

Observation b819815a-87ac-40bc-8d62-8771f3c71ef1 · outbound

This paper cites A new proof of the likelihood principle.The British Journal for the Philosophy of Science, 66(3):475–503, 2015.

Fully Offline Reinforcement Learning A new proof of the likelihood principle.The British Journal for the Philosophy of Science, 66(3):475–503, 2015

Reference 25

Resolution
verified exact
doi, observed 2026-08-07T13:16:06.339819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:15:59.057545Z digest=sha256:381c98af33a3e3b76766641c51339f3e0a518589115aaa7dce6111c888429e51

Observation ef74a39f-c140-461e-9299-238a9cacd29d · outbound

This paper cites Efficient bayes-adaptive reinforcement learn- ing using sample-based search.

Fully Offline Reinforcement Learning Efficient bayes-adaptive reinforcement learn- ing using sample-based search

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:18.234691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:15:59.158171Z digest=sha256:896abcec214d9a5c9f8d45c1b8e21730fabd586ccd83d07b87ed9db76c74b402

Observation 66c454f4-a472-49ce-9f75-be6bfcf5830a · outbound

This paper cites Scalable and efficient bayes-adaptive reinforcement learning based on monte-carlo tree search.Journal of Artificial Intelligence Research, 48:841– 883, 10 2013.

Fully Offline Reinforcement Learning Scalable and efficient bayes-adaptive reinforcement learning based on monte-carlo tree search.Journal of Artificial Intelligence Research, 48:841– 883, 10 2013

Reference 27

Resolution
verified exact
doi, observed 2026-08-07T13:16:06.134303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:15:59.266350Z digest=sha256:a27aa93414f71557e1258f9bdab41b7dd91a1c508c0de4ddc223cecce903d86e

Observation 5ae9635e-e761-4966-9f5c-475929cbdaa3 · outbound

This paper cites Bayes-adaptive simulation-based search with value function approximation.

Fully Offline Reinforcement Learning Bayes-adaptive simulation-based search with value function approximation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:18.029970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:15:59.378247Z digest=sha256:546b0160ca967d12df3d776a39d0daf10e7484c3a2bd90b4622f7781d5d9b232

Observation 0bf909f1-3bb1-4868-9454-17edfad43c83 · outbound

This paper cites an unresolved cited work.

Fully Offline Reinforcement Learning Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:16:17.800520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:15:59.456794Z digest=sha256:220c29f939c75700d11f077e5227d224f493676744e32bb781ab46b66961ae00

Observation 9d4a2e4d-9dec-438c-8374-ef6516ec5cf1 · outbound

This paper cites A Clean Slate for Offline Reinforcement Learning.

Fully Offline Reinforcement Learning A Clean Slate for Offline Reinforcement Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:15:59.584911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:15:59.584911Z digest=sha256:b72ffd0d9a8f4a8196fccb87e722ba7460ee7d063301d79dbf8734ba359dc5e2

Observation cea3e1a3-04be-4a0e-b586-9593314cd513 · outbound

This paper cites Relu to the rescue: Improve your on-policy actor-critic with positive advantages.

Fully Offline Reinforcement Learning Relu to the rescue: Improve your on-policy actor-critic with positive advantages

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:17.425672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:15:59.840841Z digest=sha256:b2d53f052b4a5781acfd4801bbf5dc708f4f084061c7dae9e96b6614f70a6362

Observation 3d7c401b-dbee-484d-97e6-317c68901b68 · outbound

This paper cites Littman, and Anthony R.

Fully Offline Reinforcement Learning Littman, and Anthony R

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:17.259390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:15:59.939273Z digest=sha256:55c374ad60d0b75759c1a7e7bba43da9bd14ac2f17a4800525aaf9af58559be6

Observation c8b21251-6035-4794-b409-1b41bf81643a · outbound

This paper cites The validity of posterior expansions based on laplace’s method.Bayesian and Likelihood Methods in Statistics and Economics, pages 473–488, 1990.

Fully Offline Reinforcement Learning The validity of posterior expansions based on laplace’s method.Bayesian and Likelihood Methods in Statistics and Economics, pages 473–488, 1990

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:17.095540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:16:00.030618Z digest=sha256:4e1c29ca330d89689098c702f9f953d87be9f2a6ebce70730fd7f4d37013ebf6

Observation 5f93167a-8333-4b84-9aa6-02e28e1791d0 · outbound

This paper cites Morel: Model-based offline reinforcement learning.

Fully Offline Reinforcement Learning Morel: Model-based offline reinforcement learning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:16.938047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:16:00.149839Z digest=sha256:e086e461fee01da238a9da45b5fd2dac2db1a60c451e38af737e420a5e820778

Observation b1048332-956b-40ae-97f3-948b0f060334 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Fully Offline Reinforcement Learning Adam: A Method for Stochastic Optimization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:00.249033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:00.249033Z digest=sha256:26d11227820a0a40c5c4499999b09d096f2ba967dd23486bc4f910a7f70e2512

Observation 07d239ca-ddc3-4b44-bd0a-7fa83874c750 · outbound

This paper cites Kleijn and A.W.

Fully Offline Reinforcement Learning Kleijn and A.W

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:00.328567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:00.328567Z digest=sha256:73fba745008de1ab2c0fec4be0301214aecc81ea906d340209500844f570df0f

Observation 60336ccb-c3d8-410e-b70c-1f30de21c1e6 · outbound

This paper cites On asymptotic properties of predictive distributions.Biometrika, 83(2):299–313, 06 1996.

Fully Offline Reinforcement Learning On asymptotic properties of predictive distributions.Biometrika, 83(2):299–313, 06 1996

Reference 37

Resolution
verified exact
doi, observed 2026-08-07T13:16:05.898060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:16:00.414859Z digest=sha256:314059606ba6fdc29aab1f172bcc248793a6d6a3b60e551617532de127795a96

Observation 5a7d85c1-7902-4a0e-b581-c368ca353447 · outbound

This paper cites Offline Reinforcement Learning with Implicit Q-Learning.

Fully Offline Reinforcement Learning Offline Reinforcement Learning with Implicit Q-Learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:00.532185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:00.532185Z digest=sha256:56b3836ed035526f096d98670c6a68474ac9954c59ebda8964d89bb91cfb9537

Observation 30ecba1b-69b7-4d9b-858a-2f4e7484e894 · outbound

This paper cites an unresolved cited work.

Fully Offline Reinforcement Learning Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:16:16.756648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:16:00.631981Z digest=sha256:3f87f035646b06d869f76dfdc4c2ba48d395cf4c6aff5417cc0743df7a1976fe

Observation ffe6b090-67c1-4b10-80c7-98f457a65e28 · outbound

This paper cites Conserva- tive q-learning for offline reinforcement learning.

Fully Offline Reinforcement Learning Conserva- tive q-learning for offline reinforcement learning

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:16.535991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:16:00.720547Z digest=sha256:28c3785d8629bec2a2c6a9fbefdbd907c3805ea99532443474c975efa527956e

Observation ca65e117-c96a-4482-946a-1488ab71b6ce · outbound

This paper cites Springer Berlin Heidelberg, Berlin, Heidelberg, 2012.

Fully Offline Reinforcement Learning Springer Berlin Heidelberg, Berlin, Heidelberg, 2012

Reference 41

Resolution
verified exact
doi, observed 2026-08-07T13:16:05.713211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:16:00.803761Z digest=sha256:6cf12db326b711b67c7868f19beb29a8b42f1d304397a8956bd1a829cf3022d6

Observation 59003146-ce29-4e1a-b6de-9a5eeb746557 · outbound

This paper cites On some asymptotic properties of maximum likelihood estimates and related bayes’ estimates.

Fully Offline Reinforcement Learning On some asymptotic properties of maximum likelihood estimates and related bayes’ estimates

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:16.367331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:16:00.913603Z digest=sha256:56e431a458953e3d38bf2beae7ce3c8ed6ebb0934e895d204d01f5ea844a70f3

Observation c3b51491-c4f4-492e-853b-98679cf6c568 · outbound

This paper cites Efficient backprop.

Fully Offline Reinforcement Learning Efficient backprop

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:16.174298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:16:01.009737Z digest=sha256:839b69a09df9a3c8fa0a9f204ffef0196c57ec3043d85c20f4b74c65504d33bd

Observation f588b94a-fc8f-49ea-8df6-56381841e2ab · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

Fully Offline Reinforcement Learning Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:01.099532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:01.099532Z digest=sha256:9008974c7fcf24b619c04c038ac0953b7dd445a7e0578caa50299c5a70f5271a

Observation 763cabbd-dca8-4f36-9bd2-80503a8a572c · outbound

This paper cites an unresolved cited work.

Fully Offline Reinforcement Learning Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:16:15.993087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:16:01.205003Z digest=sha256:0e85a0b049a7c9bf2ed6849b1f88a20b19536fe53e2a86e14dac78c77e5c7525

Observation 7f2bc730-51c6-4fd0-bd3b-dfa8cec553e6 · outbound

This paper cites Discovered policy optimisation.Advances in Neural Information Processing Systems, 35:16455–16468, 2022.

Fully Offline Reinforcement Learning Discovered policy optimisation.Advances in Neural Information Processing Systems, 35:16455–16468, 2022

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:15.846851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:16:01.369676Z digest=sha256:6e6ec4c6d0d555639fffb9a5107a84e3e87a545f25526b54361d96a6fcdcc4a6

Observation c8edee02-d5e4-4aca-8ff9-e8d4973c8c16 · outbound

This paper cites an unresolved cited work.

Fully Offline Reinforcement Learning Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:16:15.734238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:16:01.541553Z digest=sha256:5c6999a677bb747d6c0ec3b4d117ed6b58758cb0fae8b590d05ab8b5a79de228

Observation 0efa23b4-2053-4f0e-bc2e-4046f7604aa4 · outbound

This paper cites an unresolved cited work.

Fully Offline Reinforcement Learning Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:16:15.552146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:16:01.622269Z digest=sha256:cdfcce6720d232477d65f62ca8f6ab092d1fe5f40bf08813f377c5e5e86ffc24

Observation 109b40e2-9355-43a8-919c-180a8a979deb · outbound

This paper cites Reinforcement learning: An overview, 2024.

Fully Offline Reinforcement Learning Reinforcement learning: An overview, 2024

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:01.713071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:01.713071Z digest=sha256:a5da0fc3d7334a6edb3701670999662dae3f9f5651d796a50e1666de290895de

Observation faa5c574-0182-4877-94e7-37959d84b22d · outbound

This paper cites an unresolved cited work.

Fully Offline Reinforcement Learning Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:16:15.280859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:16:01.806401Z digest=sha256:9befa6a294176e303a0e233e2843769892329ac157570aea83a7c929d334b966

Observation 5ad12ac9-c509-4e25-b824-4f25a4667b3f · outbound

This paper cites Randomized prior functions for deep reinforcement learning.

Fully Offline Reinforcement Learning Randomized prior functions for deep reinforcement learning

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:15.019714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:16:01.878513Z digest=sha256:5b3c9ed20da96f973edca6c0ed4249405c1f2b8353f76741d19c0c9d08d096bb

Observation 3f69e0b2-8fea-4d05-92fa-65b796d5717b · outbound

This paper cites Hyperparameter Selection for Offline Reinforcement Learning.

Fully Offline Reinforcement Learning Hyperparameter Selection for Offline Reinforcement Learning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:02.204750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:02.204750Z digest=sha256:ca7404fd2b6c93e7bdd0daa6532e12809847c32a2e0251b2f8ee6c003359b8e1

Observation 3cdb724d-205c-4577-bfcf-229b3a585e22 · outbound

This paper cites an unresolved cited work.

Fully Offline Reinforcement Learning Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:16:14.501296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:16:02.297041Z digest=sha256:5842ebf30dd8fb30531bceb1804275b5e8c0ef15880547304a2fcc0ee5f2fc36

Observation 01135a97-f1f5-4186-8595-cbb430cf666a · outbound

This paper cites Puterman.Markov Decision Processes: Discrete Stochastic Dynamic Programming.

Fully Offline Reinforcement Learning Puterman.Markov Decision Processes: Discrete Stochastic Dynamic Programming

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:14.254455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:16:02.466920Z digest=sha256:98505a5834eb75dd525212026a0494246885d9e709188356278e91525eeb739a

Observation d3db7484-5096-47a1-8f1d-99d7ad8a2b2e · outbound

This paper cites an unresolved cited work.

Fully Offline Reinforcement Learning Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:16:13.976552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:16:02.570369Z digest=sha256:bb84173f0dec2af364b26dd765bc10e43af09255f3f78c4ab63f165f6d14a094

Observation a6c1a408-8b9c-43d6-aae5-a7adf262494e · outbound

This paper cites Roberts and Jeffrey S.

Fully Offline Reinforcement Learning Roberts and Jeffrey S

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:02.646018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:02.646018Z digest=sha256:2bcfbfa1ec99541ef4daa7a01019f65bd723deddf65f8a72ed97f6ddcba14007

Observation 6548a700-b2b2-4587-8ecf-bf30c365eb12 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Fully Offline Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:02.763290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:02.763290Z digest=sha256:b42d1e4aa759be616fcaa86367419b5f92ceb4bc48cc37f28178ccdbaec4839c

Observation c9630eac-592d-458e-962e-7de77b2b2c11 · outbound

This paper cites The edge-of-reach problem in offline model-based reinforcement learning.

Fully Offline Reinforcement Learning The edge-of-reach problem in offline model-based reinforcement learning

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:13.702817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:16:02.940355Z digest=sha256:a9d4e6296b6977fb6664c4a18189685688c1056550dc217e98d30846d5c107f1

Observation a4892ae3-3add-4651-956c-4ebe72a9e2f7 · outbound

This paper cites Smallwood and Edward J.

Fully Offline Reinforcement Learning Smallwood and Edward J

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:13.474924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:16:03.098532Z digest=sha256:2837e7ca8b06076d82431e7fc7bec2afd8c43bab146b5225222ac74c0a6e6b8b

Observation 3dbaf9de-ba74-42ea-a427-5573b308c687 · outbound

This paper cites A Strong Baseline for Batch Imitation Learning.

Fully Offline Reinforcement Learning A Strong Baseline for Batch Imitation Learning

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:16:07.261939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:16:03.231306Z digest=sha256:4c499d598f4d313f19687f7ff48ccf5b5d20e55475d63a9e4d2dd1f686dac1d6

Observation 47b7cc56-f8b3-4513-a944-01b1baf5d6da · outbound

This paper cites Sriperumbudur, Kenji Fukumizu, Arthur Gretton, Bernhard Scholkopf, and Gert R.

Fully Offline Reinforcement Learning Sriperumbudur, Kenji Fukumizu, Arthur Gretton, Bernhard Scholkopf, and Gert R

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:13.245354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:16:03.346225Z digest=sha256:326047a2319ecbdfd23c72e632f64cdbba7598960404e2c546af2e196cb00956

Observation 9d4ce3f3-a4e3-43ae-9e00-b4997b3e5697 · outbound

This paper cites Model-Bellman inconsistency for model-based offline reinforcement learning.

Fully Offline Reinforcement Learning Model-Bellman inconsistency for model-based offline reinforcement learning

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:13.080447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:16:03.495666Z digest=sha256:a485cc320dbc208d2bd83c11256126d918003f26f4dc400b339cc8ac14a93df0

Observation 6f3323e5-478d-4efc-99b4-5fca93c0204d · outbound

This paper cites Sutton and Andrew G.

Fully Offline Reinforcement Learning Sutton and Andrew G

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:12.860042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:16:03.611536Z digest=sha256:5f09cb5ccbe63ee52b7114d6996c08492e681bc5d34c047a9d2ab1dd1a9f34b0

Observation 00410bc4-e85a-44a8-81b3-59f65070875b · outbound

This paper cites Algorithms for Reinforcement Learning.Synthesis Lectures on Artificial Intelligence and Machine Learning, 4(1):1–103, 2010.

Fully Offline Reinforcement Learning Algorithms for Reinforcement Learning.Synthesis Lectures on Artificial Intelligence and Machine Learning, 4(1):1–103, 2010

Reference 64

Resolution
verified exact
doi, observed 2026-08-07T13:16:05.532252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:16:03.712384Z digest=sha256:b71734e405d56dd5505ae9ef643ed3e43b1034c6f50cc14d124f22471f4a53d2

Observation 148591fc-2b94-4a6b-95e3-f4149d0dd418 · outbound

This paper cites Revisiting the minimalist approach to offline reinforcement learning.

Fully Offline Reinforcement Learning Revisiting the minimalist approach to offline reinforcement learning

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:12.677528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:16:03.845273Z digest=sha256:3368e48fc0c808e3dd91f2809ff3d3c896a7bc678b3bf72179d2271b9c7b1056

Observation 3f1d74b1-20f4-4f2d-b848-9d4d2ddcc294 · outbound

This paper cites an unresolved cited work.

Fully Offline Reinforcement Learning Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:16:12.382724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:16:03.993853Z digest=sha256:3e79b6cef428cd6c44d3093b6aa13870bbe2d8af27ce23796d3b378837d831c6

Observation 2c8874e2-5f69-48de-8f6f-a726461790cb · outbound

This paper cites Kass, and Joseph B.

Fully Offline Reinforcement Learning Kass, and Joseph B

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:12.102869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:16:04.077806Z digest=sha256:b15720d6d1e385f81914f5bc46d5499631801418dc08f88e7f66cfa8bbba8d97

Observation 52eacba7-dd3b-435e-9019-4de3411263b5 · outbound

This paper cites an unresolved cited work.

Fully Offline Reinforcement Learning Unresolved cited work

Reference 68

Resolution
verified exact
doi, observed 2026-08-07T13:16:05.370696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:16:04.164651Z digest=sha256:e74f80bc46bfaaaee3b51e3dd85bcd9cef6bdfa3920a31be2a0e0fee1bfb13a6

Observation 200ba08e-31ef-46ab-9499-8a18109ff2ee · outbound

This paper cites Information rates of nonparametric gaussian process methods.J.

Fully Offline Reinforcement Learning Information rates of nonparametric gaussian process methods.J

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:11.719273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:16:04.247779Z digest=sha256:ed74b0c64a2150ebe531e5585b495d5e4a04d43339fcd2050e72086699bef150

Observation a5ff4e8f-5141-4868-a201-40cd77bbf3f9 · outbound

This paper cites No more pesky hyperparameters: Offline hyperparameter tuning for RL.Transactions on Machine Learning Research, 2022.

Fully Offline Reinforcement Learning No more pesky hyperparameters: Offline hyperparameter tuning for RL.Transactions on Machine Learning Research, 2022

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:11.312056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:16:04.382384Z digest=sha256:3c80e27be2d42aa0fe253db70cdf420d8b6e06ed1750ac9f94528672a864e689

Observation 6db256f6-ee66-49ff-94bc-74333915930d · outbound

This paper cites Foster, and Sham M.

Fully Offline Reinforcement Learning Foster, and Sham M

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:10.928204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:16:04.503412Z digest=sha256:7fe6f05df66f3fffbd372aa9637124b425ad1e627b919b0c74253687c29419f8

Observation 377f2112-f27e-47e5-b138-44b928cd80b0 · outbound

This paper cites Differential-space.Journal of Mathematics and Physics, 2(1-4):131–174, 1923.

Fully Offline Reinforcement Learning Differential-space.Journal of Mathematics and Physics, 2(1-4):131–174, 1923

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:04.598127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:04.598127Z digest=sha256:038f4df0770d67db10043f67e5c28e1e58777776c26aa6eee8ded224ca4fa2fa

Observation d1f9866b-e8e5-4644-b4b1-5d10a190d644 · outbound

This paper cites Information-theoretic determination of minimax rates of convergence.The Annals of Statistics, 27(5):1564 – 1599, 1999.

Fully Offline Reinforcement Learning Information-theoretic determination of minimax rates of convergence.The Annals of Statistics, 27(5):1564 – 1599, 1999

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:04.682823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:04.682823Z digest=sha256:09c2d84169d8e2b6181f9eadfe54da4086424edec7fae97075f073a49a665236

Observation 32d5bb05-0908-42bb-9fad-fc95a8b99a71 · outbound

This paper cites Mopo: Model-based offline policy optimization.

Fully Offline Reinforcement Learning Mopo: Model-based offline policy optimization

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:10.584652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:16:04.761857Z digest=sha256:cb8e6a2383b6833d6f1b6e206b0453e9d6c2f0582d05051b519987ad52718cad

Observation c848b3c1-b773-4246-9597-50e98808cd6f · outbound

This paper cites Combo: Conservative offline model-based policy optimization.

Fully Offline Reinforcement Learning Combo: Conservative offline model-based policy optimization

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:10.295053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:16:04.859926Z digest=sha256:ad84e36775727c0f77a4a782ef19d39edcfecdf9b60191b72960c2665369a7dd

Observation d95966f2-0215-4863-a16c-934f6f31a520 · outbound

This paper cites On the importance of hyperparameter optimization for model-based reinforcement learning.

Fully Offline Reinforcement Learning On the importance of hyperparameter optimization for model-based reinforcement learning

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:09.953911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:16:04.932417Z digest=sha256:0e102aa81428c38552f5095765649cf1b2d51df04f81daf6f73a219d1e59a77f

Observation 01f6eb3b-010c-4d61-ab64-ee67d6026ba2 · outbound

This paper cites Varibad: A very good method for bayes-adaptive deep rl via meta- learning.

Fully Offline Reinforcement Learning Varibad: A very good method for bayes-adaptive deep rl via meta- learning

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:09.714445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:16:05.017057Z digest=sha256:d9a0c9e226a2c145402450ffe5e0fc90cb2ab80656756dd0d418301f92207536

Observation e4a4b1c3-7783-4bf3-86a2-7dba28d76125 · outbound

This paper cites N−1X i=0 1 N D 2 log(2π) + 1 2 D−1X d=0 logσ 2 θd (xi) + (yid −µ θd (xi))2 σ2 θd (xi) !!# , =Ei∼UN.

Fully Offline Reinforcement Learning N−1X i=0 1 N D 2 log(2π) + 1 2 D−1X d=0 logσ 2 θd (xi) + (yid −µ θd (xi))2 σ2 θd (xi) !!# , =Ei∼UN

Reference 78

Resolution
verified exact
raw_fallback, observed 2026-08-07T13:16:06.856195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:16:05.082852Z digest=sha256:bda5d5e9241d12cc6895cbcd6ec3a7cf0594a50edb0ac5ff4ec628e492a5bfc6

Observation 584f4e18-5a1d-4826-9888-eb69dfcfea54 · outbound

This paper cites Eθ∼PΘ(DN ).

Fully Offline Reinforcement Learning Eθ∼PΘ(DN )

Reference 83

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T13:16:06.605584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:16:05.159203Z digest=sha256:1efeeb9030ec0e5ebbc8266e0ea08ae8e34e91c0596dfbb2e1de609d80bac48d

Observation d995e4c3-886c-44f7-9c7a-9ca3fbf81918 · outbound

This paper cites doi: 10.1093/oso/9780198504856.003.0002.

Fully Offline Reinforcement Learning doi: 10.1093/oso/9780198504856.003.0002

Reference 1999

Resolution
unresolved
no resolver link, observed 2026-08-07T13:15:57.660083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:15:57.660083Z digest=sha256:d0ae240a5a30fc85e968398034db41d2680701dd882c796e4902a0e1371291e2

Observation aa4e3f11-d59b-4bd4-a552-fd3592fb27fd · outbound

This paper cites URL https://books.google.co.uk/books?id= s6mVlgEACAAJ.

Fully Offline Reinforcement Learning URL https://books.google.co.uk/books?id= s6mVlgEACAAJ

Reference 2013

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:20.621497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:15:57.884088Z digest=sha256:8cd5730dad0ce3bce1c6ed398ab44073cae6bdc707b7322c1fc2c1348fa89400

Observation 1e434d6f-04c0-42fb-b1dc-edaa8c4dbd63 · outbound

This paper cites an unresolved cited work.

Fully Offline Reinforcement Learning Unresolved cited work

Reference 2025

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:16:17.599212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:15:59.744897Z digest=sha256:46a9b82e644dae2da9d5363bce31f270b606fb24a37db311510169a3938304c6

Observation 8c85133c-fff4-4179-b7a6-273632e7103f · outbound

This paper cites URL http://papers.nips.cc/paper/8080- randomized-prior-functions-for-deep-reinforcement-learning.pdf.

Fully Offline Reinforcement Learning URL http://papers.nips.cc/paper/8080- randomized-prior-functions-for-deep-reinforcement-learning.pdf

Reference 8629

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:14.762978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:16:02.029997Z digest=sha256:6d167ce28d9e7928311b0c3627acd91a7fa5f635dd5ee8ab10afae7ebc1e03f0

Pith citing papers

Observation 7a38f31e-998b-44d0-a06c-e8f6a48dbd1e · inbound

Long-Horizon Model-Based Offline Reinforcement Learning Without Explicit Conservatism cites this paper.

Long-Horizon Model-Based Offline Reinforcement Learning Without Explicit Conservatism Fully Offline Reinforcement Learning

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-17T00:20:38.764461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-17T01:46:16.923948Z digest=sha256:7f7ea6e85f26c7726a009355ed042f7cf29377b5fd436e464f934982db45786b

Observation 00b3f2c1-259b-40d1-9384-9e49144431e9 · inbound

Is Inter-Seed Cross-Play Enough? Evaluating the Robustness of Zero-Shot Coordination Algorithms to Implementation Details cites this paper.

Is Inter-Seed Cross-Play Enough? Evaluating the Robustness of Zero-Shot Coordination Algorithms to Implementation Details Fully Offline Reinforcement Learning

Reference 117

Resolution
unresolved
no resolver link, observed 2026-08-05T15:25:39.829506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:25:39.829506Z digest=sha256:ccd5bb32ac1e93b830e7ababde414a898e78de35dd72d235ff1d39ac8fe8654a