Pith. sign in

Paper Citation Record · LEDGER

Fully Offline Reinforcement Learning

As of 9 August 2026, this Paper Citation Record lists 83 of 83 outbound references and 2 inbound Pith citation observations for arXiv:2505.22442.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22442 v3

Coverage vector

measured 83 of 83 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:16:05.159203Z

measured 85 of 85 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T15:25:39.829506Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T01:48:51.121586Z

Reference resolution

83 of 83 outbound references displayed

  • verified exact11
  • verified fuzzy41
  • unresolved28
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7f4fa7dd-2576-4833-918f-082dad491144 · outbound

This paper cites Aitchison.

Fully Offline Reinforcement Learning Aitchison

Reference 1

Resolution
verified exact
raw_fallback, observed 2026-08-07T13:16:09.424246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:15:57.315944Z digest=sha256:339b44c9c9d3a715a4b6f988c57fda7651e6975b30f309fb578b017b08ce1058

Observation aa23fa33-78ea-43be-ad09-ec04a2ec318a · outbound

This paper cites Alaa and Mihaela van der Schaar.

Fully Offline Reinforcement Learning Alaa and Mihaela van der Schaar

Reference 2

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T13:16:09.081121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:15:57.356190Z digest=sha256:be435c5da8573dfb0114e6f20de1fadef08e08076848cc8fe6d2801a9a03c818

Observation faa3004a-74fb-466a-b59d-28e8b50b6332 · outbound

This paper cites Uncertainty-based offline reinforcement learning with diversified q-ensemble.

Fully Offline Reinforcement Learning Uncertainty-based offline reinforcement learning with diversified q-ensemble

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:15:57.413976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:15:57.413976Z digest=sha256:e4c2c90a93630f1e098d7b1fbf5f545f9851e63ce1ad718494a907e8fe3dc77e

Observation 3ad69933-f655-4d69-8dee-5f498d7ff47c · outbound

This paper cites Asymptotically minimax bayes predictive densities.The Annals of Statistics, 34(6):2921–2938, 2006.

Fully Offline Reinforcement Learning Asymptotically minimax bayes predictive densities.The Annals of Statistics, 34(6):2921–2938, 2006

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:15:57.474090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:15:57.474090Z digest=sha256:3d6541d4b2b1c240c873f89560700657d55d524489fba50a15f8889ef7a37fe0

Observation 02430707-3cef-41eb-acea-41eedfcd62cc · outbound

This paper cites Augmented world models facilitate zero-shot dynamics generalization from a single offline environment.

Fully Offline Reinforcement Learning Augmented world models facilitate zero-shot dynamics generalization from a single offline environment

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:15:57.536060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:15:57.536060Z digest=sha256:53279ddb823255a91ae555050935f008eec703bcd8c4629f02edf3b96b96018a

Observation b48111ec-9df1-470f-b256-4122791e1c06 · outbound

This paper cites Information-theoretic characterization of bayes performance and the choice of priors in parametric and nonparametric problems.

Fully Offline Reinforcement Learning Information-theoretic characterization of bayes performance and the choice of priors in parametric and nonparametric problems

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:21.255216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:15:57.590263Z digest=sha256:d4d589b0e8f4785732f7790ecbc46c1f9fe3829df5b9e8b37b8185b59cd63438

Observation ab150b87-375f-4aca-b853-765cb9f73e2f · outbound

This paper cites Barron.The Exponential Convergence of Posterior Probabilities with Implications for Bayes Estimators of Density Functions.

Fully Offline Reinforcement Learning Barron.The Exponential Convergence of Posterior Probabilities with Implications for Bayes Estimators of Density Functions

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:20.989180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:15:57.757860Z digest=sha256:06aaa6e7472bec9ec0823dd53707c2a58f718112c76637564f7dc3d687d90a76

Observation d16ea67a-4b7d-4dcf-83cb-b64d4ddd1675 · outbound

This paper cites Bass.Real Analysis for Graduate Students, chapter 21.

Fully Offline Reinforcement Learning Bass.Real Analysis for Graduate Students, chapter 21

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:20.820274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:15:57.822068Z digest=sha256:b1a9f3d38db051061ab5f4571d52a18228dd52e332408c792571d5441e15f6d3

Observation 14fff60b-933b-4eff-8483-c0f9dc237f61 · outbound

This paper cites A Tutorial on Meta-Reinforcement Learning.

Fully Offline Reinforcement Learning A Tutorial on Meta-Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:15:57.942175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:15:57.942175Z digest=sha256:ba84e251aa96ab08db4fa3cf51d2f2fa6077e062c6feb877d06444cf5bfb5bf0

Observation c275b9b1-b9bd-4ad5-a6de-e2c89d65fe86 · outbound

This paper cites A problem in the sequential design of experiments.Sankhy ¯a: The Indian Journal of Statistics (1933-1960), 16(3/4):221–229, 1956.

Fully Offline Reinforcement Learning A problem in the sequential design of experiments.Sankhy ¯a: The Indian Journal of Statistics (1933-1960), 16(3/4):221–229, 1956

Reference 10

Resolution
verified exact
raw_fallback, observed 2026-08-07T13:16:08.501815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:15:58.001783Z digest=sha256:2317647de3655ea6ceeebddc464096a3b3cb181591f955d6064f19dfbe726fff

Observation 53dbe3be-d34e-4b8f-b8c4-eeae6d108a02 · outbound

This paper cites Dynamic programming and stochastic control processes.Information and Control, 1(3):228–239, 1958.

Fully Offline Reinforcement Learning Dynamic programming and stochastic control processes.Information and Control, 1(3):228–239, 1958

Reference 11

Resolution
malformed identifier
no resolver link, observed 2026-08-07T13:15:58.052512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:15:58.052512Z digest=sha256:160fc5f25c001d4c1e31650e2ac576985e3e9142d6e4225bce6b1b4081c75771

Observation ad5273c6-cfb0-4297-8b5b-d36395eea30e · outbound

This paper cites Foster, and Daniel M.

Fully Offline Reinforcement Learning Foster, and Daniel M

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:20.427736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:15:58.116894Z digest=sha256:4b225d159baced35d8559311b240437f73f9558d4c9b3d8ba615394eb82337eb

Observation 3f3dade5-de7b-468d-8deb-ae49eabbfc43 · outbound

This paper cites On the foundations of statistical inference.Journal of the American Statistical Association, 57(298):269–306, 1962.

Fully Offline Reinforcement Learning On the foundations of statistical inference.Journal of the American Statistical Association, 57(298):269–306, 1962

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:15:58.171525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:15:58.171525Z digest=sha256:fa4bfdf6f479ea089b252dd3e0b65d3f25488c03a0ca3e1f4859b2515a861933

Observation 50ccbea6-642b-4c41-abfb-6ccd7d75f62c · outbound

This paper cites an unresolved cited work.

Fully Offline Reinforcement Learning Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:16:20.279450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:15:58.236104Z digest=sha256:bcf464cec7d65055cc9a07dc63d063a3f898ccaeaa200fb96433eefaa1a85884

Observation de0a4ee4-4b43-46b8-9c14-1c86432ce338 · outbound

This paper cites Bayes adaptive monte carlo tree search for of- fline model-based reinforcement learning, 2024.

Fully Offline Reinforcement Learning Bayes adaptive monte carlo tree search for of- fline model-based reinforcement learning, 2024

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:20.135149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:15:58.297640Z digest=sha256:be41656692ca9b33f70f1cf09e9365365e4bf312c0fd968d8b886101b9f9b8df

Observation cfe50663-3e99-4113-815f-551afb1c29ee · outbound

This paper cites Conser- vative uncertainty estimation by fitting prior networks.

Fully Offline Reinforcement Learning Conser- vative uncertainty estimation by fitting prior networks

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:19.983352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:15:58.372662Z digest=sha256:0cb8175851e5d3660f02a2530bc635c39ff965314fb4fe8428e2106289413923

Observation 016295ea-6bda-4420-99f2-2c9d7ae85950 · outbound

This paper cites Clarke and A.R.

Fully Offline Reinforcement Learning Clarke and A.R

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:19.743980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:15:58.448968Z digest=sha256:c6c2409f092cdd5302e1be129b9b022241f6486d2c83ab5be15dd457bfbaaa0a

Observation 4ba4c40f-40e6-47f5-a44a-23eb99b56a5f · outbound

This paper cites an unresolved cited work.

Fully Offline Reinforcement Learning Unresolved cited work

Reference 18

Resolution
verified exact
raw_fallback, observed 2026-08-07T13:16:07.959162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:15:58.500423Z digest=sha256:ff8358d3cdb3f14277315e75c2511cfceb2f01d6a025f7c310a0ef5ddeaf47ef

Observation 54158c56-5f57-431f-be73-7a763611224d · outbound

This paper cites Observation of a markov process through a noisy channel.PhD Thesis, 1962.

Fully Offline Reinforcement Learning Observation of a markov process through a noisy channel.PhD Thesis, 1962

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:19.551859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:15:58.563857Z digest=sha256:c1f71de12c2af93d2eee0ad4059af5a348b008a5ec06f1545764e7d15fb467e4

Observation d5475ef1-f8d4-4265-9802-3e75ba2d6176 · outbound

This paper cites Fast reinforcement learning via slow reinforcement learning.

Fully Offline Reinforcement Learning Fast reinforcement learning via slow reinforcement learning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:19.341391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:15:58.618468Z digest=sha256:1897f2aa75886781d8f232f024a64e2875d759769843a6e5ddf223887d12b215

Observation fd95ba64-4cef-4f53-a399-142fcde22945 · outbound

This paper cites PhD thesis, 2002.

Fully Offline Reinforcement Learning PhD thesis, 2002

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:19.101212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:15:58.684957Z digest=sha256:3096738bce322cb59af52e2aeae38b7412deede5ab5c8e5162489957c2745d1c

Observation b34d3302-8720-46c0-88bf-8404861c1c57 · outbound

This paper cites Bayesian exploration networks.

Fully Offline Reinforcement Learning Bayesian exploration networks

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:18.931137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:15:58.801124Z digest=sha256:57b752dd9350605ae9386a7934faff0fafa28755c2b1545402f6810049a885f7

Observation af8f42f6-eebf-41e4-a8ed-29b0c9113c05 · outbound

This paper cites D4rl: Datasets for deep data-driven reinforcement learning, 2020.

Fully Offline Reinforcement Learning D4rl: Datasets for deep data-driven reinforcement learning, 2020

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:18.703635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:15:58.874936Z digest=sha256:85b11cc2af792a37d0b14626b940d04be3e5e2af907375e6fa334915620a05fb

Observation 37eca544-1166-4b74-b6d2-90890599c766 · outbound

This paper cites A minimalist approach to offline reinforcement learn- ing.

Fully Offline Reinforcement Learning A minimalist approach to offline reinforcement learn- ing

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:18.497264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:15:58.962879Z digest=sha256:57ca71207b80623658225919c5d744e1755dc0504249effd36b216dae1442cd2

Observation b819815a-87ac-40bc-8d62-8771f3c71ef1 · outbound

This paper cites A new proof of the likelihood principle.The British Journal for the Philosophy of Science, 66(3):475–503, 2015.

Fully Offline Reinforcement Learning A new proof of the likelihood principle.The British Journal for the Philosophy of Science, 66(3):475–503, 2015

Reference 25

Resolution
verified exact
doi, observed 2026-08-07T13:16:06.339819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:15:59.057545Z digest=sha256:fc6e4c6dd84bd59d6cb8fe07f649e6fb0fee7cf3665d89dfe5d98eff4af7af1c

Observation ef74a39f-c140-461e-9299-238a9cacd29d · outbound

This paper cites Efficient bayes-adaptive reinforcement learn- ing using sample-based search.

Fully Offline Reinforcement Learning Efficient bayes-adaptive reinforcement learn- ing using sample-based search

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:18.234691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:15:59.158171Z digest=sha256:b7d0bb9d03df65fd6b18c97cf4fd56b396d65802a12e8cd9d85d5a7107b9ba99

Observation 66c454f4-a472-49ce-9f75-be6bfcf5830a · outbound

This paper cites Scalable and efficient bayes-adaptive reinforcement learning based on monte-carlo tree search.Journal of Artificial Intelligence Research, 48:841– 883, 10 2013.

Fully Offline Reinforcement Learning Scalable and efficient bayes-adaptive reinforcement learning based on monte-carlo tree search.Journal of Artificial Intelligence Research, 48:841– 883, 10 2013

Reference 27

Resolution
verified exact
doi, observed 2026-08-07T13:16:06.134303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:15:59.266350Z digest=sha256:9b199e053e5e6beab8ddcbdee1c52f431ce4da27186dc5df6515b99ee5e21e75

Observation 5ae9635e-e761-4966-9f5c-475929cbdaa3 · outbound

This paper cites Bayes-adaptive simulation-based search with value function approximation.

Fully Offline Reinforcement Learning Bayes-adaptive simulation-based search with value function approximation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:18.029970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:15:59.378247Z digest=sha256:997029c1af2b8cb3490851fb734836c4e7463f91cf340d596f0d4480f9649e59

Observation 0bf909f1-3bb1-4868-9454-17edfad43c83 · outbound

This paper cites an unresolved cited work.

Fully Offline Reinforcement Learning Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:16:17.800520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:15:59.456794Z digest=sha256:e1af9d63194e04ad294c1f25ea7f06fa60758c19b71c281303f5a8ccea144f2d

Observation 9d4a2e4d-9dec-438c-8374-ef6516ec5cf1 · outbound

This paper cites A Clean Slate for Offline Reinforcement Learning.

Fully Offline Reinforcement Learning A Clean Slate for Offline Reinforcement Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:15:59.584911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:15:59.584911Z digest=sha256:c5ad76453fee4b3814b720c188ac50fe2c5f7e7529485112f22dffb752d4a7b9

Observation cea3e1a3-04be-4a0e-b586-9593314cd513 · outbound

This paper cites Relu to the rescue: Improve your on-policy actor-critic with positive advantages.

Fully Offline Reinforcement Learning Relu to the rescue: Improve your on-policy actor-critic with positive advantages

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:17.425672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:15:59.840841Z digest=sha256:18d37eb7ae706cbc1c470c4929da74fe9322fe3e7c0907b1d72634852387be08

Observation 3d7c401b-dbee-484d-97e6-317c68901b68 · outbound

This paper cites Littman, and Anthony R.

Fully Offline Reinforcement Learning Littman, and Anthony R

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:17.259390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:15:59.939273Z digest=sha256:5372adba8fee3ef3b88e0fdd282b39a47949dde8f54f47493db3624977ef3702

Observation c8b21251-6035-4794-b409-1b41bf81643a · outbound

This paper cites The validity of posterior expansions based on laplace’s method.Bayesian and Likelihood Methods in Statistics and Economics, pages 473–488, 1990.

Fully Offline Reinforcement Learning The validity of posterior expansions based on laplace’s method.Bayesian and Likelihood Methods in Statistics and Economics, pages 473–488, 1990

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:17.095540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:00.030618Z digest=sha256:1043c87ab917259912c0d1956247b0bfae7e78bf6d28d10afeb63aae99cb9a53

Observation 5f93167a-8333-4b84-9aa6-02e28e1791d0 · outbound

This paper cites Morel: Model-based offline reinforcement learning.

Fully Offline Reinforcement Learning Morel: Model-based offline reinforcement learning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:16.938047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:00.149839Z digest=sha256:8f081eae748c1597e1fe00d1002963b4e2aa16bc3fdc52772c0658bc73716e2d

Observation b1048332-956b-40ae-97f3-948b0f060334 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Fully Offline Reinforcement Learning Adam: A Method for Stochastic Optimization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:00.249033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:00.249033Z digest=sha256:a29c7ae718cb2a9aa6117b907f7ad43863c53c696852df1b90fa90b97c5c5d88

Observation 07d239ca-ddc3-4b44-bd0a-7fa83874c750 · outbound

This paper cites Kleijn and A.W.

Fully Offline Reinforcement Learning Kleijn and A.W

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:00.328567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:00.328567Z digest=sha256:b35e59eda34c3abd56a3e858d66ad60979fdea7959d104d8442b3fc763e9b99e

Observation 60336ccb-c3d8-410e-b70c-1f30de21c1e6 · outbound

This paper cites On asymptotic properties of predictive distributions.Biometrika, 83(2):299–313, 06 1996.

Fully Offline Reinforcement Learning On asymptotic properties of predictive distributions.Biometrika, 83(2):299–313, 06 1996

Reference 37

Resolution
verified exact
doi, observed 2026-08-07T13:16:05.898060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:00.414859Z digest=sha256:dd975231b5c8aad5ac62d9a3ed99f04518b0f36e3bf4f3fb4a73e0fadda22930

Observation 5a7d85c1-7902-4a0e-b581-c368ca353447 · outbound

This paper cites Offline Reinforcement Learning with Implicit Q-Learning.

Fully Offline Reinforcement Learning Offline Reinforcement Learning with Implicit Q-Learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:00.532185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:00.532185Z digest=sha256:86d6197eec9c28ed503c4dc3f5cc73caf80fd9a9eaaf92399372793b50633993

Observation 30ecba1b-69b7-4d9b-858a-2f4e7484e894 · outbound

This paper cites an unresolved cited work.

Fully Offline Reinforcement Learning Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:16:16.756648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:00.631981Z digest=sha256:494e867de628133b9635704c254322e8fadbf615ebca562f1b0cef46a78a6e12

Observation ffe6b090-67c1-4b10-80c7-98f457a65e28 · outbound

This paper cites Conserva- tive q-learning for offline reinforcement learning.

Fully Offline Reinforcement Learning Conserva- tive q-learning for offline reinforcement learning

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:16.535991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:00.720547Z digest=sha256:6de80518d79e721517c3052ec4ca68a9655b816bc7570a00d2ade8d722f41b69

Observation ca65e117-c96a-4482-946a-1488ab71b6ce · outbound

This paper cites Springer Berlin Heidelberg, Berlin, Heidelberg, 2012.

Fully Offline Reinforcement Learning Springer Berlin Heidelberg, Berlin, Heidelberg, 2012

Reference 41

Resolution
verified exact
doi, observed 2026-08-07T13:16:05.713211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:00.803761Z digest=sha256:eaf33266a394e02decf5d552044f9488ad9cc9f6f9e60a364481e7df188ae0f0

Observation 59003146-ce29-4e1a-b6de-9a5eeb746557 · outbound

This paper cites On some asymptotic properties of maximum likelihood estimates and related bayes’ estimates.

Fully Offline Reinforcement Learning On some asymptotic properties of maximum likelihood estimates and related bayes’ estimates

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:16.367331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:00.913603Z digest=sha256:dcfccbef6b116ca05f071e9a84ea7989e93e6e8235bc9d25a56556d67e2d55c6

Observation c3b51491-c4f4-492e-853b-98679cf6c568 · outbound

This paper cites Efficient backprop.

Fully Offline Reinforcement Learning Efficient backprop

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:16.174298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:01.009737Z digest=sha256:b02a3f60bc9ada672cf9340335799ed70fd4d49660753adb718144fc12225b88

Observation f588b94a-fc8f-49ea-8df6-56381841e2ab · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

Fully Offline Reinforcement Learning Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:01.099532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:01.099532Z digest=sha256:98e3849b1f2b34b9961f2feb46b193ddf54e941ba75b01dc0acf70babd87a7ee

Observation 763cabbd-dca8-4f36-9bd2-80503a8a572c · outbound

This paper cites an unresolved cited work.

Fully Offline Reinforcement Learning Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:16:15.993087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:01.205003Z digest=sha256:b5049985901a1e4194e8c3488f54bef35662edd43b22d834eeede8071eb38222

Observation 7f2bc730-51c6-4fd0-bd3b-dfa8cec553e6 · outbound

This paper cites Discovered policy optimisation.Advances in Neural Information Processing Systems, 35:16455–16468, 2022.

Fully Offline Reinforcement Learning Discovered policy optimisation.Advances in Neural Information Processing Systems, 35:16455–16468, 2022

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:15.846851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:01.369676Z digest=sha256:a582180f4143ed7e56de730657d0cb5f515b3c010986618de1debd7580fae1b3

Observation c8edee02-d5e4-4aca-8ff9-e8d4973c8c16 · outbound

This paper cites an unresolved cited work.

Fully Offline Reinforcement Learning Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:16:15.734238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:01.541553Z digest=sha256:f30ad8960e6e2fe8da9ca33c040dae9ded76fd0b0e87124d90c7dc8e21b9356c

Observation 0efa23b4-2053-4f0e-bc2e-4046f7604aa4 · outbound

This paper cites an unresolved cited work.

Fully Offline Reinforcement Learning Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:16:15.552146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:01.622269Z digest=sha256:79ac434beab9980b86df6977055ffc74c483c65e21ef1f885a379ef109a50e97

Observation 109b40e2-9355-43a8-919c-180a8a979deb · outbound

This paper cites Reinforcement learning: An overview, 2024.

Fully Offline Reinforcement Learning Reinforcement learning: An overview, 2024

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:01.713071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:01.713071Z digest=sha256:6a6c58ea7e3796947f87a99c837c7863deacdb19aa3587748a8726be18e57d9f

Observation faa5c574-0182-4877-94e7-37959d84b22d · outbound

This paper cites an unresolved cited work.

Fully Offline Reinforcement Learning Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:16:15.280859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:01.806401Z digest=sha256:2c975d9afbb72ce7d5885053ffd47e75b917b9b4df56b8d69de3810b68b0409e

Observation 5ad12ac9-c509-4e25-b824-4f25a4667b3f · outbound

This paper cites Randomized prior functions for deep reinforcement learning.

Fully Offline Reinforcement Learning Randomized prior functions for deep reinforcement learning

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:15.019714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:01.878513Z digest=sha256:a30816e579c63fce26f908aa8a0d10bc46ba1934983ffbbb39289c6885a97c94

Observation 3f69e0b2-8fea-4d05-92fa-65b796d5717b · outbound

This paper cites Hyperparameter Selection for Offline Reinforcement Learning.

Fully Offline Reinforcement Learning Hyperparameter Selection for Offline Reinforcement Learning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:02.204750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:02.204750Z digest=sha256:4089d2661da842fc93698556087acb04c0e08a1451055cd1de5e5abbb1332491

Observation 3cdb724d-205c-4577-bfcf-229b3a585e22 · outbound

This paper cites an unresolved cited work.

Fully Offline Reinforcement Learning Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:16:14.501296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:02.297041Z digest=sha256:8a6b7204644bcdf5d101bea79324df50543890e054837f67fe8fc09ee92f7682

Observation 01135a97-f1f5-4186-8595-cbb430cf666a · outbound

This paper cites Puterman.Markov Decision Processes: Discrete Stochastic Dynamic Programming.

Fully Offline Reinforcement Learning Puterman.Markov Decision Processes: Discrete Stochastic Dynamic Programming

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:14.254455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:02.466920Z digest=sha256:1bfebef8dd878229cd0f7120e989504e849d2f62a657b081fdec399a3e947bd3

Observation d3db7484-5096-47a1-8f1d-99d7ad8a2b2e · outbound

This paper cites an unresolved cited work.

Fully Offline Reinforcement Learning Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:16:13.976552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:02.570369Z digest=sha256:c6192b8c1c17514b617a6f30bf46da1768fa29884ddba75be9446c69ed9df10a

Observation a6c1a408-8b9c-43d6-aae5-a7adf262494e · outbound

This paper cites Roberts and Jeffrey S.

Fully Offline Reinforcement Learning Roberts and Jeffrey S

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:02.646018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:02.646018Z digest=sha256:b20081cc2ae0147a76178f6ada8f22e282bc6498499e10f80ad3e849f94eab25

Observation 6548a700-b2b2-4587-8ecf-bf30c365eb12 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Fully Offline Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:02.763290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:02.763290Z digest=sha256:a34212ffa48e94f422ff0e355e9ea8ec1e1dfac8966b8f4918ef39596dac8400

Observation c9630eac-592d-458e-962e-7de77b2b2c11 · outbound

This paper cites The edge-of-reach problem in offline model-based reinforcement learning.

Fully Offline Reinforcement Learning The edge-of-reach problem in offline model-based reinforcement learning

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:13.702817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:02.940355Z digest=sha256:b052696c53aa3c9bb05ec74812b32f46d260f8138d48b1892cc8991d94099af6

Observation a4892ae3-3add-4651-956c-4ebe72a9e2f7 · outbound

This paper cites Smallwood and Edward J.

Fully Offline Reinforcement Learning Smallwood and Edward J

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:13.474924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:03.098532Z digest=sha256:5ccee3fa83d2828fd935ce606f5ed5ffecaa5d34841214a04dfca64e38f1b81e

Observation 3dbaf9de-ba74-42ea-a427-5573b308c687 · outbound

This paper cites A Strong Baseline for Batch Imitation Learning.

Fully Offline Reinforcement Learning A Strong Baseline for Batch Imitation Learning

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:16:07.261939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:03.231306Z digest=sha256:7401a9e6f20f7a0d8cdc9a1b9e3342ae7d89364b27455c39fa4ab96887345a49

Observation 47b7cc56-f8b3-4513-a944-01b1baf5d6da · outbound

This paper cites Sriperumbudur, Kenji Fukumizu, Arthur Gretton, Bernhard Scholkopf, and Gert R.

Fully Offline Reinforcement Learning Sriperumbudur, Kenji Fukumizu, Arthur Gretton, Bernhard Scholkopf, and Gert R

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:13.245354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:03.346225Z digest=sha256:069c0f1a609253fe684d87d6564d4a2e466489e62a2d428c7bdae539efbf0fb5

Observation 9d4ce3f3-a4e3-43ae-9e00-b4997b3e5697 · outbound

This paper cites Model-Bellman inconsistency for model-based offline reinforcement learning.

Fully Offline Reinforcement Learning Model-Bellman inconsistency for model-based offline reinforcement learning

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:13.080447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:03.495666Z digest=sha256:115e469ca386bd4b66585303f91f9ee8c651a3519bccb16e375bacf8c8dcea5d

Observation 6f3323e5-478d-4efc-99b4-5fca93c0204d · outbound

This paper cites Sutton and Andrew G.

Fully Offline Reinforcement Learning Sutton and Andrew G

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:12.860042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:03.611536Z digest=sha256:114c82a00ddc02c2dac364947b3b3e60029513bab8495706bf1c687c1ee5a83f

Observation 00410bc4-e85a-44a8-81b3-59f65070875b · outbound

This paper cites Algorithms for Reinforcement Learning.Synthesis Lectures on Artificial Intelligence and Machine Learning, 4(1):1–103, 2010.

Fully Offline Reinforcement Learning Algorithms for Reinforcement Learning.Synthesis Lectures on Artificial Intelligence and Machine Learning, 4(1):1–103, 2010

Reference 64

Resolution
verified exact
doi, observed 2026-08-07T13:16:05.532252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:03.712384Z digest=sha256:7b3e8699a13667c9010af4e99d1972284cf01b87ebe81031045300f686bed0a9

Observation 148591fc-2b94-4a6b-95e3-f4149d0dd418 · outbound

This paper cites Revisiting the minimalist approach to offline reinforcement learning.

Fully Offline Reinforcement Learning Revisiting the minimalist approach to offline reinforcement learning

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:12.677528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:03.845273Z digest=sha256:5ef3c5078ab47c2294223dd5db5204e00da695c3e6cbb794f4fc8289a077994c

Observation 3f1d74b1-20f4-4f2d-b848-9d4d2ddcc294 · outbound

This paper cites an unresolved cited work.

Fully Offline Reinforcement Learning Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:16:12.382724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:03.993853Z digest=sha256:cd1cc74ef00c8d59b8cf045a11b491125bf7b0731c3fced6fff0041667dede9e

Observation 2c8874e2-5f69-48de-8f6f-a726461790cb · outbound

This paper cites Kass, and Joseph B.

Fully Offline Reinforcement Learning Kass, and Joseph B

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:12.102869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:04.077806Z digest=sha256:22e1603f2ef44896441c2f993b9525e8986a4de177df9decfb8447fd0100f4bc

Observation 52eacba7-dd3b-435e-9019-4de3411263b5 · outbound

This paper cites an unresolved cited work.

Fully Offline Reinforcement Learning Unresolved cited work

Reference 68

Resolution
verified exact
doi, observed 2026-08-07T13:16:05.370696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:04.164651Z digest=sha256:f149ad50ebff8bf8a703895bad117e28d27c05d71fd5ee7064e51f6e600819d6

Observation 200ba08e-31ef-46ab-9499-8a18109ff2ee · outbound

This paper cites Information rates of nonparametric gaussian process methods.J.

Fully Offline Reinforcement Learning Information rates of nonparametric gaussian process methods.J

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:11.719273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:04.247779Z digest=sha256:c408a1b32865eb5fdae38c6571c674e3385c3c39b2dc3137a854fe04d1cb3fe2

Observation a5ff4e8f-5141-4868-a201-40cd77bbf3f9 · outbound

This paper cites No more pesky hyperparameters: Offline hyperparameter tuning for RL.Transactions on Machine Learning Research, 2022.

Fully Offline Reinforcement Learning No more pesky hyperparameters: Offline hyperparameter tuning for RL.Transactions on Machine Learning Research, 2022

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:11.312056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:04.382384Z digest=sha256:450367bb4d8cd8a7eb3618539919ccd706390b3120a4df469d6d9bbfe1ec8a11

Observation 6db256f6-ee66-49ff-94bc-74333915930d · outbound

This paper cites Foster, and Sham M.

Fully Offline Reinforcement Learning Foster, and Sham M

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:10.928204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:04.503412Z digest=sha256:55ceb8e2582beee6e6db1c8a6bb34892cd0ccf894ed519bdce9ca1e36969c386

Observation 377f2112-f27e-47e5-b138-44b928cd80b0 · outbound

This paper cites Differential-space.Journal of Mathematics and Physics, 2(1-4):131–174, 1923.

Fully Offline Reinforcement Learning Differential-space.Journal of Mathematics and Physics, 2(1-4):131–174, 1923

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:04.598127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:04.598127Z digest=sha256:4e2f3c40c3eced4b30a09df815c33ddb883fe2b2379223a0bceb20243c6bccfe

Observation d1f9866b-e8e5-4644-b4b1-5d10a190d644 · outbound

This paper cites Information-theoretic determination of minimax rates of convergence.The Annals of Statistics, 27(5):1564 – 1599, 1999.

Fully Offline Reinforcement Learning Information-theoretic determination of minimax rates of convergence.The Annals of Statistics, 27(5):1564 – 1599, 1999

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:04.682823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:04.682823Z digest=sha256:bd5f52d48b2d3f1631b5b7849d3b3743f6fe069943a90be7b5fbbabd2a6228f3

Observation 32d5bb05-0908-42bb-9fad-fc95a8b99a71 · outbound

This paper cites Mopo: Model-based offline policy optimization.

Fully Offline Reinforcement Learning Mopo: Model-based offline policy optimization

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:10.584652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:04.761857Z digest=sha256:363a68ddd6d4574e6aba0facdc4d3f7c19252d5e7b2e70439d78f703045b4c69

Observation c848b3c1-b773-4246-9597-50e98808cd6f · outbound

This paper cites Combo: Conservative offline model-based policy optimization.

Fully Offline Reinforcement Learning Combo: Conservative offline model-based policy optimization

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:10.295053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:04.859926Z digest=sha256:e2ebc0cc4878d43176aeeaafa5fed99c2c7cc3939b2289808f8c1dee0c749c91

Observation d95966f2-0215-4863-a16c-934f6f31a520 · outbound

This paper cites On the importance of hyperparameter optimization for model-based reinforcement learning.

Fully Offline Reinforcement Learning On the importance of hyperparameter optimization for model-based reinforcement learning

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:09.953911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:04.932417Z digest=sha256:ec654002f94d3e4f86f8df6c04b23cd6cbd53f840255f3c7b89198e4a90545ea

Observation 01f6eb3b-010c-4d61-ab64-ee67d6026ba2 · outbound

This paper cites Varibad: A very good method for bayes-adaptive deep rl via meta- learning.

Fully Offline Reinforcement Learning Varibad: A very good method for bayes-adaptive deep rl via meta- learning

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:09.714445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:05.017057Z digest=sha256:008f07283960f32f57e0c67568f1419b91ba2976c7c2acb669dd7ab6bdffd146

Observation e4a4b1c3-7783-4bf3-86a2-7dba28d76125 · outbound

This paper cites N−1X i=0 1 N D 2 log(2π) + 1 2 D−1X d=0 logσ 2 θd (xi) + (yid −µ θd (xi))2 σ2 θd (xi) !!# , =Ei∼UN.

Fully Offline Reinforcement Learning N−1X i=0 1 N D 2 log(2π) + 1 2 D−1X d=0 logσ 2 θd (xi) + (yid −µ θd (xi))2 σ2 θd (xi) !!# , =Ei∼UN

Reference 78

Resolution
verified exact
raw_fallback, observed 2026-08-07T13:16:06.856195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:05.082852Z digest=sha256:0dcd2102000d63b927e9459ea2076bcefc7c85df0884fa5239b1005b3b24c559

Observation 584f4e18-5a1d-4826-9888-eb69dfcfea54 · outbound

This paper cites Eθ∼PΘ(DN ).

Fully Offline Reinforcement Learning Eθ∼PΘ(DN )

Reference 83

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T13:16:06.605584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:05.159203Z digest=sha256:b53f17c229ce7f61ea570a8002170992353968e9df2d316fd8357ddd4909e413

Observation d995e4c3-886c-44f7-9c7a-9ca3fbf81918 · outbound

This paper cites doi: 10.1093/oso/9780198504856.003.0002.

Fully Offline Reinforcement Learning doi: 10.1093/oso/9780198504856.003.0002

Reference 1999

Resolution
unresolved
no resolver link, observed 2026-08-07T13:15:57.660083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:15:57.660083Z digest=sha256:0a45f8db76b3ba544d38dc436de64d58514ecedadb2681e45b8ddefbe18351c0

Observation aa4e3f11-d59b-4bd4-a552-fd3592fb27fd · outbound

This paper cites URL https://books.google.co.uk/books?id= s6mVlgEACAAJ.

Fully Offline Reinforcement Learning URL https://books.google.co.uk/books?id= s6mVlgEACAAJ

Reference 2013

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:20.621497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:15:57.884088Z digest=sha256:7f23e4f9a49c8bd4b51bf478bec3fa5f0a6144ee181ae2646eef54d12aa6df6f

Observation 1e434d6f-04c0-42fb-b1dc-edaa8c4dbd63 · outbound

This paper cites an unresolved cited work.

Fully Offline Reinforcement Learning Unresolved cited work

Reference 2025

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:16:17.599212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:15:59.744897Z digest=sha256:2cdd197ca298da5fa9a72494327658eb6503aa761927f87b1beadef3448e7431

Observation 8c85133c-fff4-4179-b7a6-273632e7103f · outbound

This paper cites URL http://papers.nips.cc/paper/8080- randomized-prior-functions-for-deep-reinforcement-learning.pdf.

Fully Offline Reinforcement Learning URL http://papers.nips.cc/paper/8080- randomized-prior-functions-for-deep-reinforcement-learning.pdf

Reference 8629

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:14.762978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:02.029997Z digest=sha256:ecc4549ccd8dac4cef71fdf5db8393f3c0d95f136ac4e0a3aa44e50d59bafb51

Pith citing papers

Observation 7a38f31e-998b-44d0-a06c-e8f6a48dbd1e · inbound

Long-Horizon Model-Based Offline Reinforcement Learning Without Explicit Conservatism cites this paper.

Long-Horizon Model-Based Offline Reinforcement Learning Without Explicit Conservatism Fully Offline Reinforcement Learning

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-17T00:20:38.764461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T01:46:16.923948Z digest=sha256:7960664fbbdf6679849dfed451e5f1d2af4437f18e65ea3012caff89865e1f6a

Observation 00b3f2c1-259b-40d1-9384-9e49144431e9 · inbound

Is Inter-Seed Cross-Play Enough? Evaluating the Robustness of Zero-Shot Coordination Algorithms to Implementation Details cites this paper.

Is Inter-Seed Cross-Play Enough? Evaluating the Robustness of Zero-Shot Coordination Algorithms to Implementation Details Fully Offline Reinforcement Learning

Reference 117

Resolution
unresolved
no resolver link, observed 2026-08-05T15:25:39.829506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:25:39.829506Z digest=sha256:27d2e75c9aec0dfb2649b638ca40060b94355bc1f1b79b03da736b389f93f626