Pith. sign in

Paper Citation Record · LEDGER

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens

As of 9 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 2 inbound Pith citation observations for arXiv:2506.06261.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.06261 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T06:03:49.520373Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-26T21:27:31.026539Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T00:09:14.926116Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact3
  • verified fuzzy25
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c8092b47-592e-4a82-b66a-cd0fce5fc801 · outbound

This paper cites write newline.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:44.662883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:44.662883Z digest=sha256:ace01847b26edbbf0758a5174970959e3fac1c888a8322edf407689b55edb255

Observation c1e2bd2b-cdcf-4cf5-8a4c-2568b5e59295 · outbound

This paper cites T., Tassa, Y., Munos, R., Heess, N., and Riedmiller, M.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens T., Tassa, Y., Munos, R., Heess, N., and Riedmiller, M

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:58.817378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:44.734527Z digest=sha256:f051713c2ffb920dd98f040c9e6dc8915b3355e7b3f5ad2851989f728dc4b1db

Observation 22a2a20c-39f7-4aec-ad69-8d7dc93f94d7 · outbound

This paper cites Deep Reinforcement Learning at the Edge of the Statistical Precipice.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Deep Reinforcement Learning at the Edge of the Statistical Precipice

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:44.808251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:44.808251Z digest=sha256:4d077a3098b541375df3a102acd7f7fbeec94c214475ad625c4a529a7a1212eb

Observation 05899d7e-9df4-4707-b93a-b9b427f6ea1e · outbound

This paper cites an unresolved cited work.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:03:58.554367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:44.892153Z digest=sha256:2e6001d518a318e93e6793d43c3c343c89311e13307fc9ea7ac6e542526d087c

Observation f78c6806-41a6-42ce-bed7-7574043fae77 · outbound

This paper cites and Dulac-Arnold, G.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens and Dulac-Arnold, G

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:58.335194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:44.989810Z digest=sha256:4d4473cea41dd9046e89c59392e1aab8df325cf6e8a341c7e65284a35430ddad

Observation 46e4276a-0f8e-48b6-a9fa-e2bc0d3300df · outbound

This paper cites Experiment tracking with weights and biases, 2020.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Experiment tracking with weights and biases, 2020

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:58.058084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:45.042792Z digest=sha256:68992f47631a25a93a6eb36216db7a908976f5c22c198e369526b87ce7b46157

Observation 9f0003b4-1283-495a-a48d-a56f1b787393 · outbound

This paper cites T., Wenjie, S., and Ye, J.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens T., Wenjie, S., and Ye, J

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:57.739705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:45.123090Z digest=sha256:5f9cbdbb3d7e0b0dd4597499e159330ae03f5679e0e49dbeba796389a956f184

Observation ee9dc7b7-43f8-491c-b5c0-6781847d8ddc · outbound

This paper cites Deep reinforcement learning in a handful of trials using probabilistic dynamics models.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Deep reinforcement learning in a handful of trials using probabilistic dynamics models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:57.533127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:45.201119Z digest=sha256:1ac83f4a2fcf2b1dcc5ef4607ab46a32c200180c8cef93fc533559b74495a0ed

Observation 9428f096-c1a4-48d4-8f2e-2aa7816e0afc · outbound

This paper cites S., Abbeel, P., Levine, S., and Finn, C.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens S., Abbeel, P., Levine, S., and Finn, C

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:57.285004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:45.285290Z digest=sha256:57d7258e2d2185804021b35c4b370e55a87dac8100bc562be59879b072fcaec1

Observation ca17b460-a715-47aa-8594-b91f28163e4d · outbound

This paper cites Offline meta reinforcement learning -- identifiability challenges and effective data collection strategies.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Offline meta reinforcement learning -- identifiability challenges and effective data collection strategies

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:56.969870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:45.365659Z digest=sha256:6c92d77d9c20cd7550a34b7ce08f6fed25916fd7bf521c5d5567da5e8a97e18e

Observation ab536021-8957-42d9-8dcb-f46a45381c51 · outbound

This paper cites an unresolved cited work.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:03:56.712428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:45.411352Z digest=sha256:e5000175b7e7a6885c9967c220f050dfd6a30169ba810074e7772c7398f67865

Observation 4c5c8bac-27bf-41dc-8dc2-08557a1ec706 · outbound

This paper cites D4RL: Datasets for Deep Data-Driven Reinforcement Learning.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens D4RL: Datasets for Deep Data-Driven Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:45.544896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:45.544896Z digest=sha256:6f6ab950de835129fc1574a780548040df569c4044b0f91053fccdc0dec0e6a7

Observation b76c7758-8e57-4861-9361-d4dd75684348 · outbound

This paper cites and Gu, S.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens and Gu, S

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:45.611725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:45.611725Z digest=sha256:813d45763268c199720b17ef661cb895e5754de41c123abd9f804aa3ffc19d0b

Observation b7b5d722-cb99-4da3-b011-02228b4b32f4 · outbound

This paper cites Off-policy deep reinforcement learning without exploration.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Off-policy deep reinforcement learning without exploration

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:56.465987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:45.696943Z digest=sha256:23895afc460830c62a6b6909bf8270f073b8ca2efbc3e5fb7d798d1281778209

Observation c639d093-f020-46bd-81ab-9e58b7ccd681 · outbound

This paper cites Bayesian reinforcement learning: A survey.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Bayesian reinforcement learning: A survey

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:56.146775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:45.765536Z digest=sha256:46fa65a96e7f92638bcd31f39ec2bfc6a4943bd8f3164e96776d8ebfda178240

Observation f6fef5e4-79ef-4c5d-ba64-f1a2869dbb01 · outbound

This paper cites P., and Levine, S.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens P., and Levine, S

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:55.861945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:45.831667Z digest=sha256:7651e2734d1cb74040f5273ce15d53b6fffb3e235555898b2eee32d802bdd11a

Observation 27827645-388a-40b6-9c8a-672b7ecd6df1 · outbound

This paper cites Offline RL policies should be trained to be adaptive.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Offline RL policies should be trained to be adaptive

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:55.550935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:45.937699Z digest=sha256:a666b8063c8bf420376bbde63a2e800294eb624a2f50364f876cf6d7799a6abe

Observation 321a86f3-abf5-44f5-8209-44aef3c51753 · outbound

This paper cites Efficient bayes-adaptive reinforcement learning using sample-based search.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Efficient bayes-adaptive reinforcement learning using sample-based search

Reference 18

Resolution
verified exact
raw_fallback, observed 2026-08-07T06:03:50.770223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:46.012409Z digest=sha256:c9b6194dae16fa0df6181f243829bc7b0c6d072babfa3cff8a394c0d751da5cf

Observation 705bd01e-1c79-4bfe-a40d-afbf8fe81c58 · outbound

This paper cites When to trust your model: Model-based policy optimization.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens When to trust your model: Model-based policy optimization

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:55.308363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:46.096762Z digest=sha256:3266137ade8603f7ef5c5f5c31c892af54ffaea18020c2d080d858abccc8ba7b

Observation 42ad2618-6304-40fe-b34f-6df7def337bc · outbound

This paper cites Planning with diffusion for flexible behavior synthesis.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Planning with diffusion for flexible behavior synthesis

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:54.959745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:46.183550Z digest=sha256:312303ce0f19278a3f1b0d166f9359953148a9fa2ad46e5e4b1c16327821e63b

Observation 0ab800cf-6a1f-4238-bbc1-dfc6983d4a9c · outbound

This paper cites Is pessimism provably efficient for offline rl? In Meila, M.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Is pessimism provably efficient for offline rl? In Meila, M

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:54.657314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:46.260650Z digest=sha256:e533d6210d4b563ccf163258dd9aa86b4f1398bff6b6e076e81356c174ab8038

Observation ca289220-e219-4c8d-96c5-209ad2b4d40e · outbound

This paper cites P., Littman, M.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens P., Littman, M

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:54.324634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:46.347403Z digest=sha256:eadfd1603036d01210b5743397a0bdad1bd94b3e7623a815768dcc296572180d

Observation 0ba19711-98f1-4007-9ad5-5c42e5c10026 · outbound

This paper cites Morel : Model-based offline reinforcement learning.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Morel : Model-based offline reinforcement learning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:53.981841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:46.452733Z digest=sha256:7bc475a39db0975e44a6b4e57f97984f5f95298d5f77faefaff214882b36528c

Observation dd60bff0-8a58-4589-850d-21e8f415fd4d · outbound

This paper cites Offline reinforcement learning with implicit q-learning.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Offline reinforcement learning with implicit q-learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:46.549158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:46.549158Z digest=sha256:8f173e53293df712316e7345c12f088b7f56b055c7b57752b74277cc324ff1f7

Observation 55e25cb3-2daf-442c-bced-130b23261ac5 · outbound

This paper cites Stabilizing off-policy q-learning via bootstrapping error reduction.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Stabilizing off-policy q-learning via bootstrapping error reduction

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:53.649910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:46.664728Z digest=sha256:c00c1f542a45550969fa7b65dca23e08f33f70710e477c9bb17abbe03d4ea124

Observation a226a1e9-50ab-4673-abec-69396cd17915 · outbound

This paper cites Conservative q-learning for offline reinforcement learning.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Conservative q-learning for offline reinforcement learning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:53.293634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:46.791841Z digest=sha256:5e281bf366b0c37068fe9fc05802b6b252e9d2e4401ceaed040d36fc417a3de0

Observation b3e38fe3-479d-4d2b-97b8-d3cc76116f40 · outbound

This paper cites Reinforcement learning and control as probabilistic inference: Tutorial and review, 2018.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Reinforcement learning and control as probabilistic inference: Tutorial and review, 2018

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:52.991334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:46.990592Z digest=sha256:8debf8ed1a1c0200178c36492591ef79db9720a26b5fd2bd03a0b94d9c600d32

Observation d0d22f6d-3053-4300-8c3b-b4cbd78b069d · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:47.167284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:47.167284Z digest=sha256:312ba825524c96911c1caf949ee6f8271ecc95361fd42f5dd3dc20813273c170

Observation e95f3e36-e84f-46ae-b63d-338187ba32a5 · outbound

This paper cites Plan Online, Learn Offline: Efficient Learning and Exploration via Model-Based Control.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Plan Online, Learn Offline: Efficient Learning and Exploration via Model-Based Control

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:47.372549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:47.372549Z digest=sha256:e5eedcd2546c4743f4608c0e6028d856b3e94f02bd0e8ecefffab3c8f5ccfbd8

Observation e17cfde7-9275-4d2c-bf81-faf758631575 · outbound

This paper cites Revisiting design choices in offline model based reinforcement learning, 2021.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Revisiting design choices in offline model based reinforcement learning, 2021

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:52.717159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:47.506550Z digest=sha256:0ec16129e4b7eb926c9fc43fef38bec66976af95980e4eca2274de067aab14c3

Observation 8de4fcca-fc38-43a3-a230-aa0ca047e3ea · outbound

This paper cites Deep Dynamics Models for Learning Dexterous Manipulation.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Deep Dynamics Models for Learning Dexterous Manipulation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:47.644501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:47.644501Z digest=sha256:0d5a5928813a526e5046984842ccf87a53cb35a1f6f03d69823d5d417452d237

Observation 3425e546-0fed-405c-9f01-b5c0be6bdda4 · outbound

This paper cites and Taniguchi, T.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens and Taniguchi, T

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:52.311287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:47.808954Z digest=sha256:29995e2b129eb55a62625363a52ade3e7c333a729b49910381f4681e3ef6da3c

Observation 76999567-9b58-4024-a5a9-a64c44821feb · outbound

This paper cites Probabilistic planning with sequential monte carlo methods.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Probabilistic planning with sequential monte carlo methods

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:52.002407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:48.015339Z digest=sha256:43b2fa536064e7433b5602732428dc723c9985b0fc01c2e914d8e1b3adeaaaed

Observation 3dcb864d-2bad-4af4-8797-ee40806379e7 · outbound

This paper cites RAMBO-RL: Robust Adversarial Model-Based Offline Reinforcement Learning.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens RAMBO-RL: Robust Adversarial Model-Based Offline Reinforcement Learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:48.178163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:48.178163Z digest=sha256:26281826772966ead94f1be807cb62e58208ac24a1dba58420424129951a9c86

Observation 6b04e500-2cc4-4f42-87c8-27ab9d7706b4 · outbound

This paper cites Learning off-policy with online planning.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Learning off-policy with online planning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:51.673944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:48.345210Z digest=sha256:f4d32e00043e4bb3eb271639cb3f99e1dd09b5dd932f526d5c1e8192ef70c6de

Observation 300b965f-9130-4fc9-bd52-e230fd1425a8 · outbound

This paper cites Improved sampling-importance resampling and reduced bias importance sampling.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Improved sampling-importance resampling and reduced bias importance sampling

Reference 36

Resolution
verified exact
raw_fallback, observed 2026-08-07T06:03:50.402148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:48.477888Z digest=sha256:ccbf193492d9bc89bca76eecd8f1a3d78b80e6bd620eff0c7153357c4961cd21

Observation 2d6978e4-bf2e-413e-9d8d-5540e253e3ac · outbound

This paper cites an unresolved cited work.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:48.624377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:48.624377Z digest=sha256:a78958bf602047f04be57857bcf9290772071e2a70238aa5b5767a07dcbb23ab

Observation 85faaf2e-b21c-4330-83d4-9a7482d1a1f0 · outbound

This paper cites Model predictive path integral control using covariance variable importance sampling, 2015.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Model predictive path integral control using covariance variable importance sampling, 2015

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:51.303992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:48.752678Z digest=sha256:86f426d8d4e0e329ff026c3e598a9205b8f0b0f698a77d2cea7904ca1ab11753

Observation 6c80e0e6-da96-4a3c-8da9-f2437ea0e9a8 · outbound

This paper cites Behavior Regularized Offline Reinforcement Learning.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Behavior Regularized Offline Reinforcement Learning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:48.860136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:48.860136Z digest=sha256:7eea119da5c6416494de3b59bcb882b0284d162be5199bd01bf5a5854e3a304d

Observation 7f0ae814-39de-4658-a455-60ce2bad1014 · outbound

This paper cites Mopo: Model-based offline policy optimization.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Mopo: Model-based offline policy optimization

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:50.976072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:49.021233Z digest=sha256:d4763bdc4cd807f3cb56e3ba5faed163e52ba61efcf08b35f590905998c2aaf4

Observation 40591d21-29e5-46d1-b52f-9a758d043e95 · outbound

This paper cites COMBO: Conservative Offline Model-Based Policy Optimization.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens COMBO: Conservative Offline Model-Based Policy Optimization

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:49.226355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:49.226355Z digest=sha256:4cc51bb656ecc634470eac11fad556ed8d5be8c36de446bbef25226d11605afd

Observation 377ac002-9335-4d18-923f-0e1bace05914 · outbound

This paper cites Model-based offline planning with trajectory pruning.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Model-based offline planning with trajectory pruning

Reference 42

Resolution
verified exact
doi, observed 2026-08-07T06:03:49.820614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:49.380210Z digest=sha256:b6d9cd290cecc712fe5527d694957a3fbc51f640d13b696b35253f40ea87e17e

Observation c4e594ae-0268-48e9-92cc-2d4f6d18b387 · outbound

This paper cites VariBAD: A Very Good Method for Bayes-Adaptive Deep RL via Meta-Learning.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens VariBAD: A Very Good Method for Bayes-Adaptive Deep RL via Meta-Learning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:49.520373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:49.520373Z digest=sha256:65b7f386999c2c37cdfea9762bb74f679f325a2e284c3fadcc734a29165a4351

Pith citing papers

Observation 26f71b9d-ad91-4084-834d-b3d50183cc9f · inbound

Agentic Reasoning for Large Language Models cites this paper.

Agentic Reasoning for Large Language Models Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:14:25.907040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T15:14:25.558878Z digest=sha256:3c083b27754e8be976527d511dccdde4b0852ce5bcebc92597016f922039e8b3

Observation c1530e55-3dc4-43ea-861d-91ed9c3c525d · inbound

UBP2: Uncertainty-Balanced Preference Planning for Efficient Preference-based Reinforcement Learning cites this paper.

UBP2: Uncertainty-Balanced Preference Planning for Efficient Preference-based Reinforcement Learning Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:09:14.928534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T21:27:31.026539Z digest=sha256:dccad7bda586b6691a1ddcae624e4a5b057929a2540bdb4abba5b81b54d7a93c