Pith. sign in

Paper Citation Record · LEDGER

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens

As of 9 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 2 inbound Pith citation observations for arXiv:2506.06261.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.06261 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T06:03:49.520373Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-26T21:27:31.026539Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T00:09:14.926116Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact3
  • verified fuzzy25
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c8092b47-592e-4a82-b66a-cd0fce5fc801 · outbound

This paper cites write newline.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:44.662883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:44.662883Z digest=sha256:137eb3a8c139298f36493f0df22e6ceaddd00deaef5cecdaff25538e4636ada7

Observation c1e2bd2b-cdcf-4cf5-8a4c-2568b5e59295 · outbound

This paper cites T., Tassa, Y., Munos, R., Heess, N., and Riedmiller, M.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens T., Tassa, Y., Munos, R., Heess, N., and Riedmiller, M

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:58.817378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:03:44.734527Z digest=sha256:7e7da91d6c8dbb545fbbbfcf66516fca0a3c73351804e89cd98b7a03ae091000

Observation 22a2a20c-39f7-4aec-ad69-8d7dc93f94d7 · outbound

This paper cites Deep Reinforcement Learning at the Edge of the Statistical Precipice.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Deep Reinforcement Learning at the Edge of the Statistical Precipice

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:44.808251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:44.808251Z digest=sha256:6bb9ec370d6cb950e9500e79cb90a1154587e42cc5b38da21fb79a1b1fa08f95

Observation 05899d7e-9df4-4707-b93a-b9b427f6ea1e · outbound

This paper cites an unresolved cited work.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:03:58.554367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:03:44.892153Z digest=sha256:1708378cfa72cbfb560495f6a21c4ede250c4f3e9902274edb6056e969e52fa8

Observation f78c6806-41a6-42ce-bed7-7574043fae77 · outbound

This paper cites and Dulac-Arnold, G.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens and Dulac-Arnold, G

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:58.335194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:03:44.989810Z digest=sha256:8df31675b5f76a76f7efa261aa1165a43bbf58f597486ddb7d415c4e85584c77

Observation 46e4276a-0f8e-48b6-a9fa-e2bc0d3300df · outbound

This paper cites Experiment tracking with weights and biases, 2020.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Experiment tracking with weights and biases, 2020

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:58.058084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:03:45.042792Z digest=sha256:fa98b1c6a144108473726b21aed77986c9e408a1a5073f66fbb598ba14a90a14

Observation 9f0003b4-1283-495a-a48d-a56f1b787393 · outbound

This paper cites T., Wenjie, S., and Ye, J.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens T., Wenjie, S., and Ye, J

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:57.739705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:03:45.123090Z digest=sha256:0e57d9baf6edd2d05531fb4e2bf6946be82e0facc29c9cb51502a5bc800fb26b

Observation ee9dc7b7-43f8-491c-b5c0-6781847d8ddc · outbound

This paper cites Deep reinforcement learning in a handful of trials using probabilistic dynamics models.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Deep reinforcement learning in a handful of trials using probabilistic dynamics models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:57.533127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:03:45.201119Z digest=sha256:f5639a4fdb26f7bb4e953f272f1b2716081d4c1396610d93008e6c0aedc03e55

Observation 9428f096-c1a4-48d4-8f2e-2aa7816e0afc · outbound

This paper cites S., Abbeel, P., Levine, S., and Finn, C.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens S., Abbeel, P., Levine, S., and Finn, C

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:57.285004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:03:45.285290Z digest=sha256:bef31da14235f1bfec4adb375a72d17469aa4085c215af377b98e3d6db619732

Observation ca17b460-a715-47aa-8594-b91f28163e4d · outbound

This paper cites Offline meta reinforcement learning -- identifiability challenges and effective data collection strategies.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Offline meta reinforcement learning -- identifiability challenges and effective data collection strategies

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:56.969870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:03:45.365659Z digest=sha256:66f789a4065dec8cb2be41c0f085b34d5b8a13e9ab34b6a146279e3b2058d225

Observation ab536021-8957-42d9-8dcb-f46a45381c51 · outbound

This paper cites an unresolved cited work.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:03:56.712428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:03:45.411352Z digest=sha256:df042886450c4e058fbfe1d4eeb20cf03f79e77627974008faf390dbaafc7b7f

Observation 4c5c8bac-27bf-41dc-8dc2-08557a1ec706 · outbound

This paper cites D4RL: Datasets for Deep Data-Driven Reinforcement Learning.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens D4RL: Datasets for Deep Data-Driven Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:45.544896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:45.544896Z digest=sha256:3fd0ba44e31393ad92ddef0fcae66b7fc4d9177c6837dc7216a22134d9d23e86

Observation b76c7758-8e57-4861-9361-d4dd75684348 · outbound

This paper cites and Gu, S.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens and Gu, S

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:45.611725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:45.611725Z digest=sha256:dffbe45038cc14d7c0bc9fb08ae9aca4c43a0d2201aa514edc3a4099288bcc27

Observation b7b5d722-cb99-4da3-b011-02228b4b32f4 · outbound

This paper cites Off-policy deep reinforcement learning without exploration.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Off-policy deep reinforcement learning without exploration

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:56.465987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:03:45.696943Z digest=sha256:e5f8c46ff9fa2a43b025137e4235a5e3b51c1143250c3907d3607f2748784d1a

Observation c639d093-f020-46bd-81ab-9e58b7ccd681 · outbound

This paper cites Bayesian reinforcement learning: A survey.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Bayesian reinforcement learning: A survey

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:56.146775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:03:45.765536Z digest=sha256:0f65af772625de52898d602175c09ea1b75baccff94ab95a540de6ca25c61f52

Observation f6fef5e4-79ef-4c5d-ba64-f1a2869dbb01 · outbound

This paper cites P., and Levine, S.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens P., and Levine, S

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:55.861945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:03:45.831667Z digest=sha256:d91670d4363d1b3326b5df283a8aa853808eecd79eb6f42282e592cdb458261c

Observation 27827645-388a-40b6-9c8a-672b7ecd6df1 · outbound

This paper cites Offline RL policies should be trained to be adaptive.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Offline RL policies should be trained to be adaptive

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:55.550935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:03:45.937699Z digest=sha256:b596d83f2d893616752ec8b0108cd1aeb0ae138c229a23e2a532bd4110ad7299

Observation 321a86f3-abf5-44f5-8209-44aef3c51753 · outbound

This paper cites Efficient bayes-adaptive reinforcement learning using sample-based search.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Efficient bayes-adaptive reinforcement learning using sample-based search

Reference 18

Resolution
verified exact
raw_fallback, observed 2026-08-07T06:03:50.770223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:03:46.012409Z digest=sha256:fa1f4bfeb0e51883c531c99161722282b2727777bb156fe2856af3e4177dcd23

Observation 705bd01e-1c79-4bfe-a40d-afbf8fe81c58 · outbound

This paper cites When to trust your model: Model-based policy optimization.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens When to trust your model: Model-based policy optimization

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:55.308363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:03:46.096762Z digest=sha256:ac38d271d660506d6bb7a2ca2cc99fa0131b1f28d281c5f56c741970c72bd310

Observation 42ad2618-6304-40fe-b34f-6df7def337bc · outbound

This paper cites Planning with diffusion for flexible behavior synthesis.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Planning with diffusion for flexible behavior synthesis

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:54.959745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:03:46.183550Z digest=sha256:00326067fbc232f32dcd9f8c14beaebac4207a55a891e38f7fd7c6aca387249d

Observation 0ab800cf-6a1f-4238-bbc1-dfc6983d4a9c · outbound

This paper cites Is pessimism provably efficient for offline rl? In Meila, M.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Is pessimism provably efficient for offline rl? In Meila, M

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:54.657314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:03:46.260650Z digest=sha256:6f642b15aa8ba070f0bde5694d43f473e0d432ae49e99d8128b531329489b037

Observation ca289220-e219-4c8d-96c5-209ad2b4d40e · outbound

This paper cites P., Littman, M.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens P., Littman, M

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:54.324634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:03:46.347403Z digest=sha256:4b73cfcabb6e5de538faad10a943bb0a388384f2027a97ad5e14732d055a0ad8

Observation 0ba19711-98f1-4007-9ad5-5c42e5c10026 · outbound

This paper cites Morel : Model-based offline reinforcement learning.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Morel : Model-based offline reinforcement learning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:53.981841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:03:46.452733Z digest=sha256:901621dfcc31cfb0ae7548a40c59f9ab2b31a943d12e586bfefa06822bbd615d

Observation dd60bff0-8a58-4589-850d-21e8f415fd4d · outbound

This paper cites Offline reinforcement learning with implicit q-learning.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Offline reinforcement learning with implicit q-learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:46.549158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:46.549158Z digest=sha256:aef462be2e5001f1800e98755259a52b076f72ee20e6c6649198322926d05316

Observation 55e25cb3-2daf-442c-bced-130b23261ac5 · outbound

This paper cites Stabilizing off-policy q-learning via bootstrapping error reduction.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Stabilizing off-policy q-learning via bootstrapping error reduction

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:53.649910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:03:46.664728Z digest=sha256:008b67b936fc91a87b235f393d3ec757260c6b9d491a29ffde2cdb851f995480

Observation a226a1e9-50ab-4673-abec-69396cd17915 · outbound

This paper cites Conservative q-learning for offline reinforcement learning.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Conservative q-learning for offline reinforcement learning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:53.293634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:03:46.791841Z digest=sha256:b1c7c8b2c80a7d2ba2ad90aa903ec7f64c685f6bf196a758f8d4e611da24cd79

Observation b3e38fe3-479d-4d2b-97b8-d3cc76116f40 · outbound

This paper cites Reinforcement learning and control as probabilistic inference: Tutorial and review, 2018.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Reinforcement learning and control as probabilistic inference: Tutorial and review, 2018

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:52.991334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:03:46.990592Z digest=sha256:6ab2dc41943a9106f7fd6c3af2d7ad5093a537f627ddaadaaeb14bd96c14cefb

Observation d0d22f6d-3053-4300-8c3b-b4cbd78b069d · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:47.167284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:47.167284Z digest=sha256:2ec2a03958161e544149094487ef80a698fb13aab7f011b7a3ce2842859c4289

Observation e95f3e36-e84f-46ae-b63d-338187ba32a5 · outbound

This paper cites Plan Online, Learn Offline: Efficient Learning and Exploration via Model-Based Control.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Plan Online, Learn Offline: Efficient Learning and Exploration via Model-Based Control

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:47.372549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:47.372549Z digest=sha256:6a6de9ff78134001108b0a41611e0e3e41a2e1cda82578a8feef6d542d582ff5

Observation e17cfde7-9275-4d2c-bf81-faf758631575 · outbound

This paper cites Revisiting design choices in offline model based reinforcement learning, 2021.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Revisiting design choices in offline model based reinforcement learning, 2021

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:52.717159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:03:47.506550Z digest=sha256:3601eb0907665f3ca0a144371d4afd770c5d4675ee7ed9689762a088776a482b

Observation 8de4fcca-fc38-43a3-a230-aa0ca047e3ea · outbound

This paper cites Deep Dynamics Models for Learning Dexterous Manipulation.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Deep Dynamics Models for Learning Dexterous Manipulation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:47.644501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:47.644501Z digest=sha256:9fb37b574ead7d1abe575b9503d920680a35a0500105cea2d13b6668d0d1017a

Observation 3425e546-0fed-405c-9f01-b5c0be6bdda4 · outbound

This paper cites and Taniguchi, T.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens and Taniguchi, T

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:52.311287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:03:47.808954Z digest=sha256:6f343f17d2228b62f4f019988834b7e7adabb083896f630042be25399df57f92

Observation 76999567-9b58-4024-a5a9-a64c44821feb · outbound

This paper cites Probabilistic planning with sequential monte carlo methods.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Probabilistic planning with sequential monte carlo methods

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:52.002407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:03:48.015339Z digest=sha256:8bfb39faa535a563d0545a60d2fbe8850406863fd309cfd242e36615640c18d1

Observation 3dcb864d-2bad-4af4-8797-ee40806379e7 · outbound

This paper cites RAMBO-RL: Robust Adversarial Model-Based Offline Reinforcement Learning.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens RAMBO-RL: Robust Adversarial Model-Based Offline Reinforcement Learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:48.178163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:48.178163Z digest=sha256:b8c677decbcf057120226631f7fef8ca8aa98c2908c30f7cbf383ec8a7c8683f

Observation 6b04e500-2cc4-4f42-87c8-27ab9d7706b4 · outbound

This paper cites Learning off-policy with online planning.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Learning off-policy with online planning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:51.673944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:03:48.345210Z digest=sha256:b042edf7f2a0db020b5ed7f49a5108a018cd0add59bc75a3d3b79b71692f7842

Observation 300b965f-9130-4fc9-bd52-e230fd1425a8 · outbound

This paper cites Improved sampling-importance resampling and reduced bias importance sampling.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Improved sampling-importance resampling and reduced bias importance sampling

Reference 36

Resolution
verified exact
raw_fallback, observed 2026-08-07T06:03:50.402148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:03:48.477888Z digest=sha256:89eca86c29e8ae98776793b9c5bf830a5bc92bc15851b14d58a2171e48ad8d3d

Observation 2d6978e4-bf2e-413e-9d8d-5540e253e3ac · outbound

This paper cites an unresolved cited work.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:48.624377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:48.624377Z digest=sha256:df9807f523302a2ccb0644349849095f24a6751ec1861d15acb3ac5701ea8daf

Observation 85faaf2e-b21c-4330-83d4-9a7482d1a1f0 · outbound

This paper cites Model predictive path integral control using covariance variable importance sampling, 2015.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Model predictive path integral control using covariance variable importance sampling, 2015

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:51.303992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:03:48.752678Z digest=sha256:8f027e8414b1d680c53a8ff603108c88c7b94eb6316329529e5a3577e63a442d

Observation 6c80e0e6-da96-4a3c-8da9-f2437ea0e9a8 · outbound

This paper cites Behavior Regularized Offline Reinforcement Learning.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Behavior Regularized Offline Reinforcement Learning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:48.860136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:48.860136Z digest=sha256:724c57843693eac75dc37f2e50e5873181373c184da5da97e22b4b63334def7c

Observation 7f0ae814-39de-4658-a455-60ce2bad1014 · outbound

This paper cites Mopo: Model-based offline policy optimization.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Mopo: Model-based offline policy optimization

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:50.976072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:03:49.021233Z digest=sha256:118eaefb30ca026df368cb7d46f663b34717c034ba7fc90b512167d88baedd7d

Observation 40591d21-29e5-46d1-b52f-9a758d043e95 · outbound

This paper cites COMBO: Conservative Offline Model-Based Policy Optimization.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens COMBO: Conservative Offline Model-Based Policy Optimization

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:49.226355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:49.226355Z digest=sha256:2154afb8a033a14790bbc15ddc4a775dd55217526c22690f7246a17c6e95ff5d

Observation 377ac002-9335-4d18-923f-0e1bace05914 · outbound

This paper cites Model-based offline planning with trajectory pruning.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens Model-based offline planning with trajectory pruning

Reference 42

Resolution
verified exact
doi, observed 2026-08-07T06:03:49.820614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:03:49.380210Z digest=sha256:0a617f2e1449efcf8b1c6e15c6bb41fb1cbd25db8745807f4964eeae14363f95

Observation c4e594ae-0268-48e9-92cc-2d4f6d18b387 · outbound

This paper cites VariBAD: A Very Good Method for Bayes-Adaptive Deep RL via Meta-Learning.

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens VariBAD: A Very Good Method for Bayes-Adaptive Deep RL via Meta-Learning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:49.520373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:49.520373Z digest=sha256:a9e057a119a730e2faa59b582fb7413bc5369f33d790f1bf99e961e0599752d3

Pith citing papers

Observation 26f71b9d-ad91-4084-834d-b3d50183cc9f · inbound

Agentic Reasoning for Large Language Models cites this paper.

Agentic Reasoning for Large Language Models Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:14:25.907040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T15:14:25.558878Z digest=sha256:a9686aaae9916a5f27b5d8b50a98dbb4ca7f493493840ae8be37964ea10ad1e5

Observation c1530e55-3dc4-43ea-861d-91ed9c3c525d · inbound

UBP2: Uncertainty-Balanced Preference Planning for Efficient Preference-based Reinforcement Learning cites this paper.

UBP2: Uncertainty-Balanced Preference Planning for Efficient Preference-based Reinforcement Learning Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:09:14.928534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T21:27:31.026539Z digest=sha256:8b4abc95e24edfd9228a512fbdbf31a07b1f1bfe43123f4ed51f23d1c06685bd