Pith. sign in

Paper Citation Record · LEDGER

Offline vs. Online Learning in Model-based RL: Lessons for Data Collection Strategies

As of 18 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 0 inbound Pith citation observations for arXiv:2509.05735.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.05735 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T05:09:29.269978Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

20 of 20 outbound references displayed

  • verified exact3
  • verified fuzzy5
  • unresolved10
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ee287cfc-e534-4381-acbc-8d2791ef3b3f · outbound

This paper cites Combating the Compounding-Error Problem with a Multi-step Model.

Offline vs. Online Learning in Model-based RL: Lessons for Data Collection Strategies Combating the Compounding-Error Problem with a Multi-step Model

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T05:09:29.205646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:09:29.205646Z digest=sha256:471b5eca6fc379ebf343321613638b92472a40f5463f30b3dbb2dc86159fade1

Observation 32071725-ddd9-4c2f-b5c2-facf1b812fed · outbound

This paper cites Below, the training of the world model M includes training all components in Eq.

Offline vs. Online Learning in Model-based RL: Lessons for Data Collection Strategies Below, the training of the world model M includes training all components in Eq

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:09:29.425275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T05:09:29.261773Z digest=sha256:13a14f58ea2b4cd022aa5d47879a9c76d90edf1acb4a60a7752a98484b6b5630

Observation 86bb3937-28c7-4ce8-90fb-f09a9540103e · outbound

This paper cites Regularized Behavior Value Estimation.

Offline vs. Online Learning in Model-based RL: Lessons for Data Collection Strategies Regularized Behavior Value Estimation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T05:09:29.215472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:09:29.215472Z digest=sha256:944c9731495ad3321ff719d51f4dfbfd2fbd4ea74a89699418a571fb3047d455

Observation 783a355f-10d1-4167-9e5b-b7b0dd5e819c · outbound

This paper cites A Survey on Offline Model-Based Reinforcement Learning.

Offline vs. Online Learning in Model-based RL: Lessons for Data Collection Strategies A Survey on Offline Model-Based Reinforcement Learning

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-05T05:09:29.343895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T05:09:29.225492Z digest=sha256:c524b1cd7a9e2cecb1c7b02fdc85db9cb06209df39fb4c6e4e195a82916d4131

Observation aeccf0ce-10ea-4c16-809e-db11a1d5e6ce · outbound

This paper cites Discor: Corrective feedback in reinforcement learning via distribution correction.

Offline vs. Online Learning in Model-based RL: Lessons for Data Collection Strategies Discor: Corrective feedback in reinforcement learning via distribution correction

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:09:29.454915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T05:09:29.234724Z digest=sha256:92d02434a9f622d263868e880682992de1c135b0ead996fa6845f9610262a9f3

Observation 49a3ba7d-8906-4510-b54e-b22e15165640 · outbound

This paper cites Andrew Bagnell.

Offline vs. Online Learning in Model-based RL: Lessons for Data Collection Strategies Andrew Bagnell

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:09:29.445196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T05:09:29.238447Z digest=sha256:e17997fdd2e8c9223ec090d0850788789b22fa4f1046feb01d2568d36ad32205

Observation d18d4844-7d4d-422d-8037-d70e1072919e · outbound

This paper cites The Edge-of-Reach Problem in Offline Model-Based Reinforcement Learning.

Offline vs. Online Learning in Model-based RL: Lessons for Data Collection Strategies The Edge-of-Reach Problem in Offline Model-Based Reinforcement Learning

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-05T05:09:29.324068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T05:09:29.241251Z digest=sha256:fabbfb2e3d7c0282a94e3a8fb1dd42ce23efa1d97c98638fa410a968109f4c53

Observation 372920bb-6144-4d2a-9de6-51c63bbeb966 · outbound

This paper cites Understanding the performance gap between online and offline alignment algorithms.

Offline vs. Online Learning in Model-based RL: Lessons for Data Collection Strategies Understanding the performance gap between online and offline alignment algorithms

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T05:09:29.245182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:09:29.245182Z digest=sha256:e89c6e7d45b6080eceedfe96c93960580840c7f5648464cb01423e3024d76096

Observation ea4bdc3f-4d47-4feb-b7de-06859784cd58 · outbound

This paper cites Don't Change the Algorithm, Change the Data: Exploratory Data for Offline Reinforcement Learning.

Offline vs. Online Learning in Model-based RL: Lessons for Data Collection Strategies Don't Change the Algorithm, Change the Data: Exploratory Data for Offline Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T05:09:29.248222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:09:29.248222Z digest=sha256:4c6a42aad9aba5d639fce2ab452d14e4cbf211e7889f5596e73b94ae1b039c5a

Observation 05cd2f78-30c3-45ae-a746-040560acf259 · outbound

This paper cites MinAtar: An Atari-Inspired Testbed for Thorough and Reproducible Reinforcement Learning Experiments.

Offline vs. Online Learning in Model-based RL: Lessons for Data Collection Strategies MinAtar: An Atari-Inspired Testbed for Thorough and Reproducible Reinforcement Learning Experiments

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T05:09:29.251166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:09:29.251166Z digest=sha256:aa66b8f211d6696dcd8b9ceb8d44a7db723c313fd5163b69a4d0b96b9d948144

Observation 8ba0897d-c032-412c-8d51-8790df24a9bc · outbound

This paper cites A Implementation Details A.1 Runtime Overview Our experiments comprised approximately 2000 runs, totaling 20000 GPU hours.

Offline vs. Online Learning in Model-based RL: Lessons for Data Collection Strategies A Implementation Details A.1 Runtime Overview Our experiments comprised approximately 2000 runs, totaling 20000 GPU hours

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:09:29.436594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T05:09:29.258053Z digest=sha256:6cde34e2dea96ac1a5201543119b9631052301db0f1d5eafe6346e397e68b008

Observation ab6cea53-7737-43db-94cd-71264f022b31 · outbound

This paper cites -same”, while the different model initialization is marked with “-diff.

Offline vs. Online Learning in Model-based RL: Lessons for Data Collection Strategies -same”, while the different model initialization is marked with “-diff

Reference 19

Resolution
malformed identifier
raw_fallback, observed 2026-08-05T05:09:29.415889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T05:09:29.265767Z digest=sha256:1589ad4008ff6acdac959692f65d576cdbf6d8d0568654a1c46d8854861c00ab

Observation a571720c-4a66-4583-aae5-55809fd18630 · outbound

This paper cites an unresolved cited work.

Offline vs. Online Learning in Model-based RL: Lessons for Data Collection Strategies Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:09:29.406556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T05:09:29.269978Z digest=sha256:c3155e9afeb608ddbbde97ea03ab60030264bfe5ddc963380414f1eb9e432acf

Observation 45390c4c-59ae-464e-8f66-d06cee0d787c · outbound

This paper cites Mastering Diverse Domains through World Models.

Offline vs. Online Learning in Model-based RL: Lessons for Data Collection Strategies Mastering Diverse Domains through World Models

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-05T05:09:29.222266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:09:29.222266Z digest=sha256:8f8dc174f3421117848f743c508c50769916964a3f7fa29d5bf402975cd1d851

Observation 5422bfd4-e554-4c47-847b-13464b51c3e6 · outbound

This paper cites Multi-task curriculum learning in a complex, visual, hard-exploration domain: Minecraft.

Offline vs. Online Learning in Model-based RL: Lessons for Data Collection Strategies Multi-task curriculum learning in a complex, visual, hard-exploration domain: Minecraft

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-05T05:09:29.228691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:09:29.228691Z digest=sha256:8fdb9655f7b7648ade8372294ef05ccc515a1c79748390cd5f7c54511940001a

Observation 970f977b-940e-4d93-b523-185692526cbc · outbound

This paper cites Exploring generalization and adaptability of offline reinforcement learning for robot manipulation.

Offline vs. Online Learning in Model-based RL: Lessons for Data Collection Strategies Exploring generalization and adaptability of offline reinforcement learning for robot manipulation

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:09:29.464306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T05:09:29.231579Z digest=sha256:0002a375d52a030cbc39481cee420381faad81a0ad34f3a5c7472ed22ed3c9c3

Observation 912e2045-4ae9-4daa-b764-dd6721502493 · outbound

This paper cites Boosting Offline Reinforcement Learning via Data Rebalancing.

Offline vs. Online Learning in Model-based RL: Lessons for Data Collection Strategies Boosting Offline Reinforcement Learning via Data Rebalancing

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-05T05:09:29.254979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:09:29.254979Z digest=sha256:5ca5f931b44846356c14a2571f5ca5eb048362e99d536fd8dfff8396773eef23

Observation e885a2ad-9441-46a0-af61-63eaf5aa4185 · outbound

This paper cites Learning Latent Dynamics for Planning from Pixels.

Offline vs. Online Learning in Model-based RL: Lessons for Data Collection Strategies Learning Latent Dynamics for Planning from Pixels

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-05T05:09:29.219130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:09:29.219130Z digest=sha256:b9f035fae5abfa05af43073abb10045a2e0921776492d799f350640b85cc3dea

Observation ac127ea4-8e9e-4825-89da-61919d9a4b69 · outbound

This paper cites Knowledge Transfer from Teachers to Learners in Growing-Batch Reinforcement Learning.

Offline vs. Online Learning in Model-based RL: Lessons for Data Collection Strategies Knowledge Transfer from Teachers to Learners in Growing-Batch Reinforcement Learning

Reference 2023

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T05:09:29.377909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T05:09:29.212419Z digest=sha256:33d4f48ff496a8ff434053c75845309fd211cde9f6e5afa46496186cb558cb81

Observation a1289718-f32d-456a-b445-3dc26cc19514 · outbound

This paper cites Behavioral Priors and Dynamics Models: Improving Performance and Domain Transfer in Offline RL.

Offline vs. Online Learning in Model-based RL: Lessons for Data Collection Strategies Behavioral Priors and Dynamics Models: Improving Performance and Domain Transfer in Offline RL

Reference 2024

Resolution
verified exact
local_arxiv, observed 2026-08-05T05:09:29.389938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T05:09:29.209188Z digest=sha256:ae4b7a7b4a1a6d68b9888f19a519729dfc50babd5f3db14ab6bfd6b653fdec47

Pith citing papers

No inbound Pith citation observations are available.