Pith. sign in

Paper Citation Record · LEDGER

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL

As of 9 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2505.19923.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19923 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:14:47.489929Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact2
  • verified fuzzy5
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 893f724a-cc24-4fe6-b9fd-1cd62d354734 · outbound

This paper cites The mean-wise best results among algorithms are highlighted in bold.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL The mean-wise best results among algorithms are highlighted in bold

Reference 3

Resolution
verified exact
raw_fallback, observed 2026-08-07T14:14:47.818399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:14:47.489929Z digest=sha256:2b66922d5959a2f8559b2b7146e2f11d8e8a2c95354efe8af3a310bf771f3974

Observation 7dfa7fc1-3245-47dd-9d6d-545b04214591 · outbound

This paper cites RvS: What is Essential for Offline RL via Supervised Learning?.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL RvS: What is Essential for Offline RL via Supervised Learning?

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:45.370476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:45.370476Z digest=sha256:d16c14506c90c77e84a802f552ded177ac088b351b20c70e39973e8a4d26f6f5

Observation 58fe6bdf-e92c-4148-80d2-f214ac4ea121 · outbound

This paper cites D4RL: Datasets for Deep Data-Driven Reinforcement Learning.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL D4RL: Datasets for Deep Data-Driven Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:45.440249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:45.440249Z digest=sha256:f204d36b60cbafafec957b464c337a37378a09ecd050d8e347c42fbc82371c93

Observation cf2ca248-7508-4b23-8b16-305983013a7a · outbound

This paper cites Planning with Diffusion for Flexible Behavior Synthesis.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Planning with Diffusion for Flexible Behavior Synthesis

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:45.772661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:45.772661Z digest=sha256:058fe582f393bb4422d1ec24337362f4417896312088470edbe89d151d1f469c

Observation 771cf11e-2c94-4336-a504-57a71276bbda · outbound

This paper cites Offline Reinforcement Learning with Implicit Q-Learning.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Offline Reinforcement Learning with Implicit Q-Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:45.838651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:45.838651Z digest=sha256:c196f097865472c43f0f836c613d40b745ea692fb83a1d3c283240a0e813e83b

Observation 56770b7a-9f6a-4c27-a856-48e770a0cb14 · outbound

This paper cites Reward-Consistent Dynamics Models are Strongly Generalizable for Offline Reinforcement Learning.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Reward-Consistent Dynamics Models are Strongly Generalizable for Offline Reinforcement Learning

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:14:48.307564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:14:45.930801Z digest=sha256:4d9b0d77ee4e6f2f81758d4d1f96d769739d4ce8dfc33784ad4827d6d4db4d56

Observation 9e633d48-ad30-4a98-b584-8c13f7e19791 · outbound

This paper cites S., Ghadirzadeh, A., Chen, X., and Finn, C.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL S., Ghadirzadeh, A., Chen, X., and Finn, C

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:50.287789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:14:46.017316Z digest=sha256:e6ace0afb62815262568228819d3e78267ef9b3d8598c4c2633b834aaa68f5e2

Observation 4b5041e9-8924-418a-a0bb-aa938abaf12b · outbound

This paper cites AWAC: Accelerating Online Reinforcement Learning with Offline Datasets.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL AWAC: Accelerating Online Reinforcement Learning with Offline Datasets

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:46.091351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:46.091351Z digest=sha256:4e64733306e67fca646a491218b9e926f5d0f6c4ca680e248e060a9d79820241

Observation 4abe55e6-075f-41d3-b5de-45378bc73e6d · outbound

This paper cites Q-Ensemble for Offline RL: Don't Scale the Ensemble, Scale the Batch Size.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Q-Ensemble for Offline RL: Don't Scale the Ensemble, Scale the Batch Size

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:46.138209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:46.138209Z digest=sha256:292511aa76b9e9750497b96b7c89dd533058f36af528d9430ba5b8167879a1d8

Observation 4f965292-2506-4043-9daa-6e2b00cb5c13 · outbound

This paper cites Hyperparameter Selection for Offline Reinforcement Learning.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Hyperparameter Selection for Offline Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:46.200212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:46.200212Z digest=sha256:94a9c25fa1813ce7ceefed18fde4a1e8cb955601be5b561addb7a180edcb66d3

Observation 8dbaed76-6543-4524-a810-85d8bc47c469 · outbound

This paper cites Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:46.292149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:46.292149Z digest=sha256:05a507b1fc2dc0cdf8dc7f8cc73f5df6eadc70d797b8c82449067dbe505fa103

Observation 7a33a587-91c8-495d-a991-4a2389b37c28 · outbound

This paper cites Diffusion Policies as an Expressive Policy Class for Offline Reinforcement Learning.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Diffusion Policies as an Expressive Policy Class for Offline Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:46.345569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:46.345569Z digest=sha256:0f08914481526430f7ef09de62246158e886dbf0865f01d6e1307ba6caae260a

Observation 5795690e-4847-4996-ba18-6a6adf07d008 · outbound

This paper cites Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:46.427182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:46.427182Z digest=sha256:e571b838f7a3909a6016a608350dfb5cf887548fde3e9788b28aeb0a412fcb13

Observation d018c99c-319e-49f4-b0e0-a31927ab0fdf · outbound

This paper cites Policy Expansion for Bridging Offline-to-Online Reinforcement Learning.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Policy Expansion for Bridging Offline-to-Online Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:46.513693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:46.513693Z digest=sha256:9c0f5c058bb95a33ce08e4c35b358897c6abffd16a8fac6ccdc9bc5869cd262f

Observation 5cd06e8b-3fc1-415c-8cfe-fe4537e8cdc8 · outbound

This paper cites ENOTO: Improving Offline-to-Online Reinforcement Learning with Q-Ensembles.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL ENOTO: Improving Offline-to-Online Reinforcement Learning with Q-Ensembles

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:46.601651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:46.601651Z digest=sha256:57a486b2f9280c30d90f043481bc9a5ea76b0db638909d7bb7126abe15900ec3

Observation 9772347e-cd21-4d34-833d-ee2ba45afc40 · outbound

This paper cites Adaptive Behavior Cloning Regularization for Stable Offline-to-Online Reinforcement Learning.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Adaptive Behavior Cloning Regularization for Stable Offline-to-Online Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:46.698844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:46.698844Z digest=sha256:e17134b4ca485e373922129add2b881ca40391fd71b07dcf7f9a2563c25d4e12

Observation e191725b-0487-48e2-a856-c8c480519051 · outbound

This paper cites an unresolved cited work.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:14:49.903609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:14:46.935078Z digest=sha256:6b0bc7562ae6b2cea1a0aaa15252312acbfbbd0e6d338a2c93ddc496417bceb0

Observation 9d17ed07-74aa-4a4a-80e6-e2d3388bac09 · outbound

This paper cites Recent works focus on explicit policy constraints for stochastic policies (Wu et al., 2022; Nair et al.,.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Recent works focus on explicit policy constraints for stochastic policies (Wu et al., 2022; Nair et al.,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:49.457619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:14:47.118568Z digest=sha256:75241a8dc6576ad435abf4bdfbcb7881983a1986a6bff258eef8b1bdc31bb285

Observation 68cea97c-3af2-4c50-b1e8-c77a93ac0da6 · outbound

This paper cites an unresolved cited work.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:14:49.244395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:14:47.219318Z digest=sha256:123dc68284e2de0aea1f7c6da996c0d172fc0b14a532e8a0fd41cfac17462daf

Observation 521b3951-9399-427c-96d8-c02efa20145c · outbound

This paper cites Moreover, the coefficients are updated by maximizing Q-values, lacking the interpretability offered by our method.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Moreover, the coefficients are updated by maximizing Q-values, lacking the interpretability offered by our method

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:48.972316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:14:47.318561Z digest=sha256:38a9acd57bced8dbf90d66cb03fa30d62dfc4e90c36cc5cd651e4f0d0af081d7

Observation cd96905c-98d3-4dec-860b-916782727a32 · outbound

This paper cites an unresolved cited work.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:14:48.759894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:14:47.389710Z digest=sha256:9dee129fc6cea14585d49ce0cf56c29a7d276d92b32e9464cf7b357a769be31f

Observation 41d6a292-3654-4631-86a1-08c3242581f0 · outbound

This paper cites Off-policy deep reinforcement learning without exploration.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Off-policy deep reinforcement learning without exploration

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:50.543479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:14:45.559317Z digest=sha256:efa007cb2690dfe2b56ca0a9879632f524124c10bac61edb295fef23383060a7

Observation 202fe6ac-af03-4756-8236-36e4ef3b38c3 · outbound

This paper cites Extreme Q-Learning: MaxEnt RL without Entropy.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Extreme Q-Learning: MaxEnt RL without Entropy

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:45.632354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:45.632354Z digest=sha256:bdd29da9deb71edcd580bf920c764ae785a3026f4e66b4b732f5de711db31661

Observation 05b7370f-d2b3-493b-ab10-8219106b87b3 · outbound

This paper cites Various CQL variants adjust the constraints or modify the regularizer to avoid excessive pessimism (Lyu et al., 2022; Nakamoto et al., 2024; Mao et al., 2024; Yu et al., 2021).

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Various CQL variants adjust the constraints or modify the regularizer to avoid excessive pessimism (Lyu et al., 2022; Nakamoto et al., 2024; Mao et al., 2024; Yu et al., 2021)

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:49.690918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:14:47.021956Z digest=sha256:49f05a88ad0d8b416c91dc4506e8a5ec96960904e4f5a8c1e583bdc5bbb26fb2

Observation e614c14e-8003-4cc3-a08d-0ac9dc9f223a · outbound

This paper cites Pessimistic Bootstrapping for Uncertainty-Driven Offline Reinforcement Learning.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Pessimistic Bootstrapping for Uncertainty-Driven Offline Reinforcement Learning

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:45.120978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:45.120978Z digest=sha256:fd8d878e6531f424ac0e3296c45004773afec5192e2c86b7ddad2f0bd08d58b7

Observation 9ca01919-e7c7-4a62-9b55-695b0f199b9b · outbound

This paper cites Efficient Online Reinforcement Learning with Offline Data.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Efficient Online Reinforcement Learning with Offline Data

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:45.188595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:45.188595Z digest=sha256:c0e3f0af5df0028e8ff7ac0f004ff049919ac262df1af938b031926c6a9335a2

Observation 563b4579-5f45-48bf-9108-6c2930b0c1e2 · outbound

This paper cites Improving TD3-BC: Relaxed Policy Constraint for Offline Learning and Stable Online Fine-Tuning.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Improving TD3-BC: Relaxed Policy Constraint for Offline Learning and Stable Online Fine-Tuning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:45.273426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:45.273426Z digest=sha256:4d0355f35bd4e19e6c5723ba7109cd73f7cc056d612135db3d2ea75e695d1d9b

Observation 4c8ab753-1a18-465f-a148-48665a553261 · outbound

This paper cites IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:45.701272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:45.701272Z digest=sha256:624b0988f6c4ea556b2302f267a6297d76aab04bb93b295d5a1bde383e8cfd68

Observation 28a30096-fa6c-4025-92f7-9aca61919fcd · outbound

This paper cites an unresolved cited work.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Unresolved cited work

Reference 2025

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:14:50.114624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:14:46.808770Z digest=sha256:a32894009d22b84cfc147a8c47cf608754e37c10f4cfe3a0b8582c80524913a1

Pith citing papers

No inbound Pith citation observations are available.