Pith. sign in

Paper Citation Record · LEDGER

Conservative Q-Learning for Offline Reinforcement Learning

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 29 inbound Pith citation observations for arXiv:2006.04779.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2006.04779 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 29 of 29 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T20:57:48.435223Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

537
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 798f5a8b-a8f1-46a2-afda-7e554d96674b · inbound

What Matters in Learning from Offline Human Demonstrations for Robot Manipulation cites this paper.

What Matters in Learning from Offline Human Demonstrations for Robot Manipulation Conservative Q-Learning for Offline Reinforcement Learning

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-13T08:51:55.904631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T08:51:55.826747Z digest=sha256:596b8f252f511cfe714cd1c585d13c1706cc075b40fbe747b516119addd69876

Observation 114f4217-f920-4a89-86b6-529d4ba6016a · inbound

Offline Reinforcement Learning with Implicit Q-Learning cites this paper.

Offline Reinforcement Learning with Implicit Q-Learning Conservative Q-Learning for Offline Reinforcement Learning

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:47:05.661694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T08:47:05.621624Z digest=sha256:16aeb3372d603d9c341e66a69884ee7c8928f254d0b09645a234f007aa38b9b1

Observation c4e0a579-b85e-4265-925d-c8e326ac58eb · inbound

Digi-Q: Learning Q-Value Functions for Training Device-Control Agents cites this paper.

Digi-Q: Learning Q-Value Functions for Training Device-Control Agents Conservative Q-Learning for Offline Reinforcement Learning

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T20:57:48.435223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:57:48.435223Z digest=sha256:a132c8f65eae2b7575a6382e3f5d6edef70ebddcecfa56ef9874841f08a9ca44

Observation 5e093172-a8ce-45df-91fc-d68450c80dea · inbound

GenAI-based Multi-Agent Reinforcement Learning towards Distributed Agent Intelligence: A Generative-RL Agent Perspective cites this paper.

GenAI-based Multi-Agent Reinforcement Learning towards Distributed Agent Intelligence: A Generative-RL Agent Perspective Conservative Q-Learning for Offline Reinforcement Learning

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T17:57:08.426907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:57:08.426907Z digest=sha256:261a70dc97b8b74365f5fb9730fa15b062964cdfa3ff2379a444924795af952a

Observation 10db63f3-0676-4291-8c36-01bea85205dc · inbound

Safe Deployment of Offline Reinforcement Learning via Input Convex Action Correction cites this paper.

Safe Deployment of Offline Reinforcement Learning via Input Convex Action Correction Conservative Q-Learning for Offline Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T11:34:19.197355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:34:19.197355Z digest=sha256:7ba59f2437d9b3b1cc2c2f194fb9e10e0c6a57a7de9edabf16d59df7fad732da

Observation aa01e485-bd3b-484b-a935-a74a66b0b451 · inbound

Generative Sequential Notification Optimization via Multi-Objective Decision Transformers cites this paper.

Generative Sequential Notification Optimization via Multi-Objective Decision Transformers Conservative Q-Learning for Offline Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T11:39:57.868349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:39:57.868349Z digest=sha256:cd384320157f326c72e74e3b2af0f3ac6c75cf752f9fc88873ebea469c8d286c

Observation 918ed758-8c20-49ea-bc32-31db1ae7d66e · inbound

DAWM: Diffusion Action World Models for Offline Reinforcement Learning via Action-Inferred Transitions cites this paper.

DAWM: Diffusion Action World Models for Offline Reinforcement Learning via Action-Inferred Transitions Conservative Q-Learning for Offline Reinforcement Learning

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T13:56:26.127974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T13:55:16.513839Z digest=sha256:364d23c5382a8057b765dbe27b62bad56c058b99584d2f13ecd8a5e3c30c0377

Observation e8e4e2d1-bdee-42ff-b814-21d6f5458937 · inbound

Value Flows cites this paper.

Value Flows Conservative Q-Learning for Offline Reinforcement Learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T11:01:31.063875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:01:31.063875Z digest=sha256:f4e48687f0ce3a63be925c07c67207fe05820559675ea50b7cd69b2258e4730d

Observation 593a8806-a6e6-448e-b68c-492efa911c48 · inbound

DVPO: Distributional Value Modeling-based Policy Optimization for LLM Post-Training cites this paper.

DVPO: Distributional Value Modeling-based Policy Optimization for LLM Post-Training Conservative Q-Learning for Offline Reinforcement Learning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:48:51.003494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T01:46:21.744857Z digest=sha256:ca63917ebc852f88f3d1397bc2ad8923a2034371c4ee54ad310904d714c16388

Observation 18072b28-e076-4023-97dd-5c4a920c8b40 · inbound

The hidden risks of temporal resampling in clinical reinforcement learning cites this paper.

The hidden risks of temporal resampling in clinical reinforcement learning Conservative Q-Learning for Offline Reinforcement Learning

Reference 62

Resolution
malformed identifier
arxiv_id, observed 2026-05-16T07:07:29.616897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T07:05:16.379996Z digest=sha256:368750fe6ec1a38a68896ec2d99f8d154010cb2232244905035a62694da15fb8

Observation df39f09e-07b6-4ef8-8e75-5515c23b8209 · inbound

Simulation Distillation: Pretraining World Models in Simulation for Rapid Real-World Adaptation cites this paper.

Simulation Distillation: Pretraining World Models in Simulation for Rapid Real-World Adaptation Conservative Q-Learning for Offline Reinforcement Learning

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-15T09:49:54.486151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T09:49:03.333757Z digest=sha256:9ebcee22307c550402a635fb58dc3d55d2aa40934f62045ffb46b303fcd6cad8

Observation 771a535d-9a5a-41fb-b930-f2cd098e19ff · inbound

JD-BP: A Joint-Decision Generative Framework for Auto-Bidding and Pricing cites this paper.

JD-BP: A Joint-Decision Generative Framework for Auto-Bidding and Pricing Conservative Q-Learning for Offline Reinforcement Learning

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:10:52.115427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:40:01.560145Z digest=sha256:8658f3b288e34df25c78f8aa5e58f62ee4883a89586448c53997a9530a77948f

Observation 13352aff-ad43-46fa-9169-02517c493e7d · inbound

JD-BP: A Joint-Decision Generative Framework for Auto-Bidding and Pricing cites this paper.

JD-BP: A Joint-Decision Generative Framework for Auto-Bidding and Pricing Conservative Q-Learning for Offline Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T16:46:55.806545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:46:55.806545Z digest=sha256:69c2d09ffbbf689ba9f8890ab39050f6dc7a58bd893045145bcb7a9235e1192f

Observation 9aeb8997-d922-4e50-aa27-2de53b38dbee · inbound

Feedback-Normalized Developer Memory for Reinforcement-Learning Coding Agents: A Safety-Gated MCP Architecture cites this paper.

Feedback-Normalized Developer Memory for Reinforcement-Learning Coding Agents: A Safety-Gated MCP Architecture Conservative Q-Learning for Offline Reinforcement Learning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-09T22:29:06.408108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T13:59:40.638085Z digest=sha256:29e99ca860c2f6c74d151929597c13221cfa93d24dc1c7df03c6f7fae22ef364

Observation f1446329-990b-4a11-bbe1-7ec26ce075f7 · inbound

An adaptive variance estimator for relative sparsity cites this paper.

An adaptive variance estimator for relative sparsity Conservative Q-Learning for Offline Reinforcement Learning

Reference 79

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T05:45:22.591920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-08T19:35:46.097113Z digest=sha256:e3e471654619863a6c87aacbcf97beb816329764b17e841efe609c553a2856a5

Observation 54b2b2da-a21f-40ea-bfd8-cebf495def48 · inbound

RankQ: Offline-to-Online Reinforcement Learning via Self-Supervised Action Ranking cites this paper.

RankQ: Offline-to-Online Reinforcement Learning via Self-Supervised Action Ranking Conservative Q-Learning for Offline Reinforcement Learning

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:37:08.274944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T02:32:16.746824Z digest=sha256:d5f0cc3c21db3b860f4a69454c5bc8424d40a231ffcf95bcccdbdf34d3cc4442

Observation 6ff6e96e-8a17-48c8-9446-e719ce29eaaa · inbound

RankQ: Offline-to-Online Reinforcement Learning via Self-Supervised Action Ranking cites this paper.

RankQ: Offline-to-Online Reinforcement Learning via Self-Supervised Action Ranking Conservative Q-Learning for Offline Reinforcement Learning

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:54:05.817616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T08:53:29.468764Z digest=sha256:0a293eea8103f4d4b8222c55403ed9f148ffb91b903e294df65a8fcf732c270d

Observation 363088f6-84cb-4d4c-a516-d248c8a40dcf · inbound

Decoupling KL and Trajectories: A Unified Perspective for SFT, DAgger, Offline RL, and OPD in LLM Distillation cites this paper.

Decoupling KL and Trajectories: A Unified Perspective for SFT, DAgger, Offline RL, and OPD in LLM Distillation Conservative Q-Learning for Offline Reinforcement Learning

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-19T20:52:46.236699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T20:49:15.724896Z digest=sha256:14b60c111e5b455c7eb98ae49cf23047d7a56e0df4e9d8d72ff6d54dd9d54485

Observation 43d4de9f-b2d2-4358-b1d2-285c45147a1c · inbound

ISEP: Implicit Support Expansion for Offline Reinforcement Learning via Stochastic Policy Optimization cites this paper.

ISEP: Implicit Support Expansion for Offline Reinforcement Learning via Stochastic Policy Optimization Conservative Q-Learning for Offline Reinforcement Learning

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:08:15.443271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T12:07:21.037980Z digest=sha256:54f04fb7dc537c8da6441f1f10b2f9808005c3a0e2f98c3764b7e27a2bbdc1f9

Observation 91af34ea-1812-483d-9e44-00cd4875c82b · inbound

Abstraction for Offline Goal-Conditioned Reinforcement Learning cites this paper.

Abstraction for Offline Goal-Conditioned Reinforcement Learning Conservative Q-Learning for Offline Reinforcement Learning

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-22T07:51:16.575399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T07:46:20.289421Z digest=sha256:ff59609ecb809dac6accf438652ccc78561d184cc37186844ee0ebfd925b77f6

Observation 1c1e04c5-22b8-4520-ac8b-0185d71fed66 · inbound

Reward-free Pretraining for Reinforcement Learning via Occupancy Coverage Maximization cites this paper.

Reward-free Pretraining for Reinforcement Learning via Occupancy Coverage Maximization Conservative Q-Learning for Offline Reinforcement Learning

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:09:37.550272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T14:48:19.683101Z digest=sha256:80cdba24c3d1c5696b8b6356c936eb62d2517032f175054c1c08b3c7d17b9240

Observation 96215305-918f-42b4-9869-48c085c7f94f · inbound

Support-Constrained RL Enables Real-World Policy Improvement without Real-World Experience cites this paper.

Support-Constrained RL Enables Real-World Policy Improvement without Real-World Experience Conservative Q-Learning for Offline Reinforcement Learning

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-01T18:25:58.742568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T01:57:04.293058Z digest=sha256:31b003d320359ddb89a4eb113d15b5bbc43ab9cc98c3dac30d44f7b4dbdd3f28

Observation d5013757-8dfe-4512-a595-49fb8945f0c8 · inbound

From Bootstrapping to Sequence Modeling: A Unified Generative Framework for Personalized Landing-Page Modeling cites this paper.

From Bootstrapping to Sequence Modeling: A Unified Generative Framework for Personalized Landing-Page Modeling Conservative Q-Learning for Offline Reinforcement Learning

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-01T17:55:51.405173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T02:56:52.303152Z digest=sha256:0fe6cdf6bbdd09c9691c61c9cd0d36de50fdca0c0bcae95764283018f8ee6aee

Observation fb1a7742-8efb-4993-a66b-a3fba610e1ad · inbound

Pessimism's Paradox: Conservative Offline Training Amplifies Reward Hacking During Online Adaptation in Reasoning Models cites this paper.

Pessimism's Paradox: Conservative Offline Training Amplifies Reward Hacking During Online Adaptation in Reasoning Models Conservative Q-Learning for Offline Reinforcement Learning

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T06:54:21.012695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T06:46:02.255137Z digest=sha256:7aa24f95fd0bcde718085a88f8c5cde3f1bb1e1e159930cece051a500ca74215

Observation 09c7d9f4-c115-4b28-88b4-ff7706ff3381 · inbound

Guided Action Flow: Q-Guided Inference for Flow-Matching Vision-Language-Action Policies cites this paper.

Guided Action Flow: Q-Guided Inference for Flow-Matching Vision-Language-Action Policies Conservative Q-Learning for Offline Reinforcement Learning

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-07-03T11:58:05.713079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-03T11:50:49.215046Z digest=sha256:b60c0261c3c7d11933824161d9146cfcba6587a9b0a4380b9ae00bb046661dbc

Observation 43654418-5787-417f-b0a6-d550dd7cc28f · inbound

Guided Action Flow: Q-Guided Inference for Flow-Matching Vision-Language-Action Policies cites this paper.

Guided Action Flow: Q-Guided Inference for Flow-Matching Vision-Language-Action Policies Conservative Q-Learning for Offline Reinforcement Learning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-12T08:26:59.428524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T08:26:59.428524Z digest=sha256:23b294fa597150bd745f38e1911aaf1138391387b67e2e75767a3d27d5f3fd4c

Observation b74f09b0-d2d0-424b-980b-f5c4a16560f2 · inbound

Reinforcement Learning: From Algorithms To Foundation Models cites this paper.

Reinforcement Learning: From Algorithms To Foundation Models Conservative Q-Learning for Offline Reinforcement Learning

Reference 172

Resolution
unresolved
no resolver link, observed 2026-08-01T17:45:13.959321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:45:13.959321Z digest=sha256:5eb520813eb9b6ebc8999219b4c24892c1e42e56a3974f54107bd93118a21aee

Observation ad055bab-729a-4f67-bdfa-70a49a77e3a4 · inbound

Good Rankers, Bad Objectives: Bilinear Contrastive Critics under Expressive Policy Search cites this paper.

Good Rankers, Bad Objectives: Bilinear Contrastive Critics under Expressive Policy Search Conservative Q-Learning for Offline Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-31T01:24:20.678068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T01:24:20.678068Z digest=sha256:daf13c3d9e25b38c1d759e57a007008575715b85609cb5a773963b6aa139f834

Observation 536637ce-d99c-4c05-b02d-56f7c329c78d · inbound

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills cites this paper.

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills Conservative Q-Learning for Offline Reinforcement Learning

Reference 126

Resolution
unresolved
no resolver link, observed 2026-08-04T19:45:34.852244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:45:34.852244Z digest=sha256:a20bfeef3feed9d1b592259a8adcc81bac66d5cc03fc6c98069bdff1be950383