Pith. sign in

Paper Citation Record · LEDGER

Deep Transfer $Q$-Learning for Offline Non-Stationary Reinforcement Learning

As of 11 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 6 inbound Pith citation observations for arXiv:2501.04870.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.04870 v2

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:32:16.181011Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:19:10.654724Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T23:49:14.996703Z

Reference resolution

20 of 20 outbound references displayed

  • verified exact0
  • verified fuzzy11
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d0ec2c96-cbe6-47d6-965c-8c40046abc08 · outbound

This paper cites Offline Multi-task Transfer RL with Representational Penalization.

Deep Transfer $Q$-Learning for Offline Non-Stationary Reinforcement Learning Offline Multi-task Transfer RL with Representational Penalization

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T21:32:16.109676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:32:16.109676Z digest=sha256:1b55d663f31e8567c8df39031006f86547c2793d5709c11a8236d40003322ca0

Observation 2caaed7d-f951-4f6e-a72a-c1d286d793d3 · outbound

This paper cites After some algebra we get that ∥ bQp t − Q∗ agg t ∥2 nM,bPagg t ≤ ∥gp t − Q∗ agg t ∥2 nM,bPagg t + 2 nM nMX i=1 (byrwt−ki t,i − Q∗ agg t,i ) · ( bQp t,i − gp t,i).

Deep Transfer $Q$-Learning for Offline Non-Stationary Reinforcement Learning After some algebra we get that ∥ bQp t − Q∗ agg t ∥2 nM,bPagg t ≤ ∥gp t − Q∗ agg t ∥2 nM,bPagg t + 2 nM nMX i=1 (byrwt−ki t,i − Q∗ agg t,i ) · ( bQp t,i − gp t,i)

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:32:16.401105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T21:32:16.161325Z digest=sha256:5757dafa11e536d440c5913b4e5b2e994d7d3014ef3dcc23690ce4a4a0788d57

Observation fcc1872c-9d52-44da-8afe-7255ce3a95a5 · outbound

This paper cites an unresolved cited work.

Deep Transfer $Q$-Learning for Offline Non-Stationary Reinforcement Learning Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:32:16.351905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T21:32:16.174866Z digest=sha256:81c28ecd9aec13e83d286a9e41f9320b36379dcabc18b9f72eef91fff5c337e6

Observation 7575b9fe-bbf0-4a72-bd2f-5b7d84f48565 · outbound

This paper cites sub-Gaussian random variables with variance parameter σ.

Deep Transfer $Q$-Learning for Offline Non-Stationary Reinforcement Learning sub-Gaussian random variables with variance parameter σ

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:32:16.413812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T21:32:16.157923Z digest=sha256:19f7bed74a13ed965cef8c10820d7b581fd4881fea7629c580e905fe24877a62

Observation 14068bac-a517-45fd-885d-c40053615633 · outbound

This paper cites copies of z, G be a b-uniformly-bounded function class satisfying log(N∞(ϵ, G, zn 1 )) ≤ v log ebn ϵ for some quantity v.

Deep Transfer $Q$-Learning for Offline Non-Stationary Reinforcement Learning copies of z, G be a b-uniformly-bounded function class satisfying log(N∞(ϵ, G, zn 1 )) ≤ v log ebn ϵ for some quantity v

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:32:16.562749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T21:32:16.154387Z digest=sha256:8e72015900ef3904faa1f31c6ebb0299ab4c2416ae011ddac3ac9421d523deef

Observation 4c3e234b-a9cc-4f49-8492-0e8e88f5e2a4 · outbound

This paper cites Robust angle-based transfer learning in high dimensions.

Deep Transfer $Q$-Learning for Offline Non-Stationary Reinforcement Learning Robust angle-based transfer learning in high dimensions

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T21:32:16.122035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:32:16.122035Z digest=sha256:39de687c441a82a91435b47122fbb89d4794e88d8120ccfb7b6ca9dcae8aa07e

Observation 322b8543-e945-4058-bdad-7302491fa76e · outbound

This paper cites an unresolved cited work.

Deep Transfer $Q$-Learning for Offline Non-Stationary Reinforcement Learning Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:32:16.388060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T21:32:16.164919Z digest=sha256:4a97d377f8b1e3bd11bee82fa49dc1ab94c2a468e0409ace4a1390655bb32057

Observation 222c81a9-84bc-42e8-a0cf-3c4dc036e6f7 · outbound

This paper cites We now use the peeling argument to extend to uniform r.

Deep Transfer $Q$-Learning for Offline Non-Stationary Reinforcement Learning We now use the peeling argument to extend to uniform r

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:32:16.376291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T21:32:16.168599Z digest=sha256:27effd7fa12a46c90725a04b0dd760a182327456e6ea4c6b1be9bef79d162f0e

Observation 13b356e3-f83d-492e-8913-7039e61e011e · outbound

This paper cites For g1, · · ·, gN being an ϵ-covering set of G, we claim that g2 1 − eg2, · · ·, g2 N − eg2 is an 2bϵ-covering set of ¯G.

Deep Transfer $Q$-Learning for Offline Non-Stationary Reinforcement Learning For g1, · · ·, gN being an ϵ-covering set of G, we claim that g2 1 − eg2, · · ·, g2 N − eg2 is an 2bϵ-covering set of ¯G

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:32:16.364756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T21:32:16.171929Z digest=sha256:9c3307d7622ed1ab287e20f388be122a740828e2f51dde7579e0a458aac73d7f

Observation 393c7740-b859-4d7e-a972-8aecdb99fde4 · outbound

This paper cites In our dataset, the mortality rate is 24.21% for female and 22.71% for male.

Deep Transfer $Q$-Learning for Offline Non-Stationary Reinforcement Learning In our dataset, the mortality rate is 24.21% for female and 22.71% for male

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:32:16.341161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T21:32:16.177942Z digest=sha256:bbf04e84a8bc243a246b7b4a732c672c19e411c0d40978aeac2c06991313779a

Observation 7e83532c-c609-4c66-9c6d-a8b9e2d5d5f1 · outbound

This paper cites Figure 5 in Chen, Li & Jordan (2022) presents mortality rates of different lengths.

Deep Transfer $Q$-Learning for Offline Non-Stationary Reinforcement Learning Figure 5 in Chen, Li & Jordan (2022) presents mortality rates of different lengths

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:32:16.329149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T21:32:16.181011Z digest=sha256:5d4eec08b3734b2ee315c999a9c7faffee12eedf471d2e7969695f9574c9826a

Observation 26bbd280-67f7-439d-b3c7-ea930d053232 · outbound

This paper cites B., Davidian, M.

Deep Transfer $Q$-Learning for Offline Non-Stationary Reinforcement Learning B., Davidian, M

Reference 114

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:32:16.585575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T21:32:16.147497Z digest=sha256:5a2a21e23e6658198870458691c3f82438bf6a012bae5e633f560f7b4d72db9b

Observation bba30a7c-7f58-49a8-8c84-50e975fb6d42 · outbound

This paper cites & Remlinger, C.

Deep Transfer $Q$-Learning for Offline Non-Stationary Reinforcement Learning & Remlinger, C

Reference 343

Resolution
unresolved
no resolver link, observed 2026-08-10T21:32:16.118274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:32:16.118274Z digest=sha256:1b40b5115930df2f5bd32a64c52ea4da03a3cd633cd7055f29e8a18d31c04afe

Observation a019a3d5-0491-4d7b-becc-8ee33a89f447 · outbound

This paper cites & Song, R.

Deep Transfer $Q$-Learning for Offline Non-Stationary Reinforcement Learning & Song, R

Reference 640

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:32:16.596606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T21:32:16.138500Z digest=sha256:e7c72d230fad7deb3516a12ffc94ead26d1a7ef2348a4cdd2bb423dba605ea78

Observation 523052cf-28b5-4f18-9179-f93475dfe160 · outbound

This paper cites an unresolved cited work.

Deep Transfer $Q$-Learning for Offline Non-Stationary Reinforcement Learning Unresolved cited work

Reference 651

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:32:16.630880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T21:32:16.114359Z digest=sha256:8ecb85a393a59e9884f42169fa890d03ec8c41c2b1cbe1d7c646434a3402be89

Observation 65c1489b-4bfd-445a-b7c5-d722c61da4b5 · outbound

This paper cites Pseudo-Labeling for Kernel Ridge Regression under Covariate Shift.

Deep Transfer $Q$-Learning for Offline Non-Stationary Reinforcement Learning Pseudo-Labeling for Kernel Ridge Regression under Covariate Shift

Reference 901

Resolution
unresolved
no resolver link, observed 2026-08-10T21:32:16.142623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:32:16.142623Z digest=sha256:fff039f2df1027a2940a0841f0ba0fe25ec60ed039fe2f8dc73ebe9f660452d2

Observation 70a0fd0f-6126-4ca7-8351-939c8091aa40 · outbound

This paper cites (2012), Transfer in reinforcement learning: A framework and a survey, in ‘Reinforcement Learning’, Springer, pp.

Deep Transfer $Q$-Learning for Offline Non-Stationary Reinforcement Learning (2012), Transfer in reinforcement learning: A framework and a survey, in ‘Reinforcement Learning’, Springer, pp

Reference 1225

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:32:16.620368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T21:32:16.126456Z digest=sha256:1731f2d9317e177adcfb5b3a2a82f41a4432bc87e551c913aeada3b0794816c1

Observation 6f06d249-26eb-4209-8c26-c81945a89489 · outbound

This paper cites Deep Transfer Q-Learning for Offline Non-Stationary Reinforcement Learning.

Deep Transfer $Q$-Learning for Offline Non-Stationary Reinforcement Learning Deep Transfer Q-Learning for Offline Non-Stationary Reinforcement Learning

Reference 1549

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:32:16.574542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T21:32:16.150696Z digest=sha256:6a9ca665db38bfc13694049aa056ab41fd8071cd03f7fea039340e56aa2da9d8

Observation d174b87b-2a88-4461-8cc9-ae65dc60a418 · outbound

This paper cites an unresolved cited work.

Deep Transfer $Q$-Learning for Offline Non-Stationary Reinforcement Learning Unresolved cited work

Reference 2014

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:32:16.609130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T21:32:16.134721Z digest=sha256:5cc1bf83c9b7a9b35882114fb29c1d99585a3f3af548b06a45e5bee36b998131

Observation a80d550b-7804-4b87-a7ab-37bdeb6d35a5 · outbound

This paper cites On the Power of Multitask Representation Learning in Linear MDP.

Deep Transfer $Q$-Learning for Offline Non-Stationary Reinforcement Learning On the Power of Multitask Representation Learning in Linear MDP

Reference 3364

Resolution
unresolved
no resolver link, observed 2026-08-10T21:32:16.130092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:32:16.130092Z digest=sha256:9c4118fd14c0ec7e6e84ccdfb999bc6b8df2903b9d2ea72c9c3b84e56e4bbe36

Pith citing papers

Observation 1c0b5eee-179a-49b9-aa6d-0e3978524b3f · inbound

Phase Transition in Nonparametric Minimax Rates for Covariate Shifts on Approximate Manifolds cites this paper.

Phase Transition in Nonparametric Minimax Rates for Covariate Shifts on Approximate Manifolds Deep Transfer $Q$-Learning for Offline Non-Stationary Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:10.654724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:19:10.654724Z digest=sha256:d376cec337d29f3a861a27c57bcfda9de5a9742a68b138e173fcb37fbc3b49c9

Observation 3f7af179-139a-430a-b8a4-f82a314dddee · inbound

One-Step Bellman Alignment Enables Provably Efficient Transfer in Online RL cites this paper.

One-Step Bellman Alignment Enables Provably Efficient Transfer in Online RL Deep Transfer $Q$-Learning for Offline Non-Stationary Reinforcement Learning

Reference 2007

Resolution
unresolved
no resolver link, observed 2026-08-03T06:51:17.592951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:51:17.592951Z digest=sha256:3e03ccf198f49e57bacc25e64959ee739fb3fbb31369fbb9e58062cf5340ee9b

Observation d3d7ee4b-295b-4114-af5d-f0fd70ad5b99 · inbound

Kernelized Advantage Estimation: From Nonparametric Statistics to LLM Reasoning cites this paper.

Kernelized Advantage Estimation: From Nonparametric Statistics to LLM Reasoning Deep Transfer $Q$-Learning for Offline Non-Stationary Reinforcement Learning

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:16:27.343602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T07:00:32.206081Z digest=sha256:1048b93b7e783803bf117f048c0559b53706a991cd524388188ae00f674f0a32

Observation 84e593e0-db75-4a6b-94c9-5b7a2ec81661 · inbound

Kernelized Advantage Estimation: From Nonparametric Statistics to LLM Reasoning cites this paper.

Kernelized Advantage Estimation: From Nonparametric Statistics to LLM Reasoning Deep Transfer $Q$-Learning for Offline Non-Stationary Reinforcement Learning

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-20T23:49:14.999844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T23:47:53.282259Z digest=sha256:763e8c17d9b0d3b5b7b595735baf40572bc836efe63877f9b51dcf120f1badee

Observation 5a3779c5-27b3-4cce-88ee-6f1bf801dfce · inbound

Dual-Channel Tensor Neural Networks: Finite-Sample Theory and Conformal Structure Selection cites this paper.

Dual-Channel Tensor Neural Networks: Finite-Sample Theory and Conformal Structure Selection Deep Transfer $Q$-Learning for Offline Non-Stationary Reinforcement Learning

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:23:06.993328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-20T07:20:43.847495Z digest=sha256:4ee49cd524d18538adc214c2b11587b153d44b653e49c12fc7a15690e3badcdc

Observation 15042302-ff5b-4033-9b07-7f2412602dbb · inbound

Learning to Hand Off: Provably Convergent Workflow Learning under Interface Constraints cites this paper.

Learning to Hand Off: Provably Convergent Workflow Learning under Interface Constraints Deep Transfer $Q$-Learning for Offline Non-Stationary Reinforcement Learning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:08:12.160228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T10:07:03.605952Z digest=sha256:b098d39fef60727150edaaadd0784f55b75dbf2863b248780b8f077170c85e75