Pith. sign in

Paper Citation Record · LEDGER

Distributionally Robust Deep Q-Learning

As of 15 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 2 inbound Pith citation observations for arXiv:2505.19058.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19058 v1

Coverage vector

measured 66 of 66 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:26:53.058645Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-10T09:56:20.351339Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T09:57:00.793817Z

Reference resolution

66 of 66 outbound references displayed

  • verified exact5
  • verified fuzzy43
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d56f98fb-9fae-4388-a4fb-338b43c8f91d · outbound

This paper cites Investigating the parameters of the beta distribution.

Distributionally Robust Deep Q-Learning Investigating the parameters of the beta distribution

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:59.823780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:26:47.148084Z digest=sha256:3dbe687694d7a3bad418c88b27479e5520d4629316a2007cc45ec8e0c4d8dcda

Observation 9ac03d49-301b-462e-aa0e-f29bd48cd45e · outbound

This paper cites Infinite dimensional analysis: a hitchhiker’s guide.

Distributionally Robust Deep Q-Learning Infinite dimensional analysis: a hitchhiker’s guide

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:59.810211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:26:47.212428Z digest=sha256:9403c5bca3703b67555696b39c6be39c9b97e7aea780a053752be89e9c96c7f9

Observation b7f522da-e813-468e-a6b7-fbdbc4005c9e · outbound

This paper cites Computational aspects of robust optimized certainty equiv- alents and option pricing.

Distributionally Robust Deep Q-Learning Computational aspects of robust optimized certainty equiv- alents and option pricing

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:59.797222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:26:47.343157Z digest=sha256:d25f90aef6c45d401ea8c12dfa27216c47c7a36cd7f5f2d6c045ccac6c5c154f

Observation c17de44c-bf89-4428-b2f5-f41e89f22035 · outbound

This paper cites On the theory of dynamic programming.

Distributionally Robust Deep Q-Learning On the theory of dynamic programming

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:59.783436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:26:47.437965Z digest=sha256:bd956011c89edad6d0d39adf679037c71a653ab043f587d2483a41736e67c9a9

Observation ae0c3ccc-57a8-4ac0-b321-2fa2e0c1c04d · outbound

This paper cites Dota 2 with Large Scale Deep Reinforcement Learning.

Distributionally Robust Deep Q-Learning Dota 2 with Large Scale Deep Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:26:47.529254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:26:47.529254Z digest=sha256:d514c1323a2197ac31a0c7535b65ebd461682650a44f0c2d66ea24ec39f186ca

Observation fc8655b3-0994-46f6-b078-75faa5cf58a5 · outbound

This paper cites Convex optimization.

Distributionally Robust Deep Q-Learning Convex optimization

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:59.768872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:26:47.603514Z digest=sha256:a6f77839d9d9da96297789a588ce6caa78fb8b9e877d671af66a3bbbbd325d94

Observation 335e8b6a-d2cc-430c-a794-e6c70c8e712c · outbound

This paper cites Distributionally robust Markov decision processes and their connection to risk measures.

Distributionally Robust Deep Q-Learning Distributionally robust Markov decision processes and their connection to risk measures

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:59.756832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:26:47.723448Z digest=sha256:66669d938569dcdd7eb541a8f38a0d49388b9cdbbbbcf368e71f70e7727d496c

Observation 1461b9f8-e872-4b1c-8a58-f98767d84ecf · outbound

This paper cites an unresolved cited work.

Distributionally Robust Deep Q-Learning Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:26:59.745726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:26:47.813404Z digest=sha256:cd08f5e494da15b51702dabdef660463bd62cc4bfeed6fb540dc2955c2f0aec1

Observation c8bb8b23-9445-48d0-bd0c-0f859f07c0a6 · outbound

This paper cites Sinkhorn distances: Lightspeed computation of optimal transport, 2013.

Distributionally Robust Deep Q-Learning Sinkhorn distances: Lightspeed computation of optimal transport, 2013

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:59.730603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:26:47.913194Z digest=sha256:66864feab5e5f45c0aa138181b764e6aa2c8d4e778885ebd06730941e5e46733

Observation 836e6293-d9a6-4848-a7b4-c4090a822da9 · outbound

This paper cites Robust Q-learning for finite ambiguity sets.

Distributionally Robust Deep Q-Learning Robust Q-learning for finite ambiguity sets

Reference 10

Resolution
verified exact
raw_fallback, observed 2026-08-07T14:26:54.193985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:26:47.992530Z digest=sha256:82e01d50f60c44e195e8d2f01438a8fb51aefa115a45bddfef8cd968d48fb71f

Observation 757a9a1a-eb1b-4fc5-b04e-5637ca04f287 · outbound

This paper cites Twice Regularized Markov Decision Processes: The Equivalence between Robustness and Regularization.

Distributionally Robust Deep Q-Learning Twice Regularized Markov Decision Processes: The Equivalence between Robustness and Regularization

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:26:53.921863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:26:48.113813Z digest=sha256:7563d4624b11adc3cf29571e9027572856b2ca73f4d0c15edb48df10e02ab5c9

Observation 016885b5-6336-4a61-83bc-6130eab112a1 · outbound

This paper cites Maximum Entropy RL (Provably) Solves Some Robust RL Problems.

Distributionally Robust Deep Q-Learning Maximum Entropy RL (Provably) Solves Some Robust RL Problems

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:26:48.211513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:26:48.211513Z digest=sha256:7864e73122ff174c704bd4f06b841293f5d121f80d462538a182aa09e228522d

Observation 6ebf42a6-9e34-4417-be0d-d482898157cd · outbound

This paper cites A theoretical analysis of deep Q-learning.

Distributionally Robust Deep Q-Learning A theoretical analysis of deep Q-learning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:59.717657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:26:48.312156Z digest=sha256:d6169486292935e8e56ea03c00c548f4bdf729b7406cf2107ee94667a76cfe4a

Observation 383dc3d1-25d8-4bb0-b7dc-22b3084143b2 · outbound

This paper cites In- terpolating between optimal transport and mmd using Sinkhorn divergences.

Distributionally Robust Deep Q-Learning In- terpolating between optimal transport and mmd using Sinkhorn divergences

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:59.703610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:26:48.411547Z digest=sha256:f6b88652c30ce859729bfff47d08a07c10ab12627e1a06cf4202d33d05df2144

Observation 742e4c8e-f8e8-491c-8a58-6d57f19be4c7 · outbound

This paper cites Sample complexity of Sinkhorn divergences.

Distributionally Robust Deep Q-Learning Sample complexity of Sinkhorn divergences

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:59.689706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:26:48.498761Z digest=sha256:6a2262408bc98f56099cb9aa18e6564996ba8f6a790a91278708e0e0bfbaa9de

Observation 61f552e5-ee19-46ee-931f-1072e7e14675 · outbound

This paper cites Stability of entropic optimal transport and Schr¨ odinger bridges.

Distributionally Robust Deep Q-Learning Stability of entropic optimal transport and Schr¨ odinger bridges

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:59.675743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:26:48.596240Z digest=sha256:b7a84917ced65a5e061223fb1c2a67a08545e0f27d4efc45ced61336a2d4e5b1

Observation 80cc6a39-69e5-4015-b210-74594bae82a9 · outbound

This paper cites Robust Markov decision processes: Beyond rectangularity.Mathematics of Operations Research, 48(1):203–226, 2023.

Distributionally Robust Deep Q-Learning Robust Markov decision processes: Beyond rectangularity.Mathematics of Operations Research, 48(1):203–226, 2023

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:59.663234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:26:48.650603Z digest=sha256:8aef6179fe81bc7dd0ff34ab4153c6837619ebb8baebf7fc7bc777a92513037d

Observation f6bfc4ed-785b-4e05-9a94-c43e1aedc790 · outbound

This paper cites Double Q-learning.

Distributionally Robust Deep Q-Learning Double Q-learning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:59.650585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:26:48.760252Z digest=sha256:d7346086cf97e87a1609595509f842e08164ac26e85aa61f8d811dd7df73f9e5

Observation 14cce011-4b45-4032-9112-e2dbd88ecc1b · outbound

This paper cites Universal approximation of an unknown mapping and its derivatives using multilayer feedforward networks.

Distributionally Robust Deep Q-Learning Universal approximation of an unknown mapping and its derivatives using multilayer feedforward networks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:26:48.845231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:26:48.845231Z digest=sha256:9bb2cd2ca725e6f049b41a69608eb560de7bd9369518b1e6879f818ac7c66f2f

Observation b10a21ee-df77-4096-b7eb-947881c8ce83 · outbound

This paper cites Learning to utilize shaping rewards: A new approach of reward shaping.

Distributionally Robust Deep Q-Learning Learning to utilize shaping rewards: A new approach of reward shaping

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:26:48.908999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:26:48.908999Z digest=sha256:7d24a9facdedcbbc952b08fb1091c37ba4aefd8323e4487d34feb51376b1fe6d

Observation c468e487-5adc-4025-a1aa-9d4ca6407773 · outbound

This paper cites Robust dynamic programming.

Distributionally Robust Deep Q-Learning Robust dynamic programming

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:59.619185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:26:48.942188Z digest=sha256:b33698967881188dfd510f74633066e1c0e16ba76c592081e18c70dbf5e46777

Observation 54768045-e007-48f4-b640-1e1cf17b974a · outbound

This paper cites Probability essentials.

Distributionally Robust Deep Q-Learning Probability essentials

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:26:49.009637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:26:49.009637Z digest=sha256:31e9c7958feda703a9d6ea81e94594009a868b22a1288da834c3f7151007e578

Observation 44bda2c0-d2e0-465a-95df-957062227227 · outbound

This paper cites Universal approximation with deep narrow networks.

Distributionally Robust Deep Q-Learning Universal approximation with deep narrow networks

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:26:49.077557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:26:49.077557Z digest=sha256:3bbe4330b77c2fea2e50b2e9d4046d1123333a5541542b7e9d8677315060a5f2

Observation 5daadba4-6eb4-4020-a685-1df0d1095287 · outbound

This paper cites Probability theory: a comprehensive course.

Distributionally Robust Deep Q-Learning Probability theory: a comprehensive course

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:59.586840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:26:49.194831Z digest=sha256:b8e9427beeca4202d11fa928129e46f4b77926f1798d35a8f237b211464e05b3

Observation 2196972d-d2b4-4108-97d5-1b9403556f1d · outbound

This paper cites An Efficient Solution to s-Rectangular Robust Markov Decision Processes.

Distributionally Robust Deep Q-Learning An Efficient Solution to s-Rectangular Robust Markov Decision Processes

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:26:53.737945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:26:49.284674Z digest=sha256:2487431a2d1ac47481ae11e2cad212de9d1839ae33f45d1b39bcb163bd92e955

Observation 7514de8d-3eea-48f6-971f-ef12ca6fd1ec · outbound

This paper cites Playing fps games with deep reinforcement learning.

Distributionally Robust Deep Q-Learning Playing fps games with deep reinforcement learning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:59.570408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:26:49.444931Z digest=sha256:124e663664b6523c9f3b6c66f8142c0fba513585f363e60b2f3a964165312c97

Observation eb124f69-56be-46e8-bf84-c9cdfc65cddb · outbound

This paper cites Policy gradient algorithms for robust MDPs with non-rectangular uncertainty sets.

Distributionally Robust Deep Q-Learning Policy gradient algorithms for robust MDPs with non-rectangular uncertainty sets

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:26:49.568230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:26:49.568230Z digest=sha256:a63d1e96ab668a096347860ac5c10f480f805be85aa0679adcd00e24ed19fd53

Observation cc11521d-e1d2-4be7-b7f0-b84df17fb853 · outbound

This paper cites On the efficiency of entropic regularized algorithms for optimal transport.

Distributionally Robust Deep Q-Learning On the efficiency of entropic regularized algorithms for optimal transport

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:59.547637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:26:49.722090Z digest=sha256:14a1ad54a75e5d3b44e625a45490028ebfe6f37ee9f86b6a3cd5b1197be3e85f

Observation c7092ac4-6b7b-4725-93a5-afb950f7fa7c · outbound

This paper cites Distri- butionally robust Q-learning.

Distributionally Robust Deep Q-Learning Distri- butionally robust Q-learning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:59.293038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:26:49.840858Z digest=sha256:ba63a0813f42a4951be3331e860f75a0cec1748ade8298bdde4734d4f9682503

Observation 2e2e4b44-ad3f-4779-b02b-8752d1864697 · outbound

This paper cites Generative model for financial time series trained with MMD using a signature kernel.

Distributionally Robust Deep Q-Learning Generative model for financial time series trained with MMD using a signature kernel

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:26:49.934758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:26:49.934758Z digest=sha256:2fce1a2f172d3144b9321a7fc190697b1d4d86a8bccbcac5cd9baf2479be17a2

Observation e7ebd7b1-035f-4e75-9410-b57014f7b327 · outbound

This paper cites Robust MDPs with k-rectangular uncertainty.Mathematics of Operations Research, 41(4):1484–1509, 2016.

Distributionally Robust Deep Q-Learning Robust MDPs with k-rectangular uncertainty.Mathematics of Operations Research, 41(4):1484–1509, 2016

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:58.911349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:26:50.014471Z digest=sha256:894fb0f5192dba0e415a925fa20116533631d0010caf54e486dc0f48266e5755

Observation 8e1c0b4b-89f1-4d19-ae29-31fa38fe1ceb · outbound

This paper cites Rusu, Joel Veness, Marc G.

Distributionally Robust Deep Q-Learning Rusu, Joel Veness, Marc G

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:58.541563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:26:50.133260Z digest=sha256:b83073308ce309d87d47e38324297e3e126029fcacbf123f9dd9ed26470a1536

Observation d33f3249-3bb9-4739-af78-1e972bea358f · outbound

This paper cites Robust SGLD algorithm for solving non-convex distributionally robust optimisation problems.

Distributionally Robust Deep Q-Learning Robust SGLD algorithm for solving non-convex distributionally robust optimisation problems

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:26:53.421784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:26:50.237029Z digest=sha256:401d675ec8c721b3c883142759e8d216c850202fa05e53d2a96d71fccf3ba559

Observation de0a5a18-4943-49b3-87de-d696a478000b · outbound

This paper cites Universal approximation results for neural networks with non-polynomial activation function over non-compact domains.

Distributionally Robust Deep Q-Learning Universal approximation results for neural networks with non-polynomial activation function over non-compact domains

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:26:50.329080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:26:50.329080Z digest=sha256:38f6736dbc84840dbddbe8e1cb8499ba336364124189b82c1116bdb3e504f81d

Observation c1a5b965-ff5d-465a-bfc0-3bcb47d4fa5e · outbound

This paper cites Neural networks can detect model-free static arbitrage strategies.

Distributionally Robust Deep Q-Learning Neural networks can detect model-free static arbitrage strategies

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:58.294955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:26:50.440013Z digest=sha256:03c981c507b8b7213e6a7e11616ca6176669ccfc3211fff7623764c2ab30f32e

Observation 8042d1d6-830d-4a55-9209-fbb9db4e4e27 · outbound

This paper cites Robust Q-learning algorithm for markov decision processes under Wasserstein uncertainty.

Distributionally Robust Deep Q-Learning Robust Q-learning algorithm for markov decision processes under Wasserstein uncertainty

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:58.082779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:26:50.518778Z digest=sha256:9447ee519e4b8e32ea4cc84cdd04518a7445c5ac1073197b7b8428c4782ec487

Observation a298fc5b-8367-49f9-a359-13bf6d74ecac · outbound

This paper cites Non-concave stochastic optimal control in finite discrete time under model uncertainty.

Distributionally Robust Deep Q-Learning Non-concave stochastic optimal control in finite discrete time under model uncertainty

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:26:53.268456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:26:50.619941Z digest=sha256:5a0ec0e2dc23bab46ce48db728000b4a5f01e6b7be23e457c0729c712f361408

Observation d9cd9e9c-8e3a-42e0-a377-cd107430d60b · outbound

This paper cites Markov decision processes under model uncertainty.Mathematical Finance, 33(3):618–665, 2023.

Distributionally Robust Deep Q-Learning Markov decision processes under model uncertainty.Mathematical Finance, 33(3):618–665, 2023

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:57.698534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:26:50.701032Z digest=sha256:667ad40c1510a42d4777b7d9fc6b8710f143f53b6cda59d075a7b85b6b0f90d1

Observation 2afa2ddc-496e-4315-bb2f-1cf669188761 · outbound

This paper cites Robust control of Markov decision processes with uncertain transition matrices.

Distributionally Robust Deep Q-Learning Robust control of Markov decision processes with uncertain transition matrices

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:57.465811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:26:50.787033Z digest=sha256:249b7ad7a4365b9e78c17950594d739576be0dffee94ecb367ba743f23d72cdd

Observation f2560ea4-7670-4b2b-a2bf-adbff400a0d5 · outbound

This paper cites Introduction to entropic optimal transport.

Distributionally Robust Deep Q-Learning Introduction to entropic optimal transport

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:26:50.872266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:26:50.872266Z digest=sha256:49108d822c86104bdcbac1280c73624ac009ebc6e53968c2b2583441204b1b76

Observation b2049884-5728-4488-b3a3-ba753efab79c · outbound

This paper cites Robust reinforcement learning using offline data.

Distributionally Robust Deep Q-Learning Robust reinforcement learning using offline data

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:57.328939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:26:50.954796Z digest=sha256:6f872b8732cc3dca44209aa2748c4aa600234d959f0e3a9a79e11b184cb1a2c1

Observation eea69139-5433-4db7-a707-11bffd83a295 · outbound

This paper cites Approximation theory of the MLP model in neural networks.

Distributionally Robust Deep Q-Learning Approximation theory of the MLP model in neural networks

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:57.165931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:26:51.087991Z digest=sha256:071777c41a334462b77824655b398b962a5088eaf3a4cdf0c75091ede4170bfc

Observation ddba481d-1872-4c3d-8d0d-08a0d80b2ec0 · outbound

This paper cites Distributionally Robust Optimization: A Review.

Distributionally Robust Deep Q-Learning Distributionally Robust Optimization: A Review

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:26:51.197695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:26:51.197695Z digest=sha256:c199389b687325be1c5be25cad838d272f0be411002d74772222dc042e885b14

Observation 8cbb9b91-566c-43a2-b474-e6d9ec509828 · outbound

This paper cites Distributionally robust model-based reinforcement learning with large state spaces.

Distributionally Robust Deep Q-Learning Distributionally robust model-based reinforcement learning with large state spaces

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:57.038771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:26:51.293275Z digest=sha256:173465f4e9bd395950f1f01d892aef616d6c20aaa41b102755a25f28a426381c

Observation c1040ac8-a747-415d-883d-d2afbf624f0f · outbound

This paper cites Principles of mathematical analysis , volume 3.

Distributionally Robust Deep Q-Learning Principles of mathematical analysis , volume 3

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:56.883665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:26:51.344975Z digest=sha256:722bc838b45a8fa4852828244515c18c6e25748229a46191276f1cacb426a53b

Observation e4022c2a-03d7-4696-aefe-c7f312d7cf30 · outbound

This paper cites Structural estimation of Markov decision processes.

Distributionally Robust Deep Q-Learning Structural estimation of Markov decision processes

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:56.753591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:26:51.495132Z digest=sha256:c6badd5a5f1b70c294ed589a7ba8855c9e0bf8951f1aee14bc2f7d87e1f208ce

Observation de461997-29b7-4fbc-b4f7-2c936034e59a · outbound

This paper cites Universal approximation using feedforward neural networks: A survey of some existing methods, and some new results.

Distributionally Robust Deep Q-Learning Universal approximation using feedforward neural networks: A survey of some existing methods, and some new results

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:56.643022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:26:51.592726Z digest=sha256:435d4bd20727a7e11d15804179d7c1b0a858bf09b658b49c76b4ed513c98d528

Observation e5d9e8fa-928a-4bf0-bce8-48391be90837 · outbound

This paper cites A relationship between arbitrary positive matrices and doubly stochastic matrices.

Distributionally Robust Deep Q-Learning A relationship between arbitrary positive matrices and doubly stochastic matrices

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:56.488784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:26:51.691369Z digest=sha256:f673d4f436568541516757e7c969b5334b806af493685e3fd078d98c0ed3cd78

Observation 66540415-4f95-4bf3-ab46-6545ddfb849a · outbound

This paper cites Distributionally Robust Reinforcement Learning.

Distributionally Robust Deep Q-Learning Distributionally Robust Reinforcement Learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T14:26:51.760331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:26:51.760331Z digest=sha256:c9845be11c71abe0ef8965e9716b05907bb1e1a7d0ec73f1f386aad2502168d5

Observation 94800049-d17f-4781-8b6f-e4ea6e729172 · outbound

This paper cites Reinforcement learning: An introduction.

Distributionally Robust Deep Q-Learning Reinforcement learning: An introduction

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T14:26:51.830034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:26:51.830034Z digest=sha256:336555bb9299b26df6d9ac1ecf5b971922de548a4d559f33539a5811a8083a6a

Observation daddacaa-5731-4c90-af39-38eaa04cb3cb · outbound

This paper cites Sinkhorn Divergences for Unbalanced Optimal Transport.

Distributionally Robust Deep Q-Learning Sinkhorn Divergences for Unbalanced Optimal Transport

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T14:26:51.926466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:26:51.926466Z digest=sha256:55734e64be1f26ef719e6151017f0cce490599a38169419df287aeefed368379

Observation 90a94789-bcdc-4323-8563-707e0907b41c · outbound

This paper cites Deep reinforcement learning: From Q-learning to deep Q-learning.

Distributionally Robust Deep Q-Learning Deep reinforcement learning: From Q-learning to deep Q-learning

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:56.330805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:26:51.997657Z digest=sha256:1fbde022e493f8a9b1ba3ea87019bb8034d6825c049794ee505bc45f14044689

Observation 9eb8acb0-7ace-4c7a-a575-38fa53266750 · outbound

This paper cites Deep reinforcement learning with double Q-learning.

Distributionally Robust Deep Q-Learning Deep reinforcement learning with double Q-learning

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:56.157590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:26:52.063423Z digest=sha256:ad28b95a73a4acc278b32332db2ce9ba895cc50e72181809ac296107819b8116

Observation c2722181-b1d1-4d24-a021-bab0144a683b · outbound

This paper cites Springer, 2009.

Distributionally Robust Deep Q-Learning Springer, 2009

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:55.978372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:26:52.124396Z digest=sha256:d2a382ee0a192e867ccb7d117e4420bef53827155a1b63a984aacc6f01712923

Observation 8a7301d0-28a0-4060-9ebc-8fe544c6c5b7 · outbound

This paper cites Sinkhorn Distributionally Robust Optimization.

Distributionally Robust Deep Q-Learning Sinkhorn Distributionally Robust Optimization

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T14:26:52.197351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:26:52.197351Z digest=sha256:a149e2e436041b20cf9e29e8efae6c70001dff36ca62006bd53622281917787b

Observation d1c9673b-498d-493f-aa91-e2d8a0a25d65 · outbound

This paper cites Policy gradient in robust MDPs with global convergence guarantee, 2023.

Distributionally Robust Deep Q-Learning Policy gradient in robust MDPs with global convergence guarantee, 2023

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:55.774906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:26:52.283186Z digest=sha256:6cc614f486e82f8f5b6b96d5ed0fb130bf28a0efcb62ba424c2898a732017eb5

Observation 5a9ea900-7541-46ae-8ad3-2da91d988d08 · outbound

This paper cites A finite sample complexity bound for distribu- tionally robust Q-learning.

Distributionally Robust Deep Q-Learning A finite sample complexity bound for distribu- tionally robust Q-learning

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:55.587589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:26:52.364238Z digest=sha256:cad30b339d79595e20caa8ea1c1f776b26b430296fe6cdcb5f5c2d6d7833ed80

Observation 1fdd2e0b-7b1d-4ce1-9e14-8bbca4e226fd · outbound

This paper cites Sample complexity of variance-reduced distribu- tionally robust Q-learning.

Distributionally Robust Deep Q-Learning Sample complexity of variance-reduced distribu- tionally robust Q-learning

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:55.405979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:26:52.436747Z digest=sha256:40a3c0f63fced6b7c0342d56778baf3e668eca9f21130a1fe9d95c397233d32f

Observation e46641e0-b583-4a1a-a953-2f322065ca27 · outbound

This paper cites Online robust reinforcement learning with model uncertainty.

Distributionally Robust Deep Q-Learning Online robust reinforcement learning with model uncertainty

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T14:26:52.508548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:26:52.508548Z digest=sha256:eebad0d6ba026139560a62645d50c8052ce052af947f5f50ebdc4e53dfbf0269

Observation a4a41a19-4422-4fd5-ae3a-fdaa96631e48 · outbound

This paper cites Policy gradient method for robust reinforcement learning, 2022.

Distributionally Robust Deep Q-Learning Policy gradient method for robust reinforcement learning, 2022

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:55.215824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:26:52.591503Z digest=sha256:5d5207469a74008ca4a6263d540fb0863fa3d2c3a518d230922e28c1e5ba16a7

Observation 0848a997-d20d-44f5-984c-55c6e2a6e9e0 · outbound

This paper cites an unresolved cited work.

Distributionally Robust Deep Q-Learning Unresolved cited work

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T14:26:52.659845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:26:52.659845Z digest=sha256:61b740d8370759c5e2c80be6ba8b135e42867f405d4c5b1db64d5c22060e3f9d

Observation 33fe3bf7-a23e-473b-bd21-36faa5913b7a · outbound

This paper cites Robust Markov decision processes.

Distributionally Robust Deep Q-Learning Robust Markov decision processes

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:55.057577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:26:52.749919Z digest=sha256:ee81a15fecce5ac1cb1397c847ac0a244eccbf515b98bf44f26f385e69dff81a

Observation 1b3e7aef-70fb-45c3-88bd-798f2076fa29 · outbound

This paper cites Distributionally robust Markov decision processes.

Distributionally Robust Deep Q-Learning Distributionally robust Markov decision processes

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:54.909557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:26:52.836265Z digest=sha256:4469a567e4032ca7a70f95ec99cd7e75e0c30f1e27f9686128b6456fd1e0843c

Observation fc240353-9188-43fd-8ea5-ac5677b2b03d · outbound

This paper cites A convex optimization approach to distributionally robust Markov decision processes with Wasser- stein distance.

Distributionally Robust Deep Q-Learning A convex optimization approach to distributionally robust Markov decision processes with Wasser- stein distance

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:54.728296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:26:52.904606Z digest=sha256:7a83e373add757e40bec0a2c8993dfa1a6fb6beb941a6f0d87ae770d2e338959

Observation 8158a332-f577-4700-bfb1-88cf205a7ebc · outbound

This paper cites Wasserstein distributionally robust stochastic control: A data-driven approach.

Distributionally Robust Deep Q-Learning Wasserstein distributionally robust stochastic control: A data-driven approach

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:54.550806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:26:52.969979Z digest=sha256:9bef64166190b4d030376b4a425f1bff7dae2cfc650cb9918df9e38c146a7685

Observation 1e5f6c2c-bf72-44e7-87f3-7a550d9b32ce · outbound

This paper cites On linear optimization over Wasserstein balls.

Distributionally Robust Deep Q-Learning On linear optimization over Wasserstein balls

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:54.369542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:26:53.058645Z digest=sha256:73621d9c4a704364e34a1d31d424dc25760d8930f36bb6293c358b95f91163b6

Pith citing papers

Observation 1aaa4507-cffe-4cfc-8c0a-8764e250a890 · inbound

Robust $Q$-learning for mean-field control under Wasserstein uncertainty in common noise cites this paper.

Robust $Q$-learning for mean-field control under Wasserstein uncertainty in common noise Distributionally Robust Deep Q-Learning

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-07-04T05:29:35.432393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T15:59:25.061726Z digest=sha256:99a3cf7d155f62b1fa0ca2472c49f69221f7e20ffb15f9868bbf16cdd91c21df

Observation 8e06f30b-2c3b-48d9-8ffd-d411166a95ae · inbound

Robustness in Sequential Decision Making under Evolving Uncertainty: Evidence from High-Frequency Market Making cites this paper.

Robustness in Sequential Decision Making under Evolving Uncertainty: Evidence from High-Frequency Market Making Distributionally Robust Deep Q-Learning

Reference 105

Resolution
verified exact
local_arxiv, observed 2026-07-10T09:57:00.795026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-07-10T09:56:20.351339Z digest=sha256:e855ca7fbbcb8d7aba73e00f9e85d32dd1e354f95dca66d4b1fe0512f282abdc