Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-12T04:23:38.033851Z
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 1 inbound Pith citation observation for arXiv:2607.03168.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-12T04:23:38.033851Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T16:09:17.938445Z
A source-named dated measurement, never combined with another source.
Source: cited_works
68 of 68 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e55c09b6-214b-41c4-8fe4-70267e612289 · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning The reality gap in robotics: Challenges, solutions, and best practices.Annual Review of Control, Robotics, and Autonomous Systems, 9:403–432, 2026
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68954cfc-3963-4741-ac1e-d62685f871fd · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning State entropy regularization for robust reinforcement learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1c3656a-ec19-47df-9b4e-17bd9f8e0751 · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning Algorithmic market making in dealer markets with hedging and market impact.Mathematical Finance, 33(1):41–79, 2023
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17b60314-984c-4c47-bc44-1d1f51d93bfb · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning Continuous-time q-learning in jump-diffusion models under Tsallis entropy, 2024
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0143cd1-6702-447f-be60-aa3e46cec213 · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning Unresolved cited work
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3eab32c3-4cf9-4e07-91eb-901d772f0714 · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning Robust multi-agent reinforcement learning via adversarial regularization: Theoretical foundation and stable algorithms
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 375f2e5c-5cde-46f8-8bf4-f4f2579de619 · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning Cambridge University Press, 2015
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93fc19e1-55bd-4ff7-9bbb-f105249d01e2 · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning Robust reinforcement learning with general utility
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05a45eaf-66a8-46bc-94b4-32e5fdc5d141 · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning Deterministic policy gradient for reinforcement learning with continuous time and state, 2026
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9871788a-9966-433f-8100-16b1431a853b · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning Dai and Mark Gluzman
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d8373fd-bd9b-41b5-bd38-c14404bc1707 · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning Twice regularized MDPs and the equivalence between robustness and regularization.Advances in Neural Information Processing Systems, 34:22274–22287, 2021
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8a486ce-ae07-45a6-a9cb-ded5fdad91b4 · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning Robustness and regularization in rein- forcement learning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9963f6f-df42-4436-ad16-87b200406be7 · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning Entropy regularization in mean-field games of optimal stopping.arXiv preprint arXiv:2509.18821, 2025
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06821e48-926d-47d9-860f-99bbed3b73f5 · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning Donsker and S
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d1741b6-4870-48e5-9b7e-1f4392276592 · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning Maximum entropy RL (provably) solves some robust RL problems
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ceaf1e3a-95e4-4c89-974f-6252c0a4ef51 · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning Actor-critic learning for mean-field control in continuous time.Journal of Machine Learning Research, 26(127):1–42, 2025
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f952601-a583-4640-a2f9-84ea97d7c283 · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning Reinforcement learning for jump-diffusions, with financial applications.Mathematical Finance, 2026
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfe9bb42-3b36-421d-9306-73fe6d92c58b · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning A theory of regularized Markov decision processes
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf25a4cd-9b13-4596-a1f4-195d8fcf92ed · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning Convergence of policy gradient methods for finite-horizon exploratory linear-quadratic control problems.SIAM Journal on Control and Optimization, 62(2):1060–1092, 2024
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4b43664-96f5-4105-8dd6-d488cb367ba9 · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning Scalable first-order methods for robust MDPs
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1da14839-0990-42c5-9ff4-278c954c8b20 · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning Continuous-time Markov decision processes
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 519eeef1-1c2a-4b6b-9844-2cc1cc5686a0 · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning Entropy regularization for mean field games with learning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75dda10d-d98a-4867-95cf-11c7bde5c433 · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning Soft Actor-Critic Algorithms and Applications
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d16ae5b1-2ed3-41e6-b042-0eec7db811c6 · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning Continuous-time reinforcement learning for optimal switching over multiple regimes, 2025
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b554bdba-10d9-4f60-a934-1bc179a4210a · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning Regularized policies are reward robust
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00812bbf-3931-488c-88f9-76d4fa584d30 · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning Unresolved cited work
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57de4f8b-0432-421c-999d-106d8dbee4ff · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning Continuous-time risk-sensitive reinforcement learning via quadratic variation penalty.Applied Mathematics & Optimization, 93(2):58, 2026
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50dd245c-0439-4ef6-9f6f-cf85d969695f · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning Accuracy of discretely sampled stochastic policies in continuous-time reinforcement learning.arXiv preprint arXiv:2503.09981, 2025
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca5b8012-018c-481a-9a3b-339c03f57ff6 · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning Policy evaluation and temporal-difference learning in continuous time and space: A martingale approach.Journal of Machine Learning Research, 23(154):1–55, 2022
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1383df64-a354-442d-a83d-f65a237b477c · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning Policy gradient and actor-critic learning in continuous time and space: Theory and algorithms.Journal of Machine Learning Research, 23(275):1–50, 2022
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f99d273-aa42-46c5-8082-2bc62f6cc01e · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning q-learning in continuous time.Journal of Machine Learning Research, 24(161):1–61, 2023
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80902972-6b15-4076-b3b6-7d024cfb8136 · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning A Fisher–Rao gradient flow for entropy-regularised Markov decision processes in Polish spaces.Foundations of Computational Mathematics, pages 1–75, 2025
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13b458a5-c8bd-4e28-80ed-eeab45b891e9 · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning Policy gradient for rectangular robust Markov decision processes
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4496fd1-8fc2-473a-b071-6806cdbaefc3 · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning Policy mirror descent for reinforcement learning: Linear convergence, new sampling complexity, and generalized problem classes.Mathematical Programming, 198:1059–1106, 2023
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7096bc84-ddec-424f-9abc-0d800e7fcb34 · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning Policy gradient algorithms for robust MDPs with nonrectan- gular uncertainty sets.SIAM Journal on Optimization, 36(1):120–151, 2026
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a7b2e0a-e9a1-4cff-b1e9-1568af924771 · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning Efficient adversarial training without attacking: Worst-case-aware robust reinforcement learning.Advances in Neural Information Processing Systems, 35:22547–22561, 2022
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad55dd71-92bb-4076-b9eb-75a7d95e6428 · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning Reinforcement learning in robust Markov decision processes
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d38b7c30-186c-4d52-a905-450a7b0c1fb0 · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning Robust value iteration for continuous control tasks
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d236428f-cfe0-44a9-b0a2-40acbdaa2310 · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning Reinforcement Learning for Intensity Control: An Application to Choice-Based Network Revenue Management
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7737300e-14dd-4916-bd0e-727fb99ea930 · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning Robust reinforcement learning.Neural Computation, 17(2):335–359, 2005
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8226e0d7-d891-4fda-8b47-8f9f08183c15 · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning Robust control of Markov decision processes with uncertain transition matrices.Operations Research, 53(5):780–798, 2005
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f10580b3-e943-4592-826e-82eefe4a3812 · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning Springer, 2007
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7388648a-f0fb-42fd-af3a-6f87f7e49b11 · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning Variational inference for Markov jump processes.Advances in Neural Information Processing Systems, 20, 2007
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f16f85a-0f9f-440c-9ef8-c57b66dbab0d · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning Robustness and risk-sensitivity in Markov decision processes
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f456ba5f-84f4-41ef-b56e-26ff18d895ce · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning Sim-to-real transfer of robotic control with dynamics randomization
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7eec0f32-c0b8-43bf-b4d2-8e51b7386498 · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning Continuous-time reinforcement learning for robust control under worst-case uncertainty.International Journal of Systems Science, 52(4):770–784, 2021
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e7c112a-2c92-46c5-ab72-5ac3a559ed11 · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning Robust adversarial reinforcement learning
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5eadbd4f-9e6c-4827-89e2-5307305ad9a2 · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning Puterman.Markov Decision Processes: Discrete Stochastic Dynamic Programming
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0975acb-b648-469c-a14a-e6a5950eedec · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning On stochastic optimal control and reinforcement learning by approximate inference
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91091be2-8e07-4505-85ca-f86e04e5149f · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning Regularity and stability of feedback relaxed controls.SIAM Journal on Control and Optimization, 59(5):3118–3151, 2021
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1554f660-0ae9-4110-b35d-4769b455d15a · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning Continuous-time q-learning for mean-field control with common noise, part-I: Theoretical foundations, 2026
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6bd3e5a-2bac-44b6-8db4-13d0b3411562 · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning Continuous-time q-learning for mean-field control with common noise, part-II: q-learning algorithms, 2026
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee18c052-cba1-4ad1-ae4f-6403bc1833d6 · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning Trust region policy optimization
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb3ea989-c6e9-4ced-88c6-f75f2888c12a · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning Proximal Policy Optimization Algorithms
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc0f0bab-6d5d-492a-8808-74f992b82ec0 · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning Optimal scheduling of entropy regularizer for continuous-time linear-quadratic reinforcement learning.SIAM Journal on Control and Optimization, 62(1):135–166, 2024
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 319dfdc7-b431-419e-a0d0-16d2c806af49 · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning Action robust reinforcement learning and applications in continuous control
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a4d9645-cfe6-4b7b-ac61-0c52c551fa53 · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning Domain randomization for transferring deep neural networks from simulation to the real world
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07434df1-89cb-4df0-a59e-143d55f9c475 · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning Linearly-solvable Markov decision problems
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 728b9346-1ba6-458b-929c-e14820ffab7e · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning Policy gradient in robust MDPs with global convergence guarantee
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f47cd47e-76b9-4c6c-9834-2e89455c1353 · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning Continuous time q-learning for mean-field control problems.Applied Mathematics & Optimization, 91(1):10, 2025
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8a690eb-adaf-4ca4-9737-ea0cfcbac422 · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning Robust Markov decision processes.Mathematics of Operations Research, 38(1):153–183, 2013
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98aeb53b-fa58-4097-b5d6-79ff899e23cf · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning Continuous-time q-learning for Markov regime switching system under Tsallis entropy, 2026
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d4dd054-e8ce-470a-a732-a683a06c4fce · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning Policy optimization for continuous reinforcement learning
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad1d66e8-f82a-4178-9e1f-1954cfefa391 · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning X s∈S ¯dπ ρ(s) X a∈As µ(a|s) exp R(s, a)−R ˜θ∗(s, a) τ # =τlog
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d47622e-dd70-4fdd-be56-2bc61b8f3730 · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning The test statistic is t= WC(πτ)−WC(π std)q SE2 τ + SE2 0 , where SEτ and SE0 are the standard errors (across seeds) at the respective worst-case grid cells
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8c45d80-4a8d-4ca2-9227-6a7782a8cc60 · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning In both market making and queueing, performance peaks at an intermediate∆t and degrades for both coarser and finer grids
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33d1054c-1eed-4d66-b8cf-0d8aa0590d72 · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning The arrival-driven implementation has no grid resolution hyperparameter
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87c9c915-756e-4036-a7b0-9f18f168bf5d · outbound
Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning Unresolved cited work
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 405b5c91-abeb-4dd3-abfe-59b019854367 · inbound
Feedback Cycles in Exploratory Equilibria Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.