Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T18:40:33.633478Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 8 of 8 outbound references and 8 inbound Pith citation observations for arXiv:2512.04697.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T18:40:33.633478Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T11:22:20.627811Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T15:38:34.106512Z
8 of 8 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation aa758c4c-9bfa-4c24-a10a-f3f9ce6b2f26 · outbound
Continuous-time reinforcement learning for optimal switching over multiple regimes Regret of exploratory policy improvement and $q$-learning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc47abb3-348a-4f1d-a86b-46ddfa7d5c89 · outbound
Continuous-time reinforcement learning for optimal switching over multiple regimes Unified continuous-time q-learning for mean-field game and mean-field control problems
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ef18a3b-e27c-4f3d-b03b-240f8d4e3257 · outbound
Continuous-time reinforcement learning for optimal switching over multiple regimes A Reinforcement Learning Framework for Some Singular Stochastic Control Problems
Reference 1965
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d62ca23-1a6c-4053-9ccc-fde9eb011d72 · outbound
Continuous-time reinforcement learning for optimal switching over multiple regimes A Two-fold Randomization Framework for Impulse Control Problems
Reference 2009
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b810dae3-1fcb-4582-ba65-5826100d48ea · outbound
Continuous-time reinforcement learning for optimal switching over multiple regimes Learning to Optimally Stop Diffusion Processes, with Financial Applications
Reference 2010
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e4c2a5f-5c00-4222-ac2e-b336fc67cad9 · outbound
Continuous-time reinforcement learning for optimal switching over multiple regimes Unresolved cited work
Reference 2012
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f534aa72-264b-4a02-a0f2-5a8175cb9bda · outbound
Continuous-time reinforcement learning for optimal switching over multiple regimes Reinforcement Learning for Jump-Diffusions, with Financial Applications
Reference 2013
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 935cf58c-9e9a-4751-a580-10f6e0482858 · outbound
Continuous-time reinforcement learning for optimal switching over multiple regimes Dianetti, G
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75640507-26e9-4183-a362-14780ded219a · inbound
Equilibrium under Time-Inconsistency: A New Existence Theory by Vanishing Entropy Regularization Continuous-time reinforcement learning for optimal switching over multiple regimes
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d5669ebd-2afc-4457-9c87-c6009bffcb29 · inbound
Continuous-time q-learning for mean-field control with common noise, part-I: Theoretical foundations Continuous-time reinforcement learning for optimal switching over multiple regimes
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 4117b598-bf81-4f0b-96b1-8232ddb200c1 · inbound
Continuous-time q-learning for mean-field control with common noise, part-II: q-learning algorithms Continuous-time reinforcement learning for optimal switching over multiple regimes
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation bedc34ca-490f-45de-8ad9-9a62f5a67d3f · inbound
Equilibrium for Time-inconsistent Mean Field Games: A Systematic Analysis by Entropy Regularization Continuous-time reinforcement learning for optimal switching over multiple regimes
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ebe6d189-a512-47f1-9bed-0c2b2bdb1baa · inbound
Mean Field Competition of Optimal Switching: The Vanishing Entropy Regularization Approach Continuous-time reinforcement learning for optimal switching over multiple regimes
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation bf738f60-25ae-434c-afb2-64351c2421dc · inbound
Deterministic Policy Gradient for Learning Equilibrium in Time-Inconsistent Control Problems Continuous-time reinforcement learning for optimal switching over multiple regimes
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e11fb221-d557-46cc-bf3c-81a7570cf5f0 · inbound
Randomized Optimal Switching Problem and Related Mirror Descent Flow Continuous-time reinforcement learning for optimal switching over multiple regimes
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e4888772-7b01-4209-aa7b-cd21a5fdc2f2 · inbound
Continuous-Time Reinforcement Learning for $N$-Player Stochastic Differential Games with Exploratory Policies Continuous-time reinforcement learning for optimal switching over multiple regimes
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.