Pith. sign in

Paper Citation Record · LEDGER

Continuous-time reinforcement learning for optimal switching over multiple regimes

As of 20 August 2026, this Paper Citation Record lists 8 of 8 outbound references and 8 inbound Pith citation observations for arXiv:2512.04697.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2512.04697 v3

Coverage vector

measured 8 of 8 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T18:40:33.633478Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T11:22:20.627811Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T15:38:34.106512Z

Reference resolution

8 of 8 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation aa758c4c-9bfa-4c24-a10a-f3f9ce6b2f26 · outbound

This paper cites Regret of exploratory policy improvement and $q$-learning.

Continuous-time reinforcement learning for optimal switching over multiple regimes Regret of exploratory policy improvement and $q$-learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T18:40:33.523529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:40:33.523529Z digest=sha256:68e66e14956121b79cd013da055a85c207170355bf8442d155e4f447e1a44130

Observation bc47abb3-348a-4f1d-a86b-46ddfa7d5c89 · outbound

This paper cites Unified continuous-time q-learning for mean-field game and mean-field control problems.

Continuous-time reinforcement learning for optimal switching over multiple regimes Unified continuous-time q-learning for mean-field game and mean-field control problems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T18:40:33.633478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:40:33.633478Z digest=sha256:6a4853baaa2f431da5b793eec4ea8161ed65f1880946f9dd0f0eab6592c3a316

Observation 8ef18a3b-e27c-4f3d-b03b-240f8d4e3257 · outbound

This paper cites A Reinforcement Learning Framework for Some Singular Stochastic Control Problems.

Continuous-time reinforcement learning for optimal switching over multiple regimes A Reinforcement Learning Framework for Some Singular Stochastic Control Problems

Reference 1965

Resolution
unresolved
no resolver link, observed 2026-08-03T18:40:33.417501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:40:33.417501Z digest=sha256:999e635a30db101476a8e71944f296716942ac31d3ee0251bdf48587ced9c49c

Observation 6d62ca23-1a6c-4053-9ccc-fde9eb011d72 · outbound

This paper cites A Two-fold Randomization Framework for Impulse Control Problems.

Continuous-time reinforcement learning for optimal switching over multiple regimes A Two-fold Randomization Framework for Impulse Control Problems

Reference 2009

Resolution
unresolved
no resolver link, observed 2026-08-03T18:40:32.937455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:40:32.937455Z digest=sha256:ec79d4614c61463091d66ca16b15324776e74434b36a4c3601c2bf6e816a54c3

Observation b810dae3-1fcb-4582-ba65-5826100d48ea · outbound

This paper cites Learning to Optimally Stop Diffusion Processes, with Financial Applications.

Continuous-time reinforcement learning for optimal switching over multiple regimes Learning to Optimally Stop Diffusion Processes, with Financial Applications

Reference 2010

Resolution
unresolved
no resolver link, observed 2026-08-03T18:40:33.058915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:40:33.058915Z digest=sha256:550c54058978ec69e915773aae7a3b81c080fae389a94ba62aa43daad5d0596c

Observation 1e4c2a5f-5c00-4222-ac2e-b336fc67cad9 · outbound

This paper cites an unresolved cited work.

Continuous-time reinforcement learning for optimal switching over multiple regimes Unresolved cited work

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-03T18:40:32.817487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:40:32.817487Z digest=sha256:f2ff5c71e4edb4840026df71d951ef8feeaaed4ea63ebd8f1224cb830439d80d

Observation f534aa72-264b-4a02-a0f2-5a8175cb9bda · outbound

This paper cites Reinforcement Learning for Jump-Diffusions, with Financial Applications.

Continuous-time reinforcement learning for optimal switching over multiple regimes Reinforcement Learning for Jump-Diffusions, with Financial Applications

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-03T18:40:33.333118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:40:33.333118Z digest=sha256:63e3b799c6c9c1a7390aefcfa96330de9703cc990ac96c0df4f6df8aeb09dbd3

Observation 935cf58c-9e9a-4751-a580-10f6e0482858 · outbound

This paper cites Dianetti, G.

Continuous-time reinforcement learning for optimal switching over multiple regimes Dianetti, G

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T18:40:33.193079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:40:33.193079Z digest=sha256:b0cbd88d385c0fbcd5e17046c89eccaabbc15ffa521dd23edddc943f7cdaa4a4

Pith citing papers

Observation 75640507-26e9-4183-a362-14780ded219a · inbound

Equilibrium under Time-Inconsistency: A New Existence Theory by Vanishing Entropy Regularization cites this paper.

Equilibrium under Time-Inconsistency: A New Existence Theory by Vanishing Entropy Regularization Continuous-time reinforcement learning for optimal switching over multiple regimes

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-31T02:03:07.292107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T14:10:21.557614Z digest=sha256:8bb9fe14e8001b4f6fa92dead15dd58e5913201fefeae5d5208e1148e4858a26

Observation d5669ebd-2afc-4457-9c87-c6009bffcb29 · inbound

Continuous-time q-learning for mean-field control with common noise, part-I: Theoretical foundations cites this paper.

Continuous-time q-learning for mean-field control with common noise, part-I: Theoretical foundations Continuous-time reinforcement learning for optimal switching over multiple regimes

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-31T02:03:07.292107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-07T09:19:03.944541Z digest=sha256:fd16f222527bf327d9bfbc960f0b9b1160ec2f91e312a8cc5ea94b02abf11ecc

Observation 4117b598-bf81-4f0b-96b1-8232ddb200c1 · inbound

Continuous-time q-learning for mean-field control with common noise, part-II: q-learning algorithms cites this paper.

Continuous-time q-learning for mean-field control with common noise, part-II: q-learning algorithms Continuous-time reinforcement learning for optimal switching over multiple regimes

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-31T02:03:07.292107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-07T08:41:31.247249Z digest=sha256:9f7d3cd12271501794c8df1d06bfb27b7ff9ddbfea707866498a8b5b7eeddfc4

Observation bedc34ca-490f-45de-8ad9-9a62f5a67d3f · inbound

Equilibrium for Time-inconsistent Mean Field Games: A Systematic Analysis by Entropy Regularization cites this paper.

Equilibrium for Time-inconsistent Mean Field Games: A Systematic Analysis by Entropy Regularization Continuous-time reinforcement learning for optimal switching over multiple regimes

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-31T02:03:07.292107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-15T02:13:18.861160Z digest=sha256:bef046e120d2539b08bba344571c415ced8c94f147412dd7eb5706c1cff1a376

Observation ebe6d189-a512-47f1-9bed-0c2b2bdb1baa · inbound

Mean Field Competition of Optimal Switching: The Vanishing Entropy Regularization Approach cites this paper.

Mean Field Competition of Optimal Switching: The Vanishing Entropy Regularization Approach Continuous-time reinforcement learning for optimal switching over multiple regimes

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-31T02:03:07.292107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-29T05:53:55.444421Z digest=sha256:e67a86757712112271615b1e7690d6f9e4bb7d840a07d7a3f4b53cae972630a8

Observation bf738f60-25ae-434c-afb2-64351c2421dc · inbound

Deterministic Policy Gradient for Learning Equilibrium in Time-Inconsistent Control Problems cites this paper.

Deterministic Policy Gradient for Learning Equilibrium in Time-Inconsistent Control Problems Continuous-time reinforcement learning for optimal switching over multiple regimes

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-31T02:03:07.292107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T07:56:11.143952Z digest=sha256:99e3f9f77445956340cc17af17a3a4a5bd13c1d8608ee7e0bb524b9fce240725

Observation e11fb221-d557-46cc-bf3c-81a7570cf5f0 · inbound

Randomized Optimal Switching Problem and Related Mirror Descent Flow cites this paper.

Randomized Optimal Switching Problem and Related Mirror Descent Flow Continuous-time reinforcement learning for optimal switching over multiple regimes

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-31T02:03:07.292107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T06:23:23.030846Z digest=sha256:73744f9d671b3b5f61f1ff2c24beaab93d7500e07c131ff37e0604b757e62fe2

Observation e4888772-7b01-4209-aa7b-cd21a5fdc2f2 · inbound

Continuous-Time Reinforcement Learning for $N$-Player Stochastic Differential Games with Exploratory Policies cites this paper.

Continuous-Time Reinforcement Learning for $N$-Player Stochastic Differential Games with Exploratory Policies Continuous-time reinforcement learning for optimal switching over multiple regimes

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T11:22:20.627811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:22:20.627811Z digest=sha256:eecf7b9a35c583353850765bbd44ad76b7e9dd7201353835f12bd7de032d1e47