Pith. sign in

REVIEW 4 cited by

Unified continuous-time q-learning for mean-field game and mean-field control problems

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.04521 v3 pith:LCFFQAKS submitted 2024-07-05 math.OC cs.LGq-fin.CP

classification math.OCcs.LGq-fin.CP
keywords mean-fieldq-learningunifieddecoupledpolicyproblemsalgorithmcontinuous-time
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper studies the continuous-time q-learning in mean-field jump-diffusion models in a setting where the environment simulator does not provide direct access to the population distribution. We propose the integrated q-function in decoupled form (decoupled Iq-function) and establish its martingale characterization, which provides a unified policy evaluation rule for both mean-field game (MFG) and mean-field control (MFC) problems. Moreover, we consider the learning procedure where population distribution is updated based on the representative agent's state values. Depending on the task to solve the MFG or MFC problem, we can employ the decoupled Iq-function differently to characterize the mean-field equilibrium policy or the mean-field optimal policy respectively. Based on these theoretical findings, we devise a unified parametric q-learning algorithm for both MFG and MFC problems by utilizing test policies and the averaged martingale orthogonality condition. In two applications within and beyond LQ framework, we illustrate the effectiveness and efficiency of our unified parametric q-learning algorithm for both MFG and MFC learning tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Reinforcement learning for irreversible reinsurance problems: the randomized singular control approach

    math.OC 2025-12 conditional novelty 7.0 of 10

    A randomized, entropy-regularized singular control law enables continuous-time reinforcement learning to solve irreversible reinsurance problems, with an explicit equilibrium boundary Γ(x)=e^{-βΦ(x)/λ}.

  2. Continuous-Time Reinforcement Learning for $N$-Player Stochastic Differential Games with Exploratory Policies

    math.OC 2026-07 conditional novelty 6.0 of 10

    For entropy-regularized N-player differential games, a Nash-type equilibrium exists exactly when the Gibbs conditional best responses are jointly compatible, checkable via a cross-partial criterion on the learned q-functions.

  3. Actor-Critic Learning for Extended Mean Field Control with Deterministic Policies

    math.OC 2026-07 accept novelty 6.0 of 10

    Model-free deterministic policy gradients and a continuous-time deep actor-critic algorithm solve extended mean-field control problems whose dynamics and rewards depend on the joint state-control law.

  4. Continuous-time reinforcement learning for optimal switching over multiple regimes

    math.OC 2025-12 conditional novelty 5.0 of 10

    An entropy-regularized exploratory formulation of multi-regime optimal switching is shown to admit well-posed HJB systems, fast-converging policy iteration, and a vanishing-entropy limit that recovers the classical problem.

Pith tools