REVIEW 4 cited by
Unified continuous-time q-learning for mean-field game and mean-field control problems
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This paper studies the continuous-time q-learning in mean-field jump-diffusion models in a setting where the environment simulator does not provide direct access to the population distribution. We propose the integrated q-function in decoupled form (decoupled Iq-function) and establish its martingale characterization, which provides a unified policy evaluation rule for both mean-field game (MFG) and mean-field control (MFC) problems. Moreover, we consider the learning procedure where population distribution is updated based on the representative agent's state values. Depending on the task to solve the MFG or MFC problem, we can employ the decoupled Iq-function differently to characterize the mean-field equilibrium policy or the mean-field optimal policy respectively. Based on these theoretical findings, we devise a unified parametric q-learning algorithm for both MFG and MFC problems by utilizing test policies and the averaged martingale orthogonality condition. In two applications within and beyond LQ framework, we illustrate the effectiveness and efficiency of our unified parametric q-learning algorithm for both MFG and MFC learning tasks.
Forward citations
Cited by 4 Pith papers
-
Reinforcement learning for irreversible reinsurance problems: the randomized singular control approach
A randomized, entropy-regularized singular control law enables continuous-time reinforcement learning to solve irreversible reinsurance problems, with an explicit equilibrium boundary Γ(x)=e^{-βΦ(x)/λ}.
-
Continuous-Time Reinforcement Learning for $N$-Player Stochastic Differential Games with Exploratory Policies
For entropy-regularized N-player differential games, a Nash-type equilibrium exists exactly when the Gibbs conditional best responses are jointly compatible, checkable via a cross-partial criterion on the learned q-functions.
-
Actor-Critic Learning for Extended Mean Field Control with Deterministic Policies
Model-free deterministic policy gradients and a continuous-time deep actor-critic algorithm solve extended mean-field control problems whose dynamics and rewards depend on the joint state-control law.
-
Continuous-time reinforcement learning for optimal switching over multiple regimes
An entropy-regularized exploratory formulation of multi-regime optimal switching is shown to admit well-posed HJB systems, fast-converging policy iteration, and a vanishing-entropy limit that recovers the classical problem.
Discussion (0). Continue with ORCID to comment.