REVIEW 3 major objections 3 minor 2 cited by
A randomisation method for mean-field control problems with common noise
T0 review · 3 major / 3 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read This paper proves that randomising a mean-field control by a Poisson point process leaves its value function unchanged, and yields a BSDE representation and a randomised dynamic programming principle.
desk verdict Genuine extension of control randomisation to mean-field control with common noise; the main theorems look sound, but Lemma 4.1 has a mis-citation and the motivating example's computations are wrong. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The engine is the control randomisation apparatus built on the time-dependent action spaces $A_s=L^0(\Omega,G\vee F^W_s,P;A)$, equivalence classes of one-time controls measurable with respect to the idiosyncratic noise and the common Brownian motion up to time $s$. A Poisson random measure $\mu$ on $(0,T]\times A_T$ with intensity $\lambda_s(d\alpha)ds$, whose topological support is $A_s$, generates a step process $\hat I^{t,\alpha_t}$; a measurable selection theorem (Proposition 3.8) and a canonical-space predictability lemma (Lemma 3.10) turn this into a genuine $A$-valued control $I^{t,\alpha_t}$ while preserving the conditional law of the state given the common noise. Optimising over strictly positive predictable intensities $\nu$ via the Girsanov tilt $d\hat P^\nu/d\hat P=\mathcal{L}^\nu$ makes the intensity the control. Penalised BSDEs (5.1) then converge to the minimal constrained BSDE (5.4), from which the supremum-over-equivalent-measures DPP is read off.
What would settle it
Check the textbook problem cited in Lemma 4.1: it establishes right-continuity of the augmented Brownian filtration, not the left-continuity used there. If the family $\lambda_s$ built from rational times cannot be shown to have support equal to $A_s$ at irrational times by some other argument, then Assumption C is unproved and the randomised state dynamics together with Theorem 4.8 lose their foundation; a concrete check is to compute the topological support of $\lambda_r$ at an irrational $r$.
Extended reading notes
Core claim
The central claim is Theorem 4.8: for all $t\in[0,T]$, initial states $\xi$, and initial actions $\alpha_t$, the randomised value function $$V^R(t,\xi,\alpha_t)=\sup_{\nu\in\mathcal{V}}\mathbb{E}^{\hat P^\nu}\big[g\big(\hat $P^{{F^{B,\mu}}$,\hat P}_T $X^{{t,\xi,\alpha_t}}$_T,$X^{{t,\xi,\alpha_t}}$_T\big)+\int_t^T f(\cdots)\,dr\big]$$ equals the original mean-field control value $V(t,\xi)$. The proof runs through an isomorphism between the original control set and the set of $F^B$-predictable processes taking values in the time-dependent spaces $A_s=L^0(\Omega,G\vee F^W_s,P;A)$, followed by a Poisson random measure with intensity $\lambda_s(d\alpha)ds$ supported on $A_s$; Girsanov tilting makes the intensity the control. From this equivalence the paper derives (Theorem 5.2) that $V$ is the minimal solution of a constrained BSDE with constrained jumps, and (Theorem 5.6) the randomised DPP $V(t,m)=\sup_{\nu}\mathbb{E}^{\hat P^\nu}[V(s,\hat P^\nu_s X)+\int_t^s f(\cdots)dr]$, which reduces to the standard DPP when $V$ is regular.
Load-bearing premise
The paper needs a family of probability-like measures on the action space whose support at each time is exactly the set of admissible one-time actions and that are absolutely continuous with respect to the later measures; the proof that such measures exist relies on a continuity property of the Brownian filtration that the cited source does not state.
Editorial extensions
If this is right
- The value function of a mean-field control problem with common noise can be computed through the randomised problem, and the result is independent of the chosen initial action, of the family $\lambda_s$, and of the probability-space extension.
- The value function admits a probabilistic representation as the minimal solution of a constrained BSDE with constrained jumps, with the value at intermediate time $s$ equal to $V(s,\hat P^{\nu}_s X)$ along the randomised state flow.
- A randomised dynamic programming principle holds for every initial law $m\in P_2(\mathbb{R}^d)$: $V(t,m)=\sup_{\nu}\mathbb{E}^{\hat P^\nu}[V(s,\hat P^{\nu}_s X)+\int_t^s f(\cdots)dr]$.
- When the value function is regular enough to serve as a terminal reward, the same randomisation machinery recovers the standard non-randomised DPP for mean-field control with common noise.
- The law-invariance of the original value function transfers to the randomised value function: $V^R$ depends only on the law of the initial state.
Reading between the lines
- This suggests a practical numerical scheme: simulate the common noise and the Poisson marks once, then optimise the intensity process $\nu$ inside the expectation; the paper does not implement such a scheme, but the DPP is already in that form.
- The $L^0$-valued reformulation is a general device, so the same randomisation method may carry over to mean-field games and to path-dependent McKean–Vlasov control, where the same measurability issue between original and reformulated controls arises.
- Because the DPP is written as a supremum over equivalent probability measures, it offers a weak-formulation counterpart to maximum-principle characterisations, potentially easing comparisons between necessary and sufficient optimality conditions in mean-field control.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops a control-randomisation approach to mean-field control (MFC) problems with common noise. It first reformulates the admissible controls as common-noise-adapted processes taking values in the spaces A_s = L^0(G ∨ F^W_s, P; A), proves an isomorphism between the original and reformulated control sets, and then replaces the reformulated control by an independent Poisson random measure whose intensity is the new control. The main results are Theorem 4.8, asserting equality V = V^R between the original and randomised value functions; Theorem 5.2, giving a constrained-BSDE representation of the value function as the minimal solution; and Theorem 5.6, a randomised dynamic programming principle expressed as a supremum over equivalent probability measures.
Significance. If the main theorems hold, this is a substantial extension of the control-randomisation method to McKean–Vlasov control with common noise, going beyond the decoupled formulation in [6] and yielding both a BSDE representation and a randomised DPP. The paper is largely self-contained and gives detailed proofs of the measurable-selection and identification results; the central equivalence is parameter-free and does not rely on fitted quantities. The main line is nevertheless supported by two points that currently need repair: the construction of the intensity family in Lemma 4.1 is justified by a mis-cited result, and the introductory example contains an incorrect optimality computation. Neither point appears to invalidate Theorem 4.8, but both must be fixed before the paper can be accepted.
major comments (3)
- [§1, Example 1.1] The numerical computation in Example 1.1 is incorrect and should be redone. For α ≡ 1, the mean satisfies m_s = e^s − 1, so X_1(ξ = −1) = e − 2 and X_1(ξ = 1) = e; the two rewards are 4.5 − e and e − 1, giving average 1.75, not (e − 1)/2. For α ≡ −1, m_s = 1 − e^s, so X_1(ξ = −1) = −e and X_1(ξ = 1) = 2 − e; the average reward is (2.5 + e + 0.5 + e)/2 = 1.5 + e, which exceeds the claimed value. Thus α* ≡ 1 is not optimal. This example is motivational and does not feed into the proof of Theorem 4.8, but the false computation must be corrected or the example replaced.
- [§4, Lemma 4.1] The proof of Lemma 4.1 says that G ∨ F^{W,P} is left-continuous and cites Karatzas–Shreve Problem 7.6. That problem concerns right-continuity of the augmented Brownian filtration, not left-continuity. The left-continuity claim itself is true for the augmented Brownian filtration because W_r is the almost-sure limit of W_{q_n} for q_n ↑ r, but a correct argument or citation needs to be supplied. Since Assumption C underpins the Poisson random measure and hence Theorem 4.8, this proof gap is load-bearing and should be repaired.
- [Appendix A.1, proof of V ≤ V^R] The inequality V ≤ V^R relies on [5, Proposition A.1] in a version adapted to the time-dependent action spaces A_s. The paper only sketches the adaptation, stating that the proof can be extended with minimal changes and modifying the kernels q^m to q^m_s. This is a key bridge between original and randomised controls, so the generalized proposition should either be stated and proved in full, or the paper should give a precise reduction to [5, Proposition A.1] that accounts for the time-dependent supports and the adapted intensity in property (iii).
minor comments (3)
- [Section 3.2, first paragraph] The sentence beginning 'This approach aligns the control set thus aligning the control set more closely...' contains a duplicated phrase and should be rewritten.
- [Notation / Section 3 and Lemma 4.1] The spaces A_s are defined using the non-augmented filtration F^W_s, while Lemma 4.1 invokes left-continuity of the augmented filtration G ∨ F^{W,P}. Please clarify whether the construction uses the augmented filtration or explain why the L^0-equivalence classes make the distinction immaterial.
- [References] Several reference entries appear incomplete as printed, for example [4], [6], and [7]; please supply the missing journal, volume, and page data.
Circularity Check
No significant circularity: the randomized and original control problems are connected by a genuine proof rather than by definition or by fitting.
full rationale
The paper's central claim, Theorem 4.8, is an equality V(t,xi)=V^R(t,xi,alpha_t) between two separately defined value functions. The randomized object V^R is constructed from the reformulated control set, a Poisson measure with intensity family (lambda_s), and admissible intensities nu in V; it is not defined as V, nor is any fitted parameter renamed as a prediction. The proof in Appendix A establishes both inequalities. The direction V <= V^R approximates a general reformulated control by marked point processes using [5, Proposition A.1], an external result whose stated assumptions do not include the target equality, and the direction V^R <= V constructs an auxiliary extended problem and compares value functions. Theorem 5.2 passes to the minimal constrained BSDE solution through [31, Theorem 2.1], a published prior theorem by two of the authors with an independent proof; invoking it as a black box is standard and does not reduce the present results to their own assumptions. Theorem 5.6 then combines Theorem 5.2's representation Y_t = sup_nu E[Y_r + integral f] with Theorem 5.5's identification Y_s = V(s, conditional law), so the DPP is derived, not assumed. No equation in the paper exhibits a quantity that is equal to its input by construction. The misattributed Karatzas-Shreve citation for left-continuity in Lemma 4.1 and the incorrect arithmetic in Example 1.1 are correctness concerns, not circularity, and neither enters the proof of Theorem 4.8 as a self-referential assumption.
Assumptions & free parameters
assumptions (5)
- domain assumption Assumption A: b, sigma, sigma0 are continuous and Lipschitz in (x,mu) uniformly in (t,a), with linear growth in a.
- domain assumption Assumption B: f and g satisfy linear growth in x and in the second moment of mu.
- ad hoc to paper Assumption C: there exists a family of finite measures lambda_s on A_T with topological support A_s, lambda_s absolutely continuous with respect to lambda_r for s<=r, and bounded total mass.
- domain assumption G is a separable sigma-algebra independent of F^{W,B} and rich enough to represent every law in P2(Rd).
- standard math Known results are used as black boxes, including [31, Theorem 2.1] on constrained BSDEs and [18, Proposition 2.4] on law invariance.
Cite this review
Pith. "Pith review of A randomisation method for mean-field control problems with common noise." pith.science (2026). https://pith.science/paper/UAPRXGEV
@misc{pith2026241220782,
author = {Pith},
title = {Pith review of: A randomisation method for mean-field control problems with common noise},
year = {2026},
howpublished = {\url{https://pith.science/paper/UAPRXGEV}},
note = {Machine review of arXiv:2412.20782}
}
read the original abstract
We study mean-field control (MFC) problems with common noise using the control randomisation framework, where we substitute the control process with an independent Poisson point process, controlling its intensity instead. To address the challenges posed by the mean-field interactions in this randomisation approach, we reformulate the admissible control as L 0 -valued processes adapted only to the common noise. We then construct the randomised control problem from this reformulated control process, and show its equivalence to the original MFC problem. Thanks to this equivalence, we can represent the value function as the minimal solution to a backward stochastic differential equation (BSDE) with constrained jumps. Finally, using this probabilistic representation, we derive a randomised dynamic programming principle (DPP) for the value function, expressed as a supremum over equivalent probability measures.
Forward citations
Cited by 2 Pith papers
-
Mean Field Control with Poissonian Common Noise: A Pathwise Compactification Approach
Mean-field control with finite-intensity Poissonian common noise admits optimal relaxed controls, and the same pathwise compactification yields strong mean-field equilibria in games.
-
The randomization method in stochastic optimal control
A survey of the randomization method proving that the value of an optimal control problem equals the value of a randomized problem and is represented by a constrained BSDE, with a complete tour of applications.
Reference graph
Works this paper leans on
-
[6]
Erhan Bayraktar, Andrea Cosso, and Huyˆ en Pham. Randomi zed dynamic programming principle and feynman-kac representation for optimal control of McKe an-vlasov dynamics. 370(3):2115–2160
-
[1]
Algorithmic trading in a microstructural limit order book model
Fr´ ed´ eric Abergel, Cˆ ome Hur´ e, and Huyˆ en Pham. Algorithmic trading in a microstructural limit order book model. Quantitative Finance, 20(8):1263–1283, August 2020
work page 2020
-
[2]
A Maximum Princi ple for SDEs of Mean-Field Type
Daniel Andersson and Boualem Djehiche. A Maximum Princi ple for SDEs of Mean-Field Type. Applied Mathematics & Optimization , 63(3):341–356, 2011
work page 2011
-
[3]
Paolo Baldi. Stochastic Calculus. Universitext. Springer International Publishing, Cham, 2017
work page 2017
-
[4]
Elena Bandini, Andrea Cosso, Marco Fuhrman, and Huyˆ en P ham. Backward SDEs for optimal control of partially observed path-dependent stochastic s ystems: A control randomization approach. 28(3):1634–1678
-
[5]
Elena Bandini, Andrea Cosso, Marco Fuhrman, and Huyˆ en P ham. Randomization method and backward SDEs for optimal control of partially observed pat h-dependent stochastic systems, 2016
work page 2016
-
[7]
A stochastic target formulation for opt imal switching problems in finite horizon
Bruno Bouchard. A stochastic target formulation for opt imal switching problems in finite horizon. 81(2):171–197. 34
-
[8]
A Genera l Stochastic Maximum Principle for SDEs of Mean-field Type
Rainer Buckdahn, Boualem Djehiche, and Juan Li. A Genera l Stochastic Maximum Principle for SDEs of Mean-field Type. Applied Mathematics & Optimization , 64(2):197–216, 2011
work page 2011
Show all 37 references
-
[9]
Mean-field stochastic differential equations and associated PDEs
Rainer Buckdahn, Juan Li, Shige Peng, and Catherine Rain er. Mean-field stochastic differential equations and associated PDEs. 45(2):824–878
-
[10]
Forward–backwar d stochastic differential equations and con- trolled McKean–Vlasov dynamics
Ren´ e Carmona and Fran¸ cois Delarue. Forward–backwar d stochastic differential equations and con- trolled McKean–Vlasov dynamics. The Annals of Probability , 43(5):2647–2700, September 2015
2015
-
[11]
Springer International Publishing, Cham, 2018
Ren´ e Carmona and Fran¸ cois Delarue.Probabilistic Theory of Mean Field Games with Applications I, volume 83 of Probability Theory and Stochastic Modelling . Springer International Publishing, Cham, 2018
2018
-
[12]
K. L. Chung and R. J. Williams. Introduction to Stochastic Integration . Birkh¨ auser Boston, Boston, MA, 1990
1990
-
[13]
A pseudo-m arkov property for controlled diffusion processes
Julien Claisse, Denis Talay, and Xiaolu Tan. A pseudo-m arkov property for controlled diffusion processes. 54(2):1017–1029
-
[14]
An Introduction to the Theory of Point Processes
Daryl Daley and David Vere-Jones. An Introduction to the Theory of Point Processes . Probability and Its Applications. Springer New York
-
[15]
Probabilities and potential , volume 29 of North-Holland mathematics studies
Claude Dellacherie and Paul-Andr´ e Meyer. Probabilities and potential , volume 29 of North-Holland mathematics studies . North-Holland, 1978
1978
-
[16]
Control randomisation approach for policy gra- dient and application to reinforcement learning in optimal switching
Robert Denkert, Huyˆ en Pham, and Xavier Warin. Control randomisation approach for policy gra- dient and application to reinforcement learning in optimal switching. 2024
2024
-
[17]
McK ean–vlasov optimal control: Limit theory and equivalence between different formulations
Mao Fabrice Djete, Dylan Possama ¨ ı, and Xiaolu Tan. McK ean–vlasov optimal control: Limit theory and equivalence between different formulations. 47(4):289 1–2930
-
[18]
McKean–vlasov optimal control: The dynamic programming principle
Mao Fabrice Djete, Dylan Possama ¨ ı, and Xiaolu Tan. McKean–vlasov optimal control: The dynamic programming principle. 50(2):791–833
-
[19]
Adding constraints t o BSDEs with jumps: an alternative to multidimensional reflections
Romuald Elie and Idris Kharroubi. Adding constraints t o BSDEs with jumps: an alternative to multidimensional reflections. 18:233–250
-
[20]
BSDE representation s for optimal switching problems with controlled volatility
Romuald Elie and Idris Kharroubi. BSDE representation s for optimal switching problems with controlled volatility. 14(3):1450003
-
[21]
Probabilistic repre sentation and approximation for coupled systems of variational inequalities
Romuald Elie and Idris Kharroubi. Probabilistic repre sentation and approximation for coupled systems of variational inequalities. 80(17):1388–1396
-
[22]
Optimal swit ching problems with an infinite set of modes: An approach by randomization and constrained backwa rd SDEs
Marco Fuhrman and Marie-Am´ elie Morlais. Optimal swit ching problems with an infinite set of modes: An approach by randomization and constrained backwa rd SDEs. 130(5):3120–3153
-
[23]
Randomized and backward SDE representation for optimal control of non-markovian SDEs
Marco Fuhrman and Huyˆ en Pham. Randomized and backward SDE representation for optimal control of non-markovian SDEs. 25(4):2134–2167
-
[24]
Represe ntation of non-markovian optimal stop- ping problems by constrained BSDEs with a single jump
Marco Fuhrman, Huyˆ en Pham, and Federica Zeni. Represe ntation of non-markovian optimal stop- ping problems by constrained BSDEs with a single jump. 21:1– 7
-
[25]
Shiryaev
Jean Jacod and Albert N. Shiryaev. Limit Theorems for Stochastic Processes , volume 288 of Grundlehren der mathematischen Wissenschaften . Springer Berlin Heidelberg, Berlin, Heidelberg, 2003
2003
-
[26]
Progressive Stochas tic Processes and an Application to the Itˆ o Integral
Svenja Kaden and J¨ urgen Potthoff. Progressive Stochas tic Processes and an Application to the Itˆ o Integral. Stochastic Analysis and Applications , 22(4):843–865, 2004
2004
-
[27]
Ioannis Karatzas and Steven E. Shreve. Brownian Motion and Stochastic Calculus , volume 113 of Graduate Texts in Mathematics . Springer New York, New York, NY, 1998. 35
1998
-
[28]
A numerical algorithm for fully nonlinear HJB equations: An approach by control randomization
Idris Kharroubi, Nicolas Langren´ e, and Huyˆ en Pham. A numerical algorithm for fully nonlinear HJB equations: An approach by control randomization. Monte Carlo Methods and Applications , 20(2):145–165, June 2014
2014
-
[29]
Discrete time approximation of fully nonlinear HJB equations via BSDEs with nonpositive jumps
Idris Kharroubi, Nicolas Langren´ e, and Huyˆ en Pham. Discrete time approximation of fully nonlinear HJB equations via BSDEs with nonpositive jumps. The Annals of Applied Probability , 25(4):2301– 2338, August 2015
2015
-
[30]
Backward SDEs with constrained jumps and quasi-variational inequalities
Idris Kharroubi, Jin Ma, Huyˆ en Pham, and Jianfeng Zhan g. Backward SDEs with constrained jumps and quasi-variational inequalities. 38(2):794–840
-
[31]
Feynman–kac represen tation for hamilton–jacobi–bellman IPDE
Idris Kharroubi and Huyˆ en Pham. Feynman–kac represen tation for hamilton–jacobi–bellman IPDE. 43(4):1823–1865
-
[32]
Th´ eorie des jeux de champ moyen et applications
Pierre-Louis Lions. Th´ eorie des jeux de champ moyen et applications. Cours au Coll` ege de France , 2006—2012
2006
-
[33]
Dynamic programming for opt imal control of stochastic McKean– vlasov dynamics
Huyˆ en Pham and Xiaoli Wei. Dynamic programming for opt imal control of stochastic McKean– vlasov dynamics. 55(2):1069–1101
-
[34]
Bellman equation and viscos ity solutions for mean-field stochastic control problem
Huyˆ en Pham and Xiaoli Wei. Bellman equation and viscos ity solutions for mean-field stochastic control problem. ESAIM: Control, Optimisation and Calculus of Variations , 24(1):437–461, 2018
2018
-
[35]
Continuous Martingales and Brownian Motion , volume 293 of Grundlehren der mathematischen Wissenschaften
Daniel Revuz and Marc Yor. Continuous Martingales and Brownian Motion , volume 293 of Grundlehren der mathematischen Wissenschaften . Springer Berlin Heidelberg, Berlin, Heidelberg, 1999
1999
-
[36]
Necessary conditions for optimal control of stochastic systems with random jumps
Shanjian Tang and Xunjing Li. Necessary conditions for optimal control of stochastic systems with random jumps. 32(5):1447–1475
-
[37]
Dynamic portfolio optimization with liquidity cost and market impa ct: a simulation-and-regression approach
Rongju Zhang, Nicolas Langren´ e, Yu Tian, Zili Zhu, Fim a Klebaner, and Kais Hamza. Dynamic portfolio optimization with liquidity cost and market impa ct: a simulation-and-regression approach. Quantitative Finance , 19(3):519–532, March 2019. 36
2019
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.