{"id":"a2385690-47a5-4f38-b964-1f648caac889","arxiv_id":"2412.20782","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"For mean-field control problems with common noise, randomizing controls via a Poisson point process preserves the value function and yields a randomized dynamic programming principle.","lead":"This paper adapts control randomization, where controls are replaced by random jumps whose rate you tune, to mean-field control problems with common noise. It proves the randomized version has the same value as the original and derives a randomized dynamic programming principle.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The cited left-continuity result is misattributed, but Brownian filtration is indeed left-continuous; Example 1.1's computations are incorrect but not load-bearing for Theorem 4.8.","rationale":"I read the proof of Theorem 4.8 as the central claim, and I do not find a fatal flaw in it. The reader's weakest assumption—that Lemma 4.1 fails because Brownian filtration is not left-continuous—is based on a misreading: the citation is indeed wrong, but the left-continuity property is true for continuous paths. The density of ⋃_{q<r} A_q in A_r follows because W_r is the limit of W_q, so the constructed measures λ_s have the stated support. Thus the randomised construction is founded. The genuine numerical error in Example 1.1 is significant as an illustrative computation, but it does not affect the proofs of the main theorems, which do not rely on that example. The paper still needs revision to correct Example 1.1 and the citation in Lemma 4.1, so the reader's CONDITIONAL verdict remains appropriate; I would not change it to ACCEPT as-is, nor to REJECT, because the central equivalence and DPP appear sound after replacing the incorrect citation with the continuity argument.","tokens_in":52265,"tokens_out":24953,"duration_ms":243046,"concrete_test":"Recompute Example 1.1 by solving dX_s = (α_s + E[X_s]) ds for α ≡ 1 and α ≡ −1, and compare the resulting rewards with the reported values 1/2(e−1) and e−5/2; this will confirm the example's numerical error, while the main theorem's proof can be checked independently via the left-continuity verification described above.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The reader's weakest assumption targets Lemma 4.1, which constructs the intensity measures (λ_s) whose full support on A_s is essential for the randomised setting. The proof cites Karatzas–Shreve Problem 7.6 for left-continuity of G ∨ F^{W,P}; that problem establishes right-continuity, not left-continuity. However, the left-continuity claim itself is true: because Brownian paths are continuous, W_r is the almost-sure limit of W_q as q↑r, so ℱ_r^W = σ(⋃_{q<r}ℱ_q^W), and the algebra generated by ⋃_{q<r}(G ∨ ℱ_q^W) is dense in L^1(G ∨ ℱ_r^W). Thus the family (λ_s) in Lemma 4.1 can be constructed; the proof merely cites the wrong result and should be repaired. The more concrete issue is Example 1.1: solving dX_s = (α_s + E[X_s])ds with α ≡ 1 and ξ ∼ 1/2δ_{-1}+1/2δ_1 gives E[X_s]=e^s−1 and rewards 4.5−e and e−1 for the two particles, averaging 1.75, not 1/2(e−1). Similarly α ≡ −1 yields average 1.5+e, so the claimed optimal control and value are incorrect. This error is confined to the motivational example and does not enter the proof of Theorem 4.8, but it should be corrected.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops a control-randomisation approach to mean-field control (MFC) problems with common noise. It first reformulates the admissible controls as common-noise-adapted processes taking values in the spaces A_s = L^0(G ∨ F^W_s, P; A), proves an isomorphism between the original and reformulated control sets, and then replaces the reformulated control by an independent Poisson random measure whose intensity is the new control. The main results are Theorem 4.8, asserting equality V = V^R between the original and randomised value functions; Theorem 5.2, giving a constrained-BSDE representation of the value function as the minimal solution; and Theorem 5.6, a randomised dynamic programming principle expressed as a supremum over equivalent probability measures.","tokens_in":52643,"tokens_out":10092,"duration_ms":105940,"significance":"If the main theorems hold, this is a substantial extension of the control-randomisation method to McKean–Vlasov control with common noise, going beyond the decoupled formulation in [6] and yielding both a BSDE representation and a randomised DPP. The paper is largely self-contained and gives detailed proofs of the measurable-selection and identification results; the central equivalence is parameter-free and does not rely on fitted quantities. The main line is nevertheless supported by two points that currently need repair: the construction of the intensity family in Lemma 4.1 is justified by a mis-cited result, and the introductory example contains an incorrect optimality computation. Neither point appears to invalidate Theorem 4.8, but both must be fixed before the paper can be accepted.","major_comments":[{"comment":"The numerical computation in Example 1.1 is incorrect and should be redone. For α ≡ 1, the mean satisfies m_s = e^s − 1, so X_1(ξ = −1) = e − 2 and X_1(ξ = 1) = e; the two rewards are 4.5 − e and e − 1, giving average 1.75, not (e − 1)/2. For α ≡ −1, m_s = 1 − e^s, so X_1(ξ = −1) = −e and X_1(ξ = 1) = 2 − e; the average reward is (2.5 + e + 0.5 + e)/2 = 1.5 + e, which exceeds the claimed value. Thus α* ≡ 1 is not optimal. This example is motivational and does not feed into the proof of Theorem 4.8, but the false computation must be corrected or the example replaced.","section":"§1, Example 1.1"},{"comment":"The proof of Lemma 4.1 says that G ∨ F^{W,P} is left-continuous and cites Karatzas–Shreve Problem 7.6. That problem concerns right-continuity of the augmented Brownian filtration, not left-continuity. The left-continuity claim itself is true for the augmented Brownian filtration because W_r is the almost-sure limit of W_{q_n} for q_n ↑ r, but a correct argument or citation needs to be supplied. Since Assumption C underpins the Poisson random measure and hence Theorem 4.8, this proof gap is load-bearing and should be repaired.","section":"§4, Lemma 4.1"},{"comment":"The inequality V ≤ V^R relies on [5, Proposition A.1] in a version adapted to the time-dependent action spaces A_s. The paper only sketches the adaptation, stating that the proof can be extended with minimal changes and modifying the kernels q^m to q^m_s. This is a key bridge between original and randomised controls, so the generalized proposition should either be stated and proved in full, or the paper should give a precise reduction to [5, Proposition A.1] that accounts for the time-dependent supports and the adapted intensity in property (iii).","section":"Appendix A.1, proof of V ≤ V^R"}],"minor_comments":[{"comment":"The sentence beginning 'This approach aligns the control set thus aligning the control set more closely...' contains a duplicated phrase and should be rewritten.","section":"Section 3.2, first paragraph"},{"comment":"The spaces A_s are defined using the non-augmented filtration F^W_s, while Lemma 4.1 invokes left-continuity of the augmented filtration G ∨ F^{W,P}. Please clarify whether the construction uses the augmented filtration or explain why the L^0-equivalence classes make the distinction immaterial.","section":"Notation / Section 3 and Lemma 4.1"},{"comment":"Several reference entries appear incomplete as printed, for example [4], [6], and [7]; please supply the missing journal, volume, and page data.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper fits the journal's scope and the central equivalence is not circular: the use of [31, Theorem 2.1] as a black box is standard, and the dependency on [5, Proposition A.1] is a proof-repair issue rather than a circularity concern. The two main repairs are local; if the authors provide a correct argument for Lemma 4.1 and correct Example 1.1, I would be willing to consider a revised version."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe thing to know: this paper extends the control randomisation method to mean-field control with common noise, and that is a real gap in the literature. The L0-valued reformulation of controls adapted only to the common noise is a genuine new idea, and it buys them a randomised DPP and a constrained BSDE representation for the value function. The paper is carefully built: the measurable selection and decomposition work in Section 3 is substantial, and the equivalence proof in Theorem 4.8 is long but structured. If the main theorems are right, this is an important within-subfield contribution, not a field reshapers, and it goes cleanly beyond Bayraktar-Cosso-Pham, which had no common noise and only deterministic initial laws.\n\nThe soft spots are local, but real. Lemma 4.1 cites Karatzas–Shreve Problem 7.6 for left-continuity of the augmented Brownian filtration, and that problem is about right-continuity. The stress-test note is right: the left-continuity claim itself is true because Brownian paths are continuous, so the construction of the family (lambda_s) is repairable with a correct argument. This is a flaw in the proof, not in the conclusion.\n\nThe motivational Example 1.1 is worse. For alpha ≡ 1, the dynamics give X_1 = x + e − 1, so the two particles get rewards roughly 1.78 and 1.72, averaging 1.75, not 0.5(e−1). And alpha ≡ −1 gives an average of 1.5 + e, which is larger. So the claimed optimal control and value are simply incorrect. The example is only motivational, and it does not enter the proof of Theorem 4.8, but it should be corrected or replaced. I would also check whether it actually demonstrates the failure of [6, Prop 2.2]: with controls allowed to depend on xi, as they are in the paper's admissible set, the equality may hold after all.\n\nWho is this for: stochastic control researchers working on mean-field control, randomisation methods, or BSDE representations. A serious referee should engage with this; it needs revision, not desk rejection. My own verdict is conditional, but the central argument appears sound and the two issues I found are fixable rather than fatal.\n\nRecommendation: send to peer review.","headline":"Genuine extension of control randomisation to mean-field control with common noise; the main theorems look sound, but Lemma 4.1 has a mis-citation and the motivating example's computations are wrong.","tokens_in":53102,"tokens_out":6094,"would_cite":true,"duration_ms":66456,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["60H30","60K35","60K37","93E20"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper proves that randomising a mean-field control by a Poisson point process leaves its value function unchanged, and yields a BSDE representation and a randomised dynamic programming principle.","keywords":["mean-field control","common noise","control randomisation","Poisson point processes","constrained backward SDEs","randomised dynamic programming principle","McKean-Vlasov dynamics","Girsanov transformation"],"falsifier":"Check the textbook problem cited in Lemma 4.1: it establishes right-continuity of the augmented Brownian filtration, not the left-continuity used there. If the family $\\lambda_s$ built from rational times cannot be shown to have support equal to $A_s$ at irrational times by some other argument, then Assumption C is unproved and the randomised state dynamics together with Theorem 4.8 lose their foundation; a concrete check is to compute the topological support of $\\lambda_r$ at an irrational $r$.","tokens_in":52130,"feed_emoji":"🎲","tokens_out":12456,"duration_ms":111761,"temperature":0.7,"pith_summary":"This paper is trying to establish that mean-field control problems with common noise, where many interacting agents share a common random environment, can be solved through control randomisation: replace the control by an independent Poisson point process and optimise its intensity instead. The main result is an exact equivalence: the value function of the original problem equals the value function of the randomised problem for every initial state and every initial action. This matters because the randomised formulation turns a mean-field optimisation into an optimisation over equivalent probability measures, which yields a representation of the value function as the minimal solution of a constrained backward stochastic differential equation and a dynamic programming principle. A key step is a reformulation of admissible controls as $L^0$-valued processes adapted only to the common noise, which keeps the mean-field interaction, namely the conditional distribution of the state given the common noise, unchanged under randomisation.","feed_headline":"Randomising the control recovers the mean-field value","feed_subtitle":"Replacing the control by a Poisson intensity gives the same value function and a dynamic programming principle.","key_machinery":"The engine is the control randomisation apparatus built on the time-dependent action spaces $A_s=L^0(\\Omega,G\\vee F^W_s,P;A)$, equivalence classes of one-time controls measurable with respect to the idiosyncratic noise and the common Brownian motion up to time $s$. A Poisson random measure $\\mu$ on $(0,T]\\times A_T$ with intensity $\\lambda_s(d\\alpha)ds$, whose topological support is $A_s$, generates a step process $\\hat I^{t,\\alpha_t}$; a measurable selection theorem (Proposition 3.8) and a canonical-space predictability lemma (Lemma 3.10) turn this into a genuine $A$-valued control $I^{t,\\alpha_t}$ while preserving the conditional law of the state given the common noise. Optimising over strictly positive predictable intensities $\\nu$ via the Girsanov tilt $d\\hat P^\\nu/d\\hat P=\\mathcal{L}^\\nu$ makes the intensity the control. Penalised BSDEs (5.1) then converge to the minimal constrained BSDE (5.4), from which the supremum-over-equivalent-measures DPP is read off.","core_discovery":"The central claim is Theorem 4.8: for all $t\\in[0,T]$, initial states $\\xi$, and initial actions $\\alpha_t$, the randomised value function $$V^R(t,\\xi,\\alpha_t)=\\sup_{\\nu\\in\\mathcal{V}}\\mathbb{E}^{\\hat P^\\nu}\\big[g\\big(\\hat $P^{{F^{B,\\mu}}$,\\hat P}_T $X^{{t,\\xi,\\alpha_t}}$_T,$X^{{t,\\xi,\\alpha_t}}$_T\\big)+\\int_t^T f(\\cdots)\\,dr\\big]$$ equals the original mean-field control value $V(t,\\xi)$. The proof runs through an isomorphism between the original control set and the set of $F^B$-predictable processes taking values in the time-dependent spaces $A_s=L^0(\\Omega,G\\vee F^W_s,P;A)$, followed by a Poisson random measure with intensity $\\lambda_s(d\\alpha)ds$ supported on $A_s$; Girsanov tilting makes the intensity the control. From this equivalence the paper derives (Theorem 5.2) that $V$ is the minimal solution of a constrained BSDE with constrained jumps, and (Theorem 5.6) the randomised DPP $V(t,m)=\\sup_{\\nu}\\mathbb{E}^{\\hat P^\\nu}[V(s,\\hat P^\\nu_s X)+\\int_t^s f(\\cdots)dr]$, which reduces to the standard DPP when $V$ is regular.","pith_inferences":["This suggests a practical numerical scheme: simulate the common noise and the Poisson marks once, then optimise the intensity process $\\nu$ inside the expectation; the paper does not implement such a scheme, but the DPP is already in that form.","The $L^0$-valued reformulation is a general device, so the same randomisation method may carry over to mean-field games and to path-dependent McKean–Vlasov control, where the same measurability issue between original and reformulated controls arises.","Because the DPP is written as a supremum over equivalent probability measures, it offers a weak-formulation counterpart to maximum-principle characterisations, potentially easing comparisons between necessary and sufficient optimality conditions in mean-field control."],"forward_implications":["The value function of a mean-field control problem with common noise can be computed through the randomised problem, and the result is independent of the chosen initial action, of the family $\\lambda_s$, and of the probability-space extension.","The value function admits a probabilistic representation as the minimal solution of a constrained BSDE with constrained jumps, with the value at intermediate time $s$ equal to $V(s,\\hat P^{\\nu}_s X)$ along the randomised state flow.","A randomised dynamic programming principle holds for every initial law $m\\in P_2(\\mathbb{R}^d)$: $V(t,m)=\\sup_{\\nu}\\mathbb{E}^{\\hat P^\\nu}[V(s,\\hat P^{\\nu}_s X)+\\int_t^s f(\\cdots)dr]$.","When the value function is regular enough to serve as a terminal reward, the same randomisation machinery recovers the standard non-randomised DPP for mean-field control with common noise.","The law-invariance of the original value function transfers to the randomised value function: $V^R$ depends only on the law of the initial state."],"supporting_citations":[{"why":"Establishes the dynamic programming principle and law-invariance for mean-field control with common noise, including the predictable conditional-law version used throughout the paper.","marker":"[18]"},{"why":"Provides the prior randomisation for McKean–Vlasov control without common noise; the paper's Example 1.1 shows the decoupled approach fails and this work extends the method.","marker":"[6]"},{"why":"Introduces the control randomisation idea of replacing a control by a Poisson process and optimising its intensity.","marker":"[7]"},{"why":"Supplies the general theorem that minimal solutions of constrained BSDEs are limits of penalised BSDEs, used in Theorem 5.2.","marker":"[31]"},{"why":"Original randomised and BSDE representation for non-Markovian control, providing the penalisation and representation technique adapted here.","marker":"[23]"},{"why":"Approximation of progressive controls by marked point processes with full support and bounded intensity, used to prove $V\\le V^R$.","marker":"[5]"},{"why":"Provides the partially-observed randomisation machinery, including Girsanov tilts, canonical extensions, and control transfer, used in the converse inequality.","marker":"[4]"},{"why":"The canonical-space lemma that progressive and predictable processes coincide, used to construct the identifying control process.","marker":"[13]"},{"why":"Cited in Lemma 4.1 for the continuity property of the augmented Brownian filtration on which the existence of $\\lambda_s$ depends.","marker":"[27]"},{"why":"Girsanov theorem for Poisson random measures, used to define the tilted measures $\\hat P^\\nu$.","marker":"[14]"}],"fun_headline_variants":["Poisson twist matches mean-field value exactly","Randomised controls hit the same mean-field value","Randomisation yields MFC value and DPP","Poisson trick equivalences MFC with common noise","Control randomisation proves MFC value equal"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper needs a family of probability-like measures on the action space whose support at each time is exactly the set of admissible one-time actions and that are absolutely continuous with respect to the later measures; the proof that such measures exist relies on a continuity property of the Brownian filtration that the cited source does not state.","fun_headline_variants_meta":{"raw":{"variants":["Poisson twist matches mean-field value exactly","Randomised controls hit the same mean-field value","Randomisation yields MFC value and DPP","Poisson trick equivalences MFC with common noise","Control randomisation proves MFC value equal"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000839,"raw_usage":{"total_tokens":3670,"prompt_tokens":971,"completion_tokens":2699,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":587,"completion_tokens_details":{"reasoning_tokens":2629}},"tokens_in":587,"tokens_out":2699,"duration_ms":16115,"temperature":1.0,"reasoning_tokens":2629,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T23:12:10.665702+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Check the textbook problem cited in Lemma 4.1: it establishes right-continuity of the augmented Brownian filtration, not the left-continuity used there. If the family $\\lambda_s$ built from rational times cannot be shown to have support equal to $A_s$ at irrational times by some other argument, then Assumption C is unproved and the randomised state dynamics together with Theorem 4.8 lose their foundation; a concrete check is to compute the topological support of $\\lambda_r$ at an irrational $r$.","supporting_citations":[{"cited_title":"McKean–vlasov optimal control: The dynamic programming principle","cited_arxiv_id":null,"evidence_quote":"Establishes the dynamic programming principle and law-invariance for mean-field control with common noise, including the predictable conditional-law version used throughout the paper."},{"cited_title":"Randomi zed dynamic programming principle and feynman-kac representation for optimal control of McKe an-vlasov dynamics","cited_arxiv_id":null,"evidence_quote":"Provides the prior randomisation for McKean–Vlasov control without common noise; the paper's Example 1.1 shows the decoupled approach fails and this work extends the method."},{"cited_title":"A stochastic target formulation for opt imal switching problems in ﬁnite horizon","cited_arxiv_id":null,"evidence_quote":"Introduces the control randomisation idea of replacing a control by a Poisson process and optimising its intensity."},{"cited_title":"Feynman–kac represen tation for hamilton–jacobi–bellman IPDE","cited_arxiv_id":null,"evidence_quote":"Supplies the general theorem that minimal solutions of constrained BSDEs are limits of penalised BSDEs, used in Theorem 5.2."},{"cited_title":"Randomized and backward SDE representation for optimal control of non-markovian SDEs","cited_arxiv_id":null,"evidence_quote":"Original randomised and BSDE representation for non-Markovian control, providing the penalisation and representation technique adapted here."},{"cited_title":"Randomization method and backward SDEs for optimal control of partially observed pat h-dependent stochastic systems, 2016","cited_arxiv_id":null,"evidence_quote":"Approximation of progressive controls by marked point processes with full support and bounded intensity, used to prove $V\\le V^R$."},{"cited_title":"Backward SDEs for optimal control of partially observed path-dependent stochastic s ystems: A control randomization approach","cited_arxiv_id":null,"evidence_quote":"Provides the partially-observed randomisation machinery, including Girsanov tilts, canonical extensions, and control transfer, used in the converse inequality."},{"cited_title":"A pseudo-m arkov property for controlled diﬀusion processes","cited_arxiv_id":null,"evidence_quote":"The canonical-space lemma that progressive and predictable processes coincide, used to construct the identifying control process."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Cited in Lemma 4.1 for the continuity property of the augmented Brownian filtration on which the existence of $\\lambda_s$ depends."},{"cited_title":"An Introduction to the Theory of Point Processes","cited_arxiv_id":null,"evidence_quote":"Girsanov theorem for Poisson random measures, used to define the tilted measures $\\hat P^\\nu$."}],"review_version":1}