{"id":"f299d173-64bb-45c2-b7e7-1403249d6197","arxiv_id":"2510.11158","paper_version":2,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Optimal policy in multi-dimensional ergodic singular control is characterized via Skorokhod reflection at free boundaries of an auxiliary Dynkin game, with two fully solved 2D inventory models.","lead":"The paper proves that in a class of multi-dimensional ergodic inventory control problems, the optimal policy keeps inventory between two time-varying barriers, and shows these barriers come from an auxiliary zero-sum stopping game. It then solves two two-dimensional inventory examples, including one with unobservable demand.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No significant objection identified: the conditional Theorems 2.2/2.3 are sound, and both case studies verify Hypothesis 2.2; only the delegated Skorokhod-reflection lemma merits an independent check.","rationale":"The paper's central claim—that the optimal policy in this class is Skorokhod reflection at Y-dependent free boundaries derived from a Dynkin game—is not contradicted by any internal inconsistency I found. Theorem 2.1 is a standard verification argument; Theorems 2.2 and 2.3 are transparently conditional on Hypothesis 2.2 and the extra regularity assumptions. The two case studies go to considerable length to verify those hypotheses in genuinely two-dimensional settings, including the degenerate partially-observed model. The only potentially fragile spot is the existence and minimality of the reflection control when the boundaries are only right/left-continuous (Section 3.1), which is delegated to earlier work. That is a normal citation practice rather than a fatal gap, and the provided construction is plausible. The reader's weakest assumption (Hypothesis 2.2) is exactly the conditional core; I agree that it is the load-bearing condition, but since it is verified in the examples and the general theorems are stated conditionally, the verdict should remain unchanged.","tokens_in":50660,"tokens_out":19262,"duration_ms":167141,"concrete_test":"Compare the construction in Lemma 3.13 (and the analogous omitted proof in Section 3.2) with [35, Section 4.3] for the specific case of a± monotone with jumps and a diffusion factor Y; simulate the reflection scheme on a one-dimensional test case (e.g., a± step functions) to confirm the process stays in [a+(Y),a-(Y)] and that the jump minimality condition in (3.65) is satisfied pathwise. If the scheme fails, the optimality claim in Theorem 3.15/3.17 would need a revised reflection construction.","verdict_should_be":"UNCHANGED","load_bearing_attack":"After a full read, I do not find a load-bearing flaw. The central claim is conditional: given Hypothesis 2.2 and the stated regularity, Theorems 2.2/2.3 construct (V,λ) and an optimal Skorokhod-reflection control. The proof of Theorem 2.1 is a standard mollifier/Itô argument and the inequalities are consistent. The two case studies verify Hypothesis 2.2: Section 3.1 establishes U∈C^1∩C^2(C) via hypoellipticity and Assumption 3.2, checks (2.23)-(2.24) via Lemma 3.4(ii), and verifies (2.35) by direct computation; Section 3.2 obtains U∈W^{2,∞}_{loc} from [39] and Lipschitz boundaries. The weakest point—existence of the reflection control with discontinuous state-dependent boundaries in Lemma 3.13—is sketched and delegated to [35,34]. This is a normal reliance on prior results, and the construction given is plausible, but it is the only step I would want to see checked independently.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper characterizes optimal policies for a class of multi-dimensional ergodic singular stochastic control problems where a one-dimensional linearly controlled process is modulated by a multi-dimensional uncontrolled factor. The main contribution is a general verification framework: given a solution (V, λ) to an auxiliary PDE with gradient constraints, a Skorokhod-reflection control at Y-dependent free boundaries is optimal (Theorem 2.1). The construction of (V, λ) is then reduced to the analysis of an auxiliary Dynkin game: under Hypothesis 2.2, the pseudo-potential V is the integral of the Dynkin game's value U, and λ is given explicitly (Theorems 2.2 and 2.3). Two inventory problems are solved: a partially observable model with a two-state Markov chain, where the separated problem yields a degenerate diffusion and Theorem 2.3 applies after proving hypoellipticity, smooth fit, and verifying the key identity (2.35); and a fully observable model with an Ornstein-Uhlenbeck factor, where uniform ellipticity gives W^{2,∞}_{loc} regularity of U and Theorem 2.2 applies with Lipschitz boundaries.","tokens_in":50931,"tokens_out":1326,"duration_ms":14338,"significance":"If the claims hold, the paper provides the first systematic connection between multi-dimensional ergodic singular stochastic control and Dynkin games, yielding explicit optimal policies in two genuinely two-dimensional settings. The general theorems are well-structured and the case studies are substantial. The verification theorems are stated as conditional results with checkable hypotheses, which is appropriate for a first paper in a new direction. The paper also contains useful methodological developments: the hypoellipticity argument for the filter dynamics in Section 3.1 and the Lipschitz regularity of free boundaries in Section 3.2. The main load-bearing steps are argued carefully, with the only notable delegation being the Skorokhod reflection lemma (Lemma 3.13), which is plausible and supported by prior work. On balance, the result is significant and the technical execution appears sound.","major_comments":[{"comment":"The existence of the reflection control with state-dependent, discontinuous boundaries a±(Π) is the key existence step for optimality in Section 3.1, but the proof is delegated to '[35, Section 4.3]' and '[34, Section 6.1]' with a sketch. In particular, Lemma 3.13 states that a solution to (3.65) exists and is admissible, but the verification of admissibility (the limit in (2.2)) is only one sentence. Since Theorem 3.15 relies on this lemma for the existence of the optimal control, I would like to see a more self-contained argument or a precise statement of which theorem in [35,34] applies to the present discontinuous-boundary setting. This is a normal reliance on prior results, but it is the one step I would want checked independently.","section":"Lemma 3.13 / Skorokhod reflection problem"},{"comment":"The identity (2.35) is load-bearing: it replaces the direct computation of L V that was possible in Theorem 2.2. In Theorem 3.15 it is verified by direct computation for the specific model, which is fine. However, the general Theorem 2.3 is stated with (2.35) as an assumption, and the paper does not discuss when (2.35) is expected to hold beyond the example. This is not an error, but the theorem would be more useful if it stated a sufficient condition (e.g., V∈C^2( C) plus the free-boundary equation for U in C) that implies (2.35).","section":"Theorem 2.3, Eq. (2.35)"},{"comment":"The inequalities (2.23)-(2.24) are assumed globally in x≥a−(y) and x≤a+(y), respectively. In the two case studies they are verified via the monotonicity of c' and the explicit bounds on a± from Lemma 3.4(ii). In the general statement, however, the hypotheses are stated as assumptions without a discussion of when they are natural or automatically satisfied. This is acceptable for a conditional theorem, but it would be helpful to mention that these inequalities are exactly what is needed to turn the free-boundary problem for U into the variational inequality for V.","section":"Hypothesis 2.2(II)"}],"minor_comments":[{"comment":"There are typographical errors and spacing issues (e.g., 'F∞', 'dP⊗dt' without spaces, missing parentheses in several displayed equations). For example, in the abstract 'F∞' appears in the definition of admissible controls, and in Eq. (3.29) the parentheses around 'sup' are missing. These should be corrected.","section":"Throughout"},{"comment":"In the proof of Lipschitz continuity of a+, the representation (3.80) is used to compute the derivative of aε+. It is not fully justified that (aε+(y), y) belongs to C for all ε>0 and that the implicit function theorem applies; this is briefly stated. A reference to a similar argument would help.","section":"Section 3.2, proof of Lemma 3.16(iv)"},{"comment":"The remark explains why the argument does not work directly for the DPE. It is useful but would benefit from a more explicit explanation of the distinction between λ(y) and the constant value λ⋆.","section":"Remark 2.4"}],"recommendation":"minor_revision","confidential_remarks":"The paper is strong and within the scope of the journal. The central claim is conditional on clear hypotheses, and the two case studies check those hypotheses in nontrivial settings. The main technical gap I identified is Lemma 3.13, which is delegated; I would encourage the editor to ask the authors to make that proof more explicit. This is not a reason to reject, but it is worth tightening."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a serious, well-built paper. The headline result—linking multi-dimensional ergodic singular control to an auxiliary Dynkin game, with the optimal policy as Skorokhod reflection at Y-dependent free boundaries—is real. The one-dimensional versions of this connection are classical (Karatzas, Taksar); carrying it to a factor-modulated multi-dimensional setting is not a trivial extension, because the value-function derivative no longer lives in a scalar ODE. The construction of the pseudo-potential V as the integral of the game value U, and λ as a boundary term, is clean and the heuristic is explained honestly.\n\nThe two case studies are the real payoff. The partial-observation inventory problem is the more demanding one: filtering reduces it to a degenerate two-dimensional diffusion, and the authors need hypoellipticity plus a custom coordinate change to get regularity; the free-boundary verification is genuinely technical. The full-observation problem with mean-reverting demand is more standard but still requires Lipschitz free boundaries, which they get via an Arzelà–Ascoli argument. Both examples verify Hypothesis 2.2, which is the load-bearing condition for the general theorems. That condition is stated honestly as a hypothesis; it is not disguised. The paper does not oversell its generality.\n\nSoft spots, in proportion: the general theorems are conditional on Hypothesis 2.2, so the 'general characterization' is really a template plus two verifications. That's still a real contribution, but readers should not expect a theorem covering arbitrary ergodic multi-dimensional problems. Lemma 3.13 (existence of the Skorokhod reflection with state-dependent barriers) is delegated to 'similar techniques' as in Federico–Pham and Federico–Ferrari–Rodosthenous. That is a normal reliance on prior work, but it is the one piece a referee should check independently, because the boundaries here are discontinuous (right/left-continuous) and the jump-reflection mechanism is subtle. The proof of Theorem 2.3 is sketched as 'analogous'—acceptable, but for a journal version I would ask for the details of identity (2.35) in the degenerate case, since it is verified explicitly in Section 3.1.4, not in the general statement.\n\nBottom line: no circularity, no hidden fitted claims. The paper is for stochastic-control and inventory-theory specialists. It deserves a serious referee and, after checking the delegated reflection lemma, should be publishable in a good journal.","headline":"Genuinely new multi-dimensional ergodic singular control / Dynkin game connection; both applications check out, with the usual rubber-meets-road gaps in delegated lemmas.","tokens_in":51388,"tokens_out":1995,"would_cite":true,"duration_ms":18703,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["93E03","60G40","49L20","35R35","90B05"],"pacs":[],"model":"deepseek-v4-flash","headline":"For a class of ergodic singular stochastic control problems with a one-dimensional controlled state and a multi-dimensional uncontrolled factor, the optimal policy is a Skorokhod reflection at factor-dependent boundaries obtained from an au","keywords":["ergodic singular stochastic control","Dynkin game","free boundary","Skorokhod reflection","verification theorem","inventory management","partial observation","value profile"],"falsifier":"Compute, for either Section 3 example, the long-run average cost of the Skorokhod reflection policy at the numerical solution of the free-boundary problem, and compare with a policy that occasionally lets X drift inside the continuation region before reflecting; if any such policy achieves strictly lower average cost, the characterization in Theorems 2.2 and 2.3 fails. A cheaper check: exhibit coefficients satisfying Assumption 2.1 for which the inequalities (2.23)-(2.24) fail at the boundary, since then Hypothesis 2.2 cannot hold.","tokens_in":50543,"feed_emoji":"📦","tokens_out":5688,"duration_ms":51781,"temperature":0.7,"pith_summary":"This paper proves that for a class of long-run-average (ergodic) singular stochastic control problems, where a one-dimensional controlled process is modulated by an uncontrolled multi-dimensional factor Y, the optimal policy has barrier form: reflect the controlled state at two Y-dependent boundaries a+(Y) and a-(Y). The value and boundaries come not from the ergodic Bellman equation but from an auxiliary zero-sum optimal stopping game, whose value function U serves as the derivative of a pseudo-potential V(x,y)=∫U(x',y)dx'. The authors give verification theorems showing that if U solves a two-obstacle free-boundary problem with the right monotonicity inequalities, then the reflection policy is optimal; they then fully solve two genuinely two-dimensional inventory problems, one with partially observable mean-reversion level and one with fully observable mean-reverting level.","feed_headline":"Optimal control emerges from a two-player stopping game","feed_subtitle":"In two solved inventory models, the policy is to keep stock between belief-dependent barriers derived from the game","key_machinery":"The load-bearing object is the auxiliary Dynkin game: a zero-sum game of optimal stopping in which one player chooses a stopping time and the other chooses a stopping time, with payoff built from the x-derivative of the running cost and the constants K±; its value U(x,y) plays the role of the derivative of a pseudo-potential. Free boundaries a+(y)<a-(y) split the state into stopping regions and the continuation region, and the optimal singular control is exactly Skorokhod reflection at these Y-dependent barriers. The verification rests on Hypothesis 2.2: U must solve the free-boundary problem classically and the inequalities (2.23)-(2.24) must hold outside the continuation region, supplying","core_discovery":"The central construction is the auxiliary Dynkin game defined in (2.22), whose value function U is used to build a pseudo-potential V(x,y)=∫_α^x U(x',y)dx' and a value profile λ(y) via (2.27). Theorems 2.2 and 2.3 show that, if U is a classical solution of the free-boundary problem (2.25) with boundaries a±(y) satisfying inequalities (2.23)-(2.24), then (V,λ) solves the auxiliary PDE (2.7), and the control that reflects X at a±(Y) is optimal. In the two inventory case studies the authors verify this hypothesis: in the partial-observation example, U∈C^1 is established via hypoellipticity and boundary regularity despite the degenerate parabolic structure; in the observable mean-reverting examp","pith_inferences":["Beyond the paper: the same construction suggests a computational route—solve a two-obstacle optimal stopping problem rather than the ergodic variational inequality; the boundaries of the stopping set are then ready-made control barriers.","Beyond the paper: in the degenerate limit where Y is constant, formula (2.40) recovers the classical smooth-fit relation c(a-)+K-b = c(a+)-K+b, so the framework contains the standard one-dimensional barrier solution as a special case.","Beyond the paper: a natural stress test is to perturb the two examples by making Y mean-reverting with coefficients that violate condition b>δ or the parameter condition (3.29); the conjecture is that reflection remains optimal but boundary regularity degrades, which would test how much of the regularity machinery is truly needed.","Beyond the paper: the value-profile viewpoint suggests that for non-ergodic factors the correct solution concept is an initial-condition-dependent value rather than a single constant, which may matter for mean-field game analogues."],"forward_implications":["If Hypothesis 2.2 holds, the optimal policy is fully characterized by the two boundaries a±(Y), with no need to solve the gradient-constrained ergodic Bellman equation.","The value of the ergodic control problem is the long-run time average of the value profile λ(Y_t); when Y is ergodic, this value is a constant independent of the initial state.","The two inventory models are solved completely: in the partial-observation case, the optimal policy reflects the inventory at belief-dependent boundaries driven by the filter; in the full-observation case, at Lipschitz-continuous boundaries driven by the mean-reversion level.","The representation remains valid even when the factor Y is not recurrent, with the value depending on the initial factor level.","Degeneracy of the state process in the partial-information example is overcome by showing the generator is hypoelliptic, so U is C^1 and the verification theorem still applies."],"fun_headline_variants":["Control policy reduced to a two-player game","Game-derived reflection barriers solve inventory control","First link between singular control and optimal stopping","Two inventory models solved via a Dynkin game","Optimal policy via game-derived free boundaries"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The whole construction depends on Hypothesis 2.2: the value U of the auxiliary Dynkin game must be a sufficiently regular classical solution of the two-obstacle free-boundary problem, with boundaries a± satisfying inequalities (2.23)-(2.24); if those fail, the constructed pseudo-potential need not satisfy the verification equation and the reflection policy may be suboptimal.","fun_headline_variants_meta":{"raw":{"variants":["Control policy reduced to a two-player game","Game-derived reflection barriers solve inventory control","First link between singular control and optimal stopping","Two inventory models solved via a Dynkin game","Optimal policy via game-derived free boundaries"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001009,"raw_usage":{"total_tokens":4101,"prompt_tokens":743,"completion_tokens":3358,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":487,"completion_tokens_details":{"reasoning_tokens":3292}},"tokens_in":487,"tokens_out":3358,"duration_ms":21631,"temperature":1.0,"reasoning_tokens":3292,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T10:11:18.753677+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute, for either Section 3 example, the long-run average cost of the Skorokhod reflection policy at the numerical solution of the free-boundary problem, and compare with a policy that occasionally lets X drift inside the continuation region before reflecting; if any such policy achieves strictly lower average cost, the characterization in Theorems 2.2 and 2.3 fails. A cheaper check: exhibit coefficients satisfying Assumption 2.1 for which the inequalities (2.23)-(2.24) fail at the boundary, since then Hypothesis 2.2 cannot hold.","supporting_citations":[],"review_version":1}