{"id":"f4d2d401-dc6f-4c4e-add1-ee1c584b9a8d","arxiv_id":"2412.17257","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A two-point communicated distribution achieves asymptotically optimal two-stage risk-averse DRO decisions under scaling of budgets and demand moments.","lead":"This paper proposes a decentralized scheme in which a forecasting team sends the operations team a simple two-point distribution, and proves that the resulting first-stage decisions become optimal for two-stage distributionally robust optimization as the problem scale grows. The result matters because two-point distributions make the operations problem a tractable linear program while keeping an asymptotic performance guarantee.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The asymptotic guarantee is confined to the Taylor-law scaling s<2; at s=2 the error terms are order-k and the ratio proof degenerates, so the central claim only covers vanishing relative uncertainty.","rationale":"I read the paper in good faith as proving asymptotic optimality of the two-point mechanism under an explicit Taylor-law scaling. The proofs are detailed, the sandwich bounds are internally coherent, and the numerical studies support the moment-based claims. I do not see an internal contradiction in the stated theorems. The most load-bearing limitation is the scaling regime itself: the proof needs s<2 so that uncertainty-induced error terms are sublinear in k, allowing the linear negative term V0 to dominate. At s=2, the error terms are the same order as V0 and the ratio bound degenerates. This is not merely a technical annoyance because s=2 describes the natural multiplicative-scaling model where demand is scaled by k while its distribution shape is preserved. The reader's weakest assumption already points at the scaling scheme, but I would sharpen it: the relevant boundary is s=2, not s<1, and the practical consequence is that the advertised optimality is an asymptotic statement about vanishing relative dispersion. This does not overturn the theorems, but it does mean the conclusion should not be read as a general resolution of two-stage DRO tractability for arbitrary large-scale problems. The reader's CONDITIONAL verdict remains appropriate.","tokens_in":47940,"tokens_out":20988,"duration_ms":220494,"concrete_test":"For the one-dimensional instance N=1, A=H=1, p=2c, b=2cμ, with the moment ambiguity set (5) and CVaR risk measure, compute the exact worst-case objective OPT(b(k),θ(k)) and the objective of the two-point mechanism's lower-level solution from (10) with (κ,η)=(1,1), under scaling (6) for s=2 and for s=1.5, over k=10^2,...,10^6. If the s=2 ratio saturates below 1 while the s=1.5 ratio approaches 1, the s<2 restriction is essential; if the s=2 ratio also approaches 1, the proof's failure is only technical and the concern would not land.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Theorems 1 and 2 are proven only under the scaling schemes (6) and (16) with s ∈ [1,2). The sandwich argument in Appendix A.1 and A.2 requires the positive error terms from Propositions 3, 4, 8, and 9 to be o(k) relative to V0(b(k),θ(k)) = kV0(b(1),θ(1)); this is exactly the condition s<2. At s=2, k^{s/2-1}=1, so the lower bound on Obj/OPT becomes 1 + C with C ≤ 0, and convergence to 1 is not established. The excluded case s=2 is not pathological: it is the natural scaling for demand of the form d(k) = k·d0, where the coefficient of variation is constant as the market grows, and for any common multiplicative shock. Thus the headline claim that the two-point mechanism yields asymptotically optimal solutions as problem magnitude increases is really a statement about uncertainty becoming relatively small compared to the mean, not about large scale per se. The theorem is internally consistent because s<2 is an explicit condition, but the practical reach of the advertised guarantee is exactly the vanishing-relative-uncertainty regime.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a decentralized, bilevel framework for two-stage risk-averse distributionally robust optimization. The forecasting team acts as leader and communicates a distribution to the operations team, which then solves a tractable two-stage stochastic program. For moment-based ambiguity sets (Assumption 2) and Wasserstein ambiguity sets (Assumption 4), the paper constructs a two-point mechanism M_{ς,τ} in (7) and (17) and proves, under the scaling schemes (6) and (16) with s∈[1,2), that the induced first-stage decision x^{(k)}_{ς,τ} satisfies Obj(x^{(k)}_{ς,τ},b(k),θ(k))/OPT(b(k),θ(k))→1. The proof uses a sandwich argument: a linear-program lower bound V0 for OPT and a truncated-linear-decision-rule upper bound for Obj, with error terms that are o(k). Numerical experiments compare the method with TLDR approximations and SAA, including a real-data case study.","tokens_in":48228,"tokens_out":14481,"duration_ms":141129,"significance":"If the main theorems are correct, the paper gives a genuinely simple and computationally tractable mechanism that is asymptotically optimal for a class of otherwise intractable two-stage DRO problems. The proof structure is explicit and self-contained: the lower bound uses only the expectation mechanism, the upper bound uses a TLDR policy, and the ratio argument is based on explicit bounds in Propositions 3, 4, 8, and 9. No fitted constants enter the asymptotic guarantee; any feasible (ς,τ) works. The numerical work is substantial, including a real sales-data experiment and comparisons against SAA and TLDR. The main reservation is that the proved regime is one where relative uncertainty vanishes as k grows, and the proof of the key upper bound contains a local but correctable error; these issues do not destroy the central idea but require revision.","major_comments":[{"comment":"The proof claims that an optimal (v,U) in problem (13a) \"must satisfy v≤U d_h, otherwise we can set v to be U d_h and then (U d_h,U) is a feasible solution with smaller objective.\" This is not correct as stated: decreasing v can only increase the pointwise loss -p^T min{v,U d} (or leave it unchanged on the support {d_l,d_h}), so by monotonicity of the risk measure the objective cannot decrease. The needed conclusion is nevertheless salvageable: for any feasible (v,U), replacing v_i by min{v_i, U_i d_h} leaves the policy unchanged on the support of M_{ς,τ}(θ(k)), so there exists an optimal solution with v≤U d_h; the subsequent estimate (v-kUμ)_+≤U(d_h-kμ) then holds for that representative. Please revise the proof to use this existence argument rather than the false necessity claim. The same issue appears in the Wasserstein analogue, Proposition 8.","section":"Section 4, equation (6) and Section 1, scaling discussion"},{"comment":"Theorems 1 and 2 are proved only for s∈[1,2). Under this scaling, the coefficient of variation of each marginal demand is of order k^{s/2-1}, which tends to zero; the asymptotic regime is therefore one of vanishing relative uncertainty, not merely of increasing problem scale with \"inherent uncertainty remain[ing] pronounced\" as the Introduction states. The excluded case s=2 is the natural scaling for a common multiplicative shock, where the coefficient of variation is constant; at s=2 the sandwich bound degenerates because the correction term k^{s/2-1} does not vanish. The paper should state this limitation explicitly in the abstract, introduction, and conclusion, and either extend the analysis to s=2 or clearly delineate that the advertised optimality applies to the regime of shrinking relative dispersion. This is a load-bearing point because it concerns the practical interpretation of the headline asymptotic-optimality claim.","section":"Section 1 and Section 4"},{"comment":"The theorem's proof first establishes the same asymptotic ratio for the expectation mechanism M0, the Dirac distribution at kμ, which is the degenerate member of the family (τ=0 or ς=0). Thus the claim that \"a two-point distribution suffices\" is not the strongest possible statement: a one-point (mean) mechanism also achieves the same limit. The paper should acknowledge that the theoretical contribution is not that two points are necessary, and should clarify that the value of the two-point mechanism lies in finite-sample/finite-k performance (as suggested by the numerical experiments) rather than in asymptotic optimality per se.","section":"Section 4.1"}],"minor_comments":[{"comment":"The phrase \"the set of all joint contributions of random vectors\" should be \"the set of all couplings (joint distributions) with the given marginals.\"","section":"Assumption 4"},{"comment":"In the first line, \"As P(k)∈A(θ(k))\" should refer to the nominal distribution \\hat P(k) used to define the Wasserstein ball; please correct the notation.","section":"Appendix A.2, proof of Proposition 7"},{"comment":"The phrase \"at an appropriate rate\" in the abstract and the Introduction's claim that the scaling reflects situations where \"inherent uncertainty remains pronounced\" should be reconciled with the fact that s<2 makes relative dispersion vanish; this is part of the major comment above, but a precise wording fix in the abstract is also needed.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":"The central idea is appealing and the asymptotic framework is a plausible way to bypass intractability. The paper is not ready in its present form because the proof of a key upper-bound proposition contains a false optimality claim, even though a WLOG repair exists, and because the advertised claim of asymptotic optimality as the problem magnitude grows is materially narrower than the abstract suggests: it covers only the vanishing-relative-uncertainty regime s<2. With those revisions, the paper could be a solid contribution. I would also encourage the authors to state plainly in the conclusion that the one-point mean mechanism is already asymptotically optimal, so that the contribution of the two-point mechanism is clearly positioned as a finite-sample/finite-k engineering choice."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the two-point mechanism result is real, and the proof is worth reading. The authors show that for two-stage risk-averse DRO with moment or Wasserstein ambiguity, a two-point distribution suffices for asymptotic optimality under scaling (6)/(16). That's a clean structural result: the forecasting team only needs to communicate two scenarios, and the operations problem becomes a linear program. The sandwich proof in Appendix A is credible, built on standard ingredients (Gallego's bound, standard risk coefficient, TLDR policies, Wasserstein outer approximation). The numerical work is extensive, with real sales data showing out-of-sample gains over SAA and competitive performance vs TLDR at roughly 100x speedup. I believe the main theorem.\n\nThe soft spots are in the scope of the guarantee, not in the math. The scaling requires s in [1,2), so variance grows sublinearly relative to mean squared; the coefficient of variation shrinks as k grows. At s=2 the proof collapses, and s=2 is the natural regime for constant CV or common multiplicative shocks. So 'large scale' in this paper really means 'large scale with relatively less uncertainty,' not large scale per se. The abstract's 'appropriate rate' is doing a lot of work. The conclusion's claim that this 'effectively resolves' the intractability of two-stage DRO is an overstatement; it resolves a specific asymptotic regime. The Wasserstein theorem is proven only by sketch and has no numerical validation. The code claim is not checkable from the text.\n\nNone of these are fatal. The theorems are internally consistent, the assumptions are explicit, and the numerical studies support the practical claims in the moment-based case. The paper is honest about the s<2 condition; it just doesn't emphasize how restrictive that is. My own verdict: conditional accept if the authors tighten the language about what the asymptotic result delivers and ideally add a numerical check for the Wasserstein case. The right reader is someone working on DRO approximations or asymptotic analysis in OM; it will get cited. I'd take it to peer review, and I'd expect it to survive with revisions.","headline":"Two-point mechanism is a genuine asymptotic simplification for two-stage DRO, but the guarantee lives in a narrow vanishing-uncertainty regime and the conclusion overstates its reach.","tokens_in":48698,"tokens_out":1801,"would_cite":true,"duration_ms":17088,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["90C15","90C47"],"pacs":[],"model":"deepseek-v4-flash","headline":"A forecasting team that sends only two demand scenarios can make the operations team's decision asymptotically optimal for two-stage distributionally robust problems.","keywords":["distributionally robust optimization","two-stage stochastic programming","bilevel optimization","asymptotic optimality","moment ambiguity set","Wasserstein ambiguity set","two-point distribution","risk measures"],"falsifier":"Set s = 2 in the scaling schemes (6) or (16) and compute the ratio Obj($x^{{(k)}}$_{ς,τ}, b(k), θ(k)) / OPT(b(k), θ(k)) for growing k; the paper's proof leaves the correction term $k^{{s/2−1}}$ = 1 non-vanishing, so observing a ratio that fails to approach one would mark the boundary of the theorem.","tokens_in":47731,"feed_emoji":"📈","tokens_out":5744,"duration_ms":51647,"temperature":0.7,"pith_summary":"This paper proposes a decentralized way to solve two-stage risk-averse distributionally robust optimization problems by splitting them into a forecasting team and an operations team. The forecasting team is treated as a leader who chooses which probability distribution to communicate, and the operations team reacts with a two-stage stochastic decision. The paper's central finding is that the best mechanism is a two-point distribution—low and high demand with carefully chosen locations and probabilities—and that this simple message is asymptotically optimal: as the scale of demand and budget grows, the cost ratio between the induced solution and the true optimum converges to one. This matters because the original problem is generally intractable, while the induced operations problem is a linear program. The same two-point structure works for both moment-based and Wasserstein ambiguity sets, and numerical experiments on assemble-to-order and real sales data show it matches or beats standard approximations.","feed_headline":"Two-point forecast is enough for asymptotically optimal decisions","feed_subtitle":"Forecasting sends just low and high demand; operations solves one LP and the cost ratio reaches one as scale grows.","key_machinery":"The load-bearing object is the two-point mechanism M_{ς,τ}(θ(k)) = (1−τ)δ_{d_l} + τδ_{d_h}, with d_l = kμ − √(τ/(1−τ)) $k^{{s/2}}$ ς and d_h = kμ + √((1−τ)/τ) $k^{{s/2}}$ ς. This distribution lives inside the ambiguity set and reduces the infinite-dimensional second-stage recourse problem to a one-stage linear program. The proof sandwiches OPT between the expectation-mechanism value V_0, a linear program with positive homogeneity, and an upper bound built from a truncated linear decision rule, using the standard risk coefficient α to control worst-case risk; the linear negative term from V_0 dominates the sublinear corrections and forces the cost ratio to one.","core_discovery":"Under the paper's Assumptions 1–4 and the scaling schemes (6) and (16), the bilevel mechanism design problem (4) has an optimal mechanism M_{ς,τ} that outputs a two-point distribution. Consequently, the first-stage decision $x^{{(k)}}$_{ς,τ} induced at the lower level satisfies lim_{k→∞} Obj($x^{{(k)}}$_{ς,τ}, b(k), θ(k)) / OPT(b(k), θ(k)) = 1. This establishes that, in the large-scale regime obeying Taylor's law with exponent s ∈ [1,2), a two-point distribution carries all the information needed for asymptotically optimal risk-averse decisions, both when ambiguity is described by marginal moments and when it is described by a 2-Wasserstein ball.","pith_inferences":["Editorial inference: because the two-point mechanism is optimal only in the limit, finite-scale users should treat the parameters (ς, τ) as tunable; cross-validation is the natural way to select them, as the paper demonstrates.","Editorial inference: the result suggests a separation principle for decentralized operations: forecasters need not report a full distribution, only a scenario pair encoding mean and spread, and this may extend to other nested stochastic programs with similar scaling.","Editorial inference: the s = 2 boundary is the natural next test; if the ratio still converges there, contrary to the proof, the practical regime would widen beyond the Taylor-law range studied here."],"forward_implications":["Two-stage distributionally robust optimization can be replaced at the operational level by solving a single linear program without losing asymptotic optimality.","In the comparison case where truncated linear decision rules are applicable, the decentralized two-point mechanism stays within a few percent of the TLDR value while running about 100 times faster at 500 products.","With cross-validated tuning of (ς, τ), the mechanism outperforms sample-average approximation out-of-sample in risk-averse scenarios, with the largest advantage when training data are scarce.","Both moment-based and data-driven Wasserstein ambiguity settings admit the same structural mechanism, so a single implementation covers two common DRO formulations."],"supporting_citations":[{"why":"Supplies the mean-variance bound on E[max{0, ζ}] used in Lemma 2 to control the risk of truncated linear decision rules.","marker":"Gallego (1992)"},{"why":"Supplies the interchangeability principle used to move the risk functional inside the minimization over recourse decisions.","marker":"Shapiro (2017)"},{"why":"Establishes the Wasserstein-to-moment outer approximation used in Proposition 5 to connect Wasserstein ambiguity with moment information.","marker":"Nguyen et al. (2021)"},{"why":"Documents the computational intractability of two-stage minimax stochastic linear programs that the paper targets.","marker":"Bertsimas et al. (2010)"},{"why":"Provides the closedness and coherence framework for risk measures and the standard risk coefficient used in the upper-bound arguments.","marker":"Rockafellar (2007)"},{"why":"Introduces the truncated linear decision rule that serves as the baseline approximation method in the numerical comparisons.","marker":"See and Sim (2010)"},{"why":"Motivates the scaling regime where variance grows as the s-th power of the mean with s in [1,2).","marker":"Taylor (1961)"}],"fun_headline_variants":["Forecast two points, decide optimally at scale","Two-point forecast achieves asymptotic optimality","Asymptotically optimal decisions from a simple two-point forecast","Two-point forecast is all you need for asymptotic optimality"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The proof holds only in the growth regime where budget and mean demand grow linearly while demand fluctuations and ambiguity radii grow as $k^{{s/2}}$ with s < 2; if variance grows as fast as the square of the mean, the correction terms no longer vanish and the optimality ratio is not established.","fun_headline_variants_meta":{"raw":{"variants":["Forecast two points, decide optimally at scale","Two-point forecast achieves asymptotic optimality","Asymptotically optimal decisions from a simple two-point forecast","Two-point forecast is all you need for asymptotic optimality"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000871,"raw_usage":{"total_tokens":3782,"prompt_tokens":966,"completion_tokens":2816,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":582,"completion_tokens_details":{"reasoning_tokens":2755}},"tokens_in":582,"tokens_out":2816,"duration_ms":19869,"temperature":1.0,"reasoning_tokens":2755,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T05:39:27.115978+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Set s = 2 in the scaling schemes (6) or (16) and compute the ratio Obj($x^{{(k)}}$_{ς,τ}, b(k), θ(k)) / OPT(b(k), θ(k)) for growing k; the paper's proof leaves the correction term $k^{{s/2−1}}$ = 1 non-vanishing, so observing a ratio that fails to approach one would mark the boundary of the theorem.","supporting_citations":[{"cited_title":"Opera- tions Research Letters 45(4):377–381","cited_arxiv_id":null,"evidence_quote":"Supplies the interchangeability principle used to move the risk functional inside the minimization over recourse decisions."}],"review_version":1}