{"id":"09245aff-51af-4c20-a007-16da8bb8a407","arxiv_id":"2507.17073","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A constant-cost, asymptotically normal estimator of Curie-Weiss interaction parameters is built from large-population moment approximations, with consistency in the double limit n, N to infinity.","lead":"This paper constructs a fast estimator for the coupling strengths of a multi-group Curie-Weiss voting model, replacing the expensive partition function with simple asymptotic moment formulas. A generalist reader might care because the estimator costs the same for any population size and supports statistically grounded voting weights in two-tier systems.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 14.4's displayed rate function in Eq. (44) is not the contraction of the empirical LDP through the estimator; as written its unique minimum is at m(tilde-beta)^4, not tilde-beta, contradicting the estimator's own consistency.","rationale":"I read the paper in good faith: the constant-cost estimator construction is coherent, and the proofs of consistency, CLT, and the exponential large-deviation bound are mostly detailed and repairable. The reader's weakest assumption concerned the unquantified constants D_high and D_low in condition (10). That is a practical obstruction to running Algorithms 24 and 25, but it does not by itself make the central theorem false. The more load-bearing defect is in the stated large-deviation rate function itself. Since the estimator is a contraction of T/N^2 through (vartheta-infinity_N)^{-1}, the contraction principle gives a rate function evaluated at vartheta-infinity_N(y); Eq. (44) instead evaluates the entropy function at points x with x^2 = y. In the low-temperature regime the printed rate function's minimizer is m(tilde-beta_N)^4, which contradicts the consistency proved in Theorem 14.1. This is an internal inconsistency in the central claim, not merely a missing constant. It is fixable by replacing (44) with the correct contraction formula, so I would not reject the paper outright; the appropriate disposition remains CONDITIONAL, matching the reader's verdict. My agreement is only partial because the reader did not flag this contraction error, instead concentrating on the unquantified constants.","tokens_in":44815,"tokens_out":17609,"duration_ms":190073,"concrete_test":"Apply Theorem 49 literally: set f = (vartheta-infinity_N)^{-1} and compute J_correct(y) = inf{Lambda*_{(S/N)^2}(x) : f(x) = y}, equivalently Lambda*_{(S/N)^2}(vartheta-infinity_N(y)) on the estimating branches. Verify that J_correct(tilde-beta_N) = 0 and J_correct(m(tilde-beta_N)^4) > 0 for beta in I_l, and analogously for the high-temperature branch, while the printed Eq. (44) gives the opposite. A simple numerical check: fix N = 100, beta = 1.5, b2 = 1.1; tabulate both functions on a fine grid over [0,1] and locate their minima.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Under Proposition 57, for fixed N the statistic T/N^2, an average of n iid copies of (S/N)^2, obeys an LDP with rate n and rate function Lambda*_{(S/N)^2}, whose unique zero is at E(S/N)^2. By Remark 33, the estimator satisfies hat-beta-infinity_N = (vartheta-infinity_N)^{-1}(T/N^2). The contraction principle (Theorem 49) therefore gives the rate function J(y) = Lambda*_{(S/N)^2}(vartheta-infinity_N(y)) on the two branches of the estimator, with infinity on the undefined or gap region. Eq. (44) instead defines J(y) = inf{Lambda*_{(S/N)^2}(x) : x^2 = y}. On beta in I_l, vartheta-infinity_N(y) = m(y)^2, so the printed formula is minimized at sqrt(y) = m(tilde-beta_N)^2, i.e. y = m(tilde-beta_N)^4, whereas consistency (Theorem 14.1) forces the LDP minimizer to be y = tilde-beta_N. Thus the statement as printed is internally inconsistent: the displayed rate function's unique minimum is not the probability limit of the estimator. The high-temperature branch has the same defect, since minimizing over x^2 = y does not track x = 1/((1-y)N). This is fixable by replacing Eq. (44), but as it stands Theorem 14.4 does not state a valid contraction. It is also more decisive than the unquantified D_high and D_low: even with explicit constants, the published rate function is wrong.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a computationally cheap estimator for the coupling parameters of a multi-group Curie-Weiss model with no inter-group interactions, replacing the exact moment E_{\\beta,N}S^2 in the maximum-likelihood equation by large-population asymptotic approximations. The main results are Proposition 11 (exponentially decaying regime-misclassification probabilities), Theorem 14 (consistency to a bias-corrected value, asymptotic normality, and a large deviation principle), and Theorem 38 on asymptotically optimal two-tier voting weights. The central statistical claims are conditional on condition (10), which requires a gap between the high- and low-temperature intervals whose verification depends on constants D_high and D_low asserted in Proposition 22.","tokens_in":45149,"tokens_out":6079,"duration_ms":64937,"significance":"If the technical issues below are fixed, the paper would make a useful contribution: it provides a constant-cost estimator with standard asymptotic guarantees in a model of social/electoral interaction, and it connects those estimates to an explicit voting-weight application. The detailed saddle-point/Laplace proofs in Section 4, the self-contained derivation of the moment asymptotics in Propositions 20 and 22, and the large deviation analysis of S/N in Proposition 30 are genuine strengths. The voting-weight theorem is standard but its proof is helpfully re-derived from the moment asymptotics. However, because the statement of Proposition 22 is incorrect as written and because the rate function in Theorem 14(4) is not the correct contraction, the significance is currently conditional on substantial corrections.","major_comments":[{"comment":"The statement of Proposition 22 is false as written: for beta < 1 it asserts E_{\\beta,N} S^{2k} \\approx (1/(1-beta))^k N^k, but the proof itself derives the limit (2k-1)!! (1/(1-beta))^k N^k (see the simplified sum over r_1 after Eq. (20)). This is load-bearing, not cosmetic: in the proof of Theorem 14(3)(a), the variance computation V_{\\beta,N}(S^2/N) \\approx 3(1-\\tilde\\beta_N)^{-2} - (1-\\tilde\\beta_N)^{-2} uses the missing factor 3 for k=2, and with the formula as printed the variance would vanish asymptotically. The statement must be corrected and all downstream variance formulas checked accordingly.","section":"Proposition 22"},{"comment":"The rate function in Eq. (44) is not the contraction of the LDP for T/N^2 through the estimator. By Remark 33, \\hat\\beta_N^\\infty = (\\vartheta_N^\\infty)^{-1}(T/N^2), so the contraction principle (Theorem 49) gives J(y) = \\Lambda^*_{(S/N)^2}((\\vartheta_N^\\infty)^{-1}(y)) on the two branches of the estimator, with infinity on the gap. The printed expression J(y) = \\inf\\{\\Lambda^*_{(S/N)^2}(x) : x^2 = y\\} instead minimizes over x = \\pm\\sqrt y. For beta in I_l, (\\vartheta_N^\\infty)^{-1}(y) = m^{-1}(\\sqrt y), so the correct rate function has its unique zero at y = m(\\tilde\\beta_N)^2, whereas the function in (44) has its unique zero at y = m(\\tilde\\beta_N)^4 because \\Lambda^*_{(S/N)^2} vanishes at x = m(\\tilde\\beta_N)^2. This contradicts the consistency statement in Theorem 14(1) and invalidates the LDP as stated. Please replace (44) with the proper contraction and update Theorem 14(4) and Remark 34 accordingly.","section":"Theorem 14(4), Eq. (44)"},{"comment":"The estimator and the definition of the intervals I_h, I_l, J_h, J_l depend on constants D_high and D_low that Proposition 22 only asserts to exist pointwise for each beta; no uniform-in-beta bounds are proved, and no explicit or computable values are supplied. As a result, a user cannot verify condition (10), cannot choose b1 and b2, and cannot certify that the advertised constant-cost algorithm is actually runnable. This is a load-bearing gap in the paper's practical claim. The authors should either prove explicit uniform bounds or reformulate the theorems with a directly checkable condition on N, b1, and b2.","section":"Definition 7, Remark 8, Algorithms 24-25"},{"comment":"In the proof of Theorem 38, the moments of the limiting normal distribution N(0, 1/(1-\\beta_\\lambda)) are stated as m_k = k!! (1/(1-\\beta_\\lambda))^{k/2} for even k. The correct formula is (k-1)!! (1/(1-\\beta_\\lambda))^{k/2}; for k=2 the printed formula gives 2 instead of 1. Since this moment sequence is used to apply the method of moments (Theorem 61), the proof as written is invalid, even though the half-normal limit quoted in the theorem is standard and the result is likely correct. Please correct the moment formula and verify that the method-of-moments hypotheses hold with the corrected sequence.","section":"Theorem 38, proof"}],"minor_comments":[{"comment":"The symbol u is introduced as a possible value of the estimator, but the target space ([−∞,∞] ∪ {u})^M is not given a topology, and the LDP in Theorem 14(4) is stated on [−∞,∞]^M; the treatment of the u/gap region in the large deviation statement should be made explicit.","section":"Definition 10 / Notation 9"},{"comment":"The sentence 'I_beta has one minimum at m(beta) = 0 if beta \\le 1' is misleading; the minimum is at x = 0, not at the point m(beta), and the phrase should be reworded.","section":"Proposition 30"},{"comment":"Figure 1 is only referenced in Remark 26 and its axes and parameter values are not described; adding a caption with the values of N and the intervals used would improve reproducibility.","section":"Figure 1 / Remark 26"},{"comment":"Proposition 57 is stated for a single-group statistic, but it is later used for the multivariate statistic T; the passage from the univariate statement to the coordinate-wise application in Theorem 14 should be spelled out.","section":"Appendix, Proposition 57"}],"recommendation":"major_revision","confidential_remarks":"The paper relies heavily on the companion manuscript [2] for Proposition 44, Lemma 48, Lemma 56, and Proposition 57; if [2] is not available or not accepted, the proofs are incomplete. The issues with Proposition 22 and Eq. (44) are independent of the reader's report and can be verified directly from the displayed equations in the manuscript."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis paper proposes an estimator for the coupling parameters of a non-interacting multi-group Curie-Weiss model based on large-population asymptotic approximations of the moments, avoiding the exponential cost of the partition function. That is a genuinely useful idea, and the estimator is new relative to the authors' earlier maximum-likelihood paper [2]. The main statistical claims — consistency at fixed N, convergence of the bias to zero as N grows, asymptotic normality, and exponential concentration — are plausible, and the proofs are mostly detailed, with the moment asymptotics in Propositions 20 and 22 derived through careful Hubbard-Stratonovich and saddle-point arguments.\n\nThe paper has real problems, though. Proposition 22, which is central, is false as stated: it drops the (2k-1)!! factor that its own proof derives and that Theorem 14.3 needs (the factor 3 for k=2). The statement must be corrected. Second, the large deviation rate function in Theorem 14.4 is not the contraction of the empirical LDP through the estimator. The proof applies the contraction principle through the map x ↦ x^2, writing Eq. (44) as inf Λ*(x) subject to x^2 = y; the correct contraction through (ϑ∞_N)^{-1} would give J(y) = Λ*_{(S/N)^2}(ϑ∞_N(y)). As printed, the unique minimum of J is at m(tildeβ)^4, not tildeβ, which contradicts the consistency statement in the same theorem. That is a concrete internal inconsistency, and it is more serious than the unquantified constants D_high and D_low, which are a practical obstruction to running the algorithm but not a mathematical contradiction. Theorem 38's auxiliary proof also uses an incorrect moment formula for the normal limit, though the result is credited to Kirsch and that proof is secondary.\n\nNone of this seems unfixable. The moment expansion in Proposition 20 and the bias analysis around tildeβ_N are sound in spirit, and the estimator itself is a real contribution for social-science applications where the MLE is computationally prohibitive. But the printed version cannot be accepted with a central proposition and a main theorem statement that are wrong as written.\n\nThe paper is for statisticians and mathematical physicists interested in mean-field inverse problems, and for social scientists using Curie-Weiss voting models. It deserves a serious referee: the idea is good, and the errors are identifiable and repairable. I would send it to review with the expectation of major revision.","headline":"Useful new constant-cost estimator for Curie-Weiss couplings, but Proposition 22 and Theorem 14.4 are wrong as printed; both are fixable, yet the paper needs major revision before acceptance.","tokens_in":45683,"tokens_out":5603,"would_cite":false,"duration_ms":57022,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F10","82B20","60F05","91B12"],"pacs":[],"model":"deepseek-v4-flash","headline":"A constant-cost estimator recovers the coupling parameters of a multi-group Curie-Weiss model from sample vote margins, with consistency, asymptotic normality, and large-deviation guarantees.","keywords":["Curie-Weiss model","coupling parameters","large population approximation","maximum likelihood estimation","large deviations","asymptotic normality","two-tier voting systems","optimal weights"],"falsifier":"For a fixed $N$, try to instantiate Algorithm 24: condition (10) needs numerical values for $D_{\\mathrm{high}}$ and $D_{\\mathrm{low}}$, and the paper provides none; demonstrating that no such constants can be extracted from the proof would block the advertised constant-cost procedure. Alternatively, simulate the estimator for a true $\\beta$ just inside $I_h$ or $I_l$ with the critical band chosen by any candidate gap and test whether the frequency of `u` responses decays exponentially in $n$ as Proposition 11 predicts; a polynomial decay rate would refute the central claim.","tokens_in":44566,"feed_emoji":"🗳️","tokens_out":7959,"duration_ms":77487,"temperature":0.7,"pith_summary":"The paper tries to estimate the coupling parameters $\\beta_\\lambda$ of a multi-group Curie-Weiss model---the strengths of within-group interaction in a population of binary voters---without computing the partition function, whose cost grows exponentially with population size. Its solution replaces the exact moment condition of maximum likelihood with sharp large-population approximations, producing an estimator that can be evaluated at constant computational cost for any group size. For parameters away from the critical value $\\beta=1$, the estimator converges as the number of samples grows to a bias-corrected value $\\tilde{\\beta}_N$, which in turn approaches the true parameter as the group grows; it is asymptotically normal and obeys a large deviation principle with exponentially small error probabilities. The authors argue this makes the coupling parameters a practical measure of social cohesion and a usable input for assigning optimal weights in two-tier voting systems.","feed_headline":"Cheap estimator recovers social-cohesion parameters from votes","feed_subtitle":"Moment asymptotics replace the exponential partition function; the estimator is normal and has exponential tails.","key_machinery":"The load-bearing object is Proposition 22's asymptotic expansion of the even moments of the group margin: for $\\beta<1$, $E S^{2k}/N^k = (1/(1-\\beta))^k + O(1/\\sqrt{N})$ with an asserted constant $D_{\\mathrm{high}}$, and for $\\beta>1$, $E S^{2k}/N^{2k} = m(\\beta)^{2k} + O((\\ln N)^{3/2}/\\sqrt{N})$ with an asserted constant $D_{\\mathrm{low}}$. These expansions are derived by representing moments as ratios of Laplace integrals and applying saddle-point analysis, and they replace the partition function in the optimality equation $E S^2 = $ sample moment. The estimator is the inverse of the piecewise function $\\vartheta^\\infty_N$ built from $1/(1-\\beta)\\cdot(1/N)$ on the high-temperature side and $m(\\beta)^2$ on the low-temperature side, with condition (10) guaranteeing a positive gap between the two regimes.","core_discovery":"The central claim is Theorem 14: for each group with $\\beta_\\lambda \\in I_h \\cup I_l$ (the high- and low-temperature intervals separated by an excluded critical band), the estimator $\\hat{\\beta}^\\infty_N$ converges in probability to the deterministic bias-corrected target $\\tilde{\\beta}_N$ as the number of i.i.d. configurations $n$ tends to infinity; $\\tilde{\\beta}_N$ converges to the true $\\beta_\\lambda$ as the group size $N$ tends to infinity; $\\sqrt{n}(\\hat{\\beta}^\\infty_N - \\tilde{\\beta}_N)$ converges in distribution to a centered normal with diagonal covariance; and $\\hat{\\beta}^\\infty_N$ satisfies a large deviation principle whose rate function has its minimum at $\\tilde{\\beta}_N$, giving exponentially decaying probabilities of large deviation from $\\beta$. The estimator is defined by inverting the asymptotic approximations: in the high-temperature regime $T = N/(1-\\hat{\\beta}^\\infty_N)$ and in the low-temperature regime $T = m(\\hat{\\beta}^\\infty_N)^2 N^2$, where $m(\\beta)$ is the largest solution of $\\tanh(\\beta x)=x$; when the sample statistic falls in the middle interval $J_c$ the estimator returns the symbol `u` rather than a number.","pith_inferences":["A natural next step the paper leaves implicit is a two-stage practical protocol: run $\\hat{\\beta}^\\infty_N$, and when it returns `u`, either enlarge the sample or switch to the exact maximum likelihood estimator, since `u` is informative about proximity to the critical point.","Because the method only needs moment asymptotics, the same template should transfer to other mean-field families---e.g., block Ising models with known large-$N$ moment limits---yielding constant-cost estimators for their interaction matrices.","The optimal-weight asymptotics imply that in large populations, almost all council weight concentrates in strongly cohesive groups; a testable consequence is that in bodies like a confederal council, estimated optimal weights should order countries by estimated cohesion, not by population alone.","Theorem 14's large deviation principle is stated for fixed $N$ as $n$ grows; a double limit in both $n$ and $N$, with the bias $\\tilde{\\beta}_N-\\beta$ controlled, would give a single finite-sample error bound and is not derived in the paper."],"forward_implications":["The coupling parameter of each non-interacting group can be estimated at cost independent of group size, so the method scales to populations of millions once the gap condition (10) holds.","Confidence intervals follow from the asymptotic normality: the high-temperature variance tends to $2(1-\\beta)^2$ and the low-temperature variance tends to $0$ as $N$ grows.","The exponential bounds of Proposition 11 let the estimator double as a classifier: the probability of mistaking a high-temperature group for a low-temperature group decays exponentially in the sample size.","Plugging the estimated parameters into Theorem 38 yields an estimator $\\hat{w}$ for optimal council weights with the same consistency, normality, and large-deviation properties (Theorem 42).","For samples that fall in the critical band, no numeric estimate is reported; the paper's guarantee is only that such samples are exponentially rare for parameters away from $\\beta=1$."],"supporting_citations":[{"why":"Supplies the maximum-likelihood optimality condition, sufficiency of T, monotonicity of $\\beta \\mapsto E S^2$, entropy-function lemmas, and Lemma 56 used in the proof of Theorem 14.","marker":"[2]"},{"why":"Provides the square-root-law form of the optimal two-tier voting weights that Theorem 38 states and that the weight estimator $\\hat{w}$ is built on.","marker":"[14]"},{"why":"Introduces the democracy-deficit criterion that defines optimal weights in Section 5, the application motivating the weight estimator.","marker":"[9]"}],"fun_headline_variants":["Fast voting-cohesion estimator bypasses partition function","Asymptotic trick turns vote moments into social cohesion","Exponential bottleneck avoided in Curie-Weiss parameter estimation","Partition function sidestepped for social cohesion estimates"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole procedure presupposes condition (10), which requires positive constants $D_{\\mathrm{high}}$ and $D_{\\mathrm{low}}$ asserted to exist in Proposition 22 but never quantified, so a user cannot verify the gap between the high- and low-temperature intervals and cannot choose $b_1$ and $b_2$ to run Algorithm 24 or 25.","fun_headline_variants_meta":{"raw":{"variants":["Fast voting-cohesion estimator bypasses partition function","Asymptotic trick turns vote moments into social cohesion","Exponential bottleneck avoided in Curie-Weiss parameter estimation","Partition function sidestepped for social cohesion estimates"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000665,"raw_usage":{"total_tokens":3078,"prompt_tokens":1028,"completion_tokens":2050,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":644,"completion_tokens_details":{"reasoning_tokens":1987}},"tokens_in":644,"tokens_out":2050,"duration_ms":16577,"temperature":1.0,"reasoning_tokens":1987,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T14:57:50.127403+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"For a fixed $N$, try to instantiate Algorithm 24: condition (10) needs numerical values for $D_{\\mathrm{high}}$ and $D_{\\mathrm{low}}$, and the paper provides none; demonstrating that no such constants can be extracted from the proof would block the advertised constant-cost procedure. Alternatively, simulate the estimator for a true $\\beta$ just inside $I_h$ or $I_l$ with the critical band chosen by any candidate gap and test whether the frequency of `u` responses decays exponentially in $n$ as Proposition 11 predicts; a polynomial decay rate would refute the central claim.","supporting_citations":[{"cited_title":"Reconstruction of the Probability Measure and the Coupling Parameters in a Curie-Weiss Model","cited_arxiv_id":"2505.21778","evidence_quote":"Supplies the maximum-likelihood optimality condition, sufficiency of T, monotonicity of $\\beta \\mapsto E S^2$, entropy-function lemmas, and Lemma 56 used in the proof of Theorem 14."},{"cited_title":"On penrose’s square-root law and beyond","cited_arxiv_id":null,"evidence_quote":"Provides the square-root-law form of the optimal two-tier voting weights that Theorem 38 states and that the weight estimator $\\hat{w}$ is built on."},{"cited_title":"Felsenthal and Mosh´ e Machover","cited_arxiv_id":null,"evidence_quote":"Introduces the democracy-deficit criterion that defines optimal weights in Section 5, the application motivating the weight estimator."}],"review_version":1}