{"id":"2d4e3fb6-04ca-49d7-9ee9-f26a083fc4cf","arxiv_id":"2505.13537","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A framework for ranking non-local games by noise-robustness using p-value based convincingness and a new analytic gapped score, applied to CHSH, 2-CHSH, Magic Square and optimized variants.","lead":"The paper proposes three measures for comparing how well different non-local games detect non-locality in noisy quantum states: noise-tolerance, convincingness, and an analytic gapped score. It reports that the CHSH game is the most noise-robust when games receive equal noisy resources, while some optimized two-copy CHSH variants can outperform it only when given far more resources.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Gapped score reliability is in-sample: its functional form was selected to match the crossing order of these 11 games, then used as a predictor on the same games; no held-out test or derivation supports Obs. 5.","rationale":"The paper makes a real contribution: it proposes a normalized, resource-aware p-value based measure (convincingness), applies it consistently to CHSH, 2-CHSH, 2-CHSH-OPT, and MSG under a clearly specified depolarizing model, and reports detailed computational comparisons. The CHSH-most-robust conclusion under equal resources is at least plausible and is supported by the simulated crossing plots, though it is conditional on the chosen channel extension, as the reader notes. I do not regard the tensor-product channel as the single most load-bearing concern: every noise-robustness measure must be relative to a noise model, and the paper is explicit that this is the model used; changing the model is a legitimate change of scope, not a hidden inconsistency. The more serious issue is that the gapped score's claimed reliability is in-sample. The paper explicitly says it 'explored multiple approaches' and found Def. V.41 'best captures' the desired behavior, using the same data to select and then to validate. That is fitting a predictor to the data and then reporting its training-set performance. The small number of games (eleven configurations) makes this especially fragile, and the need for a supplementary polynomial comparison to distinguish near-ties already signals that kappa alone is not a complete ranker. This does not invalidate the convincingness simulations or the computational ranking, but it means the analytic simplification in Def. VII.43 should be conditional on a derived or out-of-sample-validated score. I therefore keep the reader's CONDITIONAL verdict rather than accepting the analytic claim as established; I would not reject the paper, because the computational side stands on its own and the authors acknowledge several limitations. The requested check would settle whether the gapped score generalizes or is an artifact of the particular game set.","tokens_in":52447,"tokens_out":10851,"duration_ms":112896,"concrete_test":"Hold out a set of games not used in Table 2 or Fig. 8 (e.g., a 3-CHSH game, a pseudo-telepathy game other than MSG, or 2-CHSH-OPT configurations optimized at eta' values not in {0.85,...,1}). Compute their exact significance crossings eta-dagger by the paper's binomial-p-value simulation in the stable regime (Nres = 2000 and 10000), then rank them by the gapped score kappa and by the exact eta-dagger. If the Spearman or Kendall rank correlation between kappa and the actual crossing order is not high (e.g., below 0.8), Observation 5 and Def. VII.43 are unsupported. As an additional check, re-derive kappa from the large-N crossing condition Delta(eta-dagger)^2 = -ln(alpha)/Nres using the analytic Delta(eta) expressions; if the resulting score differs from Def. V.41, the chosen functional form needs explicit justification rather than empirical selection.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The analytic half of the paper rests on Observation 5: in a stable resource region, the gapped score (Def. V.41) 'reliably' ranks the significance crossings eta-dagger, and Def. VII.43 then defines noise-robustness via kappa. But the gapped score was constructed by exploring multiple candidate scores and choosing the one that best matched the observed crossing orders in Fig. 8. It is then evaluated on the same games and configurations used to select it, so the reported reliability is in-sample. No parameter-free derivation from the binomial tail, Hoeffding bound, or crossing condition Delta(eta-dagger)^2 = -ln(alpha)/Nres establishes the specific form (c1+c2)^2 / |d - omega_c|^2. The factors are plausible, but the quadratic power, the sum c1+c2, and the denominator are chosen post hoc after seeing the data. A handful of near-ties (CHSH and 2-CHSH-OPT for eta' in Q) are then resolved by a separate polynomial comparison, further showing the single-number score is not by itself a reliable ranker. Because the gapped score is the basis for the analytic definition of noise-robustness and for the claimed framework's theoretical utility, this in-sample validation is a load-bearing gap. The tensor-product channel choice is an explicit scope condition and is less damaging: the framework is noise-model-relative by design, and the paper states that other channels require re-derivation.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces three measures of noise-robustness for non-local games used as self-tests: noise-tolerance, convincingness (an upper-tailed p-value against the local bound), and the gapped score, an analytic approximation to convincingness. Under a depolarizing-noise model with a tensor-product extension from 2 to 4 qubits, the author computes scores for CHSH, 2-CHSH, several optimized 2-CHSH configurations, and the Magic Square Game, for resource counts from 10^3 to 10^6. The main claims are that convincingness is the most nuanced ranking measure, that the gapped score reliably recovers the significance-crossing order in stable resource regimes, and that CHSH is most noise-robust under equal resources while some 2-CHSH variants can surpass it with substantially more resources. The paper also proposes formal definitions of noise-robustness based on these scores.","tokens_in":52651,"tokens_out":9725,"duration_ms":96344,"significance":"The paper addresses a genuine gap: comparing self-tests of different dimensions and input-output sizes under noise. The p-value-based convincingness is an operationally meaningful normalization, and the explicit treatment of finite resources is a useful contribution. If the gapped score's reliability could be established by out-of-sample tests or by a derivation, Definition VII.43 would give a simple analytic method for ranking self-tests. The computational study is fairly extensive, and the paper is unusually honest about noise-model dependence and open questions. However, the analytic claims need repair before the framework can be accepted.","major_comments":[{"comment":"The gapped score is validated only in-sample. The text reports that multiple candidate scores were explored and that Def. V.41 was selected because it best reproduced the observed crossing order in Fig. 8; it is then evaluated on the same eleven games and configurations used for that selection. No held-out games, new noise models, or unseen resource counts are used, and no parameter-free derivation from the tail bound exp(-Nres(omega_eta_G - omega_c_G)^2) establishes the specific functional form ((c1+c2)^2 / |d - omega_c|^2) * (Nres / -ln alpha). Since Definition VII.43 defines noise-robustness through kappa_G and Observation 5 is the only evidence for its reliability, an out-of-sample prediction or an analytic derivation is needed before the analytic definition can be accepted.","section":"V.B (Def. V.41, Obs. 5, Def. VII.43)"},{"comment":"Lemma 2 asserts that Delta_G(eta) is strictly increasing for any game G, but Table 1 and App. XIV list the 2-CHSH-OPT configurations for eta' = 0.83 and 0.84 with Delta_G = 0; these configurations have constant raw score 0.53125, so strict monotonicity fails as stated. The proof's assertion that the winning state always scores above the maximally mixed state is not justified for the 4-qubit tensor-product channel and is false for those degenerate configurations. Theorem VI.6, which invokes Lemma 2 to conclude eta*_G < eta_dagger_G for every finite Nres, is therefore not proven as stated; the theorem also steps from an asymptotic equivalence CG ~ exp(-Nres Delta^2) to an exact equality at finite Nres. Please restrict the lemma and theorem to configurations with strictly positive gap and provide the explicit 4-qubit derivative computation.","section":"V (Lemma 2) and VI (Theorem VI.6)"},{"comment":"The convincingness definition and its implementation describe different statistics. Definition III.36 and Eq. (14) define C_G as the deterministic binomial tail evaluated at k = round(n*omega_v_G). Algorithm 1 first samples k ~ Binomial(n, omega_v_G), and the figure captions say the p-value was averaged over random seeds. If k is used, the plotted convincingness is a random variable and the significance crossings eta_dagger_G in Table 2 are not fixed game properties unless variances or error bars are reported; if k is not used and the tail is evaluated at round(n*omega_v_G), the random-seed averaging is superfluous and should be removed. This distinction matters because Table 2 and Observations 1-5 compare significance crossings, so the quantity being plotted must be unambiguous.","section":"III.B (Def. III.36, Algorithm 1, Figs. 4-6)"}],"minor_comments":[{"comment":"Equation (24) gives omega_eta^{2-CHSH, eta'=1} = 0.10937 eta^4 + 0.30936 eta^2(1-eta^2) + 0.21875(1-eta^2)^2, which does not match the corresponding raw expression in App. XIV (0.63748 eta^4 + 0.74686 eta^2(1-eta^2) + 0.21875(1-eta^2)^2) nor the coefficients in Table 1; please correct the typo.","section":"V.A, Eq. (24)"},{"comment":"The 4-qubit depolarizing channel is defined as a tensor product of independent 2-qubit channels, and all cross-game rankings are conditional on this extension. The paper states this in Sec. VI, but the abstract and conclusion should carry the qualifier so that readers do not interpret the equal-resource ranking as noise-model-independent.","section":"Def. II.13 and abstract"}],"recommendation":"major_revision","confidential_remarks":"This is a thesis-style manuscript with a promising core. The main risks are not the explicit tensor-product noise-model assumption, which is scoped, but the in-sample validation of the gapped score, the over-general Lemma 2, and the ambiguous p-value protocol. I would ask for an out-of-sample test or derivation for Observation 5, a corrected Lemma 2 and Theorem VI.6, and a precise statement of the plotted convincingness statistic before acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper worth knowing about is Lifshitz's Master's thesis on comparing noise-robustness of non-local games. The genuinely new piece is the systematic comparison framework: using a p-value-based 'convincingness' measure to rank self-tests of different dimensions under depolarizing noise, plus the resource-dependent analysis. The p-value formulation itself is not new — the author credits Araujo, Hirsch, and Quintino [30] — but nobody had built a comparative framework around it or applied it across CHSH, 2-CHSH, Magic Square, and optimized variants. The main empirical conclusion, that CHSH is most noise-robust under equal resource budgets while some 2-CHSH variants overtake it with far more resources, is plausible and supported by the simulations.\n\nThe paper does a lot right. The analytic expressions for the gapped gaps (Table 1) are careful, the noise-tolerance analysis matches the convincingness crossings in the asymptotic limit (a good consistency check), and the author is honest about the prior p-value work. The simulations are reasonably extensive, with averaging over seeds.\n\nSoft spots, in proportion:\n\n1. Lemma 2 is false as stated. It claims Δ_G is strictly increasing for any game, with a proof relying on the winning state scoring above the maximally mixed state. But the paper's own Table 1 shows the 0.83 and 0.84 optimized configurations have Δ_G = 0 (flat), and Theorem VI.6 uses Lemma 2. The theorem may survive for the non-flat games, but the lemma needs a narrower statement and proof.\n\n2. The gapped score κ_G is selected from several candidate predictors on the same data it is then said to 'reliably' rank. The functional form (c1+c2)^2 / |d−ω_c|^2 is plausible but not derived; Observation 5 overclaims reliability without held-out validation or a parameter-free derivation from the binomial tail. This is the clearest load-bearing gap, since Def. VII.43 makes κ the official noise-robustness measure in stable regimes. A referee should ask for either a real derivation or out-of-sample prediction on new games.\n\n3. No code or data are shipped, and the plots lack error bars. For a computational paper that limits reproducibility. Minor.\n\n4. The 'first systematic framework' wording is stronger than the paper's own acknowledgment that the p-value formulation is identical to [30]. The framework is new, but the phrasing could be toned down.\n\nThe tensor-product depolarizing channel is an explicit scope condition, so I don't count it as a flaw; comparing under other noise models is sensibly left as future work.\n\nWho this is for: experimentalists or theorists choosing self-tests for DIQKD or entanglement detection under depolarizing noise. It deserves a serious referee. Recommendation: send to peer review, but major revision — fix Lemma 2, reframe the gapped score as a heuristic with limited validation or derive it properly, release code and data. The central claim about CHSH is likely correct and the framework is a useful contribution.","headline":"A useful but overreaching comparison framework; the convincingness measure is sound, but the gapped score is an in-sample fit and Lemma 2 is false as stated.","tokens_in":53243,"tokens_out":5098,"would_cite":true,"duration_ms":53405,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["81P40","81P68","62F03"],"pacs":["03.65.Ud","03.67.-a"],"model":"deepseek-v4-flash","headline":"Convincingness, a p-value for rejecting a local explanation of a game's wins, gives the most nuanced ranking of noise-robustness among self-tests, and CHSH ranks highest on equal resources.","keywords":["noise-robustness","self-testing","non-local games","convincingness","gapped score","noise-tolerance","CHSH game","depolarizing noise"],"falsifier":"Recompute the gapped expressions and significance crossings using a global 4-qubit depolarizing channel, $\\varepsilon(\\rho) = \\eta^4\\rho + (1-\\eta^4)\\frac{I_{16}}{16}$, in place of the tensor-product channel; if any 2-CHSH variant or the Magic Square Game crosses the threshold at a lower visibility than CHSH under this channel, the reported ranking is an artifact of the channel-extension choice. A second check is to repeat the comparison under dephasing or amplitude-damping noise and test whether $\\kappa_G$ still predicts the crossing order.","tokens_in":52123,"feed_emoji":"⚛️","tokens_out":9517,"duration_ms":79704,"temperature":0.7,"pith_summary":"Non-local games are used to certify entanglement in untrusted devices, but real noise means ideal winning states are never available and there has been no agreed way to say which self-test is more robust than another. The paper proposes three comparative measures — noise-tolerance, convincingness, and an analytic gapped score — and argues that convincingness, an upper-tailed p-value for rejecting a local hidden-variable explanation of the observed win rate, is the most nuanced. Under the tested depolarizing noise model, the CHSH game is found to be the most noise-robust when all games get the same number of noisy EPR pairs, while optimized 2-CHSH variants can beat CHSH only when given roughly ten times more resources. If correct, the framework turns noise-robustness into a computable, comparable quantity that an experimentalist can match to a noise model and a resource budget.","feed_headline":"CHSH stays the most noise-robust self-test on equal resources","feed_subtitle":"A p-value 'convincingness' score ranks self-tests by noise-robustness, with CHSH first under equal resources.","key_machinery":"The load-bearing object is the gapped expression, $\\Delta_G(\\eta) = \\omega^\\eta_G - \\omega^c_G = c_1\\eta^2 + c_2\\eta^4 + d - \\omega^c_G$, which gives each game's score gap under depolarizing noise as a polynomial in the visibility $\\eta$, plus the gapped score $\\kappa_G = \\left(\\frac{c_1+c_2}{|d-\\omega^c_G|}\\right)^2 \\frac{N_{\\mathrm{res}}}{-\\ln\\alpha}$ built from its coefficients. This score tracks the convincingness p-value, $C_G \\sim \\exp(-N_{\\mathrm{res}}\\Delta_G^2)$, which measures how unlikely the observed win rate is under a local model, and it ranks games by the visibility at which their convincingness curves cross the significance threshold $\\alpha$. The machinery works by reducing an incomparable pair of statistics (different dimensions, different score scales) to a normalized, resource-aware number per game.","core_discovery":"The central claim is that noise-robustness of a self-test can be defined operationally: a game $G_1$ is more noise-robust than $G_2$ if it becomes significantly convincing of non-locality at lower visibilities (Def. VII.42), which in stable resource regimes is equivalent to comparing a single closed-form number, the gapped score (Def. VII.43). The paper reports that convincingness provides the most nuanced comparison of the tested measures, that the gapped score reliably predicts the order in which games cross the significance threshold once a few thousand noisy resources are available, and that under equal resource budgets CHSH outranks the 2-CHSH game, the Magic Square Game, and optimized 2-CHSH configurations, whereas with a roughly tenfold resource advantage the best optimized 2-CHSH variants surpass CHSH's significance crossing.","pith_inferences":["The tensor-product channel extension is the step that all rankings inherit: a global 4-qubit depolarizing model, or a dephasing model, could reorder the crossings, so the framework's main empirical claim should be re-tested under those channels.","The $\\kappa_G$ score compresses each game's noise behaviour into four numbers $(c_1, c_2, d, \\omega^c_G)$, so noise-robustness rankings could be tabulated per noise model as a lookup table — a practical step the paper leaves implicit.","The tenfold resource penalty for 2-CHSH-OPT hints at a resource-theoretic cost of parallel repetition; a rigorous lower bound on the resource amplification needed for parallel-repetition self-tests to match CHSH's crossing would settle whether this penalty is fundamental or an artifact of the specific optimization.","The finite-resource flip at lenient thresholds suggests significance levels could be tuned per application; a systematic scan over $\\alpha$ and $N_{\\mathrm{res}}$ could yield decision charts for choosing a self-test, going beyond the single-threshold analysis reported."],"forward_implications":["Given equal numbers of noisy EPR pairs, CHSH certifies non-locality at the lowest visibility among the tested games under depolarizing noise.","The gapped score $\\kappa_G$ gives a closed-form ranking that matches the exact convincingness crossing order for stable resource regimes (roughly $N_{\\mathrm{res}} \\geq 2000$).","Optimized 2-CHSH games, which are worse than CHSH at equal resources, overtake CHSH's significance crossing when given about ten times more noisy resources.","Noise-tolerance rankings coincide with convincingness rankings in the asymptotic resource limit, so noise-tolerance is the coarser measure of the two.","The framework extends to other noise models by deriving the gapped expression for the new channel and re-computing $\\kappa_G$."],"supporting_citations":[{"why":"Defines the CHSH game, the baseline self-test whose noise-robustness anchors the comparison.","marker":"[14]"},{"why":"Supplies the Tsirelson bound used for the CHSH quantum bound in the score computations.","marker":"[36]"},{"why":"Provides the local and quantum bounds for the 2-CHSH game used throughout the analysis.","marker":"[37]"},{"why":"Supplies the linear-programming method for optimizing Bell coefficients that produces the 2-CHSH-OPT configurations.","marker":"[38]"},{"why":"Contains the identical p-value formulation for single-shot rejection of local hidden-variable models, which the convincingness measure is based on.","marker":"[30]"},{"why":"Gives the parallel-repetition device-independent protocol whose resource scaling motivates the unequal-resources result for the optimized 2-CHSH games.","marker":"[18]"}],"fun_headline_variants":["CHSH most noise-robust self-test under equal resources","Convincingness measure puts CHSH first in noise-robustness","Gapped score predicts self-test robustness order, CHSH leads","Equal resources: CHSH beats complex self-tests in noise-robustness","Noise-robust self-tests: CHSH wins on convincingness measure"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The rankings assume that noise on a four-qubit system is just two independent two-qubit noise processes acting on each EPR pair separately; if the noise instead hits all four qubits collectively, the computed scores and the claimed ordering of the games could change.","fun_headline_variants_meta":{"raw":{"variants":["CHSH most noise-robust self-test under equal resources","Convincingness measure puts CHSH first in noise-robustness","Gapped score predicts self-test robustness order, CHSH leads","Equal resources: CHSH beats complex self-tests in noise-robustness","Noise-robust self-tests: CHSH wins on convincingness measure"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000307,"raw_usage":{"total_tokens":1776,"prompt_tokens":983,"completion_tokens":793,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":599,"completion_tokens_details":{"reasoning_tokens":696}},"tokens_in":599,"tokens_out":793,"duration_ms":7031,"temperature":1.0,"reasoning_tokens":696,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:32:26.589532+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the gapped expressions and significance crossings using a global 4-qubit depolarizing channel, $\\varepsilon(\\rho) = \\eta^4\\rho + (1-\\eta^4)\\frac{I_{16}}{16}$, in place of the tensor-product channel; if any 2-CHSH variant or the Magic Square Game crosses the threshold at a lower visibility than CHSH under this channel, the reported ranking is an artifact of the channel-extension choice. A second check is to repeat the comparison under dephasing or amplitude-damping noise and test whether $\\kappa_G$ still predicts the crossing order.","supporting_citations":[{"cited_title":"Proposed experiment to test local hidden- variable theories.Physical review letters, 23(15):880, 1969","cited_arxiv_id":null,"evidence_quote":"Defines the CHSH game, the baseline self-test whose noise-robustness anchors the comparison."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the Tsirelson bound used for the CHSH quantum bound in the score computations."},{"cited_title":"Quantum nonlocality, bell inequalities, and the memory loophole.Phys","cited_arxiv_id":null,"evidence_quote":"Provides the local and quantum bounds for the 2-CHSH game used throughout the analysis."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the linear-programming method for optimizing Bell coefficients that produces the 2-CHSH-OPT configurations."},{"cited_title":"Bell nonlocality with a single shot.Quantum, 4:353, 2020","cited_arxiv_id":null,"evidence_quote":"Contains the identical p-value formulation for single-shot rejection of local hidden-variable models, which the convincingness measure is based on."},{"cited_title":"Noise-tolerant testing of high entanglement of formation","cited_arxiv_id":"1712.09368","evidence_quote":"Gives the parallel-repetition device-independent protocol whose resource scaling motivates the unequal-resources result for the optimized 2-CHSH games."}],"review_version":1}