{"id":"59c00ff8-ba50-4107-bfd1-3ee4ceeeef18","arxiv_id":"2608.06940","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":1,"one_line_summary":"In LLM judge panels, a single-ballot substitution can only change decisions when the vote margin is one, and the accuracy gain from a test-suite signal is confined to those pivotal queries.","lead":"LLM judge panels almost never change their verdict when a single vote is replaced, unless the panel was split by exactly one vote; all the accuracy gain from a test-suite verification signal lands on those pivotal queries. The paper turns this into a gating rule that runs the verifier on roughly one query in six, and shows why aggregate correlation metrics miss the effect.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Positive pivotal gains are measured only for a nested test-suite proxy with one-sided errors; generalization to independent verifiers is untested, making 'verification actually helps' conditional on this signal class.","rationale":"The structural arithmetic (Propositions 1-2) is correct, and the zero gain outside the pivotal region is a theorem rather than a measurement. What needs empirical support is the positive size of the pivotal-region gain, and that is established with bootstrap CIs and replication across benchmarks and panel sizes. The soft spot is that the signal's one-sided error structure is deliberately nested in the ground-truth label, which on pivotal queries (where false rejection dominates) makes the signal correct on every full-suite pass by construction. This mechanical complementarity is disclosed honestly, and the permutation control even shows a location-random signal of equal accuracy would gain more, so the finding is not inflated by the nested design. However, the title and abstract generalize to verification as an evidence source, and no test of an independent, symmetric, or abstaining verifier is provided. This is a scope limitation rather than an internal inconsistency; the reader's CONDITIONAL verdict is appropriate. The concrete test of a disjoint held-out verifier would settle whether the concentration result is a general property of verification or an artifact of subsetting the oracle. Secondary post-hoc concerns about the majority-side replacement rule in Table 7 exist but are less load-bearing because the gating theorem holds for any single-ballot rule.","tokens_in":10732,"tokens_out":18189,"duration_ms":181359,"concrete_test":"Re-run the k=7 HumanEval+/MBPP+ uniform-random replacement analysis with an independent verifier whose signal is a 20%-coverage subset of a held-out test suite V_p disjoint from the full-suite label T_p, so the verifier has symmetric, non-nested errors relative to the operational label. If the pivotal-region accuracy gain is not significantly positive (95% task-bootstrap CI excluding zero) or drops materially below the +11.2pp nested-proxy estimate, the 'verification actually helps' claim fails to generalize beyond the nested proxy.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical claim is secure for the specific signal studied, but the headline 'verification actually helps' rests on an untested generalization. The signal S_c(p) is a nested subset of the full-suite label T_p (Section 3), so it can never reject a full-suite pass and errs only by accepting bugs outside the covered assertions. On pivotal queries the panel's false-rejection rate rises from 7.5% to 33.7% while false-acceptance rises only 1.4x (Section 5); the signal's guaranteed correctness on all full-suite passes therefore directly repairs the dominant pivotal error mode. This one-sided complementarity is a property of the proxy construction, not of verification in general. The paper's own matched-accuracy permutation control underscores the point: a synthetic signal with identical overall accuracy but randomly permuted error locations yields a larger pivotal gain (+14.6pp vs +11.2pp), so the measured magnitude is signal-specific. An independent, symmetric, or abstaining verifier could behave differently, and Section 7 explicitly declines to test this. Thus the empirical contribution is conditional on the nested-proxy signal class; the structural theorem (Propositions 1-2) is unaffected.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper separates a structural fact about majority-vote panels from an empirical claim about a verification signal. For a k-judge unweighted-majority panel with one judge's ballot replaced, Propositions 1 and 2 identify the exact set of queries whose outcome can change: for odd panels, exactly the one-vote-margin (pivotal) queries; for even panels with ties broken as incorrect, signed tallies k/2 and k/2+1. On HumanEval+/MBPP+ and LiveCodeBench, using a nested partial-test-suite signal, the paper measures accuracy gains under uniformly random single-judge substitution on pivotal queries (+10.4 to +23.3pp) and exactly zero elsewhere, while the aggregate effective-vote statistic shows no distinguishable change. It further shows that margin gating preserves the decisions of a specified substitution rule while cutting verifier calls to 12-27%, and that replacement-rule choice matters (majority-side replacement reaches 85.62% overall accuracy).","tokens_in":10880,"tokens_out":18586,"duration_ms":194485,"significance":"The paper's main contribution is the clear separation of structural from empirical claims and the demonstration that population-level dependence metrics (n_eff) and margin-conditional utility answer different questions. The proofs are elementary and correct; the empirical analysis is unusually careful, with exact averaging over replacement choices, task-level bootstrap CIs, 56 dependent judge-subset checks, and a matched-accuracy permutation control. The explicit acknowledgment that the signal is a nested one-sided proxy and that the empirical generalization to independent/symmetric/abstaining verifiers is untested is a strength, though it should be reflected in the title/abstract. If the result holds, it provides a principled call-reduction rule for single-ballot substitution policies and reframes how panel verification gains should be reported.","major_comments":[],"minor_comments":[{"comment":"The empirical claims are established only for the nested proxy S_c(p) subset of T_p with one-sided errors; Section 7 correctly states that the result may not transfer to independent, symmetric, or abstaining verifiers. The title and abstract phrase the conclusion as 'verification actually helps' without that qualifier, so I recommend adding a qualifier such as 'for nested test-suite proxies' to the title/abstract/conclusion to match the evidence.","section":"Title/Abstract; Sections 3 and 7"},{"comment":"The row 'Best judge 0% 87.16%' appears to be an oracle policy selected on the test set; please label it as an oracle/best-on-test-set judge and note in the text that it is not a deployable policy, otherwise the comparison overstates its status.","section":"Table 7"},{"comment":"The '0.0000' non-pivotal gains are consequences of Propositions 1 and 2 rather than measurements; the text says this, but a table footnote would prevent readers who see only the tables from misreading the entry as an empirical zero.","section":"Tables 2 and 6"},{"comment":"The indicator notation ⊮[·] is nonstandard; use \\mathbb{1}[·] or 1{·} for readability.","section":"Section 3"},{"comment":"In the margin-threshold sweep, 'every threshold t≥1' is potentially confusing because odd-k panels have no queries with margin 2; please state that the threshold runs over odd margins (t=1,3,5,...).","section":"Section 5"}],"recommendation":"minor_revision","confidential_remarks":"The paper is honest about its main limitation (the nested proxy) in Section 7, and the structural theorem is unaffected. The editor may wish to ask the authors to soften the title/abstract; I did not treat the untested generalization as a reason for rejection because the paper explicitly scopes its empirical contribution. The 'best judge' row in Table 7 should be corrected before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The core of this paper is right, and the empirical work is more careful than most of what we see in LLM evaluation. The structural point is elementary — for single-ballot substitution in an odd-size majority panel, only one-vote-margin queries can flip — and it is proved cleanly in Propositions 1 and 2. What is actually new is the systematic margin-stratified treatment: the even-panel asymmetry test (Proposition 2 validated on all seven 6-judge subsets, including the predicted structural zero where margin alone would mislead), the direct demonstration that the effective-vote statistic misses a concentrated effect, and the gating rule that preserves a specified substitution policy at 16% of calls. None of that is deep theory, but the application to LLM judge panels is new and the measurements back it up.\n\nThe empirical case is genuinely solid for what it tests. Bootstrap CIs, a matched-accuracy permutation control, 56 dependent subsampling checks, and full model IDs and data in the archive. The paper also earns credit for disclosing its own limitations: it says signal-only and the best judge beat every panel policy, that gating matters only when a panel must be retained, and that the replacement-policy results need held-out validation. That is rare honesty.\n\nThe soft spot is exactly the one the authors flag in Section 7 but keep out of the abstract: the verification signal is a nested subset of the full-suite operational label. It can never reject a full-suite pass, so on pivotal queries it directly repairs the panel's dominant false-rejection mode. That is a property of the proxy construction, not of verification in general. The permutation control makes the point concrete — a synthetic signal with the same overall accuracy but permuted error locations gives a bigger pivotal gain (+14.6 vs +11.2pp) — so even the magnitude is signal-specific. An independent, symmetric, or abstaining verifier could behave differently, and the paper does not test that. The claim that \"verification actually helps\" should therefore be read as \"this nested test-suite proxy helps,\" and the authors should either soften the generalization or run the follow-up.\n\nMinor quibbles: the LiveCodeBench pivotal sample is small (113 observations), and the signal-accuracy sweep covers only 87.4–88.7%, so the dose-response discussion is mostly suggestive. Neither undermines the main result.\n\nWho is this for? Practitioners running LLM judge panels with an external signal, and methodologists interested in why aggregate dependence statistics can hide conditional utility. It deserves a serious referee, and I would want the revision to address the generalization question directly rather than only acknowledging it.","headline":"A careful, honest empirical paper with a correct structural core, but the headline 'where verification actually helps' is conditional on the nested test-suite proxy it studies, not a general result about verifiers.","tokens_in":775,"tokens_out":1947,"would_cite":true,"duration_ms":35685,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"For single-ballot substitution in an unweighted-majority LLM judge panel, only one-vote-margin (pivotal) queries can change the decision, and empirically the entire accuracy gain of a test-suite verification signal is concentrated there —…","keywords":["LLM judge panels","majority voting","pivotal vote","test-suite verification","effective-vote count","margin gating","aggregate dependence","code generation evaluation"],"falsifier":"A concrete test: run the same seven-judge panel on HumanEval+/MBPP+ but replace the nested test-suite signal with an independently constructed verifier that has symmetric errors (e.g., a separate suite of generated tests that can both miss bugs and reject correct code). If the pivotal-region gain becomes indistinguishable from zero, or if non-pivotal queries show nonzero gains, then the paper's empirical claim depends on the one-sided signal structure rather than on the pivotal arithmetic; the arithmetic itself (zero change on m ≥ 3) would remain true, but its practical significance would be qualified.","tokens_in":10451,"feed_emoji":"🗳️","tokens_out":7748,"duration_ms":67125,"temperature":0.7,"pith_summary":"An LLM judge panel that decides by majority vote can be changed by substituting one judge's vote with a verification signal only when the original tally is one vote away from a tie — the pivotal queries. The paper proves this arithmetically for odd and even panels, and shows empirically on three code benchmarks that all the accuracy gain from a test-suite signal is concentrated in that pivotal region (+10.4 to +23.3 percentage points), with exactly zero gain elsewhere. Because these queries are a minority (12–27% of cases), aggregate statistics such as the effective-vote count can average the effect away and appear to show no benefit, even though the conditional effect is large. The paper thus separates a structural statement about majority arithmetic from an empirical statement about where verification helps, and derives a call-reduction rule: run the verifier only on pivotal queries.","feed_headline":"Verification helps LLM panels only when one vote decides","feed_subtitle":"Test-suite signals add up to +23 points on pivotal queries and zero elsewhere; aggregate metrics miss it.","key_machinery":"The pivotal-vote boundary: for an odd-sized panel, the majority decision has margin m = |2s − k|, and a single-ballot substitution can flip the decision if and only if m = 1 (the tally is one vote from a tie). The paper calls such queries \"pivotal\". The proof is a two-line tally bound — after replacing one vote the total moves by at most one, crossing the decision threshold ⌈k/2⌉ only from s = ⌊k/2⌋ or s = ⌈k/2⌉. For even panels with ties broken as \"incorrect\", the analogous set is s ∈ {k/2, k/2+1}, notably asymmetric in the margin. This mechanism carries the argument by fixing exactly where any single-ballot substitution can have effect, turning the empirical question into one of measuring error rates and gains inside that region. The verification signal itself is a nested subset of the full test suite (S_c(p) ⊆ T_p), which gives it a one-sided error structure: it never rejects a full-suite pass and errs only by accepting bugs outside the covered assertions.","core_discovery":"The discovery is that, for single-ballot substitution in an unweighted-majority panel, the set of decisions that can change is exactly the set of pivotal queries — those with a one-vote margin (|2s − k| = 1 for odd k; for even k with ties broken as \"incorrect\", the tally must be exactly k/2 or k/2+1, an asymmetric condition). Propositions 1 and 2 establish this by elementary tally arithmetic: replacing one vote changes the total by at most one, so it can cross the decision boundary only when the original tally sits one vote from the boundary. The empirical complement, measured across HumanEval+/MBPP+ and LiveCodeBench with panels of size 3, 5, 7, and 9, is that panel error rates rise as the margin narrows and that substituting a test-suite signal yields +10.4 to +23.3 percentage points of accuracy on pivotal queries and exactly 0.0000 elsewhere. A matched-accuracy permutation control shows that the real signal's pivotal gain is below that of synthetic signals with randomly permuted error locations, so the benefit is not explained by signal accuracy alone. As a consequence, aggregate dependence statistics like the effective-vote count (which change by −0.04, 95% CI crossing zero) can be blind to a large conditional utility, and margin-gated verification — calling the signal only on pivotal queries — reproduces the universal-substitution decisions at 12–27% of the call rate.","pith_inferences":["The structural confinement is a theorem about majority arithmetic, so it applies to any single-ballot substitution rule in any domain — not just code — including human-annotator panels, ensemble classifiers, or any scenario where one vote is replaced by another source; the empirical concentration of gains may hold more broadly, but the paper's measurements are confined to test-suite signals on cod","The one-sided error structure of the nested-subset signal suggests that a symmetric or abstaining verifier might behave differently; the paper's claim that \"verification actually helps\" is therefore best read as a claim about this class of one-sided signals, not a general property of all verification signals.","The matched-accuracy permutation control implies that the magnitude of the pivotal gain is not explained by signal accuracy alone; an independent signal of equal accuracy could, in principle, produce a different (higher or lower) gain, which is a testable prediction for future work.","The result suggests a practical design principle for evaluation pipelines: if a panel is required (e.g., for governance or audit reasons), the expensive or slow verification resource should be gated on the pivotal region rather than applied uniformly, and the replacement rule (which judge's vote to overwrite) should be chosen deliberately."],"forward_implications":["Any single-ballot substitution policy in an odd unweighted-majority panel changes predictions only on pivotal (m=1) queries; on all other queries the gain is exactly zero, regardless of the signal's accuracy.","Aggregate dependence metrics such as the effective-vote count can be statistically indistinguishable from zero (−0.04, 95% CI [−0.10, +0.02]) while the margin-conditional gain is large (+11.2pp on 16.2% of queries), so the two diagnostics answer different questions and should be reported together.","Margin gating — invoking the verifier only when m=1 — preserves a specified replacement rule's universal-substitution predictions while cutting verifier calls to 12–27% of queries.","For even panels with ties defaulting to \"incorrect\", the pivotal set is asymmetric: a bare incorrect-leaning majority (s = k/2 − 1) is not pivotal even though it has the same |m| = 2 as the bare correct-leaning majority (s = k/2 + 1); the paper verifies this zero-gain prediction on all seven 6-judge subsets of the 7-judge panel.","On HumanEval+/MBPP+ with k=7, majority-side replacement under gating raises accuracy from 82.44% to 85.62% at a 16.2% call rate, though signal-only (87.60%) and the best single judge (87.16%) remain stronger."],"supporting_citations":[{"why":"Supplies the n_eff effective-vote statistic and the correlated-errors diagnosis that the paper contrasts with margin-conditional utility.","marker":"Kohli [2026]"},{"why":"Provides the HumanEval+ and MBPP+ benchmarks, the main discovery datasets for the pivotal-gain measurements.","marker":"Liu et al. [2023]"},{"why":"Provides LiveCodeBench, the harder contamination-resistant benchmark used as a replication setting.","marker":"Jain et al. [2025]"},{"why":"Formalizes the notion of a pivotal or decisive voter that frames the paper's affected-set characterization.","marker":"Banzhaf, 1965"},{"why":"Documents prior practical use of margin-gated verification on a 3–2 ensemble split, which the paper systematically tests and isolates.","marker":"Akinfaderin and Diallo [2026]"}],"fun_headline_variants":["LLM panel help? Only when the vote is pivotal","Test-suite signals: zero gain unless one vote decides","Pivotal votes: where verification actually saves LLM panels","Aggregate metrics blind: pivotal-only gains up to +23 pts","Margin matters: LLM panel verification only on razor-thin votes"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that the verification signal has a one-sided error structure relative to the operational label — it is a nested subset of the full test suite, so it can never reject a full-suite pass and errs only by accepting bugs outside the covered assertions; if an independent, symmetric, or abstaining verifier were used, the measured concentration of gains on pivotal queries might shift or disappear.","fun_headline_variants_meta":{"raw":{"variants":["LLM panel help? Only when the vote is pivotal","Test-suite signals: zero gain unless one vote decides","Pivotal votes: where verification actually saves LLM panels","Aggregate metrics blind: pivotal-only gains up to +23 pts","Margin matters: LLM panel verification only on razor-thin votes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000231,"raw_usage":{"total_tokens":1582,"prompt_tokens":1141,"completion_tokens":441,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":757,"completion_tokens_details":{"reasoning_tokens":355}},"tokens_in":757,"tokens_out":441,"duration_ms":4601,"temperature":1.0,"reasoning_tokens":355,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T17:55:26.816611+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete test: run the same seven-judge panel on HumanEval+/MBPP+ but replace the nested test-suite signal with an independently constructed verifier that has symmetric errors (e.g., a separate suite of generated tests that can both miss bugs and reject correct code). If the pivotal-region gain becomes indistinguishable from zero, or if non-pivotal queries show nonzero gains, then the paper's empirical claim depends on the one-sided signal structure rather than on the pivotal arithmetic; the arithmetic itself (zero change on m ≥ 3) would remain true, but its practical significance would be qualified.","supporting_citations":[],"review_version":1}