{"id":"8c6fc68e-0ac3-44a7-b29d-c861d7981708","arxiv_id":"2608.06362","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Combining AIVAT variance reduction with anytime-valid confidence sequences lets poker agent evaluations stop at a median 74x fewer hands at plus or minus 1 BB, with exact finite-sample certification currently limited to games with an independent payoff bound.","lead":"An evaluation method for poker-style agents makes comparisons stop early when the evidence is clear, using variance reduction plus confidence sequences that stay valid under repeated checking. In a 71,439-hand heads-up no-limit hold'em test it cuts the hands needed for a fixed precision by a median factor of 74, while the exact certificate applies only where a structural payoff bound exists.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The HUNL 74x claim is not backed by an exact anytime-valid certificate: EB-CS there uses a data-derived BY (descriptive), and the AsympCS screen that produces the 74x ratio has finite-sample false-positive rates above 5%.","rationale":"I read the paper as honestly separating asymptotic screening from exact certification; the reader's conditional verdict reflects exactly that. The load-bearing point is that the headline quantity (74x) is produced by the asymptotic stream, while the word 'certified' is only supported by the EB-CS, which is exact only for Leduc. Since the target setting in the title is HUNL, the central claim as stated is not fully supported. I find no additional internal inconsistency: Proposition 1's proof mechanism is standard, the predictable-interface argument is correct, and the paper repeatedly labels descriptive runs as descriptive. The missing piece is an independent HUNL corrected-payoff bound or a recalibration of the AsympCS threshold to honest finite-sample levels. The proposed test would settle which of the two halves of the headline can stand.","tokens_in":26163,"tokens_out":5755,"duration_ms":72773,"concrete_test":"Derive an independent BY for the HUNL corrected stream from first principles and rerun the exact EB-CS: e.g., use Proposition 1(ii) with a verified sup-norm cap on the AIVAT value estimates and a certified maximum number of correction points per hand, or a game-tree telescoping argument in the style of Appendix D.4. Then recompute the raw-to-AIVAT EB-CS stopping-time ratio at ±1 BB over the same 200-permutation protocol. If the certified ratio is close to the descriptive 1.37x in Table 1 (as the Lemma 1 floor 4B log(2/alpha)/t suggests for large B), the 74x headline is confirmed to be an asymptotic-screen result; if it approaches 74x and the finite-horizon AsympCS false-positive rates remain above 5%, the exact-certification claim fails in HUNL either way.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Proposition 1 makes the exact EB-CS certificate conditional on a data-independent almost-sure bound BY on corrected payoffs. For HUNL, Appendix D.2 states that no such bound is available: B=200 BB is taken as the observed raw maximum (rounded upward), and the corrected stream has no matching analytic ceiling (largest observed corrected |Y|=153.79 BB). Consequently the EB-CS rows in Table 1 are explicitly descriptive, and the exact certification part of the headline holds only for Leduc (BY=117, Appendix D.4). The advertised 74x stopping-time reduction is obtained from the AsympCS screen, whose guarantee is asymptotic (Proposition 2). The paper's own finite-horizon calibration shows the screen above nominal: 7.11% average exclusion on centered P0 streams (12.43% AIVAT-only) and 10.4% standalone false-positive rate in A2. Thus the central 'certified anytime-valid stopping' claim is not exact in the setting where the 74x factor is measured; at finite horizons the user gets an approximate screen with error rates above 95% confidence, plus a descriptive EB-CS replay. This does not invalidate the method's components, but it means the headline overstates what is guaranteed for HUNL.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces AV-AIVAT, a protocol that combines AIVAT variance reduction with confidence sequences to enable anytime-valid early stopping in imperfect-information game evaluation. It proves that past-only updates of the value function preserve the conditional mean-zero property of AIVAT corrections (Proposition 1), giving an exact Empirical-Bernstein confidence sequence once an independent almost-sure bound on the corrected payoff is available. An asymptotic CLT-based CS (AsympCS) is used as the efficient primary stopping rule, and the paper characterizes when variance reduction converts into earlier stopping: a bet-capped EB-CS has a deterministic width floor proportional to the declared bound, while an asymptotic width benchmark shows the usual variance-ratio scaling. Empirically, on 71,439 paired HUNL hands from 15 LLM-agent configurations, AIVAT reduces variance by a median 54x and the AsympCS stopping-time ratio is a median 74x at the ±1 BB target; the HUNL EB-CS runs are labeled descriptive because no independent corrected-payoff bound is available. A structural bound is derived for Leduc, giving an exact finite-sample certificate there. The paper also proposes a release protocol for rechecking early-stopping claims and demonstrates the dangers of continuous monitoring with fixed-sample intervals.","tokens_in":26417,"tokens_out":5195,"duration_ms":61291,"significance":"The paper makes a valuable methodological contribution by drawing a clear distinction between asymptotic screening and exact certification, by proving that predictable online value functions preserve AIVAT validity, and by quantifying the width floor that limits how much variance reduction becomes earlier stopping for the EB-CS. The empirical corpus is substantial, and the paper is unusually transparent: it explicitly labels the HUNL EB-CS results as descriptive, reports finite-horizon calibration rates above nominal, and provides a detailed release protocol. The exact Leduc certificate is a genuine finite-sample anytime-valid result. However, the headline '74x cheaper with certified anytime-valid stopping' overstates the strength of the HUNL guarantee, since the 74x ratio is produced by the asymptotic screen and the HUNL exact certificate is not established. This is a framing and scope issue that is fixable within the manuscript's aims, but it is load-bearing for how the contribution is presented.","major_comments":[{"comment":"The title and abstract present 'Certified Anytime-Valid Stopping' as the headline contribution, but the 74x stopping-time reduction is measured under AsympCS (Eq. 5), whose guarantee is only asymptotic (Proposition 2). Appendix D.2 states that the HUNL corrected stream has no analytic ceiling; the declared B = 200 BB is the observed raw maximum rounded upward, with the largest observed corrected |Y| = 153.79 BB. Consequently the EB-CS rows in Table 1 are explicitly descriptive, and the exact anytime-valid certificate is established only for Leduc (BY = 117, Appendix D.4). The paper is transparent about this in Sections 6 and 9, but the title and abstract do not carry the qualification. This mismatch is load-bearing for the central claim and should be resolved by either supplying an independent HUNL bound or reframing the headline as an asymptotic-screen result.","section":"Abstract, Eq. (5), Table 1, Appendix D.2"},{"comment":"The AsympCS screen's finite-horizon exclusion rates exceed the nominal 5% level: P0 reports an average 7.1067% (12.43% for AIVAT-only streams) and the A2 standalone false-positive rate is 10.4%. The abstract's phrase 'At the nominal 95% level' appears immediately before the 74x claim, inviting the reader to treat the stopping rule as having 95% coverage at the observed horizons. The paper should state explicitly that AsympCS is an approximate screen with these measured finite-horizon rates, rather than implying that the 74x ratio is obtained under a guaranteed 95% confidence level.","section":"Table 3, Section 6.1"},{"comment":"Proposition 1's exact EB-CS validity is conditional on an independently justified almost-sure bound BY on the corrected payoff. The paper derives such a bound only for Leduc; for HUNL no such bound is proved, as acknowledged in Appendix D.2 and Table 1's footnote. The release protocol in Section 8 correctly requires an independent bound for exact claims, but the Contributions section and the abstract should match that standard by explicitly limiting exact certification to settings where an analytic or structural bound is available. As written, a reader may reasonably infer that the HUNL experiments enjoy the same exact anytime-valid certificate as Leduc.","section":"Proposition 1, Section 8"}],"minor_comments":[{"comment":"Reference [22] cites Peter Grünwald's discussion comment as the source of the empirical-Bernstein confidence sequence, but the construction in Eq. (4) originates from Waudby-Smith and Ramdas's paper on estimating means of bounded random variables by betting. The citation should be corrected to the primary source.","section":"References"},{"comment":"The notation 'log 2 α' is easily misread as log(2α); since the text defines it as log(2/α), the paper should use log(2/α) consistently in displayed equations.","section":"Eq. (4)"},{"comment":"Figure 2 labels the descriptive B = 22 and loose B' = 200 curves, but the analytic BY = 117 floor anchors are mentioned only in the text; adding them to the figure would make the comparison immediate.","section":"Figure 2"},{"comment":"The name 'AIVAT' is frequently typeset as 'AIV AT' with a space; this should be corrected to a single token for consistency.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is unusually honest about its limitations, which is a genuine strength. However, the title and abstract are likely to be read as claiming exact anytime-valid certification for the 74x HUNL result, while the paper's own appendix shows that the exact certificate exists only for Leduc and that the asymptotic screen's finite-horizon rates exceed nominal. The authors should be asked to align the headline with the actual guarantees, either by softening the claim or by providing an independent HUNL bound. The reference issue with [22] should also be fixed, as it may raise provenance questions. These are fixable within the manuscript's scope, so I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is a genuinely useful combination of two existing tools: AIVAT variance reduction and confidence sequences. Proposition 1 is the key structural result — past-only value function updates preserve the conditional mean-zero correction property, so you can refit online without invalidating the exact EB-CS certificate. That is a real advance, as is Lemma 1's deterministic width floor showing why a bet-capped EB-CS converts variance gains into stopping-time gains only when the declared bound is tight. The variance-regret theorem is a nice efficiency guarantee, and the Leduc structural bound BY=117 is careful, checkable work. The paper is also unusually transparent: every data-dependent bound is labeled descriptive, AsympCS is labeled asymptotic, finite-horizon calibration is reported, and the release protocol is sensible.\n\nThe soft spot is exactly where the stress-test note points. The headline \"74x cheaper with certified stopping\" holds as stated only in Leduc. The HUNL 74x number comes from the AsympCS screen, which has finite-horizon false-positive rates above nominal (7.1% overall, 12.4% AIVAT-only, 10.4% in A2). And the exact EB-CS certificate for HUNL is descriptive because B=200 is taken from the observed maximum, not an independent bound. The paper itself says this clearly in Appendix D.2 and Table 1. So the overstatement is in the marketing, not in the math: the authors separate exact from asymptotic carefully, but the abstract's \"certified anytime-valid stopping\" invites a stronger reading than the HUNL results support. The fix is straightforward: temper the abstract, or supply an independent HUNL corrected-payoff bound, and ship code/data so the stopping claims can be reconstructed. The AsympCS finite-sample rates above nominal are also a real limitation for any user who wants a 95% guarantee at a fixed horizon — the paper treats them as calibration evidence, which is fair, but it means the practical benefit is an approximate screen with inflated early rejection, not a certified 95% procedure.\n\nWho this is for: anyone doing sequential evaluation of expensive interactive agents, LLM-based or otherwise, and anyone working on anytime-valid inference for control variates. The paper deserves a serious referee: the theory is coherent, the empirical claims are mostly reproducible in structure, and the limitations are stated rather than hidden. I would accept for review with a request to fix the headline and release artifacts.","headline":"A solid anytime-valid variance-reduction paper whose advertised 74x HUNL certificate is exactly what the paper itself admits: asymptotic and descriptive, not exact finite-sample — but the components are real and the transparency is exemplary.","tokens_in":26990,"tokens_out":1607,"would_cite":true,"duration_ms":18325,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62L10","62L12","62F25","91A40"],"pacs":[],"model":"deepseek-v4-flash","headline":"AIVAT corrections plus anytime-valid confidence sequences let agent evaluation stop the moment evidence suffices, keeping the guarantee intact: a median 74x reduction in hands needed at ±1 BB in HUNL, with exact certificates in Leduc.","keywords":["AIVAT","confidence sequences","anytime-valid inference","variance reduction","imperfect-information games","agent evaluation","optional stopping","heads-up no-limit hold'em"],"falsifier":"Two checks would settle the claims. First, find any legal Leduc trajectory whose corrected payoff exceeds 117 chips, which would break the structurally derived certificate, or any AIVAT-corrected HUNL hand in the corpus whose absolute value exceeds the declared 200 BB, which would invalidate the descriptive EB-CS replay. Second, run the AsympCS as its own monitor on a zero-effect stream for 50,000 hands and count exclusions of zero: exclusion rates near the nominal 5% would support the asymptotic screen, while rates far above it would confirm that finite-horizon calibration, not the asymptotic guarantee, is its real operating characteristic.","tokens_in":25906,"feed_emoji":"🃏","tokens_out":11749,"duration_ms":120067,"temperature":0.7,"pith_summary":"The paper solves a practical matching problem: deciding which of two agents is stronger by playing games is expensive, and nobody knows in advance how many games the verdict will require. Its proposal is to let the value function that powers AIVAT corrections be refit online during the evaluation — as long as it is fixed before the hand it scores — so variance reduction and sequential monitoring work together. On 71,439 paired heads-up no-limit hold'em hands across 15 configurations, the corrected stream has median 54x lower variance than the raw stream, and at the nominal 95% level with a ±1 big blind target it stops after a median 74x fewer hands under the asymptotic confidence sequence. Exact finite-sample certification is deliberately separated from that asymptotic screen: it needs an independently justified bound on corrected payoffs, which the paper derives structurally for Leduc hold'em (|Y| ≤ 117) but not for HUNL, where the declared 200 BB bound is taken from observed data and the EB-CS run is labeled descriptive.","feed_headline":"Agent evaluation stops 74x sooner with corrected poker payoffs","feed_subtitle":"Variance-reduced payoffs plus anytime-valid confidence sequences make early stopping provable, not just cheaper.","key_machinery":"The load-bearing mechanism is the predictable AIVAT interface feeding two confidence sequences with different guarantees. At each of a fixed finite set of chance and evaluated-agent decision nodes, the evaluator must know the conditional action kernel p_{t,h}, must fix the enablement decision S_{t,h} before seeing the outgoing action, and must use a value function v_t built only from hands 1..t−1; under these conditions the correction telescopes into conditionally mean-zero summands, giving E[C_t | F_{t−1}] = 0 and hence E[Y_t | F_{t−1}] = µ. That single identity licenses refitting the value model during the run and plugs the corrected stream into the cited betting-based empirical-Bernstein construction, which becomes an exact confidence sequence once a data-independent bound B_Y on |Y_t| is supplied. The second mechanism is the width floor: because the EB-CS bets are capped at 1/2 and all numerator terms are nonnegative, the payoff-scale half-width obeys w_t ≥ 4B log(2/α)/t for every realization, so an exact interval converts a variance gain into earlier stopping only when the declared bound is tight relative to σ²/ε (Lemma 1, Theorem 1). The efficient primary interval is the asymptotic CS, whose width tracks realized variance, and Theorem 2 guarantees that a past-only value learner with sublinear variance regret reaches the oracle-value stopping horizon with no asymptotic delay.","core_discovery":"The central claim, stated on the paper's own terms, is that predictability is the only condition that matters for valid variance reduction in sequential agent evaluation: if the value function v_t used for hand t is measurable with respect to information available before hand t, the conditional action kernel is known at every enabled correction node, and correction enablement is fixed before the action is seen, then the AIVAT correction C_t has conditional mean zero (Proposition 1). The corrected payoff Y_t therefore keeps the target mean, and the bounded Empirical-Bernstein confidence sequence, rescaled by any independent almost-sure bound B_Y on |Y_t|, is an exact (1−α) time-uniform interval. The paper then shows on its HUNL corpus that this interface delivers a median 54x variance reduction and a median 74x reduction in hands needed to reach the ±1 BB target under the asymptotic CS, while the exact certificate's speed is governed by a deterministic width floor 4B log(2/α)/t set by the declared bound and the bet cap. It proves this in Leduc, where the tree structure yields the analytic bound |Y| ≤ 117 and the exact interval runs within about 1% of its floor; in HUNL the bound is taken from observed maxima, so that replay is presented as descriptive rather than certified.","pith_inferences":["The 74x factor belongs to the asymptotic screen, not the exact certificate: on the same HUNL data the descriptive EB-CS replay shows only a 1.37x ratio, so claims built on the headline number should be read as screening-grade evidence, not certified.","The single step that would promote the headline result from descriptive to certified is a sample-independent almost-sure bound on corrected HUNL payoffs, analogous to the Leduc |Y| ≤ 117 certificate; the paper's own release protocol names exactly this requirement.","The width-floor analysis implies a design target: a confidence sequence whose range dependence adapts to the realized payoff scale would convert more of the 54x variance gain into earlier exact stopping, which is the direction the paper flags as its sharpest open problem.","The continuous-monitoring audit generalizes beyond poker: any leaderboard or benchmark that inspects a fixed-sample interval repeatedly is vulnerable to the demonstrated 61% false-positive effect, and publishing the monitoring rule with a time-uniform bound before the run is the corresponding fix."],"forward_implications":["Evaluations of expensive interactive agents can be run until the evidence is in and stopped at a data-dependent time with the stated level intact, replacing fixed budgets that either overpay or underpower.","At roughly $0.07–$0.30 per evaluated hand for LLM poker agents, a median 74x reduction in hands to a ±1 BB verdict turns a budget question into a routine one for the asymptotic screen.","A published early-stopping claim is recheckable by a third party from the corrected payoff prefix and the stop metadata alone; in the constructed zero-effect audit, checking either CS at the reported stop screens out essentially all of the 61% false claims that a continuously monitored fixed interval produces.","Exact finite-sample certification is achievable in games whose structure supplies a payoff bound: in Leduc, |Y| ≤ 117 makes the EB-CS exact and its realized width sits within about 1% of the deterministic floor.","Refitting the value function on past hands only is asymptotically free: under sublinear variance regret, the stopping time matches the oracle-value benchmark while exact EB-CS validity continues to hold."],"supporting_citations":[{"why":"Supplies the original AIVAT correction construction whose conditional mean-zero property the protocol extends to online value functions.","marker":"[16]"},{"why":"The cited betting-based predictable plug-in empirical-Bernstein e-process theorem that makes the EB-CS exact under a declared bound.","marker":"[22]"},{"why":"The time-uniform CLT construction whose asymptotic confidence sequence serves as the efficient primary interval.","marker":"[47]"},{"why":"The time-uniform concentration results used to control the empirical-variance term in the stopping-time characterization.","marker":"[23]"},{"why":"Supplies the HUNL corpus of paired per-hand raw and corrected payoffs used for the main stopping-time and variance experiments.","marker":"[38]"},{"why":"Documents the spurious-significance failure mode of post hoc AIVAT tuning that motivates the predictability requirement.","marker":"[29]"},{"why":"Provides the fixed opponent against which the HUNL agent configurations are evaluated.","marker":"[42]"}],"fun_headline_variants":["AV-AIVAT: provable early stop for agent duels","74x fewer hands to settle agent skill contests","Certified anytime-valid stopping cuts eval cost 74x","Variance reduction makes agent comparison stop early","Agent evaluation halts 74x sooner with AV-AIVAT"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The exact finite-sample certificate stands on a bound on corrected payoffs that is justified independently of the data; the paper has such a bound only for Leduc (|Y| ≤ 117), while in HUNL the 200 BB bound is taken from observed maxima and the resulting EB-CS is labeled descriptive rather than certified.","fun_headline_variants_meta":{"raw":{"variants":["AV-AIVAT: provable early stop for agent duels","74x fewer hands to settle agent skill contests","Certified anytime-valid stopping cuts eval cost 74x","Variance reduction makes agent comparison stop early","Agent evaluation halts 74x sooner with AV-AIVAT"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000193,"raw_usage":{"total_tokens":1477,"prompt_tokens":1202,"completion_tokens":275,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":818,"completion_tokens_details":{"reasoning_tokens":194}},"tokens_in":818,"tokens_out":275,"duration_ms":3697,"temperature":1.0,"reasoning_tokens":194,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T04:16:55.408962+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Two checks would settle the claims. First, find any legal Leduc trajectory whose corrected payoff exceeds 117 chips, which would break the structurally derived certificate, or any AIVAT-corrected HUNL hand in the corpus whose absolute value exceeds the declared 200 BB, which would invalidate the descriptive EB-CS replay. Second, run the AsympCS as its own monitor on a zero-effect stream for 50,000 hands and count exclusions of zero: exclusion rates near the nominal 5% would support the asymptotic screen, while rates far above it would confirm that finite-horizon calibration, not the asymptotic guarantee, is its real operating characteristic.","supporting_citations":[{"cited_title":"Aivat: A new variance reduction technique for agent evaluation in imperfect information games","cited_arxiv_id":null,"evidence_quote":"Supplies the original AIVAT correction construction whose conditional mean-zero property the protocol extends to online value functions."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The cited betting-based predictable plug-in empirical-Bernstein e-process theorem that makes the EB-CS exact under a declared bound."},{"cited_title":"Time- uniform central limit theory and asymptotic confidence sequences.The Annals of Statistics, 52 (6), 2024","cited_arxiv_id":null,"evidence_quote":"The time-uniform CLT construction whose asymptotic confidence sequence serves as the efficient primary interval."},{"cited_title":"Time-uniform, nonparametric, nonasymptotic confidence sequences.The Annals of Statistics, 49(2):1055–1080, 2021","cited_arxiv_id":null,"evidence_quote":"The time-uniform concentration results used to control the empirical-variance term in the stopping-time characterization."},{"cited_title":"PokerSkill: LLMs Can Play Expert-Level Poker without Training or Solvers","cited_arxiv_id":"2605.30094","evidence_quote":"Supplies the HUNL corpus of paired per-hand raw and corrected payoffs used for the main stopping-time and variance experiments."},{"cited_title":"Heuristic Pathologies and Further Variance Reduction via Uncertainty Propagation in the AIVAT Family of Techniques","cited_arxiv_id":"2605.14261","evidence_quote":"Documents the spurious-significance failure mode of post hoc AIVAT tuning that motivates the predictability requirement."}],"review_version":1}