{"id":"cfb6fc4f-48ed-4f4c-a255-50f7005c0cb5","arxiv_id":"2607.28520","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":6.5,"correctness_risk":"low","formal_verification":"none","parameter_count":6,"one_line_summary":"CS-RNR certifies each candidate exploit by full-tree best response before atomic deployment, so model error can cost gain but not the reference-relative safety budget.","lead":"A game-playing agent can safely exploit a flawed opponent by only deploying counter-strategies whose downside it has measured itself with a full-tree best-response certificate. The method keeps every played strategy inside a user budget while earning several times the gain of prior safe-release rules on small poker-like games.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified","rationale":"Prop. 1 follows directly from the definition of exploitability and zero-sum payoffs once atomic certification is enforced; model error cannot break the bound by construction. The paper cleanly separates that strategy-level invariant from statistical detection (Assumption 1, pools, δ), and the experiments stress exactly the adaptive cases Prop. 1 covers. The only material limitation is the need for exact full-tree BR, which the manuscript scopes rather than conceals. That matches the reader’s weakest_assumption and does not justify moving off ACCEPT. A useful residual check is numerical confirmation of the η_v correction on the adversarial audit, not a challenge to soundness.","tokens_in":19002,"tokens_out":506,"duration_ms":32403,"concrete_test":"On the every-hand-BR stress trajectories (Fig. 4a), recompute exact game value v*_opp by a long CFR+ run, measure η_v = |êv_opp − v*_opp|, and verify hand-wise that realized expected gain ≥ −B̂_t − η_v within the stated 5e-3 solver tolerance; if any hand violates the corrected bound, Prop. 1’s finite-reference claim needs tightening.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central safety claim is Prop. 1: atomic deployment of a full-tree BR certificate B_t ≤ ε_max on the played strategy yields u_hero ≥ v*_hero − ε_max against any (including adaptive) opponent. That bound is nearly definitional once the executor plays the certified behavioural strategy without post-hoc modification; the finite-reference correction (ε_max + η_v) is stated explicitly. Detection, pools, δ, and Assumption 1 affect only which candidates are proposed and thus captured gain, not the deployment invariant. The reader’s scoped caveat—exact full-tree BR in solver-tractable games—is already the paper’s own limit (§2, §4.3, §7, Future Work) and does not undermine the claim as stated. Empirical support (shared-estimator ablations, 0/36k audit failures under stated tolerance, pre-registered confirmation-starvation checks) is consistent with the logic. No internal inconsistency or hidden load-bearing gap that would overturn ACCEPT was found.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper proposes CS-RNR, an online opponent-exploitation method for two-player zero-sum imperfect-information games. Anytime-valid confidence sequences flag pooled action frequencies that separate from an equilibrium reference; confirmed excesses define a conservative opponent model; restricted-response solves produce candidate strategies over a pin grid; and each complete candidate is certified by an original-game full-tree best response before atomic commit under a user budget ε_max. Proposition 1 states a runtime invariant: if the executor plays the certified behavioural strategy without post-certification modification and only admits B_t ≤ ε_max, then conditional expected hero value is at least v*_hero − ε_max against any opponent, including adaptive ones. Detection and modeling affect captured gain only. Empirically, on Leduc CS-RNR obtains 6.2× the steady-state gain of a money-verified binary gate with worst certificate ≤ 0.15; Fixed-Mix with the same estimator overshoots the budget by large factors; and all 36,000 audited hands across Leduc, Liar’s Dice, and 5-rank Leduc satisfy the reported tolerance. Confirmation starvation is identified as a distinct data-dependent limit.","tokens_in":19229,"tokens_out":1384,"duration_ms":44351,"significance":"If the result holds as stated, the paper cleanly relocates safety from response construction to a certificate on the played strategy, separating model error (which can cost gain) from deployment risk (which cannot). That separation is conceptually useful and is supported by a nearly definitional runtime invariant, shared-estimator ablations that isolate the schedule, adversarial audits including every-hand best responders, pre-registered concentrated and independently trained opponent suites, and exact full-tree evaluation of every reported gain. The work is scoped to solver-tractable games where exact certificates are available—an honest and load-bearing limit the authors state in §2, §4.3, §7, and Future Work—but within that scope the contribution is solid, reproducible in design, and of clear interest to the safe-exploitation and extensive-form game-solving communities.","major_comments":[{"comment":"§5 (Finite-reference correction) and Appendix A: the implementation certifies against a finite-iteration reference ẽv_opp, so the true-game guarantee is ε_max + η_v. The audit tolerance 5×10^{-3} covers CFR+ residue at 400 iterations, but the manuscript does not report a measured bound on |ẽv_opp − v*_opp| for the 4000-iteration Leduc reference (or the analogous references in the other games). A short quantification of η_v—or an explicit statement that the reported certificates already absorb the reference gap under the audit tolerance—would make the empirical claim “every deployed strategy within budget” line up tightly with Proposition 1’s exact-value statement.","section":"§5, Finite-reference correction; Appendix A"},{"comment":"§4.1–4.2 and Assumption 1: Proposition 1 is independent of detection, but the gain claims (6.2× gate, cross-game tables) depend on the hand-chosen 12-pool taxonomy, margin δ, and the within-pool homogeneity idealization used to inherit excesses to infosets. The horizon sweeps and 5-rank concentrated pre-registration already show confirmation starvation; it would strengthen the paper to state more explicitly that safe gain is jointly limited by (pool design, δ, T) and that alternative poolings were not systematically ablated. This does not undermine the safety invariant, but it bounds how far the headline multipliers should be read as method-intrinsic rather than taxonomy-dependent.","section":"§4.1 Detection; Assumption 1; §6.4–6.6"}],"minor_comments":[{"comment":"Figure 2 caption and §6.1: the offline frontier is central to motivating the schedule; stating the exact oracle-model protocol (same CFR+ iteration budget as online restricted solves?) in the caption would aid reproduction.","section":"Figure 2; §6.1"},{"comment":"Table 1 vs. §6.4: the main table uses stitched sub-Gaussian boundaries (2/12 diffuse releases at T=800) while the horizon sweep uses empirical-Bernstein (1/12 at T=800). The appendix notes this; a one-sentence pointer in the main comparison would prevent readers from treating the two release counts as the same detector.","section":"Table 1; §6.4; Appendix A"},{"comment":"Algorithm 3 and §4.3: the climb cap c=2 and checkpoint list are fixed without sensitivity analysis beyond the budget/pin sweeps in Table 2. A brief remark that the invariant does not require monotonicity in p or a globally maximal feasible pin (already implied by per-candidate certification) would clarify why the local search is sufficient.","section":"Algorithm 3; §4.3"},{"comment":"Related Work §2: the distinction from safe nested subgame solving (Brown & Sandholm 2017) and opponent-limited online search is present but compressed. One additional sentence stating that those methods bound re-solve loss relative to a blueprint, whereas CS-RNR certifies a whole-game deployed policy between hands, would sharpen the novelty claim.","section":"§2 Related Work"},{"comment":"Minor typography: “diffusedeviations”, “budget-constrained confidence-scheduled”, and similar missing spaces appear in the abstract/introduction PDF text; clean these in production.","section":"Abstract; §1"}],"recommendation":"minor_revision","confidential_remarks":"I agree with the reader’s ACCEPT-leaning assessment: Prop. 1 is essentially definitional once atomic play of a full-tree certificate is enforced, and the experiments are unusually careful (exact profile values, shared-estimator controls, pre-registration, 0/36k audit). I recommend minor_revision rather than accept only to force a short η_v quantification and a clearer statement that headline gain ratios are detector/taxonomy-dependent. No novelty or integrity concern. Scope is a good fit for a game-theory / multi-agent venue that accepts solver-tractable empirical work; the poker-scale certificate extension is correctly deferred."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The thing worth knowing is simple. They relocate the safety guarantee from how you build the response to a full-tree best-response certificate on the complete candidate before it is ever played. Prop 1 is nearly definitional once you enforce atomic commit of (σ_H, B_t) with B_t ≤ ε_max, but that is a feature: model error can only cost gain, never the bound. Against adaptive opponents the audit holds; Fixed-Mix with the same estimator blows the budget by 4–14× while CS-RNR stays inside. That separation is the paper’s real empirical point.\n\nWhat is new is the full online loop: anytime-valid pooled excesses → conservative model → restricted-response grid → exact BR certificate → atomic deploy. Restricted Nash responses, safe-exploitation budgets, and confidence sequences are prior; gluing them this way and measuring the certificate on the strategy that actually runs is not. The experiments are unusually careful for the area—shared-estimator ablations, pre-registered concentrated 5-rank checks that expose confirmation starvation instead of hiding it, cross-game transfer, 0/36k audited hands under stated tolerance.\n\nSoft spots are scoped and mostly owned by the authors. Exact full-tree BR and restricted solves only work in solver-tractable games; they say so repeatedly and put approximate certificates in future work. Pools, δ, and the pin grid are hand-chosen free parameters that govern whether diffuse leaks ever confirm. No public code. None of that overturns the claim as stated.\n\nThis is for people who care about safe opponent exploitation and online multi-agent methods in imperfect-information games. The math is thin but sound; the data and citation pattern are solid. I would bring it to reading group, cite it if I am working in the area, and send it to referees without hesitation.","headline":"Clean engineering move: safety lives on the played strategy via an exact BR certificate, not on the opponent model that proposed it.","tokens_in":19895,"tokens_out":467,"would_cite":true,"duration_ms":14773,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"An agent can safely exploit flawed opponents by certifying the strategy it actually plays, not the model that proposed it.","keywords":["opponent exploitation","safe exploitation","restricted Nash response","anytime-valid confidence sequences","imperfect-information games","exploitability certificate","Leduc hold'em","two-player zero-sum"],"falsifier":"Find a hand, seed, or adaptive opponent in the reported Leduc, Liar’s Dice, or 5-rank Leduc audits where the played strategy’s true expected loss exceeds its reported certificate plus the stated solver tolerance, or show that the same schedule on a game where only approximate best responses exist still claims the Proposition 1 bound.","tokens_in":19799,"feed_emoji":"🃏","tokens_out":1007,"duration_ms":19159,"temperature":0.7,"pith_summary":"In two-player zero-sum imperfect-information games, a Nash strategy locks in the game value but leaves money on the table against a flawed opponent. The hard case is diffuse deviation: slight mistakes spread across many decisions, so binary “release when sure” rules never fire, while a full best response to a half-built opponent model can be extremely exploitable. This paper introduces CS-RNR, which watches pooled action frequencies with anytime-valid confidence sequences, builds a conservative opponent model only from confirmed excesses, and turns that model into candidate counter-strategies via restricted-response solves at several pin levels. Before any candidate is played, the agent runs a full-tree best response on the complete strategy itself and commits only if the resulting certificate stays inside a user budget. Model quality then decides how much extra value is captured; the certificate alone bounds reference-relative expected loss, even against adaptive opponents. On Leduc the method earns several times the steady-state gain of a money-verified binary gate while every deployed strategy stays inside budget, and tens of thousands of audited hands across three games show no certificate violation.","feed_headline":"Agents that audit their own exploits before playing them","feed_subtitle":"A full-tree certificate on the deployed strategy keeps every hand inside a loss budget while still beating binary gates.","key_machinery":"Budget-constrained confidence-scheduled restricted responses (CS-RNR): anytime-valid confidence sequences flag pooled frequency excesses, a conservative model feeds restricted-response solves over a pin grid, and each complete candidate is admitted only after an original-game full-tree best-response certificate clears a user budget and is committed atomically with the strategy.","core_discovery":"CS-RNR is the first online opponent-exploitation method whose safety guarantee is a certificate computed on the strategy the agent actually deploys. Every committed candidate satisfies a full-tree best-response bound no larger than a chosen budget, so conditional expected value against any opponent—including adaptive ones—is at least the Nash floor minus that budget, while incomplete or wrong models can only reduce captured gain, never break the bound.","pith_inferences":["Any poker-scale extension will stand or fall on whether a sound approximate or depth-limited certificate can replace the exact full-tree best response while preserving the runtime invariant.","The same certify-the-played-strategy pattern could apply to other online adaptation settings where a planner proposes risky policies from partial models.","Pool design and observation mechanisms become first-class engineering choices: better pools shrink confirmation starvation without touching the safety argument.","Independently trained near-equilibrium opponents already show the ordering the paper predicts; larger unscripted populations would test whether the gain gap survives when leaks are not hand-designed."],"forward_implications":["Safety can be relocated from assumptions about the opponent model to a check on the played strategy, so model error costs gain rather than the loss bound.","A single budget dial interpolates continuously between pure Nash play and an unrestricted schedule while keeping every committed strategy inside the bound.","Trajectory mixtures that share the same estimator can match gain yet massively overshoot the same budget, so the behavioural restricted solve plus certificate is doing essential work.","Confirmation starvation—not certificate failure—is the binding limit on diffuse or rarely reached leaks, and longer horizons raise the number of confirmed opponents without breaking the budget.","In solver-tractable games the certificate pass is cheap enough to re-run at every checkpoint, making atomic certify-then-commit practical online."],"fun_headline_variants":["Agents that certify their own exploits before deployment","CS-RNR: full-tree certificates keep every exploit inside budget","Confidence sequences let agents audit exploits they actually play","Self-certified restricted responses beat binary gates in Leduc","Every deployed exploit carries its own full-tree safety certificate"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The safety proof needs exact full-tree best-response certificates and restricted solves to be available online, which the paper restricts to solver-tractable games.","fun_headline_variants_meta":{"raw":{"variants":["Agents that certify their own exploits before deployment","CS-RNR: full-tree certificates keep every exploit inside budget","Confidence sequences let agents audit exploits they actually play","Self-certified restricted responses beat binary gates in Leduc","Every deployed exploit carries its own full-tree safety certificate"]},"model":"grok-4.5","effort":"low","cost_usd":0.003269,"raw_usage":{"total_tokens":1198,"prompt_tokens":871,"num_sources_used":0,"completion_tokens":62,"cost_in_usd_ticks":32688000,"prompt_tokens_details":{"text_tokens":871,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":265,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":871,"tokens_out":62,"duration_ms":5664,"temperature":1.0,"reasoning_tokens":265,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-31T04:46:34.038858+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Find a hand, seed, or adaptive opponent in the reported Leduc, Liar’s Dice, or 5-rank Leduc audits where the played strategy’s true expected loss exceeds its reported certificate plus the stated solver tolerance, or show that the same schedule on a game where only approximate best responses exist still claims the Proposition 1 bound.","supporting_citations":[],"review_version":1}