{"id":"02c2128a-b048-4d98-bda2-89ca11d64b6f","arxiv_id":"2607.27035","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"Persistent randomized Weyl streams at each chance node cut MCCFR exploitability 19–34% on Kuhn and Leduc poker with local O(log N/N) frequency guarantees and no new hyperparameters.","lead":"A one-line change to how poker solvers sample card deals—using persistent low-discrepancy streams instead of fresh random draws—cuts final exploitability by roughly 20–34% on standard tabular poker benchmarks. It costs nothing extra to run and needs no new knobs, so it is immediately usable wherever Monte Carlo CFR is already the workhorse.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified beyond the reader's already-scoped open adaptive-convergence gap.","rationale":"The central empirical claim (paired 19–34% exploitability reductions on Kuhn/Leduc-family, CIs above zero, surviving to 3M touches and LCFR composition) is supported by the reported design and does not rest on a closed global adaptive proof. Local theory (marginal correctness, fixed-trajectory unbiasedness, deterministic O(log(N+1)/N) frequency error, reset guarantee) is standard and carefully scoped. The mechanism hypothesis is the weakest link, exactly as the reader states, but the paper maps its regime with revisit/symmetry/private-info diagnostics and reports nulls rather than over-generalizing. No additional concrete flaw (e.g., estimator bias that would invalidate the paired tables, or a missing assumption that breaks Thm. 1) surfaced on close reading. Verdict remains CONDITIONAL for the same reasons the reader gave; confidence in that assessment is high.","tokens_in":21522,"tokens_out":517,"duration_ms":10361,"concrete_test":"Re-run the Leduc 3M-touch paired protocol (App. H) while logging, per concrete chance node, the realized cumulative covariance between outcome indicators and the scalar regret contribution g_t along the adaptive path; if that covariance is non-negative on average in cells that still show the ~24% exploitability drop, the temporal-cancellation hypothesis would need revision even though the empirical claim stands.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The reader's weakest assumption already names the true soft spot: Theorem 1 bounds only unweighted local frequencies, while exploitability gains require favorable cumulative covariance between the persistent stream and adaptive downstream values g_t; global rates for non-reset CCS-MCCFR remain open (Prop. 3, App. A–B). That gap is real but is disclosed, bounded by the conditional scalar |B_c(t)| ≤ 2G δ_{c,t}, closed for the per-traversal reset variant (Thm. 2), and empirically stress-tested by persistence ablation, revisit/symmetry diagnostics, null HUNL/Liar's Dice/Flop boundaries, and composition with LCFR and a restricted control variate. I do not find a further load-bearing internal inconsistency, hidden hyperparameter, or unsupported overclaim that would move the verdict past the reader's CONDITIONAL.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper proposes CCS-MCCFR, a drop-in replacement for i.i.d. chance sampling in External Sampling MCCFR: each concrete chance node is assigned a persistent randomized Weyl stream whose phases are mapped through the node’s chance law. The authors prove fixed-index marginal correctness (Prop. 1), unbiasedness of External-Sampling regret estimates along fixed strategy trajectories (Prop. 2), deterministic O(log(N+1)/N) local frequency discrepancy for the first N draws at one concrete node (Thm. 1), a conditional total-variation bound on adaptive phase-selection bias (Prop. 3), and that a per-traversal phase-reset variant recovers the standard O(1/√T) External Sampling guarantee (Thm. 2). Empirically, paired experiments report 19–34% final exploitability reductions on Kuhn and a controlled Leduc deck family (all paired-bootstrap CIs above zero), a smaller significant gain on Goofspiel-4, persistence to 3M Leduc node touches, favorable composition with Linear CFR, and null or near-null endpoints on Liar’s Dice, reduced Flop, and four HUNL endgames, organized by revisit/symmetry/private-information diagnostics.","tokens_in":21661,"tokens_out":1809,"duration_ms":48380,"significance":"If the empirical pattern holds, this is a high-value practical contribution: a one-line, hyperparameter-free change to the chance sampler that yields double-digit exploitability reductions on standard tabular poker benchmarks with no measurable runtime cost, and that composes with Linear CFR to the best measured cell. The local theory is carefully scoped rather than overclaimed—classical discrepancy and External Sampling facts, no fitted free parameters, and an explicit conditional bias object δ_{c,t} for the adaptive case. Strengths include paired multi-seed designs with 10k-resample bootstrap CIs, an antithetic control, a persistence ablation, composition tests with update rules and a restricted control variate, and honest boundary diagnostics (HUNL, Liar’s Dice, low-revisit Flop). The main scientific limitation is that global rates for fully adaptive non-reset CCS-MCCFR remain open; that gap is disclosed and does not erase the local guarantees or the empirical result, but it does bound how far the paper can claim a complete theoretical replacement for i.i.d. chance sampling.","major_comments":[{"comment":"The central empirical claim rests on the temporal-cancellation hypothesis (§4–5): Thm. 1 controls only unweighted local frequencies, while exploitability reductions require favorable cumulative covariance between the persistent stream and adaptive downstream values g_t. Global convergence of non-reset CCS-MCCFR is left open (App. A–B); only fixed-trajectory unbiasedness, |B_c(t)| ≤ 2G δ_{c,t} (Prop. 3), and the reset variant (Thm. 2) are proved. This gap is disclosed, but the main text (esp. Intro/Conclusion) still frames CCS-MCCFR primarily as delivering “explicit local guarantees and large exploitability reductions” without a equally prominent main-body statement that the algorithm used in all positive experiments lacks a proven O(1/√T) guarantee. Please elevate a short, explicit caveat next to the main claims (abstract is already careful; §1 and §7 should match) so readers cannot mist","section":"§1, §4–5, §7; App. A–B"},{"comment":"Prop. 3 isolates adaptive bias in δ_{c,t}, but the manuscript never estimates or upper-bounds δ_{c,t} on the runs that produce the gains. Appendix H’s adaptive frequency diagnostic (visit-weighted max |p̂−p| after 30k iterations, seed 0) is a useful descriptive check on outcome frequencies, yet it is not a measurement of conditional phase law given g_t and reach, and ordinary count discrepancy need not control δ_{c,t} (as App. A notes). Because the load-bearing step from local balance to lower exploitability is exactly this adaptive coupling, please either (i) add a direct diagnostic of phase non-uniformity conditional on reach/downstream scalars on Kuhn/Leduc, or (ii) clearly state in §4–5 that no empirical bound on δ_{c,t} is provided and that favorable covariance is supported only indirectly (persistence ablation, Table 4; revisit table; null boundaries). Without one of these, the mec","section":"Prop. 3; §4–5; App. H; Table 4"},{"comment":"Table 5 and Figure 4 present revisit statistics as the main organizer of when CCS helps, but the diagnostic is run for 1000 vanilla iterations and then compared to final gains at very different budgets and solvers (OpenSpiel vs. standalone C++ HUNL). The reduced Flop row is explicitly a structural low-exposure probe, not full Flop Hold’em (App. A), and HUNL exploitability is within retained abstract endgames under a protocol-mixed archive (App. G). These caveats belong in the main experimental narrative when the paper claims the “operating regime” and “transfer boundary,” not only in appendices. Please qualify Table 5/Fig. 4 in §6 as descriptive correlates on this benchmark set, and avoid language that reads as a validated transfer predictor for full-scale Hold’em.","section":"§6; Table 5; Figure 4; App. A, E, G"}],"minor_comments":[{"comment":"Notation: N (per-node consumed draws), T (iterations), and node-touch budgets are distinguished in §4 but easy to conflate in figure captions and Table 1–3 headers. A single notation paragraph or consistent subscripting (N_c vs. T) in all tables would help.","section":"§4; Tables 1–3"},{"comment":"Figure 1 bands are mean±1.96 SEM over seeds, while significance is decided by paired-bootstrap intervals on differences (Fig. 2). State this dual convention once in the Fig. 1 caption to avoid readers treating SEM ribbons as paired tests.","section":"Figure 1–2; §6"},{"comment":"The antithetic baseline is a useful simple control; briefly define the exact pairing (u vs 1−u per node vs global) in §6 or App. C so the comparison is reproducible from the main text alone.","section":"§6; Table 1; App. C"},{"comment":"Related work correctly separates VR-MCCFR (estimator baselines) from CCS (temporal chance allocation). A single sentence on whether public chance sampling [Johanson et al., 2012] already correlates chance across infosets—versus CCS’s across-visit correlation at one concrete node—would sharpen the novelty claim.","section":"§2"},{"comment":"Typos/style: “V ariance” and “T emporal” appear with stray spaces in headings (§5, §4); “W eyl” similarly. Normalize heading capitalization.","section":"§4–5"},{"comment":"Appendix G’s mixed legacy/v2 HUNL protocol is handled carefully; still flag in the main HUNL paragraph that Subgame 1 mixes solver revisions so the four-endgame set is not a homogeneous scaling study.","section":"§6 HUNL paragraph; App. G"}],"recommendation":"minor_revision","confidential_remarks":"I agree with the reader/skeptic that there is no hidden inconsistency or undisclosed hyperparameter; the open adaptive-convergence gap is real but scoped. Minor revision is appropriate rather than major: the three major points are clarification and stronger mechanism diagnostics, not a rewrite of the algorithm or experiments. Fit for a serious GT/ML venue is good—practical impact on standard poker benchmarks is unusually large for a one-line sampler change. I would not block on absence of a full adaptive rate if the caveat is made unmistakable in the main body."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The thing worth knowing is simple: bind a persistent randomized Weyl stream to each concrete chance node in External Sampling MCCFR, and on Kuhn plus a controlled Leduc family you get 19–34% lower final exploitability in paired bootstrap tests, with no new hyperparameters and no timing cost. The gain holds at 3M Leduc touches and stacks with Linear CFR to the best cell they measured.\n\nWhat is new is the placement, not the math primitives. Randomized QMC and golden-ratio rotations are classical; putting one persistent stream per concrete chance node inside adaptive MCCFR, then measuring it properly, is the contribution. Local theory is careful and correctly scoped: fixed-index marginals, fixed-trajectory unbiasedness, deterministic O(log N/N) frequency error on the first N draws at a node, a conditional TV bound on the adaptive phase-selection term, and a per-traversal reset variant that recovers the usual O(1/√T) ES guarantee. Experiments are the real strength—paired seeds, bootstrap CIs, antithetic control, persistence ablation, revisit/symmetry diagnostics, and explicit nulls on Liar’s Dice, reduced Flop, and four HUNL endgames. They do not oversell transfer.\n\nThe soft spot is exactly the one they flag. Theorem 1 only controls unweighted local frequencies. Exploitability gains need favorable cumulative covariance with evolving downstream values; global rates for fully adaptive (non-reset) CCS remain open. That is a real gap, not a hidden one. They bound the scalar channel, close it for the reset variant, and map where the method helps versus where it does not. I would not call the mechanism proved; I would call the empirical claim well supported inside the revisit-heavy private-info regime they study.\n\nThis is for people who run tabular or abstracted MCCFR and care about free variance structure on the chance axis. It is not a deep-CFR or full-scale HUNL paper. Citations look appropriate; no circularity on the metric. I would send it to referees. Worth engaging if you touch poker solvers or sampling inside CFR.","headline":"A one-line persistent Weyl chance sampler that actually cuts tabular-poker exploitability 20–34% with honest local theory and clear regime boundaries.","tokens_in":22370,"tokens_out":510,"would_cite":true,"duration_ms":9841,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"A one-line change to MCCFR’s chance sampler—persistent randomized Weyl streams at each chance node—cuts final exploitability by about 19–34% across Kuhn and Leduc poker with no new hyperparameters.","keywords":["MCCFR","chance sampling","low-discrepancy sequences","Weyl sequence","exploitability","imperfect-information games","External Sampling","tabular poker"],"falsifier":"Run the same paired i.i.d.-versus-CCS protocol on a high-revisit private-information poker game at a large node-touch budget; if the paired-bootstrap interval on relative exploitability reduction includes zero or turns negative while local frequency error still tracks O(log N / N), the claimed link from local balance to lower exploitability fails.","tokens_in":22292,"feed_emoji":"🃏","tokens_out":1061,"duration_ms":22447,"temperature":0.7,"pith_summary":"Monte Carlo Counterfactual Regret Minimization repeatedly samples chance events such as card deals while strategies evolve, usually drawing independently on every visit. This paper argues that those repeated draws should instead be taken from a persistent low-discrepancy stream at each concrete chance node, so the node’s outcomes are forced to cover its distribution more evenly over time. The construction keeps every fixed-index draw correctly distributed and leaves the External Sampling estimator and regret updates unchanged, yet it guarantees much tighter local frequency error than independent sampling. In paired experiments the method lowers final exploitability by roughly one-fifth to one-third on Kuhn and several Leduc variants, with the gain holding at multi-million node-touch budgets and stacking with Linear CFR. A sympathetic reader cares because equilibrium computation in imperfect-information games is bottlenecked by sampling noise, and this is presented as a free, drop-in way to reduce that noise wherever chance nodes are revisited under private information.","feed_headline":"One-line chance sampler cuts poker exploitability up to 34%","feed_subtitle":"Persistent Weyl streams at each chance node beat i.i.d. draws in MCCFR with no extra knobs or runtime cost.","key_machinery":"Correlated Chance Sampling via persistent randomized Weyl streams: each concrete chance node keeps a phase ϕ and visit count N, draws u = (ϕ + N·g) mod 1 with g the golden-ratio conjugate, and maps u through the node’s quantile function. That stream supplies the local discrepancy bound and the cross-visit temporal structure the authors credit for lower exploitability.","core_discovery":"CCS-MCCFR replaces independent chance draws with one persistent randomized Weyl stream per concrete chance node, mapped through that node’s chance law. Each fixed-index draw remains correctly distributed, the first N draws at a node have deterministic frequency error O(log(N+1)/N) rather than the usual i.i.d. scale, and along fixed strategy trajectories the regret estimates stay unbiased. Empirically this yields large, statistically detected exploitability reductions on tabular poker and a smaller but significant gain on Goofspiel-4, with no measurable time cost and no new hyperparameters.","pith_inferences":["Any other Monte Carlo tree procedure that revisits fixed chance locations—continual resolving, search-time sampling, or model-based rollouts—could try the same per-node persistent stream without touching its value estimator.","If δ_{c,t} (adaptive phase-selection distance from uniform) can be shown to shrink with visit count, the open global-convergence gap for non-reset CCS-MCCFR would close along the paper’s own bound.","The same placement idea extends naturally to higher-dimensional chance (multi-card deals) via randomized nets or lattices once per-node visit counts stay large enough for discrepancy to matter."],"forward_implications":["A one-line chance-sampler swap can deliver double-digit exploitability cuts on Kuhn and Leduc without changing the regret estimator or adding hyperparameters.","Local frequency error at a repeatedly visited chance node can be driven to O(log(N+1)/N) instead of the i.i.d. Θ(N^{-1/2}) scale.","CCS-MCCFR stacks with Linear CFR and Discounted CFR; LCFR plus CCS reaches the lowest measured Leduc cell in the composition grid.","Per-traversal phase reset restores the standard O(1/√T) External Sampling convergence guarantee when a global adaptive proof is not yet available.","Gains concentrate where chance nodes are revisited under private-information coupling; low-revisit or highly symmetric chance structures show little or no endpoint improvement."],"fun_headline_variants":["Weyl streams at chance nodes cut poker exploitability up to 34%","Correlated sampling replaces i.i.d. draws, cuts exploitability 34%","One-line CCS-MCCFR change drops exploitability up to 34% in poker","Persistent chance streams beat i.i.d. in MCCFR by up to 34%","CCS-MCCFR trims poker exploitability 19-34% with no extra cost"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"That evening out unweighted outcome frequencies at a chance node will, under fully adaptive regret updates, produce favorable cancellation with evolving downstream values often enough to lower exploitability.","fun_headline_variants_meta":{"raw":{"variants":["Weyl streams at chance nodes cut poker exploitability up to 34%","Correlated sampling replaces i.i.d. draws, cuts exploitability 34%","One-line CCS-MCCFR change drops exploitability up to 34% in poker","Persistent chance streams beat i.i.d. in MCCFR by up to 34%","CCS-MCCFR trims poker exploitability 19-34% with no extra cost"]},"model":"grok-4.5","effort":"low","cost_usd":0.005645,"raw_usage":{"total_tokens":1602,"prompt_tokens":883,"num_sources_used":0,"completion_tokens":113,"cost_in_usd_ticks":56448000,"prompt_tokens_details":{"text_tokens":883,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":606,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":883,"tokens_out":113,"duration_ms":9891,"temperature":1.0,"reasoning_tokens":606,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-30T13:08:44.383245+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Run the same paired i.i.d.-versus-CCS protocol on a high-revisit private-information poker game at a large node-touch budget; if the paired-bootstrap interval on relative exploitability reduction includes zero or turns negative while local frequency error still tracks O(log N / N), the claimed link from local balance to lower exploitability fails.","supporting_citations":[],"review_version":1}