{"id":"3cf0dc26-0f44-47eb-baf0-510717b99609","arxiv_id":"2608.09586","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":3,"one_line_summary":"A strategy built on continuation values computed from a finite tournament model beats a strategy built on analytic ICM by $214.33 per hand in matched comparisons, and the gap persists under rollout, alternative evaluators, and opponent changes.","lead":"In a simplified three-player poker tournament, the authors show that the standard chip-to-money formula (ICM) gives up about $214 per hand in prize equity when used to choose plays, compared with a strategy built on continuation values computed from the tournament itself. The paper offers a clean, reusable protocol for measuring when an approximate value model is too inaccurate to drive decisions, a question that matters beyond poker.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The abstract's '$938.03 per hand' rollout figure is mislabeled: the full-rollout mean is per tournament, not per hand, overstating the per-hand advantage by roughly 4.4x.","rationale":"The reader's weakest_assumption concerned the external validity of the three-player jam/fold model, which the paper itself acknowledges in Appendix B.9 as an empirical question its design cannot answer. I find a different, more concrete issue: the abstract's headline rollout number is mislabeled. Section 9 and Table 11 correctly state the full-rollout mean difference of $938.03 per tournament, but the Abstract and Conclusion call it 'per hand.' The arithmetic shows the per-hand advantage in the rollout is approximately $214, matching the one-hand census mean; the 4.38x multiple equals the average number of three-player hands per tournament. This matters because readers and the reader's strongest_claim rely on the abstract's phrasing to compare magnitudes. The error does not undermine the central claim that SCO beats fixed ICM, nor the matched-comparison evidence, but it must be fixed for the paper to report its results accurately. The paper's other robustness checks, including alternative continuation endpoints, the shared-deck Monte Carlo evaluator, matched-tolerance re-solves, and the full rollout with exact pairing, are strong and support the qualitative conclusion. Conditional acceptance is appropriate pending correction of the unit label in the Abstract, Introduction, and Conclusion.","tokens_in":28328,"tokens_out":13284,"duration_ms":122188,"concrete_test":"Recompute the full-rollout mean per hero three-player hand: take the summed paired differences across all 56,760,000 replicates and divide by the number of hero three-player hands (approximately 248,217,230, half of 496,434,460). If the result is ≈$214 rather than $938, the '$938.03 per hand' phrasing in the Abstract is a mislabel and should be corrected to '$938.03 per tournament' or 'per starting state-owner unit.'","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's Abstract and Conclusion state that when the tournament is played out to a finish, the advantage is '$938.03 per hand.' But Section 9 and Table 11 report the full-rollout 'mean difference of $938.03' without per-hand units, and Table 11's caption describes it as 4.38 times the census mean. The math confirms the label is wrong: 2,838 units × 20,000 paired replicates = 56,760,000 tournaments; with 496,434,460 three-way hands total, each tournament has ~4.37 three-player hands on average. Thus $938.03 / 4.37 ≈ $214.6 per hero three-player hand, essentially the same as the one-hand census mean of $214.33. The 4.38x multiple is the average number of three-player hands, not growth in the per-hand advantage. The abstract and conclusion therefore overstate the per-hand rollout gain by ~4.4x. This is a concrete unit error in the headline, even though the direction and statistical significance of the result are not in question.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes Strategic-Continuation Optimization (SCO), which freezes a current-hand jam/fold policy obtained by optimizing against a continuation table computed from a finite three-player tournament model, and compares it with a fixed-ICM policy built with the same optimizer and analytic ICM. In a complete census of the 946 ordered stack states at T = 45, it reports $9,036 mean absolute ICM value error, a 14.08% average relative jam-range change, and a matched mean prize-equity gain of $214.33 per hand for SCO over ICM across 2,838 state-owner units. Robustness experiments include six continuation endpoints, a shared-deck Monte Carlo evaluator, two LLM opponent profiles, a family of threshold opponents, three chip depths, two prize ladders, and a full-tournament rollout. The paper's central empirical claim is that ICM's value errors change policy and that the policy change costs prize equity.","tokens_in":28474,"tokens_out":13764,"duration_ms":122505,"significance":"If the measurements hold, the paper makes a useful contribution by separating value-model accuracy from control quality, and by giving an exact, verifiable decomposition of the policy cost into positional and level terms. The complete-domain census, the machine-checked identity in Eq. (11), the reproducibility artifacts, and the multiple independent evaluators are notable strengths. The finite three-player jam/fold setting is small, and the authors correctly acknowledge in Appendix B.9 that generalization to larger games is an open empirical question. My concerns are limited to two load-bearing presentation issues: the full-rollout unit error and the overstatement about threshold opponents in the abstract.","major_comments":[{"comment":"The headline claim that the full-tournament rollout advantage is \"$938.03 per hand\" is a unit error. Section 9 reports a mean difference of $938.03 per tournament (per paired replicate starting from a state-owner unit), not per hand. The paper itself reports that the three-player stage lasts 4.37 hands on average and that the rollout mean is 4.38 times the one-hand census mean; dividing $938.03 by 4.37 gives approximately $214.6 per three-player hand, essentially the $214.33 census mean. The 4.38 factor is the average number of three-player hands per tournament, not a multiplicative gain in per-hand advantage. The Abstract, Contribution 2, Section 9, Conclusion, and Table 11 caption need to state the unit explicitly and correct the per-hand label.","section":"Abstract; Section 9; Table 11; Conclusion; Contribution 2"},{"comment":"The abstract's statement that the ordering \"survives replacing the solver-built opponent with ... a family of non-modeling threshold players\" is not supported for the full family. Table 7 reports negative mean gains of -$100.80, -$92.63, and -$25.25 per hand for the jam-top-10%, jam-top-20%, and jam-top-30% anchors, with the crossover between 30% and 40%. The main text is honest about this, but the abstract and conclusion should be qualified to say the ordering survives for the solver, LLM, and loose-to-moderate threshold opponents, with a documented negative region against tight mechanical opponents.","section":"Section 7 and Table 7; Abstract"}],"minor_comments":[{"comment":"In the paragraph \"The mean is small; the underlying gap is not,\" the paper states that $214.33 is 0.0257% of the $333,333 equal-seat value; the correct percentage is approximately 0.0643% (214.33 / 333,333).","section":"Section 5"},{"comment":"The primary census is scored with the same continuation table CSC used to construct the SCO arm, and Proposition 1 makes the sign of the matched difference a near-mathematical consequence of ε-optimality under that continuation. The paper acknowledges this in the Proposition 1 discussion, but the main-text narrative would benefit from an explicit sentence before the headline results stating that the independent external evidence is the shared-deck evaluator, the alternative endpoints, and the full-rollout experiment.","section":"Section 4, Eq. (4) discussion"},{"comment":"The caption should give the unit of the mean gain (per tournament) and provide the per-three-player-hand conversion so that the table cannot be misread as per-hand.","section":"Table 11"}],"recommendation":"major_revision","confidential_remarks":"The empirical program is careful and the reproducibility apparatus is a genuine strength. My main reservation is the headline unit error, which must be corrected; the threshold-opponent overstatement in the abstract should also be fixed. Once those are addressed, the paper is very close to acceptable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read this one. The core result is real: a complete 2,838-unit matched comparison between a policy built against strategic continuation values and one built against analytic ICM, in a three-player jam/fold tournament, with the SCO policy ahead by $214.33 per current-hand unit and favored in 85.7% of units. The design is unusually disciplined: identical opponents and evaluator, frozen policies, exact seat/level decomposition with zero residual, matched-tolerance re-solves, six continuation endpoints, a shared-deck Monte Carlo check, and a shipped verifier that recomputes the headline numbers from raw tensors. That is genuinely new and genuinely well executed.\n\nThe soft spots are real but mostly acknowledged. The game is small by design, and the paper itself says the seat/level split may not transfer to more players or streets. The main census scores both arms with the same strategic table used to construct the winning arm, but the external evaluators and the sign-preserving endpoint checks blunt that circularity; it is a weakness, not a fatal one. The negative results against tight threshold opponents are reported honestly, and the LLM section is clearly labeled as a fixed conditional sample.\n\nThe one thing that needs fixing is a concrete unit error in the headline. The abstract and conclusion say the full-rollout advantage is $938.03 per hand. Section 9 and Table 11 report a mean difference of $938.03 with no per-hand unit, and the arithmetic confirms it is per tournament, not per hand. With 2,838 units times 20,000 paired replicates, each tournament averages about 4.37 three-player hands, so $938.03 / 4.37 ≈ $214.6 per three-player hand — essentially the one-hand census mean. The 4.38x multiple is the average number of three-player hands, not growth in the per-hand advantage. The direction and statistical significance are untouched; the label is wrong by a factor of about 4.4. That must be corrected in the abstract, the conclusion, and Table 11.\n\nMy verdict: this deserves a serious referee. The evidence is strong, the artifact bundle is best-in-class, and the unit fix is straightforward. I would bring it to a reading group and cite the matched-census methodology even before revision.","headline":"Solid matched-cost census showing ICM loses money as a controller; the result survives robust checks, but the $938.03 rollout figure is per tournament, not per hand, and the abstract/conclusion need that fix.","tokens_in":29085,"tokens_out":1503,"would_cite":true,"duration_ms":27672,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"ICM, the standard poker tournament equity model, reads only stack sizes and misses seat, blind, and elimination-pressure effects; a policy built on computed continuations beats a fixed-ICM policy by $214.33 per hand in a three-player…","keywords":["Independent Chip Model","tournament poker strategy","continuation values","jam/fold equilibrium","policy evaluation","prize equity","value-based control","matched comparison"],"falsifier":"Re-run the matched comparison in a four-player jam/fold game with the same protocol at comparable depth and blinds; if the SCO-minus-ICM per-hand gain collapses to zero or reverses, or the level term turns negative, the claim that ICM is systematically inadequate as a tournament objective would fail to generalize beyond the three-player model. A second, sharper check: add a post-flop action to the three-player game and recompute the census; if the policy difference no longer favors SCO in a majority of state-owner units, the result depends on the all-in/fold restriction.","tokens_in":1986,"feed_emoji":"🃏","tokens_out":2656,"duration_ms":78061,"temperature":0.7,"pith_summary":"The paper claims that the Independent Chip Model, the standard translator from tournament chips to prize money, is reliable enough as a description of what a stack is worth but fails as an objective for choosing actions, because it sees only stack sizes and misses seat rotation, blind obligations, and the threat a covering stack exerts on short stacks. To show the failure is real, the authors build Strategic-Continuation Optimization (SCO), which prices successor states with continuation values computed from the finite tournament model, and compare the frozen policy it induces with a policy built the same way but priced by analytic ICM. In a three-player all-in/fold tournament with a $1,000,000 pool, the SCO policy earns $214.33 more prize equity per hand across all 2,838 state-owner units, and $938.03 per hand when the tournament is played out to a finish. A seat-symmetrization decomposition attributes $130.43 of the per-hand gain to knowing which seat holds which stack and $83.90 to the tournament's own chip dynamics even when seats are averaged away. The intended takeaway is that a value model used to drive decisions should be judged by the decisions it induces, not by how close its numbers land.","feed_headline":"ICM poker pricing loses $214 a hand to computed continuations","feed_subtitle":"A matched tournament experiment shows the standard chip-equity model misses seat and chip-dynamics value.","key_machinery":"The load-bearing object is the continuation table: a lookup that maps each successor stack vector to the expected prize money each surviving player will collect, standing in for all future hands. SCO optimizes the current-hand jam/fold game against the strategic continuation table $C_{SC}$ computed by value-iteration with fictitious play from the finite tournament model, while the comparison policy $\\pi_{\\mathrm{ICM}}$ solves the identical current-hand game against the analytic Malmuth-Harville table $C_{\\mathrm{ICM}}$, which reads only the stack vector. The matched-evaluation protocol replaces only the focal player's frozen policy rows, holds opponents and the common evaluator fixed, and prices both arms with the same $C_{SC}$, so the paired difference $\\Delta PE$ in Equation (4) isolates the pricing difference. A seat-symmetrization operator that averages the strategic table over the six seat assignments of each chip multiset splits the gain exactly into a positional term and a level term. The formal boundary is Proposition 1: as long as $\\pi_{SCO}$ is $\\varepsilon$-optimal in the current-hand game priced by $C_{SC}$, its matched advantage over any comparison policy is bounded below by $-\\varepsilon$; the magnitude and coverage are then empirical.","core_discovery":"The central discovery is that ICM's value error is a control error, not just an estimation error: averaging over the complete census of 946 states and three seats, fixing the policy optimizer and changing only the continuation pricing makes the SCO policy overcome the fixed-ICM policy in 2,433 of 2,838 matched units, with a mean gain of $214.33 per hand. The error chain is measured layer by layer: analytic ICM deviates from the frozen strategic-continuation benchmark by $9,036 mean absolute value error; those value differences move the induced jam frequency by an average of 14.08% relative to each decision point's own ICM jam range, and by 32.42% at the button; and the resulting matched prize-equity difference is $214.33. Because only the focal player's policy changes between the two arms while opponents and the evaluator stay fixed, the difference is attributable to the pricing, not to the scoring. The paper also shows that a seat-blind repair is not enough: a symmetrized strategic table that averages over seat assignments still beats ICM by $83.90 per hand, so the problem is not only position blindness but how chips are discounted by the tournament's own dynamics.","pith_inferences":["A testable extension suggested by the level term: add a fourth player or a post-flop street and re-run the same matched design; because ICM's proportional-elimination approximation and its blindness to future blind obligations are structural, the same positive sign should persist, but the paper's specific numbers cannot be assumed to transfer.","The seat/level decomposition is a general diagnostic recipe: for any static value function driving a controller in a sequential game, symmetrize over the state variable the model omits and price the remainder, converting 'how wrong is the value' into 'how much does acting on it cost.'","The sign reversal when both arms are scored with analytic ICM shows that an evaluator aligned with one arm biases the comparison; a neutral, simulation-based evaluator is the honest common ground for policy comparison.","Because the full-rollout gain is 4.38 times the one-hand census gain and the per-seat spread disappears, the one-hand matched protocol likely understates the practical value of the improvement over a complete tournament."],"forward_implications":["Fixed-ICM strategies leave prize equity on the table in the modeled three-player jam/fold tournament: $214.33 per hand in the matched one-hand census and $938.03 per hand when the tournament is played out to a finish.","Two separable defects in ICM are priced: seat and position blindness costs $130.43 per hand, and a seat-blind but tournament-computed table still beats ICM by $83.90 per hand, so fixing only position misses a real part of the error.","The advantage holds when opponents are replaced by two LLM-driven profiles and by a family of non-modeling threshold players, crossing zero only against very tight opponents that jam 10-30% of hands, which places the benefit in the competent-opponent regime.","Value functions used as controllers should be compared by induced policy differences and matched downstream cost, not by mean absolute value error, because a large common-level error can be action-irrelevant while a small successor-contrast error flips a decision.","The full-tournament rollout raises the gain to $938.03 per hand and flattens the per-seat asymmetry, since seats rotate over a tournament, even though the per-unit sign is less stable than the aggregate mean."],"supporting_citations":[{"why":"Supplies the sequential-ranking formula that defines the analytic ICM continuation table used to build the comparison policy.","marker":"[19]"},{"why":"Named with [19] as the source of the Malmuth-Harville table used as the fixed-ICM continuation.","marker":"[32]"},{"why":"Provides the value-iteration/fictitious-play machinery that computes the strategic continuation table from the finite tournament model.","marker":"[11]"},{"why":"Earlier three-player jam/fold equilibrium computation whose state enumeration and successor structure the paper's complete census builds on.","marker":"[10]"},{"why":"Analyzes ICM as an estimator of prize equity, the property the paper separates from its quality as a controller.","marker":"[9]"},{"why":"One of the two LLM configurations used as solver-external opponents in the robustness experiment.","marker":"[16]"}],"fun_headline_variants":["ICM chip model loses $214 per hand; new strategy wins 86% of matchups","Computed continuations beat ICM by $214 per hand in poker","ICM's $9K value error becomes $214 per-hand loss in tourneys","Seat-blind ICM costs $214 per hand; computed continuations win"],"cache_read_input_tokens":31232,"weakest_assumption_plain":"The three-player, all-in/fold, 45-chip, 169-hand-class tournament is a faithful enough stand-in for real tournament poker that the measured ICM disadvantage transfers beyond the model; the paper's own appendix flags that more players, more streets, or richer action sets could change the seat/level split.","fun_headline_variants_meta":{"raw":{"variants":["ICM chip model loses $214 per hand; new strategy wins 86% of matchups","Computed continuations beat ICM by $214 per hand in poker","ICM's $9K value error becomes $214 per-hand loss in tourneys","Seat-blind ICM costs $214 per hand; computed continuations win"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000698,"raw_usage":{"total_tokens":3250,"prompt_tokens":1140,"completion_tokens":2110,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":756,"completion_tokens_details":{"reasoning_tokens":2024}},"tokens_in":756,"tokens_out":2110,"duration_ms":17692,"temperature":1.0,"reasoning_tokens":2024,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T14:34:28.161060+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the matched comparison in a four-player jam/fold game with the same protocol at comparable depth and blinds; if the SCO-minus-ICM per-hand gain collapses to zero or reverses, or the level term turns negative, the claim that ICM is systematically inadequate as a tournament objective would fail to generalize beyond the three-player model. A second, sharper check: add a post-flop action to the three-player game and recompute the census; if the policy difference no longer favors SCO in a majority of state-owner units, the result depends on the all-in/fold restriction.","supporting_citations":[{"cited_title":"Two Plus Two Publishing LLC, 1999","cited_arxiv_id":null,"evidence_quote":"Named with [19] as the source of the Malmuth-Harville table used as the fixed-ICM continuation."},{"cited_title":"Computing equilibria in multiplayer stochastic games of imperfect information","cited_arxiv_id":null,"evidence_quote":"Provides the value-iteration/fictitious-play machinery that computes the strategic continuation table from the finite tournament model."},{"cited_title":"Computing an approximate jam/fold equilibrium for 3-player no-limit texas hold’em tournaments","cited_arxiv_id":null,"evidence_quote":"Earlier three-player jam/fold equilibrium computation whose state enumeration and successor structure the paper's complete census builds on."},{"cited_title":"Gambler’s ruin and the icm.Statistical Science, 37(3): 289–305, 2022","cited_arxiv_id":null,"evidence_quote":"Analyzes ICM as an estimator of prize equity, the property the paper separates from its quality as a controller."}],"review_version":1}