REVIEW 4 major objections 6 minor 8 references
The Curvature Shadow: An Apparent Failure of Maximum-Entropy Equilibrium Selection is a Removable Artifact
T0 review · 4 major / 6 minor · reviewed 2026-08-01 · deepseek-v4-flash
Pith's one-line read This paper claims that the apparent failure of maximum-entropy equilibrium selection in Kuhn poker is a removable artifact: the 0.0205 coordinate gap is the curvature-amplified image of a 0.00083 entropy shortfall, not a genuine selection b
desk verdict A self-aware diagnostic that explains the Kuhn gap via curvature and shortfall, but the evidence reduces to one nonzero point and the η-sweep cannot rule out a moving objective. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central identity is a second-order Taylor decomposition of the entropy landscape along the Nash segment: gap ≈ sqrt(2δ/κ), where δ = H⋆ − H(σ(c)) is the entropy shortfall from the maximum-entropy member, κ = −d²H/dc² at the peak is the local peak curvature, and gap is the coordinate distance between the solver's landing point and the maximum-entropy coordinate. Its work is to separate the two causes of any apparent selection failure: why the solver stops short of maximum entropy (δ) and how strongly that shortfall is amplified into a visible coordinate offset (κ). The identity also ensures that a solver landing exactly on the Nash segment has no off-manifold contribution to the gap.
What would settle it
Construct a sequential game whose Nash segment has a sharply curved entropy peak (κ several times larger than Kuhn's, e.g., κ ≈ 20) and where the solver leaves a measurable entropy shortfall near δ ≈ 1e-3. The law predicts a gap of approximately sqrt(2δ/κ) ≈ 0.010 times the appropriate factor, approaching exactness as the magnet weakens; observing a gap that persists near 0.02 regardless of κ or that fails to track sqrt(2δ/κ) as δ varies would settle against the paper's central claim.
Extended reading notes
Core claim
On a one-dimensional Nash segment, restricting the mean Shannon entropy to the segment and expanding around its interior maximum gives gap ≈ sqrt(2δ/κ), where δ is the solver's entropy shortfall and κ is the curvature of the entropy landscape at the peak. In Kuhn poker the measured shortfall is δ = 8.3e-4 and the measured curvature is κ ≈ 4.0, which predicts a gap of 0.0204, matching the observed 0.0205 gap to within 2e-4. The four matrix games studied have δ ≈ 0 and therefore no gap, even though two of them have flatter peaks than Kuhn's; only the sequential game leaves a nonzero shortfall. Weakening the magnet strength η drives δ toward zero and the gap down the predicted sqrt(2δ/κ) curve,
Load-bearing premise
The load-bearing premise is that the selection target is the member maximizing unweighted mean Shannon entropy over the Nash segment; if the solver's regularization actually targets a different objective—such as reach-weighted or Tsallis entropy—then δ is not a shortfall from the true target and the 'removable artifact' conclusion collapses.
Editorial extensions
If this is right
- The Kuhn poker gap is explained as a removable shortfall rather than a counterexample to maximum-entropy selection, so the I-projection account of solver-dependent equilibrium selection is upheld up to a flatness-limited residual.
- Across the five tested games, the gap law holds to within 2e-4, meaning the same decomposition applies to both matrix and sequential games with one-dimensional Nash segments.
- In matrix games the solver reaches the maximum-entropy member exactly, so flatness alone never creates a gap; a visible offset requires both a positive entropy shortfall and a sufficiently flat peak.
- A causal sweep of the magnet strength shows the gap is removable: as δ → 0, the gap follows the predicted square-root curve toward zero, with no fixed residual, until dynamics destabilize at a stability floor.
- The paper makes a falsifiable prediction: a sequential game with a sharply curved entropy peak should show a tighter match to the maximum-entropy member, scaling like sqrt(κ_Kuhn/κ) relative to Kuhn.
- The paper explicitly flags an objective-dependence limitation: the maximum-entropy target depends on the entropy used, and the analysis fixes unweighted mean Shannon entropy; if the solver's true regularizer targets a different entropy, the shortfall framing would need revision.
Reading between the lines
- The paper hypothesizes but does not prove that reach-weighted counterfactual updates cause the shortfall in sequential games; a direct test would be to construct sequential games with increasing reach skew and see whether δ grows with skew while still following the sqrt(2δ/κ) curve.
- The moving-target caveat for Tsallis entropy suggests that any family of regularizers indexed by an entropy parameter should be compared against its own matched maximum-entropy target; comparing against the Shannon target alone would manufacture a spurious curvature dependence.
- Because the stability floor prevents reaching δ = 0 directly, the zero-gap conclusion is a limit statement; an independent route would be to find a sequential game with a sharp peak and a nonzero shortfall, converting the curvature half of the law into a directly observed causal sweep.
- A practical implication for reporting solver behavior is that coordinate distance alone is misleading: two games with identical entropy shortfalls can look very different if one peak is flat and the other sharp, so reporting δ alongside coordinates would make such artifacts transparent.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper examines a single apparent counterexample to maximum-entropy equilibrium selection in regularized solvers: R-NaD in Kuhn poker converges to a Nash equilibrium at bluff coordinate 0.180 while the maximum-entropy member is at 0.201, a gap of about 0.021. The authors derive a local quadratic relation gap ≈ sqrt(2δ/κ), where δ is the entropy shortfall relative to the maximum-entropy member and κ is the curvature of the entropy landscape at its peak. They verify this relation across five games, finding that four matrix games have δ≈0 and therefore no gap, while Kuhn has δ=8.3e-4 and a predicted gap of 0.0204 against a measured 0.0205. A magnet-strength sweep on Kuhn reduces δ and the gap together with a fitted log-log slope of 0.5013, which the authors interpret as strongly supporting the hypothesis that the gap is a removable shortfall rather than a fixed selection bias. They conclude that the I-projection account survives the Kuhn counterexample up to a flatness-limited residual.
Significance. If the conclusion were established, it would resolve a nagging anomaly in the empirical study of equilibrium selection in regularized dynamics and would sharpen the distinction between shortfall-induced and curvature-induced offsets. The paper is valuable in several respects: it isolates a clean geometric quantity, the curvature κ, that converts a tiny entropy shortfall into a visible coordinate offset; it provides a fully reproducible, exact-tabular engine with no sampling noise; it honestly lists limitations, including the absence of a causal curvature sweep and the objective-dependence of the target; and it warns convincingly against a naive Tsallis-entropy experiment where the target itself moves. However, the central evidentiary weight is much weaker than the prose suggests: the relation is a Taylor identity, the nonzero part of the empirical test rests on a single game, and the objective-dependence caveat may undermine the 'removable artifact' interpretation rather than merely bound it. The paper is a useful diagnostic note, but as it stands it does not establish that the Kuhn gap is a shortfall relative to R-NaD's actual objective.
major comments (4)
- [§3, Eq. (4); §5.2, Table 2] The relation gap ≈ sqrt(2δ/κ) is a tautological Taylor expansion, not an empirically falsifiable law. δ is defined as H*−H(σ(c)) and gap as |c−c*|, both read from the same unweighted mean-entropy function on the same Nash segment. Therefore any on-manifold point near c* satisfies the relation by construction. The agreement in Table 2 is a consistency check, not independent evidence. The four matrix-game rows have δ≈0 and hence trivially satisfy the relation; only Kuhn provides a nonzero point, and that single point lies on the curve by definition. The statement in §3 that verifying (4) 'empirically confirms that its only deviation from c* is the entropy shortfall' is circular: the coordinate c was chosen precisely as the segment coordinate, and the converged profile was verified to coincide with the family member. This does not establish which objective the solver is optimizing.
- [§5.3, Table 3, Fig. 3] The magnet sweep and the fitted exponent 0.5013 do not discriminate H-flat from H-bias. If c(η) remains on the smooth segment and δ(η)→0, then continuity forces gap(η)→0 and the log-log slope must approach 1/2 regardless of whether the finite-δ deviation is a genuine optimization shortfall or a systematic difference between two distinct objectives (unweighted Shannon vs. reach-weighted entropy). A moving-target model in which the effective target approaches c* linearly in η also predicts gap∝η and δ∝η², hence the same slope. Thus the claim that the test comes out 'decisively for H-flat' is unsupported. Moreover, the alternative H-bias as defined in §1 — a nonzero gap as δ→0 — is incompatible with the quadratic expansion itself, so it is not a meaningful competing hypothesis within this framework.
- [§2, Eq. (2); §6; §7] The objective-dependence problem is load-bearing, not a mere limitation. R-NaD's update (2) uses reach-weighted counterfactual values, and Section 6 hypothesizes that this reach-weighting 'distorts the moving-reference QRE sequence away from the unweighted maximum-(mean-)entropy member.' If that hypothesis is correct, then at finite η the measured δ is not a shortfall from R-NaD's actual target; it is the systematic gap between two different selection targets. Because the η-sweep cannot reach η=0 (the stability floor at η≈0.15 is an observed obstruction, not a limit point), the paper never demonstrates that the reach-weighted target coincides with the unweighted maximum-entropy member in the limit. Without independent evidence — for example, computing the reach-weighted max-entropy member and comparing δ to the distance to that target, or testing a correctly matched Tsallis target — the
- [§5.2, §7, §8] The paper's empirical support for the general claim rests on a single sequential exemplar. Only Kuhn has δ>0; all matrix games have δ≈0 and thus provide no information about the gap law for nonzero shortfalls. The paper acknowledges this ('Single sequential exemplar') but the conclusion and title generalize from this one case. The falsifiable prediction in §8 for a sharply curved sequential game is precisely the experiment needed to give the curvature half independent support, and it is not run. A matched-target Tsallis experiment would also help. As it stands, the paper is a detailed case study of one game rather than a established general decomposition.
minor comments (6)
- [§5.3, Table 3] The 'stable' column reports NashConv≤3e-12 for all stable rows; consider stating the convergence criterion for the limit-cycle row (η=0.10) more explicitly, since NashConv=0.11 is far from equilibrium.
- [§4] The curvature-fit window is a methodological choice; the sensitivity range ±5% to ±20% is reported in §5.2, but it would help to state the exact window used for Table 1 and Figure 3 (the text says |c−c*|≤0.10, but the 'range' units are per-game and not defined until Appendix A).
- [§5.1, Figure 1] The shaded band 'within 99.7% of H*' is described verbally; a precise definition (i.e., the set of c satisfying H(σ(c)) ≥ 0.997 H*) would improve reproducibility.
- [§5.3] The phrase 'the fit's intercept independently recovers κ=3.88' could mislead: the intercept and slope are not independent in a two-parameter log-log fit; the recovered κ is a derived quantity with its own uncertainty, and the 3% agreement is consistent with the local curvature estimate but not an independent confirmation.
- [§7] The initialization-dependence paragraph reports coordinates 0.057–0.258 over ten random initializations; consider adding a small table or figure, since this is the only stochastic element and is relevant to the scope of Conjecture 1.
- [Throughout] Minor language: 'R-N aD' formatting is inconsistent (spacing after hyphen); 'numpy' should be 'NumPy'; 'Kuhn poker' is sometimes written as 'Kuhn' and sometimes 'kuhn' — unify game names.
Circularity Check
The headline gap law is a Taylor identity read off the same entropy function whose fitted curvature and measured shortfall are then called a prediction; the independent content is limited to the causal η-sweep.
-
self definitional
[Section 3, Eqs. (3)-(4)]
"H(σ(c)) = H ⋆ − 1 2 κ(c−c ⋆)2 +O((c−c ⋆)3). ... A point on the segment at entropy shortfall δ := H ⋆ −H (σ(c)) therefore satisfies 1 2 κ(c−c ⋆)2 ≈δ , i.e. gap :=|c−c ⋆| ≈ p 2δ/κ . (4)"
Eq. (4) is a rearrangement of the Taylor expansion (3), with κ defined as the second derivative of H at c⋆ and δ defined as the entropy gap from H⋆. For any point on the Nash segment, δ = (1/2)κ·gap² up to cubic terms, so the two quantities are not independent. 'Verifying' the gap law across games is therefore a tautological check of the quadratic approximation rather than an empirical discovery. The paper itself concedes this in Section 6, calling it 'a local geometric identity—a second-order Taylor readout,' yet the abstract and Section 5 present the agreement as a prediction holding across five games.
-
fitted input called prediction
[Section 4 (curvature estimation) and Section 5.2 (Table 2)]
"For the curvature we evaluate H(σ(c)) on an 801-point uniform grid over the coordinate range and estimate κ by a least-squares quadratic fit to H on the window |c−c ⋆| ≤0.10 (range)."
The 'predicted' Kuhn gap √(2δ/κ) = 0.0204 is computed using κ = 4.0 obtained by fitting the same entropy landscape H(σ(c)) from which δ = 8.3×10⁻⁴ is measured. The agreement with the observed gap 0.0205 to within 2×10⁻⁴ therefore only confirms that the quadratic fit is a good local approximation of H; it is not an independent prediction of the solver's landing point. The fitted κ carries no information beyond the shape of the landscape, so this is a self-consistency check, not a test of the removal hypothesis.
1 more flagged steps
-
self citation load bearing
[Section 2, Conjecture 1; Section 1; Conclusion]
"Conjecture 1 (I-projection selection; Conjecture 1 of Leal, 2026). R-N aDinitialized and referenced at the uniform strategy selects the I-projection of the uniform reference onto NE, i.e. the maximum-entropy equilibrium."
The theoretical account that the paper aims to rescue—'the I-projection account is upheld'—is imported as Conjecture 1 from the same author's prior work (Leal, 2026). The paper's four matrix games provide some independent support for the conjecture in those cases, but the general maximum-entropy selection claim, and the framing of Kuhn as an 'apparent failure' requiring removal, rest on this self-citation. The conclusion therefore partly states that the author's own conjecture is consistent with the new Kuhn data rather than being derived from an independent source.
full rationale
The central quantitative result, gap ≈ √(2δ/κ), is not an empirical law but a second-order Taylor identity: κ is defined as −H''(c⋆), δ is defined as H⋆ − H(σ(c)), and the relation follows by rearranging the quadratic expansion. Consequently, the five-game 'verification' is a consistency check of the quadratic approximation, and the Kuhn 'prediction' 0.0204 vs 0.0205 is a self-consistency readout from the same entropy function whose fitted κ and measured δ are the inputs. The paper transparently acknowledges this in Section 6, which lowers the severity of the circularity but does not remove the fact that the headline law is constructional rather than predictive. The genuinely independent evidence is the η-sweep: varying the magnet drives δ toward zero and the landed coordinate toward c⋆, with the gap shrinking along the predicted curve until a stability floor. That causal sweep does discriminate removability from a fixed bias, though the claimed exponent 1/2 is itself the generic Taylor prediction and cannot by itself separate H-flat from a small residual bias over the accessible δ range. A minor self-citation (Conjecture 1 from Leal, 2026) is load-bearing for the interpretation but is partially counterbalanced by the paper's own matrix-game exact matches. Overall, partial circularity: the quantitative 'prediction' reduces by construction, while the qualitative removal claim retains independent experimental content.
Assumptions & free parameters
free parameters (2)
- Peak curvature κ of the entropy landscape =
kuhn: 4.01, asym safe: 20.08, pennies safe: 2.27, two safe: 2.01, dup action: 4.02
- Curvature-fit window width =
0.10 (coordinate units)
assumptions (4)
- standard math H(σ(c)) is C² in a neighborhood of c⋆ with H′(c⋆)=0 and κ>0 (interior maximum).
- domain assumption The Nash equilibrium set of each panel game is a one-dimensional segment parameterized by c, and R-NaD's converged profile lies exactly on that segment.
- domain assumption The selection target is the unweighted mean Shannon entropy (I-projection of a uniform reference).
- ad hoc to paper The η→0 limit is the right causal handle and the stability floor at η≈0.15 does not conceal a fixed bias.
Cite this review
Pith. "Pith review of The Curvature Shadow: An Apparent Failure of Maximum-Entropy Equilibrium Selection is a Removable Artifact." pith.science (2026). https://pith.science/paper/DO4I4WLG
@misc{pith2026260717543,
author = {Pith},
title = {Pith review of: The Curvature Shadow: An Apparent Failure of Maximum-Entropy Equilibrium Selection is a Removable Artifact},
year = {2026},
howpublished = {\url{https://pith.science/paper/DO4I4WLG}},
note = {Machine review of arXiv:2607.17543}
}
abstract
In two-player zero-sum games whose Nash equilibria form a convex set, regularized solvers such as Regularized Nash Dynamics (R-NaD) empirically select the maximum-entropy member: the information projection (I-projection) of a uniform reference onto the Nash set. On a panel of small games this match is exact, with one apparent exception: in Kuhn poker R-NaD lands at bluff coordinate 0.180 while the maximum-entropy member sits at 0.201, a coordinate gap of about 0.021, even though R-NaD attains 99.7 percent of the maximum entropy. We ask whether this gap is a genuine selection bias or an artifact, and answer it quantitatively. We show that for selection on a one-dimensional Nash manifold the coordinate gap factorizes as $\mathrm{gap} \approx \sqrt{2\delta/\kappa}$, where $\delta$ is the entropy shortfall of the solver and $\kappa$ is the curvature of the entropy landscape at its peak. Across five games this relation holds to within $2 \times 10^{-4}$ (under 1 percent relative error). The four matrix games have $\delta \approx 0$ (R-NaD reaches the maximum-entropy member exactly) and therefore no gap regardless of curvature; only the sequential game (Kuhn) has $\delta > 0$. A causal sweep of the magnet strength drives $\delta \to 0$ and the gap toward zero along the predicted curve (fitted scaling exponent 0.50, $R^2 > 0.999999$, against the exact prediction of 1/2), until the dynamics destabilize at a stability floor: behavior consistent with a removable shortfall and inconsistent with a fixed bias. We quantify the curvature half of the law from measured curvatures and flag a moving-target pitfall in the natural Tsallis-entropy experiment. The Kuhn gap is thus the curvature shadow of a small, removable entropy shortfall on an unusually flat peak; the I-projection account is upheld up to a flatness-limited residual.
Figures
Figures from the paper (1 more)
Reference graph
Works this paper leans on
- [1]
-
[2]
H. W. Kuhn. A simplified two-person poker. In Contributions to the Theory of Games, vol. 1, pp. 97--103. Princeton University Press, 1950. doi:10.1515/9781400881727-010 https://doi.org/10.1515/9781400881727-010
-
[3]
L. Leal. Which Nash equilibrium? Solver-dependent selection on zero-sum Nash polytopes. arXiv:2606.28308, 2026
arXiv 2026
-
[4]
R. D. McKelvey and T. R. Palfrey. Quantal response equilibria for normal form games. Games and Economic Behavior, 10(1):6--38, 1995. doi:10.1006/game.1995.1023 https://doi.org/10.1006/game.1995.1023
arXiv 1995
-
[5]
Perolat, R
J. Perolat, R. Munos, J.-B. Lespiau, S. Omidshafiei, M. Rowland, P. Ortega, N. Burch, T. Anthony, D. Balduzzi, B. De Vylder, G. Piliouras, M. Lanctot, and K. Tuyls. From Poincar\'e recurrence to convergence in imperfect information games: Finding equilibrium via regularization. In ICML, pp. 8525--8535, 2021
2021
-
[6]
J. Perolat, B. De Vylder, D. Hennes, E. Tarassov, F. Strub, et al. Mastering the game of Stratego with model-free multiagent reinforcement learning. Science, 378(6623):990--996, 2022. doi:10.1126/science.add4679 https://doi.org/10.1126/science.add4679
-
[7]
S. Sokota, R. D'Orazio, J. Z. Kolter, N. Loizou, M. Lanctot, I. Mitliagkas, N. Brown, and C. Kroer. A unified approach to reinforcement learning, quantal response equilibria, and two-player zero-sum games. In ICLR, 2023. doi:10.48550/arxiv.2206.05825 https://doi.org/10.48550/arxiv.2206.05825
-
[8]
Zinkevich, M
M. Zinkevich, M. Johanson, M. Bowling, and C. Piccione. Regret minimization in games with incomplete information. In NeurIPS, 2007
2007
Reviewed August 1, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.