Pith. sign in

REVIEW 3 major objections 4 minor 12 references

Does a Toehold Make a Bidder Bolder? Preemption and Multiplicity in Multi-Round Takeover Auctions

T0 review · 3 major / 4 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read The toehold's deterrence effect, the reason bidders take one, does not survive a second round of bidding.

desk verdict A careful computational study with a real warning about solver multiplicity; the headline deterrence claim is plausible but rests on a single parameterization and should be read narrowly. read the letter →

arxiv 2608.08407 v1 pith:H2XZVWKC submitted 2026-08-09 cs.GT cs.AIcs.LG

classification cs.GTcs.AIcs.LG MSC 91B2691A26
keywords takeoverauctionstoeholdspreemptivebiddingequilibriummultiplicitycommon-valueimperfect-informationgamesmulti-roundascendingwinner'scurse
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether a toehold's second supposed benefit—frightening rival bidders into dropping out—survives when a takeover contest is modeled as several rounds of escalating offers rather than the one-round exchange in classical models. It solves a two-player common-value ascending auction with private noisy signals and a toehold, certifying profiles as equilibria from which neither bidder can improve its own profit by more than a tiny $\varepsilon$. The answer is that the profit benefit of a toehold survives, but the deterrence benefit does not: the same auction supports equilibria that pay the holder the same profit yet deter the rival with very different probabilities, aggressive jump bidding appears even with no toehold at all, and the tidy "bigger toehold, more deterrence" relationship holds at one round and stops responding by two. If this is right, the rarity of toeholds in practice may need no hidden cost to explain; the strategic edge they are supposed to confer may simply not be there in longer contests.

What carries the argument

The carrying object is a two-player, multi-round ascending common-value takeover auction with a toehold, cast as an extensive-form game of imperfect information: public bid history, private noisy signals, alternating raises on a discrete price ladder, and own-profit payoffs. The paper certifies computed strategies as $\varepsilon$-Nash equilibria, meaning neither bidder can improve its own profit by more than $\varepsilon$ by deviating, with $\varepsilon$ between $5\times 10^{-7}$ and $8\times 10^{-5}$. Two instruments do the work: restarted equilibrium computation to probe multiplicity, and a forced-opening best-response construction that fixes the holder's first bid at each price, lets every later decision best-respond, and reads off what each opening earns against a given rival—an exact calculation free of convergence error. The economic mechanism inside the equilibria is a winner's-curse inference: a rival who reads a high opening as evidence of a strong common value folds rather than risk winning in states where its own signal is misleadingly high.

What would settle it

Re-run the same two-round auction with a richer signal structure, for example five value levels and three private signals per bidder with lower noise; if every certified equilibrium then shows deterrence rising monotonically with the toehold, or if the low-deterrence equilibrium branch disappears, the central claim that the toehold-to-deterrence channel dies at a genuine second round is falsified.

Watch

Extended reading notes

Core claim

The paper's central discovery is that the toehold's deterrence channel is an artifact of truncating the contest at one or two moves. In the solved instance, at a zero toehold as well as at positive toeholds, independent restarts converge to equilibria that give the toehold-holder essentially the same profit but opposite conduct: one profile opens with a preemptive jump and the rival folds with probability about one-third, another opens at the bottom of the price ladder and the rival almost never folds; both are certified as $\varepsilon$-Nash equilibria with $\varepsilon$ around $10^{-5}$ to $10^{-6}$. An exact forced-opening best-response calculation confirms each conduct is a best reply to the other, so the multiplicity is not numerical slack. Preemptive jump bidding persists when the toehold is zero, tracing it to the sequential public-bid structure rather than stake ownership. Sweeping the number of rounds, deterrence rises with toehold at one round, saturates at two, and is no longer pinned down at three. The paper therefore restates preemption as a signalling phenomenon and concludes that a toehold buys profit but not identified deterrence.

Load-bearing premise

Everything economic in the paper is computed for one discrete instance—three value levels, nine bid levels, one private signal per bidder, and signal noise of 0.5—and at several toeholds no solver run converged, so the multiplicity and disappearance of deterrence could be artifacts of that parameterization or of non-convergence rather than properties of the auction itself.

Editorial extensions

If this is right

  • Larger toeholds still raise the holder's profit in every solved configuration, so the financial incentive to take a toehold stands; what falls away is the claim that it reliably deters entry.
  • Preemptive jump bidding should be attributed to turn-taking in public, not to the toehold, so eliminating toeholds would not eliminate preemption.
  • Comparative statics drawn from one-round or two-move auction models can mislead: extending the contest by a single further round can turn a clean monotone relationship into saturation or multiplicity.
  • Because equally certified profiles imply different conduct, any single reported number for what a preemptive bid is worth or for a deterrence rate is an artifact of equilibrium selection, not a property of the game.
  • Computational equilibrium analysis in general-sum games needs restart-based reporting and interval-valued output, because a single converged run can manufacture confidence.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • One testable extension: the selection between the preempting and passive equilibrium should depend on focal points or historical precedent, so observed opening-bid distributions in real takeover contests could reveal which equilibrium bidders coordinate on.
  • The two-move artifact suggests that other short-horizon auction models, not just toehold models, may have comparative statics that do not survive added rounds; re-running classic two-stage results in genuinely multi-round versions would show how widespread the fragility is.
  • At toeholds where no solver run converges, "deterrence is unidentified" conflates economic indeterminacy with computational non-convergence, so that part of the claim should be treated as provisional until alternative solvers or larger budgets either certify or refute it.
  • The forced-opening interval technique could be reused as a general diagnostic for equilibrium multiplicity in other general-sum games, since it prices unseen information-set strategies exactly and separates selection effects from solver error.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This paper models a two-bidder, multi-round ascending common-value takeover auction with a toehold as a finite extensive-form game of imperfect information and solves it numerically for approximate Bayes-Nash equilibria, certifying each profile as an epsilon-Nash equilibrium with epsilons on the order of 1e-6 to 1e-5. The authors report three economic findings: equilibrium multiplicity in which the same toehold-holder profit supports very different deterrence rates; preemptive jump bidding that persists at a zero toehold, implying the toehold does not purchase preemption; and a toehold-to-deterrence relationship that holds at one round, saturates at two rounds, and is not identified at three. They also introduce a forced-opening best-response construction to price jump bids without trusting a single solver run, benchmark several solvers on the game, and extend the analysis to an intractable regime using approximate exploitability. The paper is unusually candid: Section 10 explicitly states that all economic results come from one parameterization and that the number of values, signals, and noise were not varied, and it flags where solver non-convergence rather than theoretical multiplicity drives the three-round findings.

Significance. If the central claim is correct, the paper makes a substantive contribution to the toehold puzzle by suggesting that the deterrence rationale for toeholds is an artifact of short-horizon models, while the profit rationale survives. The methodological warning about solver-induced multiplicity and the forced-opening exact calculation are valuable independently. The paper's strengths include carefully reported epsilon certificates, exact best-response checks validated against brute force, a solver-free control with a uniform rival, and full release of code and per-restart data. However, the economic headline is supported by a single information structure, and the paper itself concedes the missing parameter variations; the conclusion is therefore not as general as the abstract and conclusion present it. With additional evidence or a suitably circumscribed claim, the paper would be a worthwhile contribution.

major comments (3)
  1. [Sections 3, 6, 10] The central claim that the toehold-to-deterrence channel is an artifact of short models rests entirely on one parameterization (num_values=3, num_bids=9, k=1, signal_noise=0.5), and Section 10 explicitly concedes that the number of values, signals, and noise were never varied. The R=2 'saturation' in Table 2 is two data points, θ=0.15 and θ=0.30, both giving deterrence 0.333; with three value levels and a signal that is uniform with probability 0.5, the rival's posterior after a jump bid is coarse, and 0.333 may be a discreteness corner rather than a robust ceiling. A sweep over signal_noise or an additional signal (e.g., k=2) would test whether a sharper jump signal restores a positive toehold response at R=2. Without such variation, the abstract's unconditional statement that 'a genuine second round flattens it' overstates the evidence; the paper should either add this variation or explicitly restrict the central claim to the solved instance throughout.
  2. [Sections 5, 6, 10] The identification language at R=3 conflates solver non-convergence with game-theoretic non-identification. Section 5 reports that no restart converges within budget at θ=0.25, 0.30, 0.40, and 0.50, and Section 6 says deterrence is 'not identified there by any route.' Non-convergence within a computational budget is a fact about the solver, not a property of the equilibrium set, and the paper itself distinguishes the two in Section 5 ('by non-convergence rather than by multiplicity'). The conclusion's statement that 'by the third the equilibria multiply and it is not identified at all' is therefore too strong: certified multiplicity exists only at the cells where restarts converged (θ=0 and θ=0.15), while the θ=0.30 cell is simply uncomputed. The text should say that deterrence is certified to be multiple at some toeholds and is not computed at others, and should not present the blank cell in Table 2 as evidence for non-identification.
  3. [Abstract, Section 5] The abstract's first finding, that 'the auction fixes what the toehold-holder earns but not how it bids,' is contradicted by the paper's own grid-refinement result in Section 5. At thirteen price levels, the two certified profiles at θ=0 yield values 0.1736 and 0.2078, a difference of 0.034 that is several hundred times the reported epsilons (1.3e-5 and 8.2e-5), so value-pinning fails in that configuration. The paper acknowledges this in Section 5 ('Value-pinning is a per-configuration observation throughout this paper rather than a theorem'), but the abstract and Section 1 present value-pinning without that caveat. Either the abstract should be qualified to the configurations where value-pinning is observed, or the first finding should be reframed as deterrence non-identification, which is the claim the paper actually rests on.
minor comments (4)
  1. [Section 6, Table 2] The label 'unique' at R=1 should read 'unique across the six restarts' rather than 'unique,' since the paper does not prove uniqueness in the game-theoretic sense; the text later acknowledges that only two positive toeholds are swept at R=1.
  2. [Section 5, Figure 1 caption] The bars spanning the certified rivals are observed extremes across the certified profiles, not confidence intervals; the caption should state this explicitly to avoid a statistical misreading.
  3. [Sections 5 and 8] The game-size numbers should be reconciled with the configuration used: Section 5 reports about 40,000 nodes for the three-round instance with num_bids=9, while Section 8 reports about 6,655 states for a three-round instance with num_bids=6; the text should state the price-grid size in each place to avoid apparent inconsistency.
  4. [Section 10] The paragraph on grid robustness says 'twelve additional starting policies' and later 'twelve of them new' for a total of eighteen solves; the arithmetic is clear, but the phrasing could be tightened to avoid confusion about how many solves come from the original six.

Circularity Check

0 steps flagged · score 2.0 of 10

No significant circularity: the economic claims are computed equilibria of a stated game, with only a non-load-bearing self-citation to the companion paper.

full rationale

The paper's derivation chain is self-contained. Every economic quantity—equilibrium values, deterrence probabilities, jump premia, and the round-sweep comparison—is computed from the game primitives stated in Section 3 (num_values=3, num_bids=9, k=1, signal_noise=0.5, toehold grid, and rounds R=1,2,3) using fictitious play and the forced-opening exact best-response construction. No parameter is fitted to any target conclusion, and no claimed prediction is defined in terms of the outcome it is supposed to predict. The forced-opening calculation is exact given a certified rival policy and is validated against brute-force enumeration, and the uniform-rival control panel is solver-free; neither reduces to the conclusion. The auhor's self-citation to the companion sealed-auction paper [Naboulsi, 2026] is used for comparison of sealed versus sequential preemption and for a standard exploitability estimator; it is not load-bearing for the central finding that the toehold's deterrence channel flattens at two rounds. The paper's own limitations—single parameterization, no variation of signal count or noise, and non-convergence at some high toeholds—are explicitly disclosed in Sections 5 and 10 and are robustness/correctness concerns, not circular reductions. The word 'unidentified' is used operationally to mean that certified equilibria or converging restarts do not pin down a value, which is an honest solver/model limitation rather than a self-referential conclusion. Therefore no circular step can be quoted or exhibited, and the only noteworthy item is the minor, non-load-bearing self-citation, which warrants a low non-circularity score rather than a finding of circularity.

Assumptions & free parameters 6 free parameters · 5 assumptions · 0 invented entities

The central economic claims rest on a single hand-chosen instance. The game parameters (value grid, price grid, signal structure, noise) are not fitted to data but are chosen by the author and not varied except for toehold, number of rounds, and price-grid resolution. The solver acceptance bound and compute budget affect which profiles are certified. No new entities are introduced; the model uses standard toehold and private-signal constructs. These choices bound the generality of the findings.

free parameters (6)
  • num_values = 3
    Number of levels in the common-value grid; chosen by hand, not estimated. Section 3.
  • price grid = num_bids = 9, bid_step = 0.375, spanning [0,3]
    Discrete price ladder; chosen by hand. Varied only in the grid-refinement study (7, 9, 13 levels). Section 3.
  • signal_noise = 0.5
    Probability that a private signal is uniform rather than equal to the common value; chosen by hand and never varied. Section 3.
  • num_signals = 1
    Number of private signals per bidder; chosen by hand and not varied in the economic results. Section 3.
  • solver acceptance bound (NashConv) = <= 1e-4
    Threshold for accepting a profile as certified; chosen by the author and affects which cells enter the tables. Section 4.
  • iteration budget and tolerance = 2e6 iterations, tolerance 1e-8
    Compute budget that ends runs; a restart enters a table only if its own-profit NashConv is at or below the acceptance bound. Section 4.
assumptions (5)
  • domain assumption The target's common value W is drawn uniformly from a grid of three levels; both bidders are risk-neutral.
    Stated in Section 3: 'W, drawn uniformly from a grid of num_values levels' and 'Both are risk-neutral'.
  • domain assumption Each bidder receives one noisy private signal; with probability 0.5 the signal equals W and otherwise is uniform on the value grid; the bid history is public and recall is perfect.
    Section 3 defines signal_noise = 0.5 and k = 1; the public history and perfect-recall structure is enforced by construction and checked by unit tests.
  • domain assumption Payoffs: if the rival wins at price p, the toehold-holder earns theta * p; if the toehold-holder wins, it pays for the non-toehold fraction only.
    Section 3 states the payoff contract, matching the companion sealed game; this is the standard Bulow-Huang-Klemperer payoff structure.
  • standard math Fictitious play and the other solvers find profiles whose own-profit NashConv certifies them as epsilon-Nash equilibria; the certificates are computed, not proven analytically for this game.
    The computational method in Section 4 relies on this; the paper reports wander and acceptance bounds to keep the certificate honest.
  • domain assumption The learned-best-response exploitability estimate is a valid lower bound, calibrated at k = 1,2 and applied at k = 6.
    Section 9 applies the estimator beyond the calibrated signal count; the paper states it is a lower bound, not a certificate.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Does a Toehold Make a Bidder Bolder? Preemption and Multiplicity in Multi-Round Takeover Auctions." pith.science (2026). https://pith.science/paper/H2XZVWKC

@misc{pith2026260808407,
  author       = {Pith},
  title        = {Pith review of: Does a Toehold Make a Bidder Bolder? Preemption and Multiplicity in Multi-Round Takeover Auctions},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/H2XZVWKC}},
  note         = {Machine review of arXiv:2608.08407}
}
read the original abstract

A bidder can quietly buy a stake in a company before making an offer for it. That stake, a toehold, is supposed to pay for itself twice: it makes the bidder willing to bid harder, and it frightens rivals into staying out of the fight. The first effect is arithmetic. The second is what would justify the cost and exposure of taking one at all. Yet toeholds are rare in practice, a standing puzzle. We ask whether that second effect is there once the contest is modelled as several rounds of escalating offers rather than the single exchange classical models assume. We turn it into a game a computer can solve, and certify the answers to an accuracy a referee can check. Three findings. The auction fixes what the toehold-holder earns but not how it bids: the same contest supports a bidder who opens aggressively against a rival who folds, and one who opens cheaply against a rival who does not, with the same profit either way. Aggressive preemptive bidding still appears when the toehold is removed entirely, so it comes from bidding in public and in turns, not from owning the stake. And the tidy "bigger toehold, more deterrence" relationship holds only in a contest cut short after one round; give it a real second round and it stops responding. So the two reasons to buy a toehold do not fare alike. The profit reason holds up; the deterrence reason does not, which suggests why toeholds may be rarer than theory predicts, alongside the procedural costs of disclosure and price impact that this model omits. A warning follows for anyone computing economics from a game solver: solve this auction once and it returns a confident figure for what a preemptive bid is worth; solve it again from a different start and it returns a different one, equally converged. We also report which solvers cope with contests of this shape, including versions too large to enumerate. Code is released.

Figures

Figures reproduced from arXiv: 2608.08407 by the authors.

Figure 1
Figure 1. Forced-opening best-response profit for the toehold-holder as a function of the open [PITH_FULL_IMAGE:figures/full_fig_p008_1.png] view at source ↗
Figure 2
Figure 2. Deep self-play on the three-round takeover auction (mean [PITH_FULL_IMAGE:figures/full_fig_p012_2.png] view at source ↗
Figure 3
Figure 3. Scaling in the number of rounds. Left: game-tree size grows steeply with rounds (the [PITH_FULL_IMAGE:figures/full_fig_p013_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Intractable multi-round, multi￾signal auction (k = 6 signals, three rounds, about 107 histories). PPO (0.004 ± 0.003) and PPG (0.003 ± 0.004) drive the learned￾best-response exploitability estimate below a naive unshaded bidder (0.083) and far below uniform play (1.77)…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

12 extracted references · 5 canonical work pages

  1. [4]

    doi: 10.1017/S0022109019001029. Kent D. Daniel and David Hirshleifer. A theory of costly sequential bidding.Review of Finance, 22(5):1631–1665,

  2. [11]

    John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov

    arXiv:2502.08938. John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal policy optimization algorithms.arXiv preprint arXiv:1707.06347,

  3. [1988]

    Deep reinforcement learning from self-play in imperfect-information games

    Johannes Heinrich and David Silver. Deep reinforcement learning from self-play in imperfect-information games. InNIPS 2016 Deep Reinforcement Learning Workshop,

  4. [2009]

    doi: 10.1016/j.jfineco.2008.02

  5. [2012]

    Michael J

    doi: 10.1016/j.econlet.2012.06.023. Michael J. Fishman. A theory of preemptive takeover bidding.RAND Journal of Economics, 19(1):88–101,

  6. [2016]

    Paul Klemperer

    arXiv:1603.01121. Paul Klemperer. Auction theory: A guide to the literature.Journal of Economic Surveys, 13 (3):227–286,

  7. [2017]

    Marc Lanctot, Edward Lockhart, Jean-Baptiste Lespiau, et al

    arXiv:1711.00832. Marc Lanctot, Edward Lockhart, Jean-Baptiste Lespiau, et al. Openspiel: A framework for reinforcement learning in games.arXiv preprint arXiv:1908.09453,

  8. [2019]

    Jeremy Bulow, Ming Huang, and Paul Klemperer

    arXiv:1811.00164. Jeremy Bulow, Ming Huang, and Paul Klemperer. Toeholds and takeovers.Journal of Political Economy, 107(3):427–454,

Show all 12 references
  1. [2021]

    Yun Dai, Sebastian Gryglewicz, and Han T

    arXiv:2009.04416. Yun Dai, Sebastian Gryglewicz, and Han T. J. Smit. Less popular but more effective toeholds in corporate takeovers.Journal of Financial and Quantitative Analysis, 56(1):283–312,

  2. [2023]

    Robert Wilson

    arXiv:2206.05825. Robert Wilson. A bidding model of perfect competition.The Review of Economic Studies, 44 (3):511–518,

  3. [2024]

    Anna Dodonova

    doi: 10.1145/3670865.3673644. Anna Dodonova. Toeholds and signalling in takeover auctions.Economics Letters, 117(2): 386–388,

  4. [2026]

    Max Rudolph, Nathan Lichtlé, Sobhan Mohammadpour, Alexandre Bayen, J

    arXiv:2606.29457. Max Rudolph, Nathan Lichtlé, Sobhan Mohammadpour, Alexandre Bayen, J. Zico Kolter, Amy Zhang, Gabriele Farina, Eugene Vinitsky, and Samuel Sokota. Reevaluating policy gradient methods for imperfect-information games. InInternational Conference on Learning Rep...

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.