Pith. sign in

REVIEW 2 major objections 3 minor 39 references

Colluding bidders can keep every agent's own prices competitive by coupling only the joint randomness of their bids, making any single-agent price audit exactly as powerful as its false-positive rate.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

Bidding agents coupled only through hidden shared randomness transfer seller rent while leaving every individual bid distribution unchanged, making single-agent price-level audits provably powerless.

T0 review reviewed 2026-08-01 challenge →

load-bearing objection The power=size theorem is a real and clean result, but the proof under-specifies the independence condition—keying the shared draw on the auditor's covariates breaks it; the fix is simple and the paper remains worth engaging. the 2 major comments →

arxiv 2607.26385 v1 pith:ZSRSMQEW submitted 2026-07-29 cs.GT cs.AIcs.CR

Collusion with Competitive Marginals: Price-Level Audits Are Blind by Construction

classification cs.GT cs.AIcs.CR
keywords algorithmic collusioncopula couplingfixed marginalspower equals sizeprice-level auditentity resolutionEthereum block-building auctionsampling temperature
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper claims that bidding agents can couple only the joint randomness of their bids while keeping each agent's own bid distribution exactly at the competitive law, and that this profitable conduct is invisible to any test reading a single agent's price or bid history. Such tests have power exactly equal to their false-positive rate, no matter the sample size. The construction uses a shared pseudorandom draw that is refreshed each round, and it is profitable because it lowers the seller's expected maximum bid. The authors demonstrate the coupling in real language-model agents and show that a live auction market already sits in a dependence regime where detection is impossible. Their positive proposal is to stop trying to detect the conduct and instead count distinct decision-makers.

Core claim

Any test whose input is a single agent's marginal bid or price distribution, compared against a competitive benchmark, has power exactly equal to its size against conduct that couples only the unexplained components of bids. Because a shared draw, refreshed independently each round, can correlate bids while preserving every agent's own marginal law, the entire single-agent process distribution remains unchanged—even at the comonotone limit. No measurable function of one agent's history can reject the conspiracy more often than it rejects the honest market. This is an identity, not a low-power phenomenon; no sample size repairs it.

What carries the argument

Lemma 1 gives a zero-communication realisation: conspirators compute a shared Gaussian latent z_t from a pseudorandom function on public round data, then form each bid's unexplained component as F_i^{-1}(Φ(√ρ z_t + √(1−ρ) e_it)). Because the inner term is standard normal for every ρ, each bidder's own component has the prescribed law F_i exactly; ρ=0 reproduces independence and ρ=1 gives the comonotone coupling that minimises the expected maximum bid. The independence of z_t across rounds is what makes the entire single-bidder process law identical to the competitive one, which is the load-bearing step for the power-equals-size result.

Load-bearing premise

The entire impossibility result rests on the shared random draw being freshly independent each round; if it is persistent, one agent's own bids become serially correlated and the conduct becomes detectable.

What would settle it

Run the Lemma 1 construction with the shared draw replaced by an AR(1) common signal: at persistence 0.5 a single bidder's residual series shows lag-one autocorrelation of about 0.40, and at 0.9 even a marginal goodness-of-fit test rejects. Observing serial dependence in one agent's residuals would therefore falsify the claim that the conduct leaves every single-agent test powerless.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Single-agent price-level audits—the standard empirical tool for algorithmic collusion—have power exactly equal to their false-positive rate against this conduct, at any sample size.
  • The conduct is profitable at meaningful magnitudes: a five-member coalition takes roughly 4.6% to 10.6% of seller revenue in the Ethereum block-building market, comparable to the price-lift effects reported in the literature.
  • Regret-based transcript audits also fail, but for a different reason: the coupling does not raise calibrated swap regret above the honest market's own off-best-response noise.
  • In deployed data, the honest population of bidder pairs is so dependent that the smallest detectable coupling is bounded below by a constant 20 to 32 times the family-wise sampling threshold, so longer audits do not help.
  • Shifting from detecting conduct to counting independent operators raises the measured Herfindahl index by 247.5% to 324.5%, a decidable regulatory target from public bid data.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the shared draw were persistent rather than independent across rounds, the invisibility would collapse: one bidder's own series would show autocorrelation and a single-agent test would gain power. Auditors could look for serial dependence in residuals as a first check.
  • The temperature result suggests a deployable mitigation—mandating higher sampling temperatures for bidding language models could mechanically weaken the coupling the conduct requires—but the paper does not quantify rent at the higher end, so that trade-off is untested.
  • The paper's dual-use framing implies the coupling arises by default from shared model supply chains; a regulatory response aimed at source concentration (how many independent models underpin a market) might be more tractable than screening bids.
  • Entity counting from public bid streams is a portable technique: any market with persistent identity and public bid data—not just block building—could use pairwise residual correlation to estimate the number of independent decision-makers.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 3 minor

Summary. The paper proposes a conduct in which bidding agents preserve each agent's own marginal bid distribution exactly while coupling the unexplained components through an iid shared draw (Lemma 1). The central theorem (Proposition 1) states that any test using a single agent's bid history has power exactly equal to its size against this conduct, for all coupling strengths up to comonotonicity. The paper then argues that price-level detection methods are therefore blind by construction, presents a language-model-agent replication showing same-base residual correlation (+0.053 under the strictest auditor at temperature 0.3) and a temperature trend, calibrates an ambient dependence floor on MEV-Boost data, and proposes entity counting as a more tractable regulatory target.

Significance. If the formal claim is repaired, the paper makes a meaningful and non-obvious point: a profitable conspiracy can be statistically invisible to single-agent marginal audits, and the impossibility is not a sample-size issue. The simulation checks, the reproducible scripts, and the transparent reporting of swept parameters are notable strengths, as is the attempt to measure the honest-population dependence floor in a real auction market. The empirical calibration to LLM agents and MEV-Boost is novel and policy-relevant. However, the current proof of Proposition 1 has a gap concerning the shared draw's dependence on auditor covariates, and the paper's broader 'price-level audits are blind' claim is narrower than the title suggests. These issues are fixable but need to be addressed before the central result is fully established.

major comments (2)
  1. [Lemma 1, Proposition 1; Figure 1(a)] The proof of Proposition 1 asserts that the single-bidder process law is unchanged because z_t are independent across rounds. This is insufficient: the test input is {(b_it, x_t)}, so the conditional law of ε_it given x_t must also be unchanged. Lemma 1 sets z_t = Φ^{-1}(PRF_k(id_t)) with id_t public round data, and Figure 1 labels the shared PRF as being 'on x_t'. If id_t includes or equals x_t, then ε_it is a deterministic function of x_t plus noise; e.g., with F_i = N(0,1), z_t = x_t, the coefficient on x_t in b_it becomes β_i + σ√ρ, so a t-test on that coefficient has power strictly greater than α for any ρ>0. The theorem is repairable by requiring z_t independent of x_t (e.g., an independent shared seed), but as written the construction and Proposition 1 are inconsistent. Appendix C checks persistence of the shared draw but not its dependence on covariates.
  2. [Proposition 3; 'Why the Empirical Literature Cannot See This'] Proposition 3 defines a 'price-level test' as a test whose input is one agent's marginal price or bid distribution. The title and abstract, however, claim that 'price-level audits' and the published detection methodology are blind by construction. This is a scope mismatch: tests based on market-level or aggregate price series are not covered by Proposition 3. Under Lemma 1, the cross-sectional average bid has the same mean but a different variance under coupling, so a test exploiting that variance would have power greater than α. Please either restrict the 'published methodology' claim to single-agent marginal tests or extend the analysis to aggregate price-level statistics.
minor comments (3)
  1. [Abstract and Table 2] The headline +0.053 is one temperature (T=0.3) under auditor (C); at T=0.0 and T=0.6 the clustered 95% CI includes zero. Consider leading with the monotone temperature trend or accounting for multiple comparisons across temperatures.
  2. [Proposition 4 and Conclusion] The conclusion states that regret-based audits 'do not separate' the conduct, relying on Proposition 4(iii), which the authors themselves describe as simulation-verified and 'left as a conjecture in general.' This should be stated more conditionally in the conclusion, as the main text already does.
  3. [Appendix F and Limitations] The claim that 'No parameter in this paper is tuned to maximise a reported quantity' is hard to reconcile with the acknowledgement in the Limitations that thresholds are not derived from a loss function but chosen from monotone sweeps. Please reconcile these statements.

Circularity Check

0 steps flagged

No significant circularity: the formal core is a self-contained construction and the empirical sections are measurements, not fitted predictions.

full rationale

Proposition 1 follows directly from Lemma 1, which explicitly constructs a coupling with i.i.d. shared draws and fixed marginals, so any single-bidder test has power equal to its size by a standard law-preservation argument. Proposition 2 cites external classical results (Mueller-Stoyan; Puccetti-Wang). The empirical sections are measurements and sweeps, and the paper explicitly states that thresholds are not derived from a loss function and reports all swept ranges. No load-bearing self-citation appears; the shared-model coupling evidence is external (Kleinberg-Raghavan; Spiro). The definition of 'price-level test' makes Proposition 3 a corollary, but it is grounded in Lemma 1's construction rather than assuming the conclusion. The only caveat is a proof-completeness concern: Lemma 1 lets the shared draw be a PRF on public round data, and Proposition 1 only shows the unconditional single-bidder law is unchanged; if the shared draw is a function of auditor covariates x_t, the joint law of (b_it, x_t) may differ. This is an assumption gap, not a circular reduction, so it does not affect the circularity score.

Axiom & Free-Parameter Ledger

3 free parameters · 5 axioms · 1 invented entities

The central theorem has no fitted parameters; it is a derivation from the stated iid-shared-draw construction. The free parameters listed are operating points in the empirical sections, all reported as sweeps. The main load-bearing assumptions are the independence of the shared draw and the first-price auction model.

free parameters (3)
  • submission-timing gap filter = swept 0, 1, 2, 4 s; floor drops 0.81 to 0.50
    Choosing the gap changes the cleaned floor by 0.31; the paper assumes co-submission implies shared infrastructure, stated as an assumption rather than established.
  • detector threshold = 0.80, 0.85, 0.90, 0.93, 0.95 (sweep)
    Operating point for same-operator detection affects the recall/false-positive trade-off; reported as a sweep, not derived from a loss function.
  • clustering threshold = 0.85, 0.90, 0.95 (sweep)
    Affects behavioural-cluster HHI lifts; the paper itself notes the result is order-dependent and has no point estimate.
axioms (5)
  • standard math Comonotone coupling is extremal for supermodular expectations (Fréchet bounds).
    Used in Proposition 2 to say coupling lowers expected maximum bid; cited to Müller & Stoyan 2002 and Puccetti & Wang 2015.
  • domain assumption Shared draw z_t is independent across rounds.
    Necessary for Proposition 1; the paper proves tightness by showing persistence would make single-bidder tests detect. Technical Appendix C.
  • domain assumption First-price auction with private values, bidders use equilibrium bid 1/2 v + σ ε; coalition faces bidders with unexplained dispersion.
    Used to compute coalition surplus and regret; Appendix D and Table 3. The paper acknowledges the implication fails when outsiders have no dispersion.
  • ad hoc to paper Co-submission in MEV-Boost implies shared infrastructure.
    The timing filter that lowers the floor from 0.81 to 0.50 rests on this; the paper says it assumes rather than establishes it.
  • domain assumption Auditor (C) residualization correctly captures the unexplained component.
    All LM residual correlation claims depend on the regression basis and out-of-sample cross-fitting removing all explainable common response.
invented entities (1)
  • Shared pseudorandom coupling draw z_t (PRF keyed on public round data) independent evidence
    purpose: Realizes dependent residuals across coalition members without a communication channel, leaving each bidder's marginal law unchanged.
    The paper supplies falsifiable handles: same-base LM deployments should show positive residual correlation and that correlation should decrease with sampling temperature; the Ethereum market is used to calibrate the required floor. These are handles outside the construction itself.

reviewed 2026-08-01 · how reviews work

0 comments
Cite this review

Pith. "Pith review of Collusion with Competitive Marginals: Price-Level Audits Are Blind by Construction." pith.science (2026). https://pith.science/paper/ZSRSMQEW

@misc{pith2026260726385,
  author       = {Pith},
  title        = {Pith review of: Collusion with Competitive Marginals: Price-Level Audits Are Blind by Construction},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ZSRSMQEW}},
  note         = {Machine review of arXiv:2607.26385}
}
Share X Bluesky LinkedIn Reddit HN
abstract

Empirical work on algorithmic collusion asks one question of the data: are prices supracompetitive? We show this can be answered "no" by a conspiracy that is nonetheless profitable. Consider bidding agents that couple only through the joint distribution of their unexplained bid components, leaving every agent's own bid law exactly at the competitive law. Any test whose input is a single agent's price or bid history then has power exactly equal to its false-positive rate, for every coupling strength up to comonotonicity. The published detection methodology is therefore blind to this conduct by construction rather than underpowered, and no sample size repairs it. Three empirical results follow. First, the mechanism appears in real language-model agents: twenty models from nineteen independent developers, three deployment prompts each, show residual correlation of $+0.053$ between two deployments of one model against $+0.0001$ across models, with a 95% interval clustered by developer of $[0.030, 0.078]$, under an auditor that sees every order feature and is fitted out of sample. Second, the coupling falls monotonically as sampling temperature rises ($p=0.002$), turning a deployment parameter into a candidate mitigation. Third, on 24 days of Ethereum block-building auction data covering 77,684 bids from 39 bidders, the honest population of bidder pairs is itself so dependent that a screen held at a 5% false-positive rate must sit above a floor of $+0.50$ to $+0.81$, which is 20 to 32 times the family-wise sampling threshold and does not fall as the audit window grows. Since lawful multi-identity operation and conspiracy are behaviourally indistinguishable here, the tractable regulatory target is not detection but counting: resolving 40 bidding identities into 23 operators raises the Herfindahl index by 247.5%, and adding behavioural clusters from public bid streams reaches 324.5%.

Figures

Figures reproduced from arXiv: 2607.26385 by Chengrui Wu, Hanzhe Hong, Jiayu Lu, Kaizhen Tan, Siru Tao, Xin Xu.

Figure 1
Figure 1. Figure 1: The mechanism. (a) Two agents deployed by different firms share an upstream origin, either one base model under two deployment prompts or a pseudorandom function keyed on public round data, and never exchange a message. (b) Each agent’s own bid distribution is identical under competitive and collusive conduct, so every test that reads one agent’s marginal has power equal to its size. (c) The conspiracy liv… view at source ↗
Figure 2
Figure 2. Figure 2: Marginal audits have power equal to their size. Pair [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Coupling is base-specific and weakens with sam [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: More data does not lower the detection floor. Ob [PITH_FULL_IMAGE:figures/full_fig_p006_4.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

39 extracted references · 4 linked inside Pith

  1. [1]

    Assad, S.; Calvano, E.; Calzolari, G.; Clark, R.; Denicol \`o , V.; Ershov, D.; Johnson, J.; Pastorello, S.; Rhodes, A.; Xu, L.; and Wildenbeest, M. 2021. Autonomous Algorithmic Collusion: Economic Research and Policy Implications. Oxford Review of Economic Policy, 37(3): 459--478

  2. [2]

    Calvano, E.; Calzolari, G.; Denicol \`o , V.; and Pastorello, S. 2020. Artificial Intelligence, Algorithmic Pricing, and Collusion. American Economic Review, 110(10): 3267--3297

  3. [3]

    Chassang, S.; Kawai, K.; Nakabayashi, J.; and Ortner, J. 2022. Robust Screens for Noncompetitive Bidding in Procurement Auctions. Econometrica, 90(1): 315--346

  4. [4]

    A.; and Shorrer, R

    Fish, S.; Gonczarowski, Y. A.; and Shorrer, R. I. 2024. Algorithmic Collusion by Large Language Models. arXiv:2404.00806

  5. [5]

    Flashbots . 2024. MEV -Boost Specifications. Flashbots Documentation. Accessed 2026-07-27

  6. [6]

    Harrington, J. E., Jr. 2008. Detecting Cartels. In Buccirossi, P., ed., Handbook of Antitrust Economics. Cambridge, MA: MIT Press

  7. [7]

    D.; Long, S.; and Zhang, C

    Hartline, J. D.; Long, S.; and Zhang, C. 2024. Regulation of Algorithmic Collusion. In Proceedings of the Symposium on Computer Science and Law (CSLAW), 98--108. ACM

  8. [8]

    D.; Wang, C.; and Zhang, C

    Hartline, J. D.; Wang, C.; and Zhang, C. 2025. Regulation of Algorithmic Collusion, Refined: Testing Pessimistic Calibrated Regret. In Proceedings of the Symposium on Computer Science and Law (CSLAW), 108--120. ACM

  9. [9]

    Keppo, J.; Li, Y.; Tsoukalas, G.; and Yuan, N. 2026. On the Fragility of AI Agent Collusion. arXiv:2603.20281

  10. [10]

    Kleinberg, J.; and Raghavan, M. 2021. Algorithmic Monoculture and Social Welfare. Proceedings of the National Academy of Sciences, 118(22): e2018340118

  11. [11]

    Y.; Ojha, S.; Cai, K.; and Chen, M

    Lin, R. Y.; Ojha, S.; Cai, K.; and Chen, M. F. 2024. Strategic Collusion of LLM Agents: Market Division in Multi-Commodity Competitions. arXiv:2410.00031

  12. [12]

    M \"u ller, A.; and Stoyan, D. 2002. Comparison Methods for Stochastic Models and Risks. Wiley Series in Probability and Statistics. Chichester: Wiley

  13. [13]

    Nelsen, R. B. 2006. An Introduction to Copulas. Springer Series in Statistics. New York: Springer, 2nd edition

  14. [14]

    OECD . 2022. Data Screening Tools for Competition Investigations. OECD Roundtables on Competition Policy Papers No. 284, OECD Publishing, Paris

  15. [15]

    Puccetti, G.; and Wang, R. 2015. Extremal Dependence Concepts. Statistical Science, 30(4): 485--517

  16. [16]

    Spiro, T. 2026. The Oracle's Fingerprint: Correlated AI Forecasting Errors and the Limits of Bias Transmission. arXiv:2605.00844

  17. [18]

    Wu, F.; Thiery, T.; Leonardos, S.; and Ventre, C. 2024. To Compete or Collude: Bidding Incentives in E thereum Block Building Auctions. In Proceedings of the 5th ACM International Conference on AI in Finance (ICAIF) , 813--821. ACM

  18. [19]

    Zhang, A.; Liu, Y.; Zhang, R.; Shan, Y.; and Wu, Y. 2026. Order Flow Exclusivity and Value Extraction Mechanisms: An Analysis of E thereum Builder Centralization. arXiv:2605.04471

  19. [20]

    Artificial Intelligence, Algorithmic Pricing, and Collusion , journal =

    Calvano, Emilio and Calzolari, Giacomo and Denicol. Artificial Intelligence, Algorithmic Pricing, and Collusion , journal =

  20. [21]

    Proceedings of the National Academy of Sciences , volume =

    Kleinberg, Jon and Raghavan, Manish , title =. Proceedings of the National Academy of Sciences , volume =

  21. [22]

    2026 , eprint =

    Spiro, Theodor , title =. 2026 , eprint =

  22. [23]

    and Shorrer, Ran I

    Fish, Sara and Gonczarowski, Yannai A. and Shorrer, Ran I. , title =. 2024 , eprint =

  23. [24]

    and Ojha, Siddhartha and Cai, Kevin and Chen, Maxwell F

    Lin, Ryan Y. and Ojha, Siddhartha and Cai, Kevin and Chen, Maxwell F. , title =. 2024 , eprint =

  24. [25]

    2026 , eprint =

    Keppo, Jussi and Li, Yuze and Tsoukalas, Gerry and Yuan, Nuo , title =. 2026 , eprint =

  25. [26]

    and Long, Sheng and Zhang, Chenhao , title =

    Hartline, Jason D. and Long, Sheng and Zhang, Chenhao , title =. Proceedings of the Symposium on Computer Science and Law (CSLAW) , pages =

  26. [27]

    and Wang, Chang and Zhang, Chenhao , title =

    Hartline, Jason D. and Wang, Chang and Zhang, Chenhao , title =. Proceedings of the Symposium on Computer Science and Law (CSLAW) , pages =

  27. [28]

    Proceedings of the 5th

    Wu, Fei and Thiery, Thomas and Leonardos, Stefanos and Ventre, Carmine , title =. Proceedings of the 5th

  28. [29]

    2026 , eprint =

    Zhang, Ao and Liu, Yunwen and Zhang, Ren and Shan, Yingdi and Wu, Yongwei , title =. 2026 , eprint =

  29. [30]

    Time to Bribe: Measuring Block Construction Market , year =

    Wahrst. Time to Bribe: Measuring Block Construction Market , year =. 2305.16468 , archivePrefix =

  30. [31]

    Comparison Methods for Stochastic Models and Risks , series =

    M. Comparison Methods for Stochastic Models and Risks , series =

  31. [32]

    Statistical Science , volume =

    Puccetti, Giovanni and Wang, Ruodu , title =. Statistical Science , volume =

  32. [33]

    Econometrica , volume =

    Chassang, Sylvain and Kawai, Kei and Nakabayashi, Jun and Ortner, Juan , title =. Econometrica , volume =

  33. [34]

    , title =

    Harrington, Jr., Joseph E. , title =. Handbook of Antitrust Economics , editor =

  34. [35]

    2022 , url =

    Data Screening Tools for Competition Investigations , howpublished =. 2022 , url =

  35. [36]

    2026 , eprint =

    Xie, Yi and Zhou, Zhanke and Cao, Chentao and Liu, Bo and Han, Bo , title =. 2026 , eprint =

  36. [37]

    , title =

    Nelsen, Roger B. , title =

  37. [38]

    van der Vaart, A. W. , title =

  38. [39]

    Autonomous Algorithmic Collusion: Economic Research and Policy Implications , journal =

    Assad, Stephanie and Calvano, Emilio and Calzolari, Giacomo and Clark, Robert and Denicol. Autonomous Algorithmic Collusion: Economic Research and Policy Implications , journal =

  39. [40]

    and Savage, Stefan , title =

    Meiklejohn, Sarah and Pomarole, Marjori and Jordan, Grant and Levchenko, Kirill and McCoy, Damon and Voelker, Geoffrey M. and Savage, Stefan , title =. Proceedings of the 13th

This paper was first reviewed by deepseek-v4-flash on August 1, 2026.