REVIEW 2 major objections 3 minor 39 references
Colluding bidders can keep every agent's own prices competitive by coupling only the joint randomness of their bids, making any single-agent price audit exactly as powerful as its false-positive rate.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
Bidding agents coupled only through hidden shared randomness transfer seller rent while leaving every individual bid distribution unchanged, making single-agent price-level audits provably powerless.
T0 review reviewed 2026-08-01 challenge →
load-bearing objection The power=size theorem is a real and clean result, but the proof under-specifies the independence condition—keying the shared draw on the auditor's covariates breaks it; the fix is simple and the paper remains worth engaging. the 2 major comments →
Collusion with Competitive Marginals: Price-Level Audits Are Blind by Construction
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
Any test whose input is a single agent's marginal bid or price distribution, compared against a competitive benchmark, has power exactly equal to its size against conduct that couples only the unexplained components of bids. Because a shared draw, refreshed independently each round, can correlate bids while preserving every agent's own marginal law, the entire single-agent process distribution remains unchanged—even at the comonotone limit. No measurable function of one agent's history can reject the conspiracy more often than it rejects the honest market. This is an identity, not a low-power phenomenon; no sample size repairs it.
What carries the argument
Lemma 1 gives a zero-communication realisation: conspirators compute a shared Gaussian latent z_t from a pseudorandom function on public round data, then form each bid's unexplained component as F_i^{-1}(Φ(√ρ z_t + √(1−ρ) e_it)). Because the inner term is standard normal for every ρ, each bidder's own component has the prescribed law F_i exactly; ρ=0 reproduces independence and ρ=1 gives the comonotone coupling that minimises the expected maximum bid. The independence of z_t across rounds is what makes the entire single-bidder process law identical to the competitive one, which is the load-bearing step for the power-equals-size result.
Load-bearing premise
The entire impossibility result rests on the shared random draw being freshly independent each round; if it is persistent, one agent's own bids become serially correlated and the conduct becomes detectable.
What would settle it
Run the Lemma 1 construction with the shared draw replaced by an AR(1) common signal: at persistence 0.5 a single bidder's residual series shows lag-one autocorrelation of about 0.40, and at 0.9 even a marginal goodness-of-fit test rejects. Observing serial dependence in one agent's residuals would therefore falsify the claim that the conduct leaves every single-agent test powerless.
If this is right
- Single-agent price-level audits—the standard empirical tool for algorithmic collusion—have power exactly equal to their false-positive rate against this conduct, at any sample size.
- The conduct is profitable at meaningful magnitudes: a five-member coalition takes roughly 4.6% to 10.6% of seller revenue in the Ethereum block-building market, comparable to the price-lift effects reported in the literature.
- Regret-based transcript audits also fail, but for a different reason: the coupling does not raise calibrated swap regret above the honest market's own off-best-response noise.
- In deployed data, the honest population of bidder pairs is so dependent that the smallest detectable coupling is bounded below by a constant 20 to 32 times the family-wise sampling threshold, so longer audits do not help.
- Shifting from detecting conduct to counting independent operators raises the measured Herfindahl index by 247.5% to 324.5%, a decidable regulatory target from public bid data.
Where Pith is reading between the lines
- If the shared draw were persistent rather than independent across rounds, the invisibility would collapse: one bidder's own series would show autocorrelation and a single-agent test would gain power. Auditors could look for serial dependence in residuals as a first check.
- The temperature result suggests a deployable mitigation—mandating higher sampling temperatures for bidding language models could mechanically weaken the coupling the conduct requires—but the paper does not quantify rent at the higher end, so that trade-off is untested.
- The paper's dual-use framing implies the coupling arises by default from shared model supply chains; a regulatory response aimed at source concentration (how many independent models underpin a market) might be more tractable than screening bids.
- Entity counting from public bid streams is a portable technique: any market with persistent identity and public bid data—not just block building—could use pairwise residual correlation to estimate the number of independent decision-makers.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a conduct in which bidding agents preserve each agent's own marginal bid distribution exactly while coupling the unexplained components through an iid shared draw (Lemma 1). The central theorem (Proposition 1) states that any test using a single agent's bid history has power exactly equal to its size against this conduct, for all coupling strengths up to comonotonicity. The paper then argues that price-level detection methods are therefore blind by construction, presents a language-model-agent replication showing same-base residual correlation (+0.053 under the strictest auditor at temperature 0.3) and a temperature trend, calibrates an ambient dependence floor on MEV-Boost data, and proposes entity counting as a more tractable regulatory target.
Significance. If the formal claim is repaired, the paper makes a meaningful and non-obvious point: a profitable conspiracy can be statistically invisible to single-agent marginal audits, and the impossibility is not a sample-size issue. The simulation checks, the reproducible scripts, and the transparent reporting of swept parameters are notable strengths, as is the attempt to measure the honest-population dependence floor in a real auction market. The empirical calibration to LLM agents and MEV-Boost is novel and policy-relevant. However, the current proof of Proposition 1 has a gap concerning the shared draw's dependence on auditor covariates, and the paper's broader 'price-level audits are blind' claim is narrower than the title suggests. These issues are fixable but need to be addressed before the central result is fully established.
major comments (2)
- [Lemma 1, Proposition 1; Figure 1(a)] The proof of Proposition 1 asserts that the single-bidder process law is unchanged because z_t are independent across rounds. This is insufficient: the test input is {(b_it, x_t)}, so the conditional law of ε_it given x_t must also be unchanged. Lemma 1 sets z_t = Φ^{-1}(PRF_k(id_t)) with id_t public round data, and Figure 1 labels the shared PRF as being 'on x_t'. If id_t includes or equals x_t, then ε_it is a deterministic function of x_t plus noise; e.g., with F_i = N(0,1), z_t = x_t, the coefficient on x_t in b_it becomes β_i + σ√ρ, so a t-test on that coefficient has power strictly greater than α for any ρ>0. The theorem is repairable by requiring z_t independent of x_t (e.g., an independent shared seed), but as written the construction and Proposition 1 are inconsistent. Appendix C checks persistence of the shared draw but not its dependence on covariates.
- [Proposition 3; 'Why the Empirical Literature Cannot See This'] Proposition 3 defines a 'price-level test' as a test whose input is one agent's marginal price or bid distribution. The title and abstract, however, claim that 'price-level audits' and the published detection methodology are blind by construction. This is a scope mismatch: tests based on market-level or aggregate price series are not covered by Proposition 3. Under Lemma 1, the cross-sectional average bid has the same mean but a different variance under coupling, so a test exploiting that variance would have power greater than α. Please either restrict the 'published methodology' claim to single-agent marginal tests or extend the analysis to aggregate price-level statistics.
minor comments (3)
- [Abstract and Table 2] The headline +0.053 is one temperature (T=0.3) under auditor (C); at T=0.0 and T=0.6 the clustered 95% CI includes zero. Consider leading with the monotone temperature trend or accounting for multiple comparisons across temperatures.
- [Proposition 4 and Conclusion] The conclusion states that regret-based audits 'do not separate' the conduct, relying on Proposition 4(iii), which the authors themselves describe as simulation-verified and 'left as a conjecture in general.' This should be stated more conditionally in the conclusion, as the main text already does.
- [Appendix F and Limitations] The claim that 'No parameter in this paper is tuned to maximise a reported quantity' is hard to reconcile with the acknowledgement in the Limitations that thresholds are not derived from a loss function but chosen from monotone sweeps. Please reconcile these statements.
Circularity Check
No significant circularity: the formal core is a self-contained construction and the empirical sections are measurements, not fitted predictions.
full rationale
Proposition 1 follows directly from Lemma 1, which explicitly constructs a coupling with i.i.d. shared draws and fixed marginals, so any single-bidder test has power equal to its size by a standard law-preservation argument. Proposition 2 cites external classical results (Mueller-Stoyan; Puccetti-Wang). The empirical sections are measurements and sweeps, and the paper explicitly states that thresholds are not derived from a loss function and reports all swept ranges. No load-bearing self-citation appears; the shared-model coupling evidence is external (Kleinberg-Raghavan; Spiro). The definition of 'price-level test' makes Proposition 3 a corollary, but it is grounded in Lemma 1's construction rather than assuming the conclusion. The only caveat is a proof-completeness concern: Lemma 1 lets the shared draw be a PRF on public round data, and Proposition 1 only shows the unconditional single-bidder law is unchanged; if the shared draw is a function of auditor covariates x_t, the joint law of (b_it, x_t) may differ. This is an assumption gap, not a circular reduction, so it does not affect the circularity score.
Axiom & Free-Parameter Ledger
free parameters (3)
- submission-timing gap filter =
swept 0, 1, 2, 4 s; floor drops 0.81 to 0.50
- detector threshold =
0.80, 0.85, 0.90, 0.93, 0.95 (sweep)
- clustering threshold =
0.85, 0.90, 0.95 (sweep)
axioms (5)
- standard math Comonotone coupling is extremal for supermodular expectations (Fréchet bounds).
- domain assumption Shared draw z_t is independent across rounds.
- domain assumption First-price auction with private values, bidders use equilibrium bid 1/2 v + σ ε; coalition faces bidders with unexplained dispersion.
- ad hoc to paper Co-submission in MEV-Boost implies shared infrastructure.
- domain assumption Auditor (C) residualization correctly captures the unexplained component.
invented entities (1)
-
Shared pseudorandom coupling draw z_t (PRF keyed on public round data)
independent evidence
Cite this review
Pith. "Pith review of Collusion with Competitive Marginals: Price-Level Audits Are Blind by Construction." pith.science (2026). https://pith.science/paper/ZSRSMQEW
@misc{pith2026260726385,
author = {Pith},
title = {Pith review of: Collusion with Competitive Marginals: Price-Level Audits Are Blind by Construction},
year = {2026},
howpublished = {\url{https://pith.science/paper/ZSRSMQEW}},
note = {Machine review of arXiv:2607.26385}
}
abstract
Empirical work on algorithmic collusion asks one question of the data: are prices supracompetitive? We show this can be answered "no" by a conspiracy that is nonetheless profitable. Consider bidding agents that couple only through the joint distribution of their unexplained bid components, leaving every agent's own bid law exactly at the competitive law. Any test whose input is a single agent's price or bid history then has power exactly equal to its false-positive rate, for every coupling strength up to comonotonicity. The published detection methodology is therefore blind to this conduct by construction rather than underpowered, and no sample size repairs it. Three empirical results follow. First, the mechanism appears in real language-model agents: twenty models from nineteen independent developers, three deployment prompts each, show residual correlation of $+0.053$ between two deployments of one model against $+0.0001$ across models, with a 95% interval clustered by developer of $[0.030, 0.078]$, under an auditor that sees every order feature and is fitted out of sample. Second, the coupling falls monotonically as sampling temperature rises ($p=0.002$), turning a deployment parameter into a candidate mitigation. Third, on 24 days of Ethereum block-building auction data covering 77,684 bids from 39 bidders, the honest population of bidder pairs is itself so dependent that a screen held at a 5% false-positive rate must sit above a floor of $+0.50$ to $+0.81$, which is 20 to 32 times the family-wise sampling threshold and does not fall as the audit window grows. Since lawful multi-identity operation and conspiracy are behaviourally indistinguishable here, the tractable regulatory target is not detection but counting: resolving 40 bidding identities into 23 operators raises the Herfindahl index by 247.5%, and adding behavioural clusters from public bid streams reaches 324.5%.
Figures
Reference graph
Works this paper leans on
-
[1]
Assad, S.; Calvano, E.; Calzolari, G.; Clark, R.; Denicol \`o , V.; Ershov, D.; Johnson, J.; Pastorello, S.; Rhodes, A.; Xu, L.; and Wildenbeest, M. 2021. Autonomous Algorithmic Collusion: Economic Research and Policy Implications. Oxford Review of Economic Policy, 37(3): 459--478
2021
-
[2]
Calvano, E.; Calzolari, G.; Denicol \`o , V.; and Pastorello, S. 2020. Artificial Intelligence, Algorithmic Pricing, and Collusion. American Economic Review, 110(10): 3267--3297
2020
-
[3]
Chassang, S.; Kawai, K.; Nakabayashi, J.; and Ortner, J. 2022. Robust Screens for Noncompetitive Bidding in Procurement Auctions. Econometrica, 90(1): 315--346
2022
-
[4]
Fish, S.; Gonczarowski, Y. A.; and Shorrer, R. I. 2024. Algorithmic Collusion by Large Language Models. arXiv:2404.00806
arXiv 2024
-
[5]
Flashbots . 2024. MEV -Boost Specifications. Flashbots Documentation. Accessed 2026-07-27
2024
-
[6]
Harrington, J. E., Jr. 2008. Detecting Cartels. In Buccirossi, P., ed., Handbook of Antitrust Economics. Cambridge, MA: MIT Press
2008
-
[7]
D.; Long, S.; and Zhang, C
Hartline, J. D.; Long, S.; and Zhang, C. 2024. Regulation of Algorithmic Collusion. In Proceedings of the Symposium on Computer Science and Law (CSLAW), 98--108. ACM
2024
-
[8]
D.; Wang, C.; and Zhang, C
Hartline, J. D.; Wang, C.; and Zhang, C. 2025. Regulation of Algorithmic Collusion, Refined: Testing Pessimistic Calibrated Regret. In Proceedings of the Symposium on Computer Science and Law (CSLAW), 108--120. ACM
2025
-
[9]
Keppo, J.; Li, Y.; Tsoukalas, G.; and Yuan, N. 2026. On the Fragility of AI Agent Collusion. arXiv:2603.20281
arXiv 2026
-
[10]
Kleinberg, J.; and Raghavan, M. 2021. Algorithmic Monoculture and Social Welfare. Proceedings of the National Academy of Sciences, 118(22): e2018340118
2021
-
[11]
Y.; Ojha, S.; Cai, K.; and Chen, M
Lin, R. Y.; Ojha, S.; Cai, K.; and Chen, M. F. 2024. Strategic Collusion of LLM Agents: Market Division in Multi-Commodity Competitions. arXiv:2410.00031
Pith/arXiv arXiv 2024
-
[12]
M \"u ller, A.; and Stoyan, D. 2002. Comparison Methods for Stochastic Models and Risks. Wiley Series in Probability and Statistics. Chichester: Wiley
2002
-
[13]
Nelsen, R. B. 2006. An Introduction to Copulas. Springer Series in Statistics. New York: Springer, 2nd edition
2006
-
[14]
OECD . 2022. Data Screening Tools for Competition Investigations. OECD Roundtables on Competition Policy Papers No. 284, OECD Publishing, Paris
2022
-
[15]
Puccetti, G.; and Wang, R. 2015. Extremal Dependence Concepts. Statistical Science, 30(4): 485--517
2015
-
[16]
Spiro, T. 2026. The Oracle's Fingerprint: Correlated AI Forecasting Errors and the Limits of Bias Transmission. arXiv:2605.00844
Pith/arXiv arXiv 2026
-
[18]
Wu, F.; Thiery, T.; Leonardos, S.; and Ventre, C. 2024. To Compete or Collude: Bidding Incentives in E thereum Block Building Auctions. In Proceedings of the 5th ACM International Conference on AI in Finance (ICAIF) , 813--821. ACM
2024
-
[19]
Zhang, A.; Liu, Y.; Zhang, R.; Shan, Y.; and Wu, Y. 2026. Order Flow Exclusivity and Value Extraction Mechanisms: An Analysis of E thereum Builder Centralization. arXiv:2605.04471
Pith/arXiv arXiv 2026
-
[20]
Artificial Intelligence, Algorithmic Pricing, and Collusion , journal =
Calvano, Emilio and Calzolari, Giacomo and Denicol. Artificial Intelligence, Algorithmic Pricing, and Collusion , journal =
-
[21]
Proceedings of the National Academy of Sciences , volume =
Kleinberg, Jon and Raghavan, Manish , title =. Proceedings of the National Academy of Sciences , volume =
-
[22]
2026 , eprint =
Spiro, Theodor , title =. 2026 , eprint =
2026
-
[23]
and Shorrer, Ran I
Fish, Sara and Gonczarowski, Yannai A. and Shorrer, Ran I. , title =. 2024 , eprint =
2024
-
[24]
and Ojha, Siddhartha and Cai, Kevin and Chen, Maxwell F
Lin, Ryan Y. and Ojha, Siddhartha and Cai, Kevin and Chen, Maxwell F. , title =. 2024 , eprint =
2024
-
[25]
2026 , eprint =
Keppo, Jussi and Li, Yuze and Tsoukalas, Gerry and Yuan, Nuo , title =. 2026 , eprint =
2026
-
[26]
and Long, Sheng and Zhang, Chenhao , title =
Hartline, Jason D. and Long, Sheng and Zhang, Chenhao , title =. Proceedings of the Symposium on Computer Science and Law (CSLAW) , pages =
-
[27]
and Wang, Chang and Zhang, Chenhao , title =
Hartline, Jason D. and Wang, Chang and Zhang, Chenhao , title =. Proceedings of the Symposium on Computer Science and Law (CSLAW) , pages =
-
[28]
Proceedings of the 5th
Wu, Fei and Thiery, Thomas and Leonardos, Stefanos and Ventre, Carmine , title =. Proceedings of the 5th
-
[29]
2026 , eprint =
Zhang, Ao and Liu, Yunwen and Zhang, Ren and Shan, Yingdi and Wu, Yongwei , title =. 2026 , eprint =
2026
-
[30]
Time to Bribe: Measuring Block Construction Market , year =
Wahrst. Time to Bribe: Measuring Block Construction Market , year =. 2305.16468 , archivePrefix =
-
[31]
Comparison Methods for Stochastic Models and Risks , series =
M. Comparison Methods for Stochastic Models and Risks , series =
-
[32]
Statistical Science , volume =
Puccetti, Giovanni and Wang, Ruodu , title =. Statistical Science , volume =
-
[33]
Econometrica , volume =
Chassang, Sylvain and Kawai, Kei and Nakabayashi, Jun and Ortner, Juan , title =. Econometrica , volume =
-
[34]
, title =
Harrington, Jr., Joseph E. , title =. Handbook of Antitrust Economics , editor =
-
[35]
2022 , url =
Data Screening Tools for Competition Investigations , howpublished =. 2022 , url =
2022
-
[36]
2026 , eprint =
Xie, Yi and Zhou, Zhanke and Cao, Chentao and Liu, Bo and Han, Bo , title =. 2026 , eprint =
2026
-
[37]
, title =
Nelsen, Roger B. , title =
-
[38]
van der Vaart, A. W. , title =
-
[39]
Autonomous Algorithmic Collusion: Economic Research and Policy Implications , journal =
Assad, Stephanie and Calvano, Emilio and Calzolari, Giacomo and Clark, Robert and Denicol. Autonomous Algorithmic Collusion: Economic Research and Policy Implications , journal =
-
[40]
and Savage, Stefan , title =
Meiklejohn, Sarah and Pomarole, Marjori and Jordan, Grant and Levchenko, Kirill and McCoy, Damon and Voelker, Geoffrey M. and Savage, Stefan , title =. Proceedings of the 13th
This paper was first reviewed by deepseek-v4-flash on August 1, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.