REVIEW 5 major objections 4 minor 2 cited by
Fill-Side Behavioral Concentration on Polymarket: Identification Limits under Record-Level Attribution
T0 review · 5 major / 4 minor · reviewed 2026-08-02 · deepseek-v4-flash
Pith's one-line read Polymarket's public fill record yields one behavioral cluster, not four-to-five archetypes, and cannot identify quote-based market making.
desk verdict The identification-limits claim is real and worth keeping; the empirical headline claims are, by the paper's own abstract, not stable, and the body hasn't caught up with that fact. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is a validity-gate split between fill attribution and quote-lifecycle attribution. The paper separates 'can we see who executed a fill?' (yes, via maker and taker on every OrderFilled log) from 'can we see who posted and withdrew quotes?' (no, because those events are off-chain), and forces all downstream analysis through the second gate's failure. The active object is then a six-feature fill-side behavioral vector per address — log trade intensity, log average notional, directional ratio, market Herfindahl concentration, intraday entropy, and log market breadth — standardized and fed to DBSCAN; the one-cluster output is interpreted as a null. Identification itself
What would settle it
Re-run the pipeline on the same logs after normalizing fills by matched maker-taker pairs and excluding or separately coding mint/burn executions. If the single dense cluster splits into multiple density-separated clusters, or if the 12.6%-of-addresses / 81.4%-of-notional split moves materially, the headline results are artifacts of the attribution rule rather than properties of the venue.
Extended reading notes
Core claim
The paper claims a bounded empirical result about Polymarket's public fills. Quote placement and cancellation events are off-chain and absent from logs, so market-making, spoofing, and quote withdrawal cannot be identified at address level. Density-based clustering over a six-feature fill-side vector (intensity, notional, directional ratio, market concentration, entropy, breadth) returns one dense cluster with zero noise in the sampled week, refuting the pre-registered four-to-five archetype hypothesis — kept as a null; the abstract cautions both headline numbers are conditional on crediting maker and taker on every fill. Pre-registered tier thresholds isolate 12.6% of addresses holding 81.4
Load-bearing premise
The analysis assumes each fill record can be credited to two meaningful addresses with the full notional counted twice; if order splitting or contract mint/burn executions break that assumption, the feature vector, the single-cluster null, and the headline concentration shares can change without any real change in trader behavior.
Editorial extensions
If this is right
- Quote-based participant types — market makers, passive liquidity providers, spoofers — cannot be labeled at address level from Polymarket public data; future studies must either obtain off-chain order data or restrict themselves to fill-side proxies.
- The four-to-five archetype view of retail versus non-retail behavior is not supported for the sampled week; high-end operators sit in the tails of one continuous fill-behavior distribution.
- Tier-based stratification gives a threshold-stable separation: roughly one in eight addresses accounts for over four-fifths of fill notional, so retail and non-retail are separable without clustering.
- Engine-calibration studies should replace fixed synthetic retail order sizes with the measured retail per-fill notional near $4.77, a distribution-aware correction the paper argues is strong.
- The paper's negative results document that Polymarket-class venues are structurally opaque on the quote side: absence of market-making evidence reflects absence of data, not absence of market makers.
Reading between the lines
- If a match-normalized attribution were applied (crediting one execution per matched pair rather than both maker and taker), the per-address notional ranking and the 81.4% share would likely change; the paper's own abstract concedes this, but the body's headline table remains maker-plus-taker arithmetic.
- The one-cluster null may be an artifact of the feature set: the six fill-side features exclude spread width, quote lifetime, and two-sidedness, which are precisely the dimensions that typically separate market makers from directional traders. A venue exposing quote events might still show multiple archetypes.
- The retail per-fill scale discovered here (about $4.77) suggests synthetic-trader evaluation grids in event-linked perpetual designs are calibrated an order of magnitude too large; if real users transact in such small slices, leverage and liquidation thresholds behave differently in simulation than assumed.
- A testable extension: run the same tier thresholds on a second, non-overlapping week and on a match-normalized fill set; movement of the 81.4% figure by more than a few points would indicate the concentration is an attribution convention rather than a venue property.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper claims to characterize non-retail participation on Polymarket from 13,356,931 OrderFilled events on the CTFExchange contract over 2026-04-21 to 2026-04-27. Its central empirical claims are: (i) a six-feature fill-side behavioral vector is uni-modal under DBSCAN (one dense cluster, zero noise, across fifteen sensitivity configurations), refuting a pre-registered four-to-five archetype hypothesis; and (ii) feature-tier stratification shows that whale, high-frequency-operator, and power-trader tiers — 12.6% of addresses — hold 81.4% of fill notional. The paper also documents a structural validity-gate failure (G-QUOTE-LIFE) because Polymarket's off-chain CLOB does not expose OrderPlaced/OrderCancelled events, and reports a range of descriptive microstructure correlations, manipulation-pattern candidates, and Paper 1 feedback tests. An appended abstract, however, corrects the empirical scope to approximately 25–28 April 2026, concedes that the maker+taker attribution convention is not invariant to match fragmentation, states that mint/burn executions do not admit a universal buyer/seller interpretation, and downgrades the one-cluster result to 'a null under the original record-level representation' rather than evidence of intrinsic unimodality.
Significance. If the empirical results were supportable, the paper would make a useful supply-side contribution to prediction-market microstructure: it provides reproducible infrastructure, a three-gate validity framework, honest handling of the quote-lifecycle data limitation, and a clear separation of fill-side from quote-side claims. The associated repository, manifests, and derived dataset are constructive. However, the load-bearing empirical claims are not established as stated. The body of the manuscript presents the single-cluster 'unimodality' finding and the 81.4% concentration table as substantive behavioral results, while the appended abstract — part of the same record — concedes that both are representation-dependent artifacts of the record-level attribution convention. The window correction also undermines the comparability and specificity of every reported statistic. These are not presentation issues; they affect the paper's main conclusions. The durable methodological content (G-QUOTE-LIFE as a structural limit, the need for match-normalized and mint-aware replication) is real but does not rescue the headline empirical claims.
major comments (5)
- [Abstract vs. Sections 1, 3.1, 4, 12] The appended abstract states the archived extraction covers approximately 25–28 April 2026, not the 21–27 April window used throughout the body (e.g., Section 1, Section 3.1, Table 2, Figure 2, Section 12). All population counts, tier shares, cluster results, and microstructure panels are window-specific; the body has not been reconciled with this correction. This is a load-bearing inconsistency, not a typo, because the sample composition and all headline numbers change with the window.
- [Sections 3.2–3.5, 4.2, 4.3] The data construction credits both maker and taker addresses on every OrderFilled record and appears to include mint/burn executions. The appended abstract concedes that this convention 'is not invariant to match fragmentation' and that mint/burn executions 'do not admit a universal buyer/seller interpretation.' Every central feature — f2 (fill counts), f3 (per-fill notional), f5 (directional ratio), and therefore the tier shares and the DBSCAN null — depends on this convention. The body nevertheless reports the single-cluster and 81.4% results as stable behavioral findings. Since the numbers can change without any change in real trader behavior, the headline claims are not identified as venue properties.
- [Section 4.3] A single DBSCAN cluster with zero noise is a property of the chosen density partition and feature scaling; it is not a statistical test of unimodality. The abstract itself says the result is 'retained only as a null under the original record-level representation.' The body's language in Sections 4.5 and 12 ('the behavioral space is uni-modal', 'pre-registered hypothesis ... empirically refuted') overstates the evidentiary content. The correct statement — that one density cluster was observed under one attribution convention — does not support the paper's claimed refutation of behavioral archetypes.
- [Section 4.2, Tables 2 and 3] The 81.4%-of-notional / 12.6%-of-addresses concentration is arithmetic given the tier definitions, whose P75/P95 thresholds are computed from the same 77,203-address sample. It therefore does not independently demonstrate 'robust retail-vs-non-retail separation'; it restates the quantile cutoffs in notional terms. Table 3 shows the strict non-retail tier size varies from roughly 9,700 to 16,500 across threshold variants, but notional shares for those variants are not reported, so the robustness of the 81.4% figure is not established.
- [Section 4.5 vs Table 2] Section 4.5 states 'approximately 91% of total notional concentrates in the top four tiers (≈39% of addresses)', while Table 2's extended non-retail subtotal, which includes five non-retail tiers, is 93.2% of notional across 17.93% of addresses; the top four tiers (whale, high-frequency operator, power trader, active retail) sum to 92.0%. These internal inconsistencies make the headline concentration claims difficult to audit and should be corrected regardless of the larger attribution issue.
minor comments (4)
- [Section 3.7 / Table 2] The address count after CTFExchange exclusion is given as 77,204 in Section 3.7 and 77,203 in Table 2 and Figure 3. Please reconcile.
- [Section 2.2] The reference to 'Dubach (2026)' is incomplete: the bibliography states 'Complete citation pending venue identification at camera-ready.' This is not acceptable in a submitted manuscript.
- [Title] The arXiv title 'Fill-Side Behavioral Concentration on Polymarket: Identification Limits under Record-Level Attribution' differs from the internal title 'Fill-Side Non-Retail Trading on Polymarket...'. Align them.
- [Sections 5.1 and 9.8] Several analyses are described as deferred to follow-up work (e.g., per-address 5-minute post-fill price moves, per-class wash-volume breakdown, realized-spread distributions), yet related correlations are reported in the bilateral tables. Please clarify which numbers are final and which are placeholders.
Circularity Check
Partial definitional circularity in the tier-stratification 'retail vs non-retail' finding and in the T3 retail-notional refutation; the DBSCAN unimodality result is not circular but is scope-limited by the paper's own appended abstract.
-
self definitional
[Section 4.2, Table 2, and appended abstract]
"Whale-tier (notional overlay; locked in r0.4.4 Section 8.1): total notional ≥ $1,000,000 ... Power trader: f2 ≥ P75 and total notional ≥ P75 ... Episodic retail: total notional < $10,000 ... The 81.4%-of-notional concentration in 12.6% of addresses ... is the substantive empirical finding: retail-vs-non-retail separation is robust on these fill-side data."
The non-retail tiers are defined by thresholds on total notional and trade intensity computed from the same 77,203-address sample. The conclusion that these tiers hold 81.4% of notional is therefore a within-sample summary of the very variables used to define the tiers, not an independent test of a 'non-retail' population. The exact share is arithmetic once the tiers are set; the qualitative 'separation' is entailed by construction because the tiers select high-notional/intensity addresses. The appended abstract concedes: 'The concentration table is likewise attribution-weighted arithmetic under that convention.'
-
fitted input called prediction
[Section 4.3 Table 4 and Section 10 (T3)]
"K5-Retail-sell-skew ... f3 = 0.679 ... K5-Retail-buy-skew ... f3 = 0.709 ... The fill-side empirical run measured the mean per-fill notional of the retail-proximate k-means partitions (K5-Retail-sell-skew and K5-Retail-buy-skew) at ≈ $4.77 USDC. Paper 1's E2/E3 evaluation parameterized synthetic retail traders with fixed notional $1,000 per fill ... Strong empirical refutation."
The k-means input vector includes f3 = log average notional, and the clusters labeled 'Retail' were assigned that mnemonic partly because their f3 is low. T3 then uses the same f3 (restated as ≈$4.77 per fill) to 'refute' Paper 1's $1,000 retail notional parameter. The comparison is therefore not an independent estimate of retail trade size; it restates the low-f3 input of the very clusters chosen as 'retail-proximate.' The paper's caveat that labels are mnemonic softens but does not remove the circularity, since the load-bearing contrast to Paper 1 still relies on the same feature.
full rationale
The DBSCAN unimodality result is a genuine empirical output of a clustering algorithm applied to the six-feature vector; it does not reduce to its inputs by construction. The appended abstract's statement that the one-cluster result is 'retained only as a null under the original record-level representation' is a scope/validity caveat rather than a circularity, and it weakens the body's stronger interpretation without making the derivation circular. Heavy self-citation to ForesightFlow and Papers 1-3 is present, but it is not load-bearing for the paper's central empirical claims: the G-QUOTE-LIFE failure is documented from the paper's own data, and the ForesightFlow metrics are either reproduced or cited as methodology rather than used to force the present conclusions. The main circularity is in the tier-stratification 'finding': the non-retail tiers are defined by pre-registered thresholds on notional and intensity computed on the same sample, so the 81.4%-of-notional concentration is arithmetic under that definition, and the qualitative 'separation' is entailed by the selection rule. A secondary, less central circularity appears in T3, where the 'retail' k-means partitions are labeled using low f3 and then their low f3 is presented as a refutation of Paper 1's retail notional parameter. These are partial rather than total circularities because the exact percentages and magnitudes are empirical, and the DBSCAN result retains independent content. Score 4 reflects this partial, definitional circularity without treating the whole paper as a self-citation chain.
Assumptions & free parameters
free parameters (6)
- Activity threshold =
≥5 fills per address
- Whale-tier notional threshold =
$1,000,000 total fill notional
- High-frequency / power-trader percentile thresholds =
f2≥P95 & f9≥P75 (HFO); f2≥P75 & notional≥P75 (power trader)
- Episodic retail notional threshold =
total notional < $10,000
- DBSCAN sensitivity grid =
ε∈[1.15,3.44], minPts∈{10,20,30}
- Winsorization bounds =
features at p99.5; Kyle's λ at [P01,P99]=[-4.2042,+0.1052]
assumptions (5)
- ad hoc to paper Each OrderFilled log with maker+taker addresses can be treated as one economically comparable fill for both counterparties.
- domain assumption The PMXT v2 archive faithfully and completely records the executed fills used in the analysis.
- ad hoc to paper A single DBSCAN cluster with zero noise can be interpreted as behavioral uni-modality.
- domain assumption Sports-dominant single-week data are adequate for venue-level behavioral characterization.
- domain assumption Polymarket quote-lifecycle events (OrderPlaced, OrderCancelled) are truly absent from all public channels used here.
Cite this review
Pith. "Pith review of Fill-Side Behavioral Concentration on Polymarket: Identification Limits under Record-Level Attribution." pith.science (2026). https://pith.science/paper/XH2I5WFO
@misc{pith2026260511640,
author = {Pith},
title = {Pith review of: Fill-Side Behavioral Concentration on Polymarket: Identification Limits under Record-Level Attribution},
year = {2026},
howpublished = {\url{https://pith.science/paper/XH2I5WFO}},
note = {Machine review of arXiv:2605.11640}
}
read the original abstract
This paper studies behavioral concentration in Polymarket's public executed-fill record and formalizes what that record can and cannot identify. A pre-publication reconciliation corrects the empirical scope: the archived extraction covers the legacy CTF Exchange over Polygon blocks 86,008,447-86,107,178, approximately 25 April 2026 17:09 UTC through 28 April 2026 00:00 UTC, rather than the full 21-27 April week stated previously. It contains 13,356,931 OrderFilled records, 77,204 addresses with at least five attributed records, and 43,116 token identifiers; negative-risk markets are absent. The archived feature construction credits both maker and taker addresses on each record. This convention is not invariant to match fragmentation, and mint/burn executions do not admit a universal buyer/seller interpretation. The reported one-cluster result is therefore retained only as a null under the original record-level representation, not as evidence that the participant population is intrinsically unimodal. The concentration table is likewise attribution-weighted arithmetic under that convention. Two methodological results remain durable: public fills do not identify the quote lifecycle required to infer market making, spoofing, or strategic withdrawal; and a one-cluster result rejects density separation in an observed feature space, not latent economic heterogeneity. A match-normalized, mint-aware, multi-window replication is required before treating the cluster null or exact cohort shares as stable venue properties.
Figures
Figures from the paper (3 more)
Forward citations
Cited by 2 Pith papers
-
Axient: Debt-Free Finality for Leveraged Binary Event Markets
Axient clears leveraged event-position debt before finality by selling the smallest quantity whose worst-case settled proceeds cover an upper debt bound, removing terminal payout risk from the lender's loan channel wi...
-
Axient: On-Chain Credit and Loss Allocation for Leveraged Event Markets: A Venue-Agnostic Protocol for Traders, Credit Providers, Market Makers, and Liquidation Backstops
A formal on-chain credit architecture for leveraged event markets, with a synthetic stress test showing layered protection reduces but does not eliminate senior-lender loss and bonded market-maker capacity can expand risk.
Reference graph
Works this paper leans on
-
[1]
The Adoption of Blockchain-based Decentralized Exchanges
Capponi, Agostino and Ruizhe Jia (2021). “The Adoption of Blockchain-based Decentralized Exchanges”. In:Working paper. 50 Dubach (2026). “Polymarket Anatomy”. Working paper / preprint, 2026; cited in Paper 1 for depth profile geometric grid distribution. Complete citation pending venue identification at camera-ready
2021
-
[2]
A Density-Based Algorithm for Discovering Clusters in Large Spatial Databases with Noise
Ester, Martin, Hans-Peter Kriegel, Jörg Sander, and Xiaowei Xu (1996). “A Density-Based Algorithm for Discovering Clusters in Large Spatial Databases with Noise”. In:Proceedings of the Second International Conference on Knowledge Discovery and Data Mining (KDD- 96), pp. 226–231. ForesightFlow (2026).ForesightFlow Datasets. Online repository, accessed 2026...
1996
-
[3]
Strategic Trading When Agents Forecast the Forecasts of Others
Foster, F. Douglas and S. Viswanathan (1996). “Strategic Trading When Agents Forecast the Forecasts of Others”. In:Journal of Finance51.4, pp. 1437–1478
1996
-
[4]
Bid, Ask and Transaction Prices in a Specialist Market with Heterogeneously Informed Traders
Glosten, Lawrence R. and Paul R. Milgrom (1985). “Bid, Ask and Transaction Prices in a Specialist Market with Heterogeneously Informed Traders”. In:Journal of Financial Economics14.1, pp. 71–100
1985
-
[5]
Combinatorial Information Market Design
Hanson, Robin (2003). “Combinatorial Information Market Design”. In:Information Systems Frontiers5.1, pp. 107–119
2003
-
[6]
Oxford University Press
Hasbrouck, Joel (2007).Empirical Market Microstructure: The Institutions, Economics, and Econometrics of Securities Trading. Oxford University Press
2007
-
[7]
(2020).BlockSci: Design and Applications of a Blockchain Analysis Platform
Kalodner, Harry et al. (2020).BlockSci: Design and Applications of a Blockchain Analysis Platform
2020
-
[8]
Continuous Auctions and Insider Trading
Kyle, Albert S. (1985). “Continuous Auctions and Insider Trading”. In:Econometrica53.6, pp. 1315–1335
1985
Show all 16 references
-
[9]
Decentralized Exchanges
Lehar, Alfred and Christine A. Parlour (2022). “Decentralized Exchanges”. In:Working paper
2022
-
[10]
Interpreting the Predictions of Prediction Markets
Manski, Charles F. (2006). “Interpreting the Predictions of Prediction Markets”. In:Economics Letters91.3, pp. 425–429
2006
-
[11]
A Fistful of Bitcoins: Characterizing Payments Among Men with No Names
Meiklejohn, Sarah, Marjori Pomarole, Grant Jordan, Kirill Levchenko, Damon McCoy, Geoffrey M. Voelker, and Stefan Savage (2013). “A Fistful of Bitcoins: Characterizing Payments Among Men with No Names”. In:Internet Measurement Conference (IMC)
2013
-
[12]
A Taxonomy of Event-Linked Perpetual Futures: Variant Designs Beyond the Single-Market Binary Case
Nechepurenko, Maksym (2026a). “A Taxonomy of Event-Linked Perpetual Futures: Variant Designs Beyond the Single-Market Binary Case”. Paper 2, four-paper Event-Linked Perpet- uals programme. Working paper, Devnull Research. Available at SSRN:https://papers. ssrn.com/abstract=674...
-
[13]
Resolution-Aware Perpetual Futures on Binary Prediction Markets: An Empirical Risk-Design Framework Using Polymarket Data
Nechepurenko, Maksym (2026h). “Resolution-Aware Perpetual Futures on Binary Prediction Markets: An Empirical Risk-Design Framework Using Polymarket Data”. Paper 1, four- paper Event-Linked Perpetuals programme. Working paper, Devnull Research. Available at SSRN: https://papers...
-
[14]
Silhouettes: A Graphical Aid to the Interpretation and Validation of Cluster Analysis
Rousseeuw, Peter J. (1987). “Silhouettes: A Graphical Aid to the Interpretation and Validation of Cluster Analysis”. In:Journal of Computational and Applied Mathematics20, pp. 53–65
1987
-
[15]
Inferring the Components of the Bid-Ask Spread: Theory and Empirical Tests
Stoll, Hans R. (1989). “Inferring the Components of the Bid-Ask Spread: Theory and Empirical Tests”. In:Journal of Finance44.1, pp. 115–134
1989
-
[16]
Prediction Markets
Wolfers, Justin and Eric Zitzewitz (2004). “Prediction Markets”. In:Journal of Economic Perspectives18.2, pp. 107–126. 52
2004
Reviewed August 2, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.