REVIEW 3 major objections 5 minor 26 references
On Devon Allen's Disqualification at the 2022 World Track and Field Championships
T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read The paper argues that Devon Allen's 0.099-second reaction at the 2022 World Championships was not statistically impossible and that the 0.1-second disqualification threshold is too strict for men.
desk verdict A solid applied analysis of the 2022 reaction-time anomaly; the timing-system claim is too strong, and the 1-in-362 tail probability needs uncertainty bounds. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The argument is carried by two statistical tools. The first is a rank-sum test for clustered data, in which each athlete is a cluster and each observed reaction time is a subunit; this test decides whether the distribution of reaction times shifts between competitions while respecting within-athlete dependence. The second is a generalized Gamma distribution $GG(\mu,\sigma,\nu)$ with random effects for venue and for heat, fitted inside the GAMLSS framework (generalized additive models for location, scale, and shape). The generalized Gamma shape parameter lets the model capture asymmetry in reaction times, the venue random effect absorbs year-to-year differences such as 2022, and the heat random effect absorbs race-to-race variability; simulating 10 million draws from the fitted model supplies the tail probabilities on which the threshold argument rests.
What would settle it
Hold out a recent championship, refit the model on the remaining years, and compare the predicted number of sub-0.1-second starts in the held-out meet with the actual count; a large mismatch would overturn the threshold recommendation, as would a year-by-year comparison showing that the left-tail counts do not track the model's predictions.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is that the 2022 timing data and the 25-year historical record together undermine both an implicit assumption about the 2022 meet and the justification for the 0.1-second barrier. Within-athlete matched comparisons show faster 2022 reaction times across all three control groups, pointing to a systematic venue- or equipment-level difference rather than individual improvement. A generalized Gamma model with random effects for venue and heat estimates the probability of a sub-0.1-second reaction at $2.76\times 10^{-3}$ (about one in 362 starts) when 2022 is included, and at $1.94\times 10^{-3}$ (about one in 515) when it is excluded. The same model yields a men's barrier of 0.094 seconds at a one-in-a-thousand false-start rate, and the paper concludes that sub-0.1-second reactions are physiologically plausible, that the 2022 meet stands out statistically, and that the uniform 0.1-second threshold is not well grounded for men.
Load-bearing premise
The whole threshold recommendation depends on trusting the fitted generalized Gamma model's extrapolation into the extreme left tail, even though the model is only checked with an overall Q-Q plot and density overlay, with no direct validation of the tail and no uncertainty interval on the one-in-362 estimate.
Editorial extensions
If this is right
- Under the fitted model, the current rule will keep producing sub-0.1-second reactions at a rate of about one in 362 men's starts, so the next Devon Allen case is a matter of when, not whether.
- A men's threshold near 0.094 seconds would place the false-start probability near one in 1,000 starts, and a threshold near 0.082 seconds would place it near one in 10,000.
- The matched comparisons single out the 2022 World Championships as statistically anomalous relative to national meets, 2019, and 2023, which points to a systematic timing-system difference rather than athlete improvement.
- Women's fitted model gives a lower probability of sub-0.1-second reactions at every threshold, so a uniform barrier does not place the same disqualification burden on men and women.
- Excluding 2022 lowers the one-in-362 estimate to one in 515, so the conclusion that 0.1 seconds is not a statistically grounded barrier is not an artifact of the single controversial year.
Reading between the lines
- The paper does not claim the 2022 timing equipment malfunctioned; an extension it leaves implicit is that comparing raw starting-block sensor traces from the same athletes at 2022 and other meets could directly test that explanation.
- The one-in-362 figure is a population-average risk across starts, not Devon Allen's personal probability; his own reaction-time distribution is likely shifted faster than the average.
- The analysis prices the cost of wrongly disqualifying a genuine fast reaction but not the cost of more false starts and restarts, so the socially optimal threshold could differ from the statistically symmetric one.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper analyses reaction times (RTs) from elite sprint events to address two questions raised by Devon Allen's disqualification at the 2022 World Championships: whether RTs at that meet were systematically faster than at comparable competitions, and whether the 0.1-second disqualification threshold is statistically justified. For the first question, the authors use rank-sum tests for clustered data (Datta–Satten and permutations) to compare the same athletes' RTs across the 2022 World Championships versus 2022 national championships, 2019 World Championships, and 2023 World Championships. For the second, they fit a generalized Gamma (GG) GAMLSS model with venue- and heat-level random effects to men's 110m hurdles and 100m dash RTs from 1999–2023 (semifinals and finals only), estimate tail probabilities below 0.08, 0.09, and 0.10 seconds, and propose alternative thresholds based on nominal tail probabilities. The results show significantly faster RTs in 2022 in all six comparisons, and model-based estimates of P(RT<0.10) of about 1/362 including 2022 data, leading the authors to recommend standardized timing protocols and reconsideration of the threshold.
Significance. If the findings are robust, the paper would make a useful contribution to an ongoing debate in athletics governance, providing statistical evidence on timing-system variability and a quantitative framework for setting false-start thresholds. The paper's strengths include the use of a clustered-data rank-sum test with exact permutation inference, a flexible parametric model for the full RT distribution, inclusion of sensitivity analyses in the supplement, and public availability of data and code. However, the two main conclusions rest on analyses that currently have important gaps: the rank-sum comparisons do not condition on race round, and the tail probabilities are extrapolations from a single parametric model without uncertainty intervals or tail-specific validation. These gaps should be addressed before the claims can be regarded as established.
major comments (3)
- [Section 2.2, Table 1] The rank-sum comparisons pool RTs across heats, semifinals, and finals without conditioning on race round. If the distribution of rounds differs between the 2022 World Championships and the comparison competitions within the same athletes, the significant p-values may reflect round effects (e.g., more final-round observations in 2022) rather than a systematic timing-system difference. The paper should either restrict to a single round (e.g., semifinals only) or include round as a stratum in the analysis to support the claim that the 2022 timing system produced faster RTs.
- [Section 3.3, Table 3] The headline probability P(RT < 0.10) = 2.76e-3 is a left-tail extrapolation from a single fitted generalized Gamma model, with no standard error or confidence interval reported. Figure 3 validates the overall fit via a density overlay and Q-Q plot but does not assess the y < 0.1 region where data are sparse. Furthermore, the sensitivity analysis in Supplement Section 3 shows the estimate drops from 2.76e-3 to 1.97e-3 when 17 disqualified RTs are excluded, indicating that the estimate is not robust to plausible model/data choices. The paper should present tail-focused diagnostics (e.g., empirical exceedance rates in subsets, alternative distributions, bootstrap intervals) before using this probability to support threshold recommendations.
- [Supplement Section 2, Tables 2–3] The comparison of men's and women's tail probabilities and the resulting claim that the uniform 0.1-second threshold unfairly penalizes men rely on model-based extrapolations without uncertainty quantification. The women's fitted model has a much larger shape parameter ν and the probability at threshold 0.08 is reported as 1e-7, an extremely small value; no tail diagnostics are provided for the women's model. The paper should report confidence intervals or at least a sensitivity analysis for the gender comparison before drawing conclusions about differential fairness.
minor comments (5)
- [Section 3.2] The manuscript contains an embedded editing note: "EDS: (Is Fig 3 for the analysis with or without 2022?) OF: With 2022. I added that to the caption of the figure". This note should be removed before publication.
- [Section 2.1.1 vs Section 3.1] The data inclusion criteria differ between the rank-sum analysis (which includes heats, semifinals, and finals, as stated in Section 2.1.1) and the GAMLSS analysis (which uses only semifinals and finals, as stated in Section 3.1). These differing choices should be explicitly justified in both places.
- [Table 2] Table 2 reports identical values for β0 and γ0 for the excluding-2022 and including-2022 fits (both −1.910 and −2.200, respectively), which suggests that more decimal places are needed to see the actual differences; consider reporting additional digits or the estimated differences.
- [Abstract] The abstract contains a typo: "RTs be low 0.1 seconds" should read "RTs below 0.1 seconds".
- [Throughout] The paper alternates between "IAAF" and "World Athletics" when referring to the governing body; since the organization changed its name, the authors should use "World Athletics" consistently after the first mention.
Circularity Check
No circular derivation chain: the 1-in-362 tail probability is a standard transform of an explicitly fitted model, and the sole self-citation (the clusrank package, co-created by an author) is non-load-bearing implementation support.
full rationale
The central threshold analysis (Section 3) fits a generalized Gamma GAMLSS model with venue- and heat-level random effects to World Championship RT data, then reports P(RT < 0.10) = 2.76e-3 by simulating 10 million draws from the fitted mixture (Table 3 and Section 3.3). This is ordinary fitted-distribution inference: the tail probability is a functional of the fitted parameters, not an identity with any input value. The estimate is sensitive to the inclusion of the 17 disqualified RTs (2.76e-3 with DQs vs. 1.97e-3 without, Supplement Table 4), but sensitivity to model/data changes is not circularity; the reported probability is not the empirical frequency of sub-0.1s RTs in the sample and is not forced to equal any fitted parameter. The choice of the GG family is stated in-text ('Based on an exploratory analysis') and defended by AIC comparison against competing models in Section 3.2, not imported through an authorship citation, so no ansatz-smuggling or imported-uniqueness pattern applies. The rank-sum analysis (Section 2) uses the clusWilcox.test() function from the clusrank package (Jiang et al., 2020, co-authored by Jun Yan), but the underlying method is the external Datta and Satten (2005) clustered rank-sum test, and the authors corroborate every p-value with a 1-million-permutation exact test, so the self-citation is purely instrumental and not load-bearing. No self-definitional step, no fitted-input-called-prediction reduction, and no renaming of a known result were found. The skeptic's concern that the left tail is an unvalidated parametric extrapolation without uncertainty intervals is a legitimate statistical robustness critique, but the instructions explicitly direct such model-dependence concerns to correctness risk rather than circularity. Accordingly, the derivation chain is self-contained against its stated assumptions, and the score reflects only the single minor, non-load-bearing self-citation.
Assumptions & free parameters
free parameters (5)
- beta0 (log-location intercept) =
-1.910
- gamma0 (log-scale intercept) =
-2.200
- nu (shape parameter) =
-1.178
- tau_v (venue random effect SD) =
0.058
- tau_h (heat random effect SD) =
0.320
assumptions (6)
- domain assumption Reaction times are positive and can be modeled by the generalized Gamma distribution
- domain assumption Reaction times from 100m dash and 110m hurdles have the same distribution and can be pooled
- domain assumption Negative reaction times (starting before the gun) are excluded as errors
- domain assumption Athletes who competed in both 2022 WC and another competition provide valid matched clusters for comparison
- standard math The Datta-Satten clustered rank-sum test correctly accounts for within-athlete dependence in this setting
- domain assumption Random effects for venue (year) and heat adequately capture the correlation structure
Cite this review
Pith. "Pith review of On Devon Allen's Disqualification at the 2022 World Track and Field Championships." pith.science (2026). https://pith.science/paper/SXQE42RZ
@misc{pith2026250611460,
author = {Pith},
title = {Pith review of: On Devon Allen's Disqualification at the 2022 World Track and Field Championships},
year = {2026},
howpublished = {\url{https://pith.science/paper/SXQE42RZ}},
note = {Machine review of arXiv:2506.11460}
}
read the original abstract
Devon Allen's disqualification at the men's 110-meter hurdle final at the 2022 World Track and Field Championships, due to a reaction time (RT) of 0.099 seconds-just 0.001 seconds below the allowable threshold-sparked widespread debate over the fairness and validity of RT rules. This study investigates two key issues: variations in timing systems and the justification for the 0.1-second disqualification threshold. We pooled RT data from men's 110-meter hurdles and 100-meter dash, as well as women's 100-meter hurdles and 100-meter dash, spanning national and international competitions. Using a rank-sum test for clustered data, we compared RTs across multiple competitions, while a generalized Gamma model with random effects for venue and heat was applied to evaluate the threshold. Our analyses reveal significant differences in RTs between the 2022 World Championships and other competitions, pointing to systematic variations in timing systems. Additionally, the model shows that RTs be low 0.1 seconds, though rare, are physiologically plausible. These findings highlight the need for standardized timing protocols and a re-evaluation of the 0.1-second disqualification threshold to promote fairness in elite competition.
Figures
Reference graph
Works this paper leans on
-
[1]
Almeida, A., Loy, A., and Hofmann, H. (2018). ggplot2 compatible quantile-quantile plots in R . The R Journal , 10(2):248--261
work page 2018
-
[2]
Babi c , V. and Delalija, A. (2009). Reaction time trends in the sprint and hurdle events at the 2004 O lympic G ames: Differences between male and female athletes. New Studies in Athletics , 24(1):59--68
work page 2009
-
[3]
C., Hayes, K., and Harrison, A
Brosnan, K. C., Hayes, K., and Harrison, A. J. (2017). Effects of false-start disqualification rules on response-times of elite-standard sprinters. Journal of Sports Sciences , 35(10):929--935
work page 2017
-
[4]
Collet, C. (1999). Strategic aspects of reaction time in world-class sprinters. Perceptual and Motor Skills , 88(1):65--75
work page 1999
-
[5]
Datta, S. and Satten, G. A. (2005). Rank-sum tests for clustered data. Journal of the American Statistical Association , 100(471):908--915
work page 2005
-
[6]
Dunn, P. K. and Smyth, G. K. (1996). Randomized quantile residuals. Journal of Computational and Graphical Statistics , 5(3):236--244
work page 1996
-
[7]
A., Shalfawi, S., and T nnessen, E
Haugen, T. A., Shalfawi, S., and T nnessen, E. (2013). The effect of different starting procedures on sprinters’ reaction time. Journal of Sports Sciences , 31(7):699--705
work page 2013
-
[8]
IAAF (2009). Comparison of false starts. https://worldathletics.org/download/download?filename=58540761-210b-4685-8b38-21fd68f70430.pdf&urlSlug=comparison-of-false-starts
work page 2009
Show all 26 references
-
[9]
Ishikawa, M., Komi, P., and Salmi, J. (2009). IAAF sprint start research project: Is the 100 ms limit still valid? IAAF New Studies in Athletics , 24:37--47
2009
-
[10]
T., Rosner, B., and Yan, J
Jiang, Y., He, X., Lee, M.-L. T., Rosner, B., and Yan, J. (2020). Wilcoxon rank-based tests for clustered data with R package clusrank . Journal of Statistical Software , 96(6):1--26
2020
-
[11]
Johnson, R. (2022a). The data keeps pouring in and it continues to look bad for World Athletics and great for D evon A llen. https://www.letsrun.com/news/2022/07/the-data-keeps-pouring-in-and-it-continues-to-look-bad-for-world-athletics-and-great-for-devon-allen/
2022
-
[12]
Johnson, R. (2022b). Was D evon A llen screwed? T here's at least a 99.9\ he was. https://www.letsrun.com/news/2022/07/was-devon-allen-screwed-theres-at-least-a-99-9-chance-that-he-was/
2022
-
[13]
B., Galecki, A
Lipps, D. B., Galecki, A. T., and Ashton-Miller, J. A. (2011). On the implications of a sex difference in the reaction times of sprinters at the beijing olympics. PLOS One , 6(10):e26141
2011
-
[14]
and Komi, P
Mero, A. and Komi, P. V. (1990). Reaction time and electromyographic activity during a sprint start. European Journal of Applied Physiology and Occupational Physiology , 61(1-2):73--80
1990
-
[15]
Milloz, M., Hayes, K., and Harrison, A. J. (2021). Sprint start regulation in athletics: A critical review. Sports Medicine , 51:21--31
2021
-
[16]
Pain, M. T. and Hibbs, A. (2007). Sprint starts and the minimum auditory reaction time. Journal of Sports Sciences , 25(1):79--86
2007
-
[17]
S., Kotzamanidou, M
Panoutsakopoulos, V., Theodorou, A. S., Kotzamanidou, M. C., Fragkoulis, E., Smirniotou, A., and Kollias, I. A. (2020). Gender and event specificity differences in kinematical parameters of a 60 m hurdles race. International Journal of Performance Analysis in Sport , 20(4):668--682
2020
-
[18]
Rigby, R. A. and Stasinopoulos, D. M. (2005). Generalized additive models for location, scale and shape (with discussion). Journal of the Royal Statistical Society Series C: Applied Statistics , 54(3):507--554
2005
-
[19]
A., Stasinopoulos, M
Rigby, R. A., Stasinopoulos, M. D., Heller, G. Z., and De Bastiani, F. (2019). Distributions for Modeling Location, Scale, and Shape: U sing GAMLSS in R . Chapman and Hall/CRC
2019
-
[20]
Stasinopoulos, D. M. and Rigby, R. A. (2008). Generalized additive models for location scale and shape (GAMLSS) in R . Journal of Statistical Software , 23(7):1--46
2008
-
[21]
D., Kneib, T., Klein, N., Mayr, A., and Heller, G
Stasinopoulos, M. D., Kneib, T., Klein, N., Mayr, A., and Heller, G. Z. (2024). Generalized Additive Models for Location, Scale and Shape: A Distributional Regression Approach, with Applications , volume 56. Cambridge University Press
2024
-
[22]
T nnessen, E., Haugen, T., and Shalfawi, S. A. (2013). Reaction time aspects of elite sprinters in athletic world championships. The Journal of Strength & Conditioning Research , 27(4):885--892
2013
-
[23]
Willwacher, S., Feldker, M.-K., Zohren, S., Herrmann, V., and Br \"u ggemann, G.-P. (2013). A novel method for the evaluation and certification of false start apparatus in sprint running. Procedia Engineering , 60:124--129
2013
-
[24]
World athletics: World athletics championships: Results
World Athletics (2023). World athletics: World athletics championships: Results. https://www.worldathletics.org/results/world-athletics-championships
2023
-
[25]
Zhang, J., Lin, X.-Y., and Zhang, S. (2021). Correlation analysis of sprint performance and reaction time based on double logarithm model. Complexity , 2021(1):6633326
2021
-
[26]
and Delalija, A
Babi c , V. and Delalija, A. (2009). Reaction time trends in the sprint and hurdle events at the 2004 olympic games: Differences between male and female athletes. New Studies in Athletics , 24(1):59--68
2009
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.