{"id":"88b8eb08-0d3b-4bc6-9634-5656f1a174cc","arxiv_id":"2506.11460","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"The paper shows 2022 World Championship reaction times were statistically faster than other major meets and estimates a 1-in-362 chance of a sub-0.1s reaction, prompting a call to revisit the 0.1s disqualification rule.","lead":"A statistical analysis of sprint reaction times finds that the 2022 World Championships produced unusually fast starts, and that sub-0.1 second reactions, though rare, occur often enough to question the current disqualification threshold. The paper argues for lower reaction time barriers and standardized timing protocols in elite sprinting.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 1-in-362 sub-0.1s probability is a left-tail extrapolation from a single generalized Gamma fit, with no tail validation or uncertainty interval; a parametric bootstrap or alternative distribution could shift it substantially.","rationale":"The paper's most consequential claim is not the rank-sum comparison (which is well supported by consistently tiny p-values and a matched design) but the assertion that a sub-0.1s RT is 'approximately one in 362 starts' and that the 0.1s barrier may be too strict. That number is a tail functional of a parametric generalized Gamma model fit to 776 men's RTs. The left tail below 0.1s is extremely sparse in the observed data; the only real empirical mass comes from a few 2022 observations, including the highly contested 0.099s RT. Thus the estimate is dominated by the functional form of the GG distribution and by inclusion/exclusion decisions. The paper provides an overall Q-Q plot and density overlay but no tail calibration, no comparison with other plausible left-tail models, and no confidence interval for the tail probability. The supplement's DQ-exclusion sensitivity changes the estimate by about 30%, which suggests the uncertainty is non-negligible. A parametric bootstrap and an alternative-distribution refit would directly test whether the 1-in-362 figure is an artifact of model choice. Because the reader already assigned CONDITIONAL based on the same load-bearing assumption, my read does not change the verdict, but it sharpens the required condition: the authors should supply tail-focused uncertainty quantification before the threshold recommendation is accepted.","tokens_in":13742,"tokens_out":5239,"duration_ms":53013,"concrete_test":"Run B=500 parametric bootstrap fits of the Section 3.2 model: resample venue and heat random effects, refit, and for each refit simulate 10 million RTs to approximate P(RT<0.10); report the 2.5-97.5 percentile interval. Independently, refit the same data with a Box-Cox t or shifted log-normal distribution (same random-effects structure) and recompute P(RT<0.10). If the bootstrap interval spans more than a factor of ~3, or the alternative distribution changes the estimate by more than 50%, the 1-in-362 claim should be downgraded to model-dependent.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The threshold recommendation in Section 3.3 rests on P(RT < 0.10) = 2.76e-3 computed by simulating from one fitted generalized Gamma GAMLSS model. The diagnostic support (Figure 3) is an overall density overlay and Q-Q plot; neither validates the y < 0.1 region where observations are sparse (e.g., Allen's 0.099 and a handful of 2022 DQ times). Because the left tail is determined by the parametric GG family, the choice of distribution is load-bearing, and the reported probability carries no standard error or confidence interval despite available parameter standard errors. The paper's own sensitivity analysis (Supplement Section 3) moves the estimate from 2.76e-3 to 1.97e-3 when 17 DQ observations are excluded, illustrating that the headline number is not robust to modest model changes. Without a tail-focused check or uncertainty quantification, the 'approximately one in 362 starts' claim is not an established empirical rate.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper analyses reaction times (RTs) from elite sprint events to address two questions raised by Devon Allen's disqualification at the 2022 World Championships: whether RTs at that meet were systematically faster than at comparable competitions, and whether the 0.1-second disqualification threshold is statistically justified. For the first question, the authors use rank-sum tests for clustered data (Datta–Satten and permutations) to compare the same athletes' RTs across the 2022 World Championships versus 2022 national championships, 2019 World Championships, and 2023 World Championships. For the second, they fit a generalized Gamma (GG) GAMLSS model with venue- and heat-level random effects to men's 110m hurdles and 100m dash RTs from 1999–2023 (semifinals and finals only), estimate tail probabilities below 0.08, 0.09, and 0.10 seconds, and propose alternative thresholds based on nominal tail probabilities. The results show significantly faster RTs in 2022 in all six comparisons, and model-based estimates of P(RT<0.10) of about 1/362 including 2022 data, leading the authors to recommend standardized timing protocols and reconsideration of the threshold.","tokens_in":13957,"tokens_out":4604,"duration_ms":41733,"significance":"If the findings are robust, the paper would make a useful contribution to an ongoing debate in athletics governance, providing statistical evidence on timing-system variability and a quantitative framework for setting false-start thresholds. The paper's strengths include the use of a clustered-data rank-sum test with exact permutation inference, a flexible parametric model for the full RT distribution, inclusion of sensitivity analyses in the supplement, and public availability of data and code. However, the two main conclusions rest on analyses that currently have important gaps: the rank-sum comparisons do not condition on race round, and the tail probabilities are extrapolations from a single parametric model without uncertainty intervals or tail-specific validation. These gaps should be addressed before the claims can be regarded as established.","major_comments":[{"comment":"The rank-sum comparisons pool RTs across heats, semifinals, and finals without conditioning on race round. If the distribution of rounds differs between the 2022 World Championships and the comparison competitions within the same athletes, the significant p-values may reflect round effects (e.g., more final-round observations in 2022) rather than a systematic timing-system difference. The paper should either restrict to a single round (e.g., semifinals only) or include round as a stratum in the analysis to support the claim that the 2022 timing system produced faster RTs.","section":"Section 2.2, Table 1"},{"comment":"The headline probability P(RT < 0.10) = 2.76e-3 is a left-tail extrapolation from a single fitted generalized Gamma model, with no standard error or confidence interval reported. Figure 3 validates the overall fit via a density overlay and Q-Q plot but does not assess the y < 0.1 region where data are sparse. Furthermore, the sensitivity analysis in Supplement Section 3 shows the estimate drops from 2.76e-3 to 1.97e-3 when 17 disqualified RTs are excluded, indicating that the estimate is not robust to plausible model/data choices. The paper should present tail-focused diagnostics (e.g., empirical exceedance rates in subsets, alternative distributions, bootstrap intervals) before using this probability to support threshold recommendations.","section":"Section 3.3, Table 3"},{"comment":"The comparison of men's and women's tail probabilities and the resulting claim that the uniform 0.1-second threshold unfairly penalizes men rely on model-based extrapolations without uncertainty quantification. The women's fitted model has a much larger shape parameter ν and the probability at threshold 0.08 is reported as 1e-7, an extremely small value; no tail diagnostics are provided for the women's model. The paper should report confidence intervals or at least a sensitivity analysis for the gender comparison before drawing conclusions about differential fairness.","section":"Supplement Section 2, Tables 2–3"}],"minor_comments":[{"comment":"The manuscript contains an embedded editing note: \"EDS: (Is Fig 3 for the analysis with or without 2022?) OF: With 2022. I added that to the caption of the figure\". This note should be removed before publication.","section":"Section 3.2"},{"comment":"The data inclusion criteria differ between the rank-sum analysis (which includes heats, semifinals, and finals, as stated in Section 2.1.1) and the GAMLSS analysis (which uses only semifinals and finals, as stated in Section 3.1). These differing choices should be explicitly justified in both places.","section":"Section 2.1.1 vs Section 3.1"},{"comment":"Table 2 reports identical values for β0 and γ0 for the excluding-2022 and including-2022 fits (both −1.910 and −2.200, respectively), which suggests that more decimal places are needed to see the actual differences; consider reporting additional digits or the estimated differences.","section":"Table 2"},{"comment":"The abstract contains a typo: \"RTs be low 0.1 seconds\" should read \"RTs below 0.1 seconds\".","section":"Abstract"},{"comment":"The paper alternates between \"IAAF\" and \"World Athletics\" when referring to the governing body; since the organization changed its name, the authors should use \"World Athletics\" consistently after the first mention.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The embedded editing note in Section 3.2 suggests the manuscript is not fully cleaned, but that is a minor fix. The substantive concerns are the lack of round-stratified analysis for the 2022 comparison and the lack of uncertainty quantification and tail validation for the GAMLSS extrapolations. The authors should be encouraged to provide these, as the paper's conclusions depend on them. The paper fits the scope of a statistical applications journal, but its impact would be strengthened by addressing these points."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read the paper. The main result — that reaction times at the 2022 World Championships were anomalously fast relative to 2019, 2023, and national meets for the same athletes — mostly holds up. The clustering-aware rank tests are appropriate, and the data collection effort is genuinely impressive. The threshold analysis is more fragile.\n\nWhat's new: this is the first pooled statistical treatment of the Devon Allen disqualification, and the matched-athlete comparisons quantify the 2022 anomaly cleanly. The GAMLSS model with venue and heat random effects is reasonable, and the supplement does real sensitivity work (excluding DQs, excluding 2022, women's data). Code and data are provided, which is more than most papers in this area offer.\n\nSoft spots, in order of importance. First, the abstract says the results point to \"systematic variations in timing systems,\" but the discussion is appropriately cautious. That's an overstatement — the analysis shows faster RTs in 2022, not directly that the timing hardware was at fault. Second, the 1-in-362 sub-0.1s probability is a left-tail extrapolation from a single fitted generalized Gamma model. The diagnostics in Figure 3 are overall fits; they don't validate the y<0.1 region, and there's no uncertainty interval on the tail probability. The sensitivity analysis shows the estimate moves from 2.76e-3 to 1.94e-3 (excluding 2022) or 1.97e-3 (excluding DQs) — a roughly 30% swing. So treat the precise probability as illustrative, not established. Third, the Section 2 rank-sum comparisons don't condition on race round (heats vs. semifinals vs. finals). RTs are typically slower in earlier rounds, and a mismatch in round mix between competitions for the same athletes could bias results. The p-values are so small that I doubt this flips the finding, but it should be addressed. Finally, there's a leftover editorial note in the Figure 3 caption (\"EDS: (Is Fig 3...?) OF: With 2022...\"). That needs to be cleaned up before any submission.\n\nWho this is for: people working on sports statistics, false-start rules, or GAMLSS applications. It's also a good case study for an applied statistics class.\n\nRecommendation: send it to peer review. The central anomaly claim is supported; the threshold recommendation needs stronger tail diagnostics and clearer uncertainty, and the abstract should be toned down. With those revisions it would be a solid applied paper.","headline":"A solid applied analysis of the 2022 reaction-time anomaly; the timing-system claim is too strong, and the 1-in-362 tail probability needs uncertainty bounds.","tokens_in":14476,"tokens_out":2631,"would_cite":true,"duration_ms":26347,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that Devon Allen's 0.099-second reaction at the 2022 World Championships was not statistically impossible and that the 0.1-second disqualification threshold is too strict for men.","keywords":["reaction time","false start","disqualification threshold","generalized Gamma distribution","GAMLSS","clustered rank-sum test","sprint starts","sports statistics"],"falsifier":"Hold out a recent championship, refit the model on the remaining years, and compare the predicted number of sub-0.1-second starts in the held-out meet with the actual count; a large mismatch would overturn the threshold recommendation, as would a year-by-year comparison showing that the left-tail counts do not track the model's predictions.","tokens_in":13557,"feed_emoji":"⏱️","tokens_out":8488,"duration_ms":77049,"temperature":0.7,"pith_summary":"This paper tries to settle two questions raised by Devon Allen's disqualification: whether reaction times measured at the 2022 World Championships were systematically faster than elsewhere, and whether the 0.1-second false-start limit is statistically defensible. For the first question, it compares reaction times of the same athletes across competitions and finds consistent, significant evidence that 2022 times were faster than at national meets, at the 2019 World Championships, and at the 2023 World Championships. For the second, it fits a generalized Gamma model to 25 years of men's World Championship reaction times and estimates that a reaction below 0.1 seconds occurs roughly once in every 362 starts, a rare event but not an impossible one. If the results are right, the current uniform threshold penalizes men more than women and should be reconsidered, and timing certification should be standardized.","feed_headline":"Model says sub-0.1-second reaction times are a 1-in-362 event","feed_subtitle":"A 25-year model of men's sprint starts says Devon Allen's 0.099 reaction was rare but not impossible.","key_machinery":"The argument is carried by two statistical tools. The first is a rank-sum test for clustered data, in which each athlete is a cluster and each observed reaction time is a subunit; this test decides whether the distribution of reaction times shifts between competitions while respecting within-athlete dependence. The second is a generalized Gamma distribution $GG(\\mu,\\sigma,\\nu)$ with random effects for venue and for heat, fitted inside the GAMLSS framework (generalized additive models for location, scale, and shape). The generalized Gamma shape parameter lets the model capture asymmetry in reaction times, the venue random effect absorbs year-to-year differences such as 2022, and the heat random effect absorbs race-to-race variability; simulating 10 million draws from the fitted model supplies the tail probabilities on which the threshold argument rests.","core_discovery":"On the paper's own terms, the central discovery is that the 2022 timing data and the 25-year historical record together undermine both an implicit assumption about the 2022 meet and the justification for the 0.1-second barrier. Within-athlete matched comparisons show faster 2022 reaction times across all three control groups, pointing to a systematic venue- or equipment-level difference rather than individual improvement. A generalized Gamma model with random effects for venue and heat estimates the probability of a sub-0.1-second reaction at $2.76\\times 10^{-3}$ (about one in 362 starts) when 2022 is included, and at $1.94\\times 10^{-3}$ (about one in 515) when it is excluded. The same model yields a men's barrier of 0.094 seconds at a one-in-a-thousand false-start rate, and the paper concludes that sub-0.1-second reactions are physiologically plausible, that the 2022 meet stands out statistically, and that the uniform 0.1-second threshold is not well grounded for men.","pith_inferences":["The paper does not claim the 2022 timing equipment malfunctioned; an extension it leaves implicit is that comparing raw starting-block sensor traces from the same athletes at 2022 and other meets could directly test that explanation.","The one-in-362 figure is a population-average risk across starts, not Devon Allen's personal probability; his own reaction-time distribution is likely shifted faster than the average.","The analysis prices the cost of wrongly disqualifying a genuine fast reaction but not the cost of more false starts and restarts, so the socially optimal threshold could differ from the statistically symmetric one."],"forward_implications":["Under the fitted model, the current rule will keep producing sub-0.1-second reactions at a rate of about one in 362 men's starts, so the next Devon Allen case is a matter of when, not whether.","A men's threshold near 0.094 seconds would place the false-start probability near one in 1,000 starts, and a threshold near 0.082 seconds would place it near one in 10,000.","The matched comparisons single out the 2022 World Championships as statistically anomalous relative to national meets, 2019, and 2023, which points to a systematic timing-system difference rather than athlete improvement.","Women's fitted model gives a lower probability of sub-0.1-second reactions at every threshold, so a uniform barrier does not place the same disqualification burden on men and women.","Excluding 2022 lowers the one-in-362 estimate to one in 515, so the conclusion that 0.1 seconds is not a statistically grounded barrier is not an artifact of the single controversial year."],"supporting_citations":[{"why":"Provides the limited Finnish national-level data originally used to justify the 0.1-second threshold.","marker":"Mero and Komi (1990)"},{"why":"Controlled experiments showing sub-0.1-second auditory reaction times are physiologically possible and that simple force-threshold sensors can delay detection by up to 26 ms.","marker":"Pain and Hibbs (2007)"},{"why":"IAAF sprint start research questioning whether the 100 ms limit is still valid; the threshold evaluation is compared against its recommendations.","marker":"Ishikawa et al. (2009)"},{"why":"Documents effects of false-start rules on elite sprinters' response times and advocates gender-specific barriers, which the paper's threshold analysis aligns with.","marker":"Brosnan et al. (2017)"},{"why":"Supplies the rank-sum test for clustered data used in the within-athlete comparisons of 2022 versus other competitions.","marker":"Datta and Satten (2005)"},{"why":"Introduces the GAMLSS framework in which the generalized Gamma model with random effects is fitted.","marker":"Rigby and Stasinopoulos (2005)"},{"why":"Shows how starting procedures and rule changes affect reaction times, used to interpret year-to-year variation in the historical data.","marker":"Haugen et al. (2013)"},{"why":"Public data analysis documenting that the 2022 World Championships had 25 reaction times under 0.115 seconds versus 3 in 2019, motivating the systematic-timing investigation.","marker":"Johnson (2022b)"}],"fun_headline_variants":["Sub-0.1-second reactions are plausible, model finds","2022 track meet timing differs from 25-year historical record","Study questions 0.1-second reaction time rule in sprints","Devon Allen's 0.099 reaction: rare but physiologically possible"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole threshold recommendation depends on trusting the fitted generalized Gamma model's extrapolation into the extreme left tail, even though the model is only checked with an overall Q-Q plot and density overlay, with no direct validation of the tail and no uncertainty interval on the one-in-362 estimate.","fun_headline_variants_meta":{"raw":{"variants":["Sub-0.1-second reactions are plausible, model finds","2022 track meet timing differs from 25-year historical record","Study questions 0.1-second reaction time rule in sprints","Devon Allen's 0.099 reaction: rare but physiologically possible"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000196,"raw_usage":{"total_tokens":1386,"prompt_tokens":993,"completion_tokens":393,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":609,"completion_tokens_details":{"reasoning_tokens":318}},"tokens_in":609,"tokens_out":393,"duration_ms":4018,"temperature":1.0,"reasoning_tokens":318,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T04:05:05.649125+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Hold out a recent championship, refit the model on the remaining years, and compare the predicted number of sub-0.1-second starts in the held-out meet with the actual count; a large mismatch would overturn the threshold recommendation, as would a year-by-year comparison showing that the left-tail counts do not track the model's predictions.","supporting_citations":[{"cited_title":"and Komi, P","cited_arxiv_id":null,"evidence_quote":"Provides the limited Finnish national-level data originally used to justify the 0.1-second threshold."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Controlled experiments showing sub-0.1-second auditory reaction times are physiologically possible and that simple force-threshold sensors can delay detection by up to 26 ms."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"IAAF sprint start research questioning whether the 100 ms limit is still valid; the threshold evaluation is compared against its recommendations."},{"cited_title":"C., Hayes, K., and Harrison, A","cited_arxiv_id":null,"evidence_quote":"Documents effects of false-start rules on elite sprinters' response times and advocates gender-specific barriers, which the paper's threshold analysis aligns with."},{"cited_title":"and Satten, G","cited_arxiv_id":null,"evidence_quote":"Supplies the rank-sum test for clustered data used in the within-athlete comparisons of 2022 versus other competitions."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the GAMLSS framework in which the generalized Gamma model with random effects is fitted."},{"cited_title":"A., Shalfawi, S., and T nnessen, E","cited_arxiv_id":null,"evidence_quote":"Shows how starting procedures and rule changes affect reaction times, used to interpret year-to-year variation in the historical data."}],"review_version":1}