{"id":"76c037d8-aa11-4c9e-92ac-46bd3f15f76a","arxiv_id":"2608.01882","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":9,"one_line_summary":"Inter-record gaps in Test, ODI, and T20 cricket follow a truncated power law with exponent about 0.8, while shuffled careers give about 0.95, implying temporal career structure drives record statistics.","lead":"A study of cricket batting records finds that the time between successive personal best scores follows a broad truncated power law with exponent around 0.8, fatter than the classic 1/g prediction for random performances. Randomly shuffling each player's innings pushes the exponent close to 1, suggesting career ordering, not scoring ability, shapes when records fall.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No significance test supports the claimed difference between empirical and shuffled exponents; a permutation test is needed.","rationale":"The reader's weakest_assumption points to several factors that might bias the shuffled comparison (finite career lengths, discrete scores, T20 selection, within-player dependence). I have narrowed the single most load-bearing concern to the absence of a proper null distribution for the shuffled exponents. The entire inference that temporal ordering matters hinges on the difference between real and shuffled α. The paper provides only point estimates and profile-likelihood errors that assume independence, and it does not report the variance of α across shuffles. This is a specific, testable gap in the statistical argument. I agree with the reader's overall conditional verdict; the paper presents an interesting observation but needs a permutation test or block-bootstrap confidence interval before the central claim can be accepted. My concern is a subset of the reader's broader concern, hence 'partial' agreement. I do not recommend a more severe verdict than the reader's, because the observed difference is large and the shuffle design is sound in principle; it simply needs a proper significance test.","tokens_in":8899,"tokens_out":8245,"duration_ms":103211,"concrete_test":"For each format (Test, ODI, T20), generate B=1000 independent random shuffles of each player's innings, recompute the full gap sequence, and fit the truncated power law by MLE to obtain a distribution of shuffled α. Compute a one-sided p-value as the fraction of shuffled α less than or equal to the empirical α. Additionally, perform a player-level bootstrap (resample players with replacement) to obtain a confidence interval for the difference Δ = α_empirical − α_shuffled that properly accounts for within-player dependence. Report both the p-value and the CI; if the p-value exceeds 0.05 or the CI includes zero, the central claim loses its empirical support.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that 'temporal ordering of innings' drives the deviation from α≈1 rests entirely on the contrast between empirical exponents (0.799–0.843) and bootstrap-shuffled exponents (0.939–0.979). The paper reports only a single shuffled fit per format, with profile-likelihood errors that treat each gap as an independent observation. This is problematic for two reasons. First, gaps are clustered within players; profile-likelihood error bars ignore this, so the uncertainty on the shuffled exponents is likely underestimated. Second, a single shuffle is a single Monte Carlo realization; the variance across shuffles is not reported. If the distribution of shuffled α is broad (e.g., ±0.1 or more), then the observed gap of ~0.15 may not be statistically significant. Furthermore, the shuffled exponents are themselves below the classical α=1, which may reflect the discrete, finite-support nature of cricket scores; the paper does not establish what the correct null is for this discrete process. Without a permutation test that generates the null distribution of α under shuffle, the claim that temporal ordering 'drives the deviation' is not supported by the reported statistics.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies inter-record gaps (innings between successive personal best scores) in the careers of leading Test, ODI, and T20 cricketers, using data scraped from ESPN Cricinfo. For each format, the authors aggregate gaps across players and fit a truncated power law P(g) ∝ g^{-α} e^{-λg} by maximum likelihood, reporting α ≈ 0.799–0.843, well below the classical i.i.d. prediction α = 1. They then shuffle each player's innings to preserve the marginal score distribution and career length while destroying temporal ordering; the shuffled exponents increase to 0.939–0.979. Reversing careers gives intermediate exponents. The authors conclude that temporal ordering within careers, not heterogeneity in ability or career length, is the primary driver of the deviation from classical record statistics, and they argue that simple synthetic models do not reproduce the empirical behavior.","tokens_in":9220,"tokens_out":7267,"duration_ms":94657,"significance":"If established, the main claim is of genuine interest to record statistics and sports analytics: it would show that the temporal organization of performance, rather than the marginal distribution of scores, controls the tail of inter-record waiting times. The design has clear strengths: the shuffle is a direct within-data control, the comparison across three formats is informative, the fitting is likelihood-based with AIC model comparison, and the paper explicitly acknowledges the discrete/finite-support caveat of classical theory. However, the headline inference is not yet statistically supported. The reported error bars treat pooled gaps as independent, only a single shuffled realization is presented, no permutation test or null distribution of shuffled exponents is given, and the T20 dataset is a selected subset. These issues are fixable but are load-bearing for the central claim.","major_comments":[{"comment":"The central inference rests on the contrast between the empirical exponents (0.799–0.843) and the bootstrap-shuffled exponents (0.939–0.979), but the paper reports only a single shuffling realization. No sampling distribution of the shuffled exponent is provided, so the statement that shuffling 'significantly' increases α is not supported. The quoted profile-likelihood intervals are conditional on one shuffled dataset and ignore shuffle-to-shuffle variability and within-player dependence. A permutation test is needed: repeat the within-player shuffle many times, refit α each time, and report the distribution, ideally with player-block resampling. If the null distribution overlaps the empirical α, the claim that temporal ordering drives the deviation must be withdrawn. Moreover, since the shuffled exponents are themselves below the classical α = 1, the correct null for finite, discrete, b","section":"Section 4.4.1 / Table 1"},{"comment":"The maximum-likelihood fit treats every observed gap as an independent observation. Gaps from the same career are not independent: after a new record is set, the distribution of the next gap depends on the value of that record and on the remaining career length. In addition, the final gap after the last record in each finite career is right-censored and is silently discarded; this length-biased sampling can affect the fitted exponent. Because the empirical-versus-shuffled difference is only about 0.1–0.18, the conclusion requires cluster-robust confidence intervals (e.g., block bootstrap by player) and a treatment of censoring. The current profile-likelihood intervals likely understate the uncertainty and do not by themselves establish a significant difference.","section":"Section 4.3, Eqs. (10)–(11)"},{"comment":"The T20 dataset is restricted to the top 101 run scorers, and the Test/ODI datasets are also taken from career-runs rankings. Selection on career runs and career length can bias the gap distribution, especially in T20, where the average career length is only 91 innings. The paper acknowledges the T20 limitation but does not quantify its effect. A sensitivity analysis is needed: vary the inclusion threshold (e.g., all players with at least 50/100/200 innings; top 50/150/200 run scorers), and compare fitted exponents. Without this, the T20 result and its contribution to the '0.799–0.843' range remain fragile.","section":"Section 3.1"},{"comment":"The synthetic-model claim is not quantified. The text states that homogeneous, heterogeneous, and correlated models 'produce results qualitatively similar to bootstrap-shuffled data' and 'do not reproduce the deviations observed in the empirical records,' but no parameter values, sample sizes, fitted exponents, or error bars are reported. One of the Highlights asserts 'Synthetic null models explain broad trends but not all empirical scaling behavior.' This claim needs either quantitative support (with simulation details and distributions of fitted exponents) or removal from the paper.","section":"Section 3.2D and Section 4.4.3"}],"minor_comments":[{"comment":"Typos and wording: 'cuttoff' (Section 4.3), 'mentined' (Section 4.4), extra comma in Eq. (3), and 'as mentined above' should be corrected.","section":"Throughout"},{"comment":"The symbol R(n) is used for the record probability and R(g) for the CCDF. These are different quantities; use distinct notation to avoid confusion.","section":"Eqs. (1) and (12)"},{"comment":"The table should report the number of gaps and the number of players underlying each fit, not only log-likelihood and AIC. This is important for assessing the reliability of the profile-likelihood intervals.","section":"Table 1"},{"comment":"Reference [12] is the ESPN Cricinfo data source, but the text refers to a 'python-based scraping pipeline'; no code repository is given. The data availability statement says data are available 'upon request.' For reproducibility, the processed datasets and analysis code should be deposited in a public repository.","section":"References and reproducibility"},{"comment":"The caption says 'solid lines represent maximum-likelihood fits' but the text does not clearly state whether the fit is to the CCDF or to the PMF and then transformed. Clarify the fitting and plotting procedure.","section":"Figure 3"}],"recommendation":"major_revision","confidential_remarks":"The paper's main empirical claim is plausible and interesting, but the statistical support is currently incomplete: a single shuffle, no permutation null, unquantified selection effects, and no data/code deposit. The editors may wish to require a reproducibility package and a genuine permutation test before considering acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the empirical record-gap exponents for cricket are a real, new observation, and the shuffle/reversal comparison is a reasonable way to show that temporal ordering matters. But the paper's central claim is not yet statistically supported as written: the empirical-vs-shuffled exponent gap (~0.15) is presented without a permutation test or multiple shuffles, and the error bars ignore within-player dependence. I would send it to review, not desk reject, and ask for code/data plus a proper null distribution.\n\nWhat's new: the inter-record gap distribution for personal bests in cricket, fitted with truncated power laws, and the contrast with shuffled careers. That is a clean application of known record theory, not a new theory, and the authors are appropriately modest. The 1/g baseline derivation is fine. The decision to restrict T20 to top run scorers is disclosed, and the paper flags the finite-support discrete-score caveat in the conclusion. Those are points in its favor.\n\nSoft spots, in rough order of importance. First, no code or data is shipped; a 'reasonable request' is not enough for a paper whose entire claim is an empirical exponent. Second, the shuffle analysis appears to be a single shuffle per format. The stress-test note is right: you need the distribution of shuffled α to know whether 0.94–0.98 is really far from 0.80–0.84. Profile-likelihood error bars on pooled gaps are too narrow because gaps from the same player are dependent. Third, the synthetic models are described but never quantified — no exponents, no error bars — so the claim that they 'do not reproduce' the empirical behavior is uncheckable. Fourth, the T20 sample is short-career by nature; the authors admit this, and the T20 reversed career exponent is oddly close to 1, which weakens the temporal-ordering story for that format. Minor point: reversed careers are not a clean null for 'direction of progression' because the endpoint record structure is asymmetric, though the authors don't overclaim there.\n\nOn balance, the central argument — temporal ordering, not the marginal score distribution, drives the deviation — is plausible and probably right, but the paper currently proves it with two fitted points instead of a significance test. Whoever gets this in review should ask for the permutation test and shared code. It deserves referee time.","headline":"New empirical exponents for cricket personal-best gaps, but the ordering claim needs a permutation test and shared data before it is fully supported.","tokens_in":9674,"tokens_out":1746,"would_cite":true,"duration_ms":20787,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Cricket personal-best gaps follow a truncated power-law with exponent 0.799–0.843, well below the classical 1/g prediction, and the deviation is driven by the ordering of innings within a career.","keywords":["record statistics","cricket analytics","truncated power law","temporal correlations","personal bests","nonstationary processes","heavy-tailed distributions","bootstrap shuffle"],"falsifier":"Simulate careers with the same length distribution and the same empirical marginal score distributions as the real data, but with scores generated independently in each inning (no temporal correlations), then fit the gap distribution with the same MLE procedure; if the simulated exponents match the empirical 0.8 rather than the shuffled 0.95, the temporal-ordering interpretation would be wrong and the deviation would stem from something else, such as discretization or selection effects.","tokens_in":8787,"feed_emoji":"🏏","tokens_out":2123,"duration_ms":28087,"temperature":0.7,"pith_summary":"This paper tries to establish that the waiting times between successive personal-best scores in real cricket careers are not described by classical record theory, which predicts a 1/g gap distribution. Instead, the authors find that these gaps follow a truncated power law with exponents around 0.8 in Test, ODI, and T20 cricket. They then show that randomly shuffling each player's innings—preserving the exact set of scores and career lengths but destroying temporal ordering—restores the exponent close to 1. The paper argues this contrast is direct evidence that career evolution, aging, and changing conditions leave a measurable fingerprint in record statistics, beyond what a player's score distribution alone would predict.","feed_headline":"Cricket record gaps follow exponent 0.8, not the classic 1","feed_subtitle":"Shuffling innings restores the classical exponent, so career ordering—not score mix—drives record timing.","key_machinery":"The analysis rests on two quantitative tools: (i) a maximum-likelihood fit of the aggregated gap distribution to a truncated power law, defined as P(g) ∝ g^(−α) e^(−λg) with normalization over finite gap values; and (ii) a bootstrap-shuffle null model that randomly permutes each player's innings, thereby preserving the full marginal score distribution and career length while eliminating all serial correlations. The contrast between the fitted exponent for the real data and the shuffled data is the mechanism that isolates temporal ordering as the cause of the heavy tails.","core_discovery":"The central discovery is that inter-record gap distributions in professional cricket careers are broad, heavy-tailed, and well described by a truncated power law P(g) ∝ g^(−α) e^(−λg) with fitted exponents α in the range 0.799–0.843 across all three formats. This is substantially below the classical i.i.d. value of α = 1. When the temporal ordering of each player's innings is destroyed by random shuffling, the exponents rise to 0.939–0.979, close to the classical prediction. Because the shuffle preserves each player's score distribution, batting average, and career length, the paper concludes that the deviation from classical record statistics arises primarily from the temporal organization","pith_inferences":["A natural extension would be to test whether the same gap-exponent shift appears in other sports with different career-length distributions, such as tennis or baseball, which would generalize the claim beyond cricket.","If the temporal-ordering interpretation holds, one would expect that career phases marked by rule changes, coaching changes, or injury returns might locally alter the gap exponent, a prediction the authors do not test but which follows from their reasoning.","The paper's T20 results rest on much shorter careers and may be more sensitive to the selection of top run-scorers; a logical next step is to apply the same analysis to an expanded T20 dataset as the format matures, or to women's cricket, to check whether α≈0.8 persists."],"forward_implications":["If the claim is correct, personal-best gap statistics can serve as a probe of nonstationarity and path dependence in any long performance sequence, not just cricket.","The empirical exponents below 1 imply that the probability of a new personal best remains relatively high even after long gaps, suggesting that career progression is not a simple monotonic improvement but includes late-career peaks and long-term fluctuations.","The bootstrap results indicate that any adequate generative model of cricket careers must include temporal correlations or nonstationarity; models with only heterogeneous abilities, variable career lengths, or weak correlations fail to reproduce the observed exponents.","The finding that reversed careers still yield exponents below 1 (except T20) shows that the effect is not merely early-career improvement, pointing to more complex career dynamics that future work could model explicitly."],"fun_headline_variants":["Shuffling innings restores cricket's classical 1/g record-gap law","Cricket personal-bests gap exponent is 0.8, shuffle raises it to ~1","Why cricket record gaps deviate from 1/g: career evolution","Temporal order, not score mix, sets cricket inter-record gaps","Cricket records: gap distribution exponent 0.8 beats classic 1"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The claim that temporal ordering drives the observed scaling assumes that the maximum-likelihood fit of pooled gaps and the comparison with bootstrap-shuffled careers are not biased by finite career lengths, discrete scores, and the restriction to top run-scorers in T20; if those factors systematically lower the empirical exponent or raise the shuffled exponent, the inferred ordering effect could shrink or vanish.","fun_headline_variants_meta":{"raw":{"variants":["Shuffling innings restores cricket's classical 1/g record-gap law","Cricket personal-bests gap exponent is 0.8, shuffle raises it to ~1","Why cricket record gaps deviate from 1/g: career evolution","Temporal order, not score mix, sets cricket inter-record gaps","Cricket records: gap distribution exponent 0.8 beats classic 1"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000338,"raw_usage":{"total_tokens":1730,"prompt_tokens":796,"completion_tokens":934,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":540,"completion_tokens_details":{"reasoning_tokens":833}},"tokens_in":540,"tokens_out":934,"duration_ms":10407,"temperature":1.0,"reasoning_tokens":833,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T18:57:10.084550+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate careers with the same length distribution and the same empirical marginal score distributions as the real data, but with scores generated independently in each inning (no temporal correlations), then fit the gap distribution with the same MLE procedure; if the simulated exponents match the empirical 0.8 rather than the shuffled 0.95, the temporal-ordering interpretation would be wrong and the deviation would stem from something else, such as discretization or selection effects.","supporting_citations":[],"review_version":1}