{"id":"6c65cb13-87df-4858-ab69-b2efb487eebf","arxiv_id":"2501.03500","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Most isolated pulsars show significant spin-down rate variability, with amplitude scaling as spin-down rate to the power 0.85 and little dependence on spin frequency.","lead":"Using three decades of Parkes telescope timing, the authors find that 238 of 259 isolated pulsars show measurable changes in their spin-down rate, and 52 also change their radio pulse shape. The result suggests pulsar clocks are far less stable than often assumed, with direct consequences for gravitational wave searches that rely on millisecond pulsar timing.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The K>1 variability criterion in Eq. 4 is uncalibrated; residual glitch recoveries and annual positional offsets are known to mimic spin-down variations, so the 238/259 count (92%) and the scaling relation in Eq. 7 may be inflated without injection-based false-positive tests.","rationale":"The paper is a large, careful population study with clear methods and extensive appendices. I read the central claim as twofold: (1) 92% of the 259 isolated pulsars show significant spin-down variability, and (2) the variability amplitude scales as Eq. 7 with no strong spin-frequency dependence (Eq. 9). Both claims depend entirely on the K>1 criterion in Eq. 4. That criterion is applied to a GP-derived nu_dot timeseries, not to direct measurements, and no calibration is provided. The paper itself flags two mechanisms--glitch-recovery residuals (Section 3.1) and unresolved annual positional offsets (Section 3.3)--that can imprint exactly the kind of smooth nu_dot excursions the GP would fit. The BIC-based model selection is a heuristic, and the authors admit making 'by-eye judgement calls' when the BIC disagrees with visual inspection. Under these conditions, an unvalidated threshold invites systematic false positives. The scaling relation in Eq. 7 is then fit only to the 238 'variable' pulsars, so contamination of that sample propagates directly into the reported power-law index and normalization, and into the PTA forecast in Section 5.2. The sample-selection bias (high-Edot pulsars favored by Fermi support) is a real but secondary limitation affecting the word 'ubiquitous'; it does not by itself invalidate the scaling relation, which is conditional on the variable subsample. My proposed injection test isolates the classification step and would settle whether the count and scaling are trustworthy. I therefore agree with the reader's weakest assumption and see no reason to move the CONDITIONAL verdict.","tokens_in":66235,"tokens_out":4371,"duration_ms":41210,"concrete_test":"For each of the 259 pulsars, generate N (e.g., 100) synthetic timing-residual realizations consisting of the best-fit deterministic timing model (including fitted glitches) plus Gaussian white noise with the observed per-ToA uncertainties. Run the exact Section 3.3 pipeline on these simulations: GP fits with one/two squared-exponential kernels and the optional fixed 1-yr sinusoidal kernel, BIC model selection, then compute K from Eq. 4. The fraction of simulations with K>1 gives the false-positive rate of the variability classification under the null hypothesis of no spin-down variability. If the aggregate false-positive rate is >5%, or if simulated pulsars with low ToA counts or high glitch activity preferentially exceed K>1, the 238 count and the Eq. 7 scaling fit are uncalibrated and the central claim needs revision.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.3 defines K (Eq. 4) from the extrema of the Gaussian-process second derivative of the timing residuals divided by the mean GP uncertainty, and declares a pulsar variable when K>1. No injection, null, or false-positive test is presented for this threshold, and the paper explicitly acknowledges two processes that can generate spurious nu_dot variability: imperfectly removed glitch recoveries (Section 3.1) and annual positional offsets that remain degenerate with quasi-periodic timing noise for 'a substantial number of pulsars' (Section 3.3). Because K is computed from the GP model rather than from directly measured nu_dot values, the model can create smooth excursions even in pure noise, while the GP predictive uncertainty (Brook et al. 2016 Eqs. 9-10) may understate the true variance arising from model selection and hyperparameter uncertainty. If the false-positive rate at K>1 is, say, 20-50% (plausible given the demonstrated 1-yr artifacts), the headline 238/259 (92%) claim and the scaling relation Eq. 7, which is fit only to the 'variable' pulsars, would both be biased. This is the most load-bearing assumption because the population ubiquity statement, the PTA-noise implication, and the quoted power-law all rest on that single uncalibrated threshold.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript applies Gaussian-process regression and Bayesian inference to roughly 30 years of Parkes (Murriyang) timing data for 259 isolated, non-recycled pulsars, and claims that 238 of them show significant spin-down rate variability under a K>1 criterion, with 52 also showing profile shape changes. It derives an empirical scaling relation δν̇ = 10^(-4.5±0.5) |ν̇_weak|^(0.85±0.04) with only marginal spin-frequency dependence, and discusses implications for pulsar timing array searches, quasi-periodic variability, and planetesimal interaction scenarios.","tokens_in":66447,"tokens_out":6408,"duration_ms":62446,"significance":"The paper is potentially important because it assembles the largest catalogue of variable pulsars to date and, if the central claims hold, establishes that spin-down instability is common across the isolated pulsar population rather than confined to a few notable objects. The long-baseline Parkes data, the use of public data products, and the explicit Bayesian framework for the scaling fit are strengths. However, the headline detection rate rests entirely on an uncalibrated K-metric threshold; without a false-positive quantification, the ubiquity claim and the fitted power law are not yet established. The paper is therefore promising but requires a validation step before the central claims can be accepted.","major_comments":[{"comment":"The K>1 criterion is not calibrated against a null hypothesis. The paper states 'We used a threshold of K > 1' but does not report the false-positive rate of this threshold for noise-only data, for data with imperfectly removed glitch recoveries, or for data with annual positional offsets, both of which are acknowledged as sources of spurious ν̇ variability in Sections 3.1 and 3.3. Because K is computed from the extrema of the Gaussian-process second derivative divided by the GP predictive uncertainty, smooth excursions can appear in pure noise while model-selection and hyperparameter uncertainties are not included in σ_ν̇,mean. An injection study that reports the fraction of K>1 pulsars expected by chance, for white noise and for simulated glitch-recovery and annual-position signals, is required to support the 238/259 claim and the scaling relation in Eq. (7), which is fit only to the selected pulsars.","section":"3.3, Eq. (4)"},{"comment":"The kernel selection procedure combines BIC with unspecified visual inspection. The text says 'on occasion we had to make by-eye judgement calls when one model visually matched the data better than another in spite of the reported BIC.' This subjective step is not quantified: the number of pulsars affected, the criteria used, and the reproducibility of the decisions are not reported. Since the choice of one-kernel, two-kernel, and annual-sinusoid models directly shapes the ν̇ timeseries and hence K, this selection uncertainty should be propagated or at least enumerated.","section":"3.3, model selection"},{"comment":"The scaling relation is fitted only to the 238 K>1 pulsars and treats δν̇ as known data. However, δν̇ = |ν̇_min| − |ν̇_max| and |ν̇_weak| are both outputs of the same Gaussian-process fit, so their uncertainties are correlated and model-dependent; the likelihood in Eq. (6) adds a single scatter σ_Q but does not propagate the GP posterior covariance. Given the selection-threshold issue in Eq. (4), the reported index 0.85±0.04 and spin-frequency exponent −0.18±0.17 in Eqs. (7) and (9) should be presented as conditional on the detection method, with an additional sensitivity analysis excluding marginal K values.","section":"5.1, Eqs. (5)–(9)"}],"minor_comments":[{"comment":"The rate calculation uses 27/260 pulsars, but the sample is 259 and Section 4 reports 238/259; please reconcile the denominator.","section":"5.4, Eq. (13)"},{"comment":"The tables contain inconsistent notation (e.g., missing minus signs in several exponents and 'e' notation such as '7.1𝑒+ 01'), which makes verification difficult.","section":"Tables A1 and A2"},{"comment":"Section 4.2 mentions 'another 28 pulsars' while the Conclusions state '29 pulsars for which we describe the links ... for the first time'; the counting should be clarified.","section":"4.2 and Conclusions"},{"comment":"The equations for the GP predictive variance from Brook et al. (2016) are not reproduced; since the K-metric denominator is central, a brief summary of those equations would help the reader assess the significance metric.","section":"3.3"},{"comment":"The labels give δν̇/|ν̇| without error bars; adding typical uncertainties or a note on how σ_ν̇,mean varies would make the K>1 selection easier to evaluate.","section":"Figure 3"}],"recommendation":"major_revision","confidential_remarks":"To the editor: the manuscript is within scope for MNRAS and the dataset is valuable. The main weakness is not the GP machinery itself but the absence of any false-positive calibration for the variability classification; I believe this is fixable with injection tests and should be required before acceptance. I did not find citation or novelty concerns, though the paper would benefit from a short validation subsection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The headline result is the size of the variable-pulsar catalogue: 238 of 259 isolated pulsars with K > 1 spin-down variability, plus a population scaling delta_nu_dot = 10^-4.5 |nu_dot_weak|^0.85. That is a genuine step up from the few dozen previously known objects, and the paper does the field a service by assembling it and laying out the GP methodology clearly. The timing solutions, glitch handling, and the appendix tables are thorough. The scaling relation is a real empirical result, not a circular derivation, and the spin-frequency dependence is sensibly left as marginal.\n\nThe weak spot is exactly where the stress-test puts it: the K > 1 threshold in Eq. 4 is not calibrated. K is computed from the extrema of the GP second derivative divided by a mean uncertainty, and there is no injection, null, or false-positive test anywhere in the paper. That would matter less if the known artifacts were minor, but the text itself says imperfectly removed glitch recoveries will introduce unwanted artefacts, and that annual positional offsets remain degenerate with quasi-periodic timing noise for a substantial number of pulsars. Those are precisely the processes that can create spurious spin-down variations in a GP model. With a threshold that is essentially 'the model moved more than one sigma', it is entirely possible that a meaningful fraction of the 238 are noise-driven excursions. The 92% figure and the power-law fit, which is performed only on the 'variable' sample, both inherit any inflation.\n\nA second, softer concern: the sample is explicitly biased toward high-spin-down-energy, Fermi-target pulsars, so the word 'ubiquitous' in the abstract outruns the evidence. That is an interpretive overreach rather than a technical error, and it is easy to fix in revision.\n\nThe central argument, that spin-down variability is common and scales with spin-down rate, probably survives even if the exact fraction is lower. But the paper needs false-positive calibration before the headline count can be trusted. I would send it to a serious referee, with the explicit request that the authors add injection-based null tests and either re-derive the scaling relation or report how robust it is to a plausible contamination rate. As it stands, I would not cite the 92% or the quoted power-law without checking the robustness myself.","headline":"Large, carefully built catalogue whose headline 92% variability rate rests on an uncalibrated GP threshold; the scaling relation is plausible but inherits the risk.","tokens_in":67155,"tokens_out":2681,"would_cite":false,"duration_ms":26689,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Most isolated pulsars do not spin down steadily: 238 of 259 show significant spin-down variability, and the fluctuation amplitude grows with spin-down rate.","keywords":["pulsars","spin-down variability","radio emission variability","Gaussian process regression","pulsar timing arrays","mode switching","timing noise","gravitational waves"],"falsifier":"Run the same Gaussian-process and K-metric pipeline on simulated timing residuals that contain only white noise and the same observation times as the 259 pulsars, and count how many simulated objects cross K > 1; a non-negligible false-positive rate would mean the 92% variable fraction is inflated by the method rather than by the pulsars.","tokens_in":1592,"feed_emoji":"🔄","tokens_out":1907,"duration_ms":70257,"temperature":0.7,"pith_summary":"The paper sets out to establish that rotational instability is the norm, not the exception, for isolated pulsars: it reports that 238 of 259 monitored pulsars show significant variation in spin-down rate, with 52 also changing radio pulse shape. Because the sample is the largest yet assembled, the claim moves this behaviour from a collection of curiosities to a population-wide property. The authors also derive a quantitative scaling between the amplitude of spin-down fluctuations and the mean spin-down rate, with little dependence on spin frequency. If correct, the same process may operate in millisecond pulsars and must be modelled in pulsar timing array searches for nanohertz gravitational waves. The paper further argues that quasi-periodic spin-down modulations do not follow free-precession scaling, and that transient spin-down events look consistent with asteroid impacts.","feed_headline":"Nine in ten isolated pulsars show unstable spin-down","feed_subtitle":"A three-decade survey finds 238 of 259 pulsars wobble in spin-down rate, complicating gravitational-wave searches.","key_machinery":"The load-bearing object is the K-metric, $$K = \\frac{|\\dot{\\nu}_{\\rm min}| - |\\dot{\\nu}_{\\rm max}|}{2\\sigma_{\\dot{\\nu},\\,{\\rm mean}}},$$ computed from Gaussian process fits to timing residuals; K > 1 marks a pulsar as variable. Gaussian process regression with squared-exponential kernels (and Matérn kernels for profile variability maps) produces continuous spin-down and profile models from unevenly sampled observations, while Bayesian information criterion model selection decides between one or two kernels and a fixed one-year sinusoidal kernel used to absorb positional offsets. This machinery lets the authors measure fluctuation amplitudes, search for correlations between profile changes and spin-down, and test scaling relations against spin, spin-down rate, characteristic age, and magnetic field.","core_discovery":"On its own terms, the paper claims that 238 of 259 isolated, non-recycled pulsars display significant spin-down variability as measured by the K-metric (K > 1), and 52 of those also show substantial changes in pulse profile shape. The fluctuation amplitude follows the relation $\\delta\\dot{\\nu} = 10^{-4.5 \\pm 0.5}\\,|\\dot{\\nu}_{\\rm weak}|^{0.85 \\pm 0.04}$, with only a marginal dependence on spin frequency ($\\nu^{-0.18 \\pm 0.17}$). This is the largest catalogue of variable pulsars to date, and the authors interpret it as evidence that these behaviours are ubiquitous among the broader pulsar population. They also find that quasi-periodic spin-down modulations in 45 pulsars do not follow the scaling expected from free precession, and that 68 transient spin-down events in 26 pulsars imply frequent interactions with small bodies if interpreted as asteroid impacts.","pith_inferences":["If the scaling relation extends to millisecond pulsars, timing-array analyses should include an explicit spin-down fluctuation term; standard red-noise models will absorb part of it and can bias the inferred gravitational-wave background amplitude.","The 52/238 profile-change fraction is likely sensitivity-limited by per-epoch signal-to-noise and pulse jitter, so longer integrations should reveal more shape-changing pulsars even if the spin-down result is unchanged.","Because K > 1 was set without injection-based false-positive calibration, re-running the pipeline on synthetic noise-only data would directly test the 92% detection rate; until then, the ubiquity claim rests on the assumption that glitch-recovery and positional artefacts are negligible.","A natural next test is to monitor the 45 quasi-periodic pulsars for phase drift or state changes; stable periods over decades would keep a geometric clock such as precession viable, while drift and switching would favour magnetospheric reconfiguration."],"forward_implications":["Spin-down variability is common enough that the steady-clock assumption for isolated pulsars needs revision in population studies.","The amplitude scaling $\\delta\\dot{\\nu} \\propto |\\dot{\\nu}_{\\rm weak}|^{0.85}$ predicts detectable spin-down fluctuations in millisecond pulsars, where timing-array noise models currently use red power laws.","The same Gaussian-process pipeline can be applied to other long-term timing data sets to enlarge the variable-pulsar catalogue.","Quasi-periodic modulation periods scattered across $P$, $\\tau_c$, and $\\dot{E}$ fail the free-precession scaling relations, strengthening magnetospheric state-switching as the driver.","Transient spin-down events, if caused by asteroid impacts, imply that debris discs and asteroid belts around pulsars should be common."],"supporting_citations":[{"why":"Establishes the Gaussian-process methodology for spin-down and profile variability maps that this paper adapts.","marker":"Brook et al. (2016)"},{"why":"Provides the gp_nudot.py fitting approach and the earlier scaling relation that this work extends.","marker":"Shaw et al. (2022)"},{"why":"The original 17-pulsar sample of correlated profile/spin-down switching whose $\\delta\\dot{\\nu} \\propto 0.01|\\dot{\\nu}|$ relation is compared against.","marker":"Lyne et al. (2010)"},{"why":"Independent large-sample profile variability catalogue used for comparison and context.","marker":"Basu et al. (2024)"},{"why":"Evidence for a common 'red' signal in pulsar timing arrays, the measurement the paper says spin-down variability could contaminate.","marker":"Reardon et al. (2023)"},{"why":"Shows red power-law noise models are insufficient for non-stationary timing behaviour, supporting the claim that spin-down variations need explicit modelling.","marker":"Keith & Niţu (2023)"},{"why":"Derives free-precession scaling relations that the 45 quasi-periodic pulsars are tested against and found inconsistent.","marker":"Jones (2012)"}],"fun_headline_variants":["238 of 259 pulsars show unstable spin","Spin-down wobble is the norm for pulsars","Nine in ten pulsars spin down erratically","92% of pulsars have erratic spin-downs","Largest pulsar variability catalogue yet"],"cache_read_input_tokens":69120,"weakest_assumption_plain":"The classification of 238 pulsars as variable assumes that the K-metric, with its K > 1 threshold and no injected-noise false-positive test, separates genuine spin-down variability from artefacts caused by imperfect glitch recovery, annual positional sinusoids, and receiver changes.","fun_headline_variants_meta":{"raw":{"variants":["238 of 259 pulsars show unstable spin","Spin-down wobble is the norm for pulsars","Nine in ten pulsars spin down erratically","92% of pulsars have erratic spin-downs","Largest pulsar variability catalogue yet"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000955,"raw_usage":{"total_tokens":4074,"prompt_tokens":947,"completion_tokens":3127,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":563,"completion_tokens_details":{"reasoning_tokens":3055}},"tokens_in":563,"tokens_out":3127,"duration_ms":21277,"temperature":1.0,"reasoning_tokens":3055,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T21:52:32.948406+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same Gaussian-process and K-metric pipeline on simulated timing residuals that contain only white noise and the same observation times as the 259 pulsars, and count how many simulated objects cross K > 1; a non-negligible false-positive rate would mean the 92% variable fraction is inflated by the method rather than by the pulsars.","supporting_citations":[],"review_version":1}