{"id":"ae128bfd-9832-4e20-90f6-26c814a3d45d","arxiv_id":"2506.22697","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A 173-pulsar timing survey confirms known timing-noise scaling, shows that second-frequency-derivative measurements are sensitive to the low-frequency cutoff in the timing noise model, and reports 17 glitches including one previously unknown.","lead":"A refurbished Australian radio telescope monitored 173 pulsars for about two years, and the team measured how pulsars' rotation wanders and jumps. The results confirm that timing noise can masquerade as a change in a pulsar's spin-down rate, and set limits on glitches the survey might have missed.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The analytic δ\\ddot{\\nu} scaling in §5.1 assumes a pure power-law timing noise PSD with no low-frequency turnover; if real noise mean-reverts on timescales ≤ T_span, the claimed three-fold cutoff sensitivity and the conclusion that only the Crab's \\ddot{\\nu} is real would be overstrong.","rationale":"The reader's weakest_assumption identifies exactly the same load-bearing concern: the analytic δ\\ddot{\\nu} estimate assumes a pure power-law timing noise PSD with no mean reversion on timescales comparable to the observing span. My analysis confirms this is the most critical point because it directly impacts both the factor-of-three sensitivity claim and the model-selection conclusion that only the Crab pulsar has a significant \\ddot{\\nu}. The paper does not independently test for low-frequency turnover; its consistency checks in §4.2 only cover frequencies ≥ 1/T_span. The injection-recovery test I propose would settle whether the pure power-law assumption is adequate for the conclusions. Since the reader already conditioned the verdict on this issue, no change is needed.","tokens_in":45571,"tokens_out":5091,"duration_ms":76698,"concrete_test":"Run injection-recovery simulations with the same enterprise pipeline on synthetic ToAs generated with a timing noise PSD that includes a low-frequency turnover (e.g., a broken power law with a knee at f_knee ~ 1/(2T_span), or a mean-reverting Ornstein-Uhlenbeck process in \\ddot{\\nu} with timescale ≈ T_span). For a sample of pulsars with |β−6|<1, inject a known secular \\ddot{\\nu} just below the current 95% upper limit, and compare (a) the empirical spread of recovered \\ddot{\\nu} across realizations with the analytic δ\\ddot{\\nu} from equation (20), and (b) the ratio of \\ddot{\\nu} uncertainties between TNF2 and TNLONGF2. If the ratio drops below ~2 or the recovered \\ddot{\\nu} is biased relative to the injected value, the paper's interpretation that most non-zero \\ddot{\\nu} are timing-noise artifacts is not robust to low-frequency PSD turnover.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim—that the \\ddot{\\nu} uncertainty varies three-fold with the low-frequency cutoff and that most of the 39 non-zero \\ddot{\\nu} measurements are timing-noise artifacts—rests on the analytic estimate in §5.1, equations (14)–(20). This estimate assumes timing noise is a pure random walk in \\ddot{\\nu} (for |β−6|<1) with no mean reversion on timescales comparable to T_span, as stated in the footnote after equation (16). The PSD is then taken to be a pure power law extending to arbitrarily low frequency, and the enterprise-fitted A_red and β are used to set σ_\\ddot{\\nu} by equating PSDs. If the true PSD has a turnover or mean-reverting behavior below 1/T_span (or even between 1/T_span and 1/(2T_span)), the expected spread δ\\ddot{\\nu} would be smaller than predicted. Consequently, the median ratio δ\\ddot{\\nu}/Δ\\ddot{\\nu} ≈ 2.6 for TNF2 could be an overestimate, weakening the inference that most non-zero \\ddot{\\nu} detections are noise-dominated. The same pure-power-law assumption underpins the TNLONGF2 model used for model selection: adding a 1/(2T_span) mode without allowing a turnover may misestimate the evidence difference between TNLONGF2 and TNLONG, potentially biasing the conclusion that only the Crab pulsar retains a significant \\ddot{\\nu}. Section 4.2's consistency checks only probe frequencies ≥ 1/T_span and cannot exclude lower-frequency turnover, so this is not empirically settled by the paper.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents the first scientific results from the UTMOST-NS pulsar timing programme, which monitored 173 pulsars at high cadence from April 2021 to June 2023, combined with earlier UTMOST-EW data. The authors use enterprise with a power-law red-noise model and nested sampling to estimate timing noise parameters, investigate correlations between timing noise strength and spin parameters, and study second frequency derivatives. They report 39 pulsars with 95% credible intervals excluding zero for ddot nu under the per-pulsar preferred model, find that this number drops to 17 when a low-frequency cutoff 1/(2Tspan) is imposed, and that only the Crab pulsar survives a further Bayes-factor test against a model with ddot nu fixed to zero. They also derive analytic conditions for when timing noise should produce anomalous braking indices. The glitch analysis uses an HMM detector: three online detections, a newly reported glitch in PSR J1902+0615, 17 glitches analysed with enterprise, and systematic upper limits on undetected glitch sizes with mean Delta nu^90/nu = 6.3e-9 offline.","tokens_in":46008,"tokens_out":10792,"duration_ms":114512,"significance":"If the central claims hold, the paper is significant for two reasons. First, it provides a high-cadence, multi-year timing dataset from a newly refurbished instrument and gives the community a transparent Bayesian treatment of timing noise and ddot nu, including a demonstration that the low-frequency cutoff of the red-noise model changes the inferred uncertainty on ddot nu by roughly a factor of three and that most 95% detections of ddot nu do not survive a stricter model comparison. Second, the systematic HMM-based glitch upper limits, with the full injection procedure described, are a useful and falsifiable product for population studies. The paper is commendably explicit about its model assumptions and about the ambiguity of the surviving ddot nu detections. The main caveat is that the quantitative ddot-nu conclusions rest on a pure power-law, no-mean-reversion timing-noise model; the paper should either test that assumption or qualify the headline claims accordingly.","major_comments":[{"comment":"The derivation of delta ddot nu assumes that timing noise is a pure random walk in nu_dot with no mean reversion on timescales comparable to Tspan, and the fitted power-law PSD is taken to extend to arbitrarily low frequency. This assumption is load-bearing for the three-fold cutoff claim and for the model comparison in §5.3: the median delta ddot nu / Delta ddot nu values (2.59 for TNF2 versus 0.91 for TNLONGF2) and the conclusion that only the Crab survives a strict Bayes-factor test would be altered if the true PSD turns over or mean-reverts below 1/Tspan. The consistency checks in §4.2 only probe frequencies at or above 1/Tspan and therefore do not constrain this regime. I request a robustness test with a low-frequency turnover or a mean-reverting model (e.g., the OU model in Vargas & Melatos 2023), or with injections, and the abstract and conclusion should be conditioned on the pure-power-law assumption.","section":"§5.1 (Eqs. 14–16; footnote 3)"},{"comment":"The paper says the analytic scaling is 'validated', but the validation is not external: delta ddot nu is computed from the same A_red and beta fitted to the same datasets, and the comparison with the enterprise uncertainties is an internal consistency check of the power-law noise model. Equations (24)–(27) and Table 5 follow from the same model rather than being tested against independent data. I recommend relabelling this as a self-consistency check, or adding an independent test such as posterior predictive simulations, before the word 'validated' is used in the abstract and conclusion.","section":"Abstract; §5.1–5.2"}],"minor_comments":[{"comment":"The sentence 'We measure 39 non-zero values of ddot nu when considering both models with and without low-frequency modes' is ambiguous, because Table 4 reports detections under the per-pulsar selected model while Table 6 reports the 17 that survive the TNLONGF2-only analysis; please state explicitly that 39 is the union over the two model choices.","section":"Abstract and §5.3"},{"comment":"For the |beta-4| case the text reports 'the mean delta ddot nu / Delta ddot nu in our sample is 0.50 and 0.40', whereas the |beta-6| results are quoted as medians; please specify which statistic is used in each case.","section":"§5.1"},{"comment":"Several rows in Table 4 have malformed uncertainty formatting, notably J1605-5257 ('8.4+7.4 -7.1 x 10^4') and J1703-3241 ('8.2+2.5 -.3 x 10^-27'); these should be corrected to match the convention used elsewhere.","section":"Table 4"},{"comment":"The RWF0 model approximates away the correlation between nu and nu_dot (Eqs. 35–37) to avoid a singular covariance matrix; the paper should state whether this approximation is conservative for glitch detection and, if possible, quantify its effect with injected signals.","section":"§6.2.2"},{"comment":"The sentence concluding that the PSD 'makes good sense' to model as a simple power law should include the qualifier 'over the frequencies probed by our datasets' in the same sentence, since the text immediately acknowledges that lower-frequency behaviour is not tested.","section":"§4.2"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is a solid dataset paper. You get a new high-cadence timing sample for 173 pulsars, a previously unreported glitch in J1902+0615, and the first systematic glitch sensitivity limits at the 6e-9 level from this kind of HMM search. The timing-noise scaling results are consistent with earlier work, and the Bayesian inference is carefully done and transparent. The glitch parameters and upper limits are the cleanest part of the paper and stand on their own.\n\nThe second-derivative story is interesting but the abstract oversells it. The 'analytic scaling' in Sections 5.1 and 5.2 is not an independent validation: it uses the same A_red and beta fitted by enterprise to compute an expected ddot-nu spread, then compares that to the enterprise uncertainty. That is a consistency check, not a validation. The stress-test note is right: the analytic estimate assumes a pure power-law PSD with no low-frequency turnover or mean reversion. If real timing noise mean-reverts on timescales near T_span, the predicted delta-ddot-nu shrinks and the median ratio of 2.6 could be too high. The paper does flag the assumption in a footnote, and the empirical difference between the TNF2 and TNLONGF2 models still shows that the cutoff matters, but the specific claim about 'approximately three-fold' sensitivity is conditional on the no-turnover model.\n\nI would not treat the 'only the Crab survives' conclusion as definitive. It depends on a particular model-selection setup, and the authors themselves are appropriately cautious in Section 5.3. The result is suggestive and worth reporting, but it should be framed as model-dependent.\n\nMinor points: the data products are promised through a portal but are not public yet, and the full analysis code is not released. That is worth fixing before acceptance. The HMM code is on GitHub, which helps.\n\nWho should read this: pulsar timing people, glitch population folks, and anyone interpreting braking indices. It deserves a serious referee. Send it to review, but ask for a softer 'validated' wording, a direct discussion of low-frequency turnover (or a test with a mean-reverting noise model), and public data and code.\n\nYes, I would cite it for the dataset and the glitch upper limits, and I would take it to a reading group.","headline":"A valuable new dataset and careful timing-noise analysis, but the 'validated' analytic scaling is an internal consistency check that inherits the pure power-law assumption, so the cutoff-sensitivity claim should be read with that caveat.","tokens_in":46534,"tokens_out":1701,"would_cite":true,"duration_ms":23572,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["97.60.Gb"],"model":"deepseek-v4-flash","headline":"The paper argues that the choice of low-frequency cutoff in the timing-noise model changes the inferred uncertainty on pulsar second frequency derivatives by about a factor of three, and that of 39 apparent non-zero detections only the…","keywords":["pulsar timing","timing noise","second frequency derivative","braking index","glitches","hidden Markov model","Molonglo telescope","neutron stars"],"falsifier":"Take any pulsar in this sample with $\\beta \\approx 6$ and a reported non-zero $\\ddot{\\nu}$, extend its timing baseline several-fold, and check whether the posterior uncertainty on $\\ddot{\\nu}$ grows as $T_{\\mathrm{span}}^{-1/2}$ as predicted by equation (20) and whether the recovered noise spectrum remains a single power law below $1/T_{\\mathrm{span}}$. A measured low-frequency turnover or mean reversion in the noise spectrum, or a $\\ddot{\\nu}$ that stays stable and non-zero as the baseline grows, would falsify the noise-dominated interpretation.","tokens_in":45352,"feed_emoji":"🦀","tokens_out":17390,"duration_ms":148465,"temperature":0.7,"pith_summary":"This paper reports the first science results of a two-year, high-cadence pulsar timing campaign at the Molonglo telescope, covering 173 southern pulsars with combined baselines of up to seven years. Its central claim is that timing noise masquerades as a secular second frequency derivative: the inferred uncertainty on $\\ddot{\\nu}$ changes by roughly a factor of three depending on whether the noise model includes Fourier modes below the reciprocal of the observing timespan. Of the 39 pulsars whose $\\ddot{\\nu}$ excludes zero at 95% confidence under either model, only 17 survive when the low-frequency modes are always included, and only the Crab pulsar survives when the model containing $\\ddot{\\nu}$ must also beat a model without it by a Bayes factor of 5, giving $n = 2.54 \\pm 0.03$. The paper validates an analytic scaling that predicts when anomalously large braking indices are expected from timing noise alone, and it catalogues 17 glitches, one previously unreported, with systematic upper limits on the size of undetected glitches.","feed_headline":"39 apparent spin-down signals drop to one under a stricter noise test","feed_subtitle":"The Crab is the only pulsar whose second frequency derivative survives the timing-noise check.","key_machinery":"The load-bearing object is an assumed power-law timing-noise spectrum $P(f) = A_{\\mathrm{red}}^2 (12\\pi^2 f_{\\mathrm{yr}}^3)^{-1}(f/f_{\\mathrm{yr}})^{-\\beta}$, implemented as a set of harmonically related sinusoids whose lowest frequency is either $1/T_{\\mathrm{span}}$ (model TNF2) or $1/(2T_{\\mathrm{span}})$ (model TNLONGF2). The $\\ddot{\\nu}$ argument runs through an analytic estimate of the spread in $\\ddot{\\nu}$ produced by a random walk in $\\dot{\\nu}$ (spectral index $\\beta \\approx 6$), $\\delta\\ddot{\\nu} \\sim \\sqrt{f_{\\mathrm{yr}}^3\\nu^2(2\\pi)^6/24\\pi^2}\\, A_{\\mathrm{red}}^{\\beta=6}\\, T_{\\mathrm{span}}^{-1/2}$, and its analogue $\\delta\\ddot{\\nu} \\propto T_{\\mathrm{span}}^{-3/2}$ for a random walk in $\\nu$ ($\\beta \\approx 4$). Comparing this predicted spread with the uncertainty returned by the inference engine shows that the $1/T_{\\mathrm{span}}$ cutoff underestimates $\\ddot{\\nu}$ errors by a median factor of 2.59, while the $1/(2T_{\\mathrm{span}})$ model tracks the expected variance; the same comparison yields the condition $\\delta n < 1$ that separates measurable braking indices from noise-dominated ones. For the glitch half of the paper, the central object is a hidden Markov model over discretised $(\\nu, \\dot{\\nu})$ states, whose transition matrix encodes one of two random-walk noise models (RWF0 for $\\beta \\approx 4$, RWF1 for $\\beta \\approx 6$) and whose log-Bayes-factor maps flag glitch candidates.","core_discovery":"The paper's central claim is that most reported 'anomalous' braking indices — values of $n = \\nu\\ddot{\\nu}/\\dot{\\nu}^2$ orders of magnitude outside the canonical range $1 \\lesssim n_{\\mathrm{pl}} \\lesssim 7$ — are artifacts of timing noise that the standard noise model fails to absorb. In the standard model (TNF2), timing noise is a sum of sinusoids with lowest frequency $1/T_{\\mathrm{span}}$; in the alternative model (TNLONGF2), modes down to $1/(2T_{\\mathrm{span}})$ are always present. The inferred uncertainty on $\\ddot{\\nu}$ differs by about a factor of three between the two models, and the number of pulsars with $\\ddot{\\nu} \\neq 0$ at 95% confidence drops from 39 to 17 when the standard model is discarded. Requiring the alternative model to beat a model with $\\ddot{\\nu}$ fixed at zero by $\\ln B > 5$ leaves exactly one detection, the Crab pulsar, whose measured $n = 2.54 \\pm 0.03$ matches its interglitch braking index. The paper also derives explicit analytic conditions, in terms of the noise amplitude $A_{\\mathrm{red}}$, spectral index $\\beta$, spin frequency, first derivative, and observing span, for when timing noise alone is expected to produce $|n| \\gg 1$, and shows that all but two pulsars in the sample satisfy them.","pith_inferences":["Timing-noise model selection of this kind should routinely include a second noise model that always retains the lowest Fourier mode; pulsar timing arrays searching for a gravitational-wave background face the same degeneracy between low-frequency noise and a low-frequency signal.","A testable prediction follows from the two scalings: pulsars with $\\beta \\approx 4$ should settle onto a stable $\\ddot{\\nu}$ as baselines grow over decades, while $\\beta \\approx 6$ pulsars should keep re-drawing their apparent $\\ddot{\\nu}$, a distinction future long-baseline timing can check.","Treating pre-baseline glitch recovery as one possible origin of residual $\\ddot{\\nu}$, the sample's own glitch amplitudes could be combined with the $\\delta\\ddot{\\nu}$ scalings to estimate how many of the 17 surviving detections are recovered glitches rather than secular braking."],"forward_implications":["Any reported $\\ddot{\\nu}$ detection from a fit that drops low-frequency noise modes must be re-checked: the $1/T_{\\mathrm{span}}$ cutoff underestimates the $\\ddot{\\nu}$ uncertainty by a typical factor of about three for pulsars with $\\beta \\approx 6$.","Pulsars whose timing noise behaves like a random walk in $\\dot{\\nu}$ will need median baselines of about $4 \\times 10^7$ yr before a canonical braking index can outlive the noise, while random-walk-in-$\\nu$ pulsars need only about 280 yr.","The Crab pulsar's braking index $n = 2.54 \\pm 0.03$ is the sample's only surviving measurement and agrees with its interglitch value.","The recovered timing-noise parameters $A_{\\mathrm{red}}$ and $\\beta$ do not shift systematically when the data set is extended or the inference pipeline is changed, supporting power-law noise modelling over the probed frequency range.","The hidden Markov model pipeline places a mean 90% upper limit of $\\Delta\\nu^{90\\%}/\\nu = 6.3 \\times 10^{-9}$ on undetected glitches across the sample, an order of magnitude tighter than the online glitch search."],"supporting_citations":[{"why":"Supplies the stochastic timing-noise model whose $\\dot{\\nu}$-random-walk limit appears in equations (14)–(16), and the general condition for when a secular $\\ddot{\\nu}$ can be measured.","marker":"Vargas & Melatos (2023)"},{"why":"Shows that the standard $1/T_{\\mathrm{span}}$ noise cutoff can mis-estimate $\\ddot{\\nu}$ uncertainties, the effect quantified here as a threefold sensitivity.","marker":"Keith & Niţu (2023)"},{"why":"Provides the Crab interglitch braking index $2.519(2)$ against which the sample's only surviving detection, $n = 2.54 \\pm 0.03$, is checked.","marker":"Lyne et al. (2015)"},{"why":"Gives the earlier timing-noise scaling relation whose parameters and log-normal likelihood this work compares with its own estimates.","marker":"Shannon & Cordes (2010)"},{"why":"Supplies the UTMOST-EW data set, earlier $\\ddot{\\nu}$ values, and the $\\sigma_{\\mathrm{RN}}$ figure of merit that the combined-sample analysis builds on.","marker":"Lower et al. (2020)"},{"why":"Defines the hidden Markov model glitch detector, its transition matrices, and the greedy Bayes-factor selection used in both glitch searches.","marker":"Melatos et al. (2020)"},{"why":"Introduced the $\\sigma_{\\mathrm{RN}}$ timing-noise-strength measure used to select the pulsars with significant timing noise.","marker":"Parthasarathy et al. (2019)"}],"fun_headline_variants":["Only the Crab survives a timing-noise audit of pulsar spin-down","Stricter noise model leaves just one pulsar with true spin-down change","Most pulsar 'anomalous braking indices' are timing noise, study finds","Crab pulsar's spin-down change withstands timing-noise scrutiny","Timing-noise test trims 39 pulsar spin-down signals to 1"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that timing noise is a pure power-law random walk with no mean reversion on timescales comparable to the observing span; if the true noise spectrum bends over or reverts to a mean at low frequencies, the predicted spread in $\\ddot{\\nu}$ and the conclusion that nearly all anomalous braking indices are timing-noise artifacts would both be wrong.","fun_headline_variants_meta":{"raw":{"variants":["Only the Crab survives a timing-noise audit of pulsar spin-down","Stricter noise model leaves just one pulsar with true spin-down change","Most pulsar 'anomalous braking indices' are timing noise, study finds","Crab pulsar's spin-down change withstands timing-noise scrutiny","Timing-noise test trims 39 pulsar spin-down signals to 1"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00083,"raw_usage":{"total_tokens":3741,"prompt_tokens":1174,"completion_tokens":2567,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":790,"completion_tokens_details":{"reasoning_tokens":2468}},"tokens_in":790,"tokens_out":2567,"duration_ms":20060,"temperature":1.0,"reasoning_tokens":2468,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T22:00:49.123775+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take any pulsar in this sample with $\\beta \\approx 6$ and a reported non-zero $\\ddot{\\nu}$, extend its timing baseline several-fold, and check whether the posterior uncertainty on $\\ddot{\\nu}$ grows as $T_{\\mathrm{span}}^{-1/2}$ as predicted by equation (20) and whether the recovered noise spectrum remains a single power law below $1/T_{\\mathrm{span}}$. A measured low-frequency turnover or mean reversion in the noise spectrum, or a $\\ddot{\\nu}$ that stays stable and non-zero as the baseline grows, would falsify the noise-dominated interpretation.","supporting_citations":[{"cited_title":"J., Ni t u I","cited_arxiv_id":null,"evidence_quote":"Shows that the standard $1/T_{\\mathrm{span}}$ noise cutoff can mis-estimate $\\ddot{\\nu}$ uncertainties, the effect quantified here as a threefold sensitivity."},{"cited_title":"M., Suvorova S., Moran W., Evans R","cited_arxiv_id":null,"evidence_quote":"Defines the hidden Markov model glitch detector, its transition matrices, and the greedy Bayes-factor selection used in both glitch searches."}],"review_version":1}