{"id":"5771ccce-25e9-4fab-8a43-155a8552bc60","arxiv_id":"2505.06763","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A 45-day campaign compared ten optical clocks across six countries and produced 38 frequency ratios, including four measured directly for the first time, with a full correlation analysis.","lead":"Ten optical clocks in six countries were compared at the same time over 45 days using fiber and satellite links, yielding 38 precise frequency ratios and the first direct measurements of four clock transitions. The results feed the international effort to redefine the second and to build optical time scales.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Satellite-link uncertainties are partly self-calibrated, and Table 2 warns some ratios may have significantly larger uncertainties; the headline GNSS precision is not yet settled.","rationale":"The reader's weakest assumption identifies the GNSS/IPPP uncertainty model and the campaign-calibrated maser models. This is exactly the load-bearing point. The paper's own Table 2 warning and Section 5 discussion of unidentified offsets in INRIM, SYRTE Sr, and PTB Sr make the concern concrete rather than hypothetical. A secondary issue is the post-hoc exclusion of all GNSS ratios involving INRIM; the exclusion is transparently reported, but it means the published '38-ratio subset' is a selected set, and the selection rule is partly based on the very comparisons the paper uses to claim consistency. This does not invalidate the campaign, and the fiber/local ratios are much less affected because their uncertainties are dominated by clock systematics with documented budgets. The empirical FTU itself has external validation up to about 100 days, so the concern is not that the link model is invented; it is that the maser extrapolation and correlation contributions, which are substantial for many GNSS ratios, are estimated using the same dataset. The proposed test would break that circularity by re-anchoring the noise models to independent long-term data. Because the reader already assigned CONDITIONAL and the suggested test is a modest re-analysis rather than a demonstrated failure, the verdict should remain unchanged rather than move to ACCEPT or REJECT. Data availability is a limitation for independent verification, but it is not the decisive technical weakness.","tokens_in":23964,"tokens_out":5504,"duration_ms":59792,"concrete_test":"Recompute the extrapolation uncertainties for all GNSS ratios using maser power spectral densities taken from independent long-term hydrogen-maser-vs-hydrogen-maser comparisons at each institute, rather than from the optical-clock-vs-maser fits used to build Table S2. If any ratio's total uncertainty increases by more than about 30 percent, or if any ratio value shifts by more than one quoted total uncertainty, the stated satellite-link precision is not supported and the affected conclusions should be relaxed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is a dense set of 38 frequency ratios with trustworthy uncertainties. That requires the GNSS uncertainty budget to be reliable, but the budget is not independently anchored. In Section 4 the IPPP link uncertainty is set to the empirical FTU = 1e-15/(T/d), which has external validation, but the maser extrapolation uncertainties that often dominate are computed from noise models (Table S2) fitted partly with the campaign's own optical-clock-vs-maser data; Section 3.4 likewise tunes the IPPP autocorrelation model to reproduce the FTU. More directly, Table 2 contains the authors' caveat that 'it is likely ... some of the frequency ratios have significantly larger uncertainties than the estimates shown here,' naming GNSS ratios with INRIM and ratios involving SYRTE Sr and PTB Sr. Section 5 then identifies a 4e-16 offset in the INRIM GNSS chain, a possible uncontrolled shift in SYRTE Sr, and a possible few-1e-17 offset in PTB Sr. If these effects leak into the quoted statistical uncertainties, the 30 satellite-based ratios, the 'lowest satellite comparison uncertainty' claim, and the reference-offset interpretations would all shift by more than the stated error bars. The issue is not formal circularity, but that the uncertainty model and the measured data are not fully independent, so an underestimate cannot be excluded.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript reports a 45-day coordinated campaign in 2022 in which ten optical clocks at six institutes were compared simultaneously, using a mix of local, international optical-fiber, and GNSS IPPP links. The paper presents 38 optical frequency ratios, of which four are claimed as first direct measurements (Yb+(E3)/Yb, In+/Yb, Sr+/Sr, Sr+/Yb), and it provides a 38x38 correlation analysis. The authors identify several anomalies, most notably an unexplained ~4e-16 offset in all GNSS comparisons involving the INRIM Yb clock, a possible uncontrolled shift in the SYRTE Sr clock, and a possible few-1e-17 offset in the PTB Sr clock. The abstract and conclusions use these results to argue that coordinated multi-lab optical clock networks are feasible and that the data will support the redefinition of the second.","tokens_in":24206,"tokens_out":8075,"duration_ms":72523,"significance":"If the uncertainty estimates are accepted, this is a landmark dataset: it is the densest simultaneous cross-checked comparison of optical clocks to date, it demonstrates the operation of a multi-link network with two independent transfer techniques, and it adds genuinely new direct frequency ratios that will feed into the next CIPM adjustment. The paper's strengths are its internal cross-checks (fiber versus GNSS, same-transition ratios, double GNSS receivers), the explicit correlation analysis, and the unusually candid reporting of anomalies. The main risk is that the uncertainty budget for the satellite-based ratios is not fully independent of the measured data, and the paper itself warns that some quoted uncertainties may be too small. Because the reliability of the quoted error bars is load-bearing for the '38 ratios', the 'lowest satellite comparison uncertainty' claim, and the offsets interpreted in Section 5, the manuscript needs revision before the results can be used with confidence.","major_comments":[{"comment":"The uncertainties of the 26 GNSS-based ratios rest on noise models that are calibrated, at least in part, with the campaign's own data. The hydrogen-maser noise models in Table S2 are stated to be 'estimated using optical clock vs HM data from the campaign and/or prior information' (Section 4), and the IPPP phase-noise model in Eq. (5) has its amplitude k^-1 adjusted 'to make the FTU agree with the conservative estimate' and its low cutoff f_l = 1/(30 d) chosen to reproduce the autocorrelation from the same IPPP-fiber comparison (Supplementary Sec. 3.4). This makes the uncertainty budget not fully independent of the measured ratios, so an unmodeled common excess noise during this particular campaign would be invisible and would propagate into all GNSS ratios. Please add a sensitivity analysis (e.g., doubling the maser noise coefficients and the IPPP amplitude, or using independent long-term maser characterizations) and report how the claimed uncertainties and the 'lowest satellite comparison uncertainty' statement change.","section":"Section 4 and Supplementary Material Sec. 3.4"},{"comment":"The paper declares 'we consider the results of all the frequency ratios via GNSS to INRIM to be unreliable' but still lists these seven ratios (rows 14, 16, 24, 25, 27, 28 and 29) in Table 2 with quantitative uncertainties and includes them in Fig. 2 and in the '38 frequency ratios' claims of the abstract and Section 7. This ambiguity is load-bearing because a reader combining these ratios with future data or with the CIPM adjustment could use uncertainties that the authors themselves state are wrong by about 4e-16. Please either mark these rows explicitly as unreliable (and exclude them from the headline count of usable ratios), or add a systematic uncertainty term that accounts for the observed offset.","section":"Section 5.1 and Table 2"},{"comment":"The paper reports a fractional frequency difference of 1.46(21)e-16 between PTB Sr and SYRTE Sr on the fiber link and states that SYRTE Sr had an uncontrolled shift at the 1e-16 level and PTB Sr possibly a few 1e-17, with Birge ratios of 3.3-5.3 for ratios involving SYRTE Sr (Supplementary Fig. S1). Nevertheless, Table 2 does not add any correlated systematic uncertainty for these clocks, so the 'agreement within 1-2 sigma' statements for the Sr/Sr and Yb+(E3)/Sr comparisons are likely understated. The Birge-ratio inflation acts on each ratio independently and cannot capture a common-mode clock shift that affects all ratios sharing a clock. Please either assign an additional, explicitly correlated uncertainty to the suspect clocks or clearly reclassify the affected ratios as consistency checks rather than precision results.","section":"Section 5.2 and the Birge-ratio treatment in Section 3"}],"minor_comments":[{"comment":"The caveat that 'some of the frequency ratios have significantly larger uncertainties than the estimates shown here' should be made actionable: list the affected row numbers and state explicitly whether users should treat those ratios as upper limits, as invalid, or as needing an extra uncertainty.","section":"Table 2 caption"},{"comment":"The ratio ID numbers and the 'inv' markers are difficult to read at publication size; please enlarge the labels or split the figure so that the GNSS and fiber/local panels are legible.","section":"Fig. 2"},{"comment":"For a paper whose central contribution is a dataset, 'Data underlying the results ... may be obtained from the authors upon reasonable request' limits reproducibility; please consider releasing at least the daily-binned ratios and the correlation/covariance matrix.","section":"Data Availability Statement"},{"comment":"The sentence 'We have demonstrated agreement between GNSS and optical fiber links over a continental scale' should be qualified, since Section 5.1 removes the INRIM GNSS data from that conclusion.","section":"Section 7, first paragraph"}],"recommendation":"major_revision","confidential_remarks":"This is a strong metrology paper with a unique dataset, and the anomalies are reported honestly. My recommendation is driven by the need to make the uncertainty budget and the status of the suspect ratios unambiguous before the 38-ratio dataset is used by the community. I do not see a circularity problem in the frequency values themselves, but the self-calibrated link/maser noise models need a sensitivity test."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is the real thing—ten clocks, six countries, 38 ratios, four first direct measurements, and the first full correlation matrix for a campaign like this. The paper deserves a serious referee, but the GNSS uncertainties are not fully independent of the data they describe, so treat the headline satellite numbers as provisional.\n\nWhat is new: the scale and the cross-checking. Five ratios were measured by both fiber and GNSS, same-transition ratios give null checks, and dual receivers per institute caught phase problems. The fiber-based results, including the first In+/Yb, Sr+/Sr, and Sr+/Yb direct measurements, are solid and directly useful. The correlation matrix is a real methodological contribution, especially the generalization of the Fourier-transform extrapolation covariance to shared masers and IPPP link noise.\n\nThe soft spots are where the reader's report says. The IPPP FTU is an empirical conservative estimate, and the maser noise models in Table S2 are partly fitted using the campaign's own clock-vs-maser data. Section 3.4 tunes the IPPP phase PSD to reproduce the FTU. That is not formal circularity—these are standard practices—but it means part of the satellite-link uncertainty budget is self-calibrated. Table 2 itself warns that some ratios may have significantly larger uncertainties than shown, and Section 5 then identifies a 4e-16 offset in the INRIM GNSS chain, a likely uncontrolled shift in SYRTE Sr, and a possible few-1e-17 offset in PTB Sr. The authors handle this honestly: they exclude the INRIM GNSS ratios, note the Sr issues, and refrain from applying corrections. But it does mean the 30 satellite-based ratios should not be taken at face value until those effects are understood. Data are not public, which limits independent verification.\n\nThat said, the central claim—that this is the largest coordinated optical clock comparison and provides the densest cross-checked dataset—holds up. The paper is transparent about its caveats, and the fiber/local subset alone is a significant contribution.\n\nWho it is for: the CIPM least-squares adjustment community, optical clock metrologists, and anyone building optical clock networks. I would send it to a careful referee, expecting that referee to push for a fuller treatment of the maser and IPPP model uncertainties and ideally for public data. The paper deserves peer review, not desk rejection.","headline":"Largest multi-lab optical clock comparison to date, with a genuinely useful dataset and correlation analysis, but the satellite-link uncertainty budget is partly self-calibrated and needs scrutiny before the numbers feed the least-squares adjustment.","tokens_in":25335,"tokens_out":1762,"would_cite":true,"duration_ms":17668,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper reports the largest coordinated comparison of optical clocks to date, simultaneously comparing ten clocks in six countries over fiber and satellite links and presenting 38 frequency ratios, including four measured directly for…","keywords":["optical clocks","frequency ratios","optical clock comparisons","integer precise point positioning","satellite frequency transfer","optical fiber links","correlation analysis","redefinition of the second"],"falsifier":"A direct comparison of the same two clocks over both a satellite link and a fiber link for at least 30 days would settle whether the transfer-uncertainty model holds: if the satellite-versus-fiber disagreement exceeds the combined uncertainties consistently, the model is optimistic. For the specific $4\\times10^{-16}$ offset seen with the Italian Yb clock, feeding the same clock signal to two independent receivers through separate distribution chains during a GNSS-versus-fiber comparison would confirm whether the offset originates in the signal distribution.","tokens_in":23711,"feed_emoji":"🕰️","tokens_out":7786,"duration_ms":73364,"temperature":0.7,"pith_summary":"The paper reports the largest coordinated international comparison of optical clocks carried out to date: ten clocks in six countries ran simultaneously for 45 days, and frequency ratios between them were measured over phase-stabilized fiber and satellite (GNSS-IPPP) links. The authors present 38 frequency ratios, including the first direct measurements of the Yb+(E3)/Yb, In+/Yb, Sr+/Sr, and Sr+/Yb ratios, and compute the correlations among all of them. The campaign is designed to test whether optical clocks from different laboratories agree within their claimed uncertainties, to expose hidden systematic errors, and to build the cross-checked dataset needed before the SI second can be redefined on an optical transition. The redefinition of the second, targeted for the 2030s, requires such validation of clock uncertainty budgets, so this campaign is a concrete step toward making optical clocks the basis of international timekeeping.","feed_headline":"38 optical clock ratios measured across six countries at once","feed_subtitle":"First direct measurements of four ratios sharpen the case for redefining the second.","key_machinery":"The central object is the optical frequency ratio between two clock transitions, with a total fractional uncertainty made of clock systematic uncertainties ($u_B$), relativistic redshift uncertainty ($u_{RRS}$), link statistics, and maser extrapolation. The comparison machinery is a star-shaped phase-stabilized fiber network, whose links contribute under $10^{-18}$ after 1000 s, plus satellite links using Integer Precise Point Positioning (IPPP), with hydrogen masers acting as flywheels so that clock data gaps can be bridged. The load-bearing model is the IPPP frequency transfer uncertainty $FTU = 1\\times10^{-15}/(T/\\mathrm{d})$, and the new analytical tool is the covariance calculation: the Fourier-transform method for extrapolation uncertainty is generalized to compute the covariance between two ratios that share a common maser, yielding correlation coefficients up to 0.94 for fiber ratios and 0.80 for GNSS ratios.","core_discovery":"On its own terms, the paper establishes that a coordinated multi-clock, multi-link campaign can produce a dense, redundant set of optical frequency ratios whose cross-checks verify uncertainty budgets and expose inconsistencies. The 38 ratios include several GNSS links with total uncertainties below $1.8\\times10^{-16}$, improving on the best previous satellite comparison, and the first direct measurements of Yb+(E3)/Yb, In+/Yb, Sr+/Sr, and Sr+/Yb. The redundancy revealed that all satellite-based ratios involving the Italian Yb clock are offset by about $4\\times10^{-16}$ from fiber-based results, likely due to an unidentified problem in its signal distribution; the French Sr clock shows excess scatter and offsets near $1\\times10^{-16}$; and the German Sr clock may have been a few $10^{-17}$ low. Because several discrepancies are ambiguous—a clock may be wrong, or the reference value derived from previous data may be wrong—the paper presents the results with a full $38\\times38$ correlation matrix so future adjustments can use them correctly.","pith_inferences":["If the FTU model continues to hold, the same satellite-plus-maser procedure could make routine intercontinental comparisons of optical clocks at the $10^{-16}$ level, and the covariance formalism could be extended to networks that mix fiber and future free-space optical links.","A testable extension would be to operate the Italian Yb clock with two independent radio-frequency distribution paths during a simultaneous fiber-and-GNSS comparison; if the $4\\times10^{-16}$ offset reproduces only on the satellite side, the problem is in the distribution chain rather than the clock itself.","The correlation machinery for common-maser extrapolation could be reused for optical time scales, where a single flywheel oscillator bridges gaps in real time; the same covariance formulas would quantify how much of the time-scale noise is shared between successive clock comparisons."],"forward_implications":["The first direct measurements of the Yb+(E3)/Yb, In+/Yb, Sr+/Sr, and Sr+/Yb ratios will feed into the next least-squares adjustment of recommended optical frequencies, with the fiber-based Yb+(E3)/Yb value (about $5\\times10^{-17}$ uncertainty) expected to pull the optimized value significantly.","Several satellite-based ratios with total uncertainties below $1.8\\times10^{-16}$ show that IPPP with maser flywheels can audit optical clocks across continents at a level previously reached only by fiber or local comparisons.","The full correlation matrix, with 242 non-zero coefficients, allows all 38 ratios to be combined in future multivariate adjustments without double-counting shared clocks, links, and masers.","The campaign's identified inconsistencies—the Italian Yb satellite offset, the French Sr scatter, and the possible German Sr offset—demonstrate that redundant multi-clock, multi-link measurements can uncover problems that pairwise comparisons would leave hidden."],"supporting_citations":[{"why":"Supplies the Integer Precise Point Positioning method used for all satellite-based clock comparisons.","marker":"[26]"},{"why":"Provides the empirical basis for the frequency transfer uncertainty model $1\\times10^{-15}/(T/\\mathrm{d})$ used for the GNSS links.","marker":"[43–45]"},{"why":"Supplies the Fourier-transform method for maser extrapolation uncertainty, generalized here to covariances between ratios sharing a maser.","marker":"[42]"},{"why":"Provides the universal data-processing formalism for combining clock and link data onto a common grid.","marker":"[40]"},{"why":"Provides the 2021 recommended frequency values and reference frequency ratios used as the comparison baseline.","marker":"[20]"},{"why":"Establishes the performance of the phase-stabilized fiber network used for the continental comparisons.","marker":"[22–25]"},{"why":"Provides the guidelines and formulas used to compute correlation coefficients from shared systematic and statistical uncertainties.","marker":"[56]"}],"fun_headline_variants":["Global clock network yields 38 frequency ratios","Six countries, ten clocks, one coordinated comparison","Optical clock ratios sharpen case for new second","First direct ratios for four clock pairs","Multi-nation clock test uncovers offset mysteries"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The satellite-based results rest on the assumption that the noise added by the satellite link and by the auxiliary hydrogen clocks used to fill gaps in the optical-clock data is no larger than the model used to compute the uncertainties; if the real noise is larger, the reported uncertainties on the thirty satellite-based ratios are too small and the comparisons built on them shift.","fun_headline_variants_meta":{"raw":{"variants":["Global clock network yields 38 frequency ratios","Six countries, ten clocks, one coordinated comparison","Optical clock ratios sharpen case for new second","First direct ratios for four clock pairs","Multi-nation clock test uncovers offset mysteries"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000153,"raw_usage":{"total_tokens":1162,"prompt_tokens":852,"completion_tokens":310,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":468,"completion_tokens_details":{"reasoning_tokens":242}},"tokens_in":468,"tokens_out":310,"duration_ms":3203,"temperature":1.0,"reasoning_tokens":242,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T22:34:31.480497+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A direct comparison of the same two clocks over both a satellite link and a fiber link for at least 30 days would settle whether the transfer-uncertainty model holds: if the satellite-versus-fiber disagreement exceeds the combined uncertainties consistently, the model is optimistic. For the specific $4\\times10^{-16}$ offset seen with the Italian Yb clock, feeding the same clock signal to two independent receivers through separate distribution chains during a GNSS-versus-fiber comparison would confirm whether the offset originates in the signal distribution.","supporting_citations":[{"cited_title":"Petit, A","cited_arxiv_id":null,"evidence_quote":"Supplies the Integer Precise Point Positioning method used for all satellite-based clock comparisons."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the Fourier-transform method for maser extrapolation uncertainty, generalized here to covariances between ratios sharing a maser."},{"cited_title":"Lodewyck, R","cited_arxiv_id":null,"evidence_quote":"Provides the universal data-processing formalism for combining clock and link data onto a common grid."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the 2021 recommended frequency values and reference frequency ratios used as the comparison baseline."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the guidelines and formulas used to compute correlation coefficients from shared systematic and statistical uncertainties."}],"review_version":1}