{"id":"40d6cb53-9610-4b42-b507-9c81b85703ec","arxiv_id":"2608.00910","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":9,"one_line_summary":"An all-optical time scale built from cryogenic silicon cavities and an iodine clock, steered by a Sr optical clock, maintains sub-1e-16 relative instability and <100 ps total difference over 20 days.","lead":"Three light-based 'optical flywheel' clocks, steered by a strontium atomic clock, ran together continuously for over 20 days and kept mutual time differences under 100 picoseconds. This suggests future time scales can be built entirely from optical systems, replacing the microwave masers that currently limit UTC.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Simulation-based superiority over AT1 is fitted from the same measured data; the pairwise comparison cannot validate individual time-scale stability.","rationale":"The reader's weakest_assumption focuses on the steering extrapolation during Sr gaps. That concern is not the most load-bearing for the central measured claim, because the accumulated time differences during gaps are directly measured rather than inferred from the drift model; if the drift changed on short timescales, the measured distribution would already show it. The stronger issue is the simulation circularity: the paper uses the same measured data both to fit the noise model and to validate the simulated individual time-scale stability versus AT1. This is explicitly described in the Methods and in Fig. 4(c), and it is the main correctness risk for the paper's broader significance claim. The reader did identify this issue in the rationale, so agreement is partial. The measured pairwise relative instability and final <100 ps time difference are well supported by the phase-continuous comparison, so the verdict should remain CONDITIONAL as the reader proposed; no change to the verdict is needed, but the condition should emphasize the need for out-of-sample or independent validation of the AT1 comparison.","tokens_in":9135,"tokens_out":17022,"duration_ms":202179,"concrete_test":"Split the 27-day record into a fitting window (e.g., first 17 days) and a holdout window (last 10 days). Fit the flywheel noise models using only the fitting window, then simulate the holdout window with the actual Sr uptime pattern and steering algorithms. Compare the simulated pairwise time differences and overlapping Allan deviations to the measured holdout data; if the out-of-sample simulation falls outside the measured confidence intervals, the model is overfit and the AT1 comparison is unsupported. Alternatively, directly compare one all-optical time scale (e.g., TS(Si3)) to an independent optical clock or a calibrated UTC(NIST) transfer link over 20 days and check whether the measured individual instability actually reaches <1e-16 after a few days.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The directly measured 20-day comparison (Fig. 4a,b) is internally consistent and supports sub-1e-16 relative instability and a final <100 ps time difference. However, the paper's broader claim that all three all-optical time scales outperform NIST AT1 and are viable replacements for maser-based time scales is not directly measured. The Methods (Extended Data Fig. 2) state that the noise parameters for Si3, Si6, and VA21 were estimated from the drift-removed overlapping Allan deviations of the same flywheels (Fig. 2c). These fitted parameters are then used to simulate individual time-scale stability versus AT1 (Fig. 4c). The 'validation' of the simulation by reproducing the measured relative time differences (Fig. 4b) is a circular consistency check: a model fitted to the same noise data is expected to reproduce it. Moreover, because all three time scales share the same Sr steering reference, pairwise comparisons cancel common-mode Sr noise and do not constrain individual time-scale stability against an external reference. Thus the AT1-superiority claim rests on two unverified assumptions: the fitted noise models are representative out-of-sample, and the shared Sr reference contributes negligible common-mode error. This does not invalidate the central measured comparison, but it is load-bearing for the paper's forward-looking thesis and for the quantitative claim in Fig. 4(c).","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a 27-day phase-continuous comparison of three all-optical time scales, each generated by steering an optical flywheel (two cryogenic silicon cavities and one iodine clock) with a shared Sr optical lattice clock. The authors measure pairwise time differences and overlapping Allan deviations, reporting relative instability below 1e-16 after a few days of averaging and a total time difference below 100 ps over 20 days. They also verify the fidelity of optical-to-rf conversion with a residual slope consistent with zero. A simulation based on measured flywheel noise characteristics is used to project individual time-scale stability and to compare with the NIST AT1 maser ensemble.","tokens_in":9476,"tokens_out":7305,"duration_ms":77066,"significance":"The central measurement is a significant technical milestone. It demonstrates that optical flywheels can sustain continuously operating time scales with sub-nanosecond accumulated errors over 20 days, and that pairwise relative stability reaches the low-1e-16 regime. The direct, phase-continuous comparison is a strong result, and the explicit correction of comb cycle slips and the verification of optical-to-rf conversion fidelity add credibility. However, the broader claim of superiority over AT1 rests on a simulation whose noise parameters are fitted from the same data, and the common Sr reference means the pairwise measurements do not constrain individual time-scale stability. The paper would be strengthened by out-of-sample validation or a more careful qualification of the simulated comparison.","major_comments":[{"comment":"The noise parameters for Si3, Si6, and VA21 are estimated from the drift-removed overlapping Allan deviations of the same flywheels (Fig. 2c, Extended Data Fig. 2). The agreement between simulation and measurement in Fig. 4(b) is therefore a consistency check, not an out-of-sample validation. The claim that all three time scales outperform AT1 (Fig. 4c) rests on the unverified assumption that these noise models are representative outside the fitted data. Please provide an out-of-sample validation (e.g., fit on a subset of the 7-month record and test on the 20-day campaign) or explicitly reframe Fig. 4(c) as a model-based projection rather than a measurement.","section":"Methods (Simulation of all-optical time scales); Fig. 4(c)"},{"comment":"Because all three time scales are steered by the same Sr optical frequency standard, the pairwise comparisons in Fig. 4(a,b) cannot detect common-mode errors in the Sr reference. The individual time-scale stability curves in Fig. 4(c) rely on the simulation's implicit assumption that the Sr reference contributes negligibly, which is not tested by this experiment. This limitation should be stated explicitly in the main text, and the statement 'All three time scales perform better than NIST AT1' should be qualified accordingly.","section":"Main text, Fig. 4"}],"minor_comments":[{"comment":"The concern that the reported ~20 ps per 6-hour gap is an underestimate because of drift extrapolation does not land: the time differences in Fig. 4(a) are directly measured, so they already reflect the actual drift behavior. However, the paper should distinguish this empirical average from a worst-case bound; a brief discussion of the distribution of gap errors would be helpful.","section":"Steering description (main text)"},{"comment":"The text mentions a gap in the rf comparison during MJD 61055-61060 due to a 'numerical precision issue' (Ref. 37). Please clarify what is meant by this issue and confirm that the optical comparison remains continuous for the full 20 days.","section":"Fig. 5(a)"},{"comment":"Minor grammar: 'superior short-term stability than hydrogen masers' should be 'superior short-term stability to hydrogen masers' or 'better than'.","section":"Abstract / Introduction"}],"recommendation":"major_revision","confidential_remarks":"This is a strong experimental paper from a leading group. The direct measurement is impressive and will likely be of great interest. The main issue is the circularity of the AT1 comparison; the authors should be pushed to either validate the simulation out-of-sample or downgrade the claim. The manuscript may also benefit from a more explicit discussion of the common-mode Sr reference limitation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here is the take: this is the first multi-flywheel all-optical time scale comparison running phase-continuously for 20+ days. Two cryogenic silicon cavities and a commercial iodine clock, all steered by one Sr lattice, reach sub-1e-16 relative instability after a few days and accumulate <100 ps mutual time difference. The central measurement is direct, internally consistent, and looks solid. The optical-to-rf conversion fidelity check with a residual slope consistent with zero is a nice, honest addition, and the cycle-slip tracking in post-processing is the kind of detail that makes the result believable.\n\nWhat's new: Milner et al. demonstrated a single optical-carrier timescale; this paper scales it to three parallel optical flywheels, includes a turnkey iodine clock, and runs the comparison long enough to make the stability statement interesting. The 20-day phase-continuous record, with the gap-length distribution and time-difference correlation, is real evidence.\n\nThe soft spots are real but contained. First, the claim that each of the three time scales individually beats NIST AT1 is not measured; it comes from a simulation whose noise parameters were estimated from the same overlapping Allan deviations that the simulation then reproduces (Extended Data Fig. 2 and Methods). The agreement between simulated and measured relative time differences is a consistency check, not an independent validation. That overstates the strength of the 'outperforms AT1' claim. Second, all three time scales share the same Sr steering reference, so pairwise comparisons cancel common-mode Sr errors. The quoted <10^-16 is a relative statement; absolute accuracy of any single all-optical time scale is not demonstrated. The paper is mostly careful about saying 'when compared with each other,' but the forward-looking part leans on the unvalidated simulation.\n\nOne more minor concern: the Si3 extrapolation uses the prior 10-hour average drift. If the drift changes on shorter timescales, the per-gap prediction error grows. In this data set that behavior would show up in the measured time differences, so the <100 ps final value is not an underestimate for this run; it just doesn't generalize automatically to other operating conditions.\n\nWho this is for: time/frequency metrology people, especially those thinking about optical redefinition of the second and UTC realization. It deserves a serious referee; the measured comparison is a milestone even if the simulation claims need to be tempered or independently verified.","headline":"A strong direct 20-day comparison of three optical-flywheel time scales; the headline numbers are measured and credible, but the simulation-based claim to beat AT1 is fitted from the same data.","tokens_in":10036,"tokens_out":2562,"would_cite":true,"duration_ms":27947,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Three all-optical time scales, steered by a strontium clock, kept <1e-16 instability and <100 ps total difference over 20+ days.","keywords":["all-optical time scale","optical flywheel","cryogenic silicon cavity","iodine optical clock","strontium optical lattice clock","hydrogen maser","Dick effect","optical frequency comb"],"falsifier":"Run the same three-flywheel comparison with an additional continuously operating optical frequency standard that is independent of the Sr clock; if the absolute time error of any all-optical time scale exceeds the pairwise differences reported here, the pairwise comparison is not representative of true timekeeping error.","tokens_in":9036,"feed_emoji":"⏱️","tokens_out":4596,"duration_ms":51698,"temperature":0.7,"pith_summary":"The paper claims that a time scale can be kept entirely in the optical domain by steering optical flywheels with a strontium optical lattice clock, removing the hydrogen-maser bottleneck that currently limits timekeeping. Three such time scales, based on two cryogenic silicon cavities and one commercial iodine clock, ran continuously for over 20 days and were compared phase-continuously. They reached relative fractional instability below 1e-16 after a few days of averaging and a total time difference under 100 ps. This matters because current time scales need weeks of averaging to reach performance that optical flywheels reach in days, and the result points toward a practical path for realizing the future optical second.","feed_headline":"Optical time scales hit sub-1e-16 stability in days","feed_subtitle":"Cryogenic silicon cavities and an iodine clock, steered by a strontium standard, outperform hydrogen-maser time scales.","key_machinery":"The optical flywheel: a continuously running laser oscillator whose short-term stability is orders of magnitude better than a hydrogen maser, so that a time scale can bridge gaps in optical-clock operation without a large Dick noise penalty. This paper uses two cryogenic silicon cavities (Si3 and Si6) and an iodine optical clock (VA21), each steered by a Sr optical lattice clock. During Sr-off gaps, steering uses linear drift extrapolation for Si3 (slope from the past 10 hours), a zero-drift assumption for Si6, and a Kalman filter for VA21. Optical frequency combs convert each flywheel to the rf domain for phase-continuous comparison.","core_discovery":"The central claim is that replacing hydrogen masers with optical flywheel oscillators removes the main bottleneck in time scale performance. The paper realizes three independent all-optical time scales—two from cryogenic silicon cavities and one from a commercial iodine optical clock—each steered by a high-uptime 87Sr optical lattice clock, and compares them phase-continuously for 27 days. The measured relative fractional instability reaches below 1e-16 after a few days of averaging, and the peak-to-peak time difference between pairs stays around 200 ps, ending below 100 ps. The authors also verify that conversion of the optical signals to 100 MHz rf via optical frequency combs adds less tha","pith_inferences":["If flywheel drift during gaps is the limiting error, then improving drift prediction (e.g., integrating cavity temperature or using a second flywheel as a drift reference) should reduce the roughly 20 ps per 6-hour gap in proportion to prediction error, not to gap length alone.","Because the reported comparison is pairwise, the <100 ps total difference is a strict bound on each time scale's error only if the three flywheels' noises are independent; correlated errors such as common Sr steering effects would be invisible.","A direct test of the gap-error model would be to add a fourth, continuously operating optical clock and measure each all-optical time scale's absolute time error through a 10+ hour gap; the Si3 prediction error should equal the extrapolation error of its 10-hour average drift.","The approach transfers naturally to optical clock networks: pairing each optical clock with an optical flywheel and comparing over fiber links would let the network realize a distributed all-optical time scale without masers."],"forward_implications":["Maser-free time scales are realistic now: with roughly 70% Sr uptime, three independent all-optical time scales reached sub-1e-16 instability in days rather than weeks.","A cryogenic silicon cavity steered only one hour per day would still outperform a hydrogen maser steered twelve hours a day, according to the paper's simulation.","All-optical time scales can be downconverted to conventional 100 MHz rf signals with less than 2 ps added timing error, so existing rf infrastructure can distribute the benefit.","The same steering algorithms and flywheel hardware can be adopted by national timing laboratories to improve UTC and TAI once optical frequency standards become more continuous.","Commercially available iodine clocks and fiber lasers make the flywheel stack realistic outside specialized laboratories."],"supporting_citations":[{"why":"Prior demonstration of an optical-clock-steered maser time scale, the baseline that this work's performance is compared against.","marker":"[8]"},{"why":"Supplies the cryogenic silicon cavity frequency stability (2.5e-17) that qualifies Si3 as an optical flywheel.","marker":"[9]"},{"why":"Earlier demonstration of a timescale based on a stable optical carrier, which this work extends to three parallel optical flywheels.","marker":"[10]"},{"why":"Shows the commercial iodine optical clock operating reliably at sea, supporting VA21's use as an optical flywheel.","marker":"[22]"},{"why":"Provides the Sr optical lattice clock and silicon cavity integration that anchor the steering measurements.","marker":"[30]"},{"why":"Supplies the cryogenic silicon cavity with semiconductor crystalline coatings used for the Si6 flywheel.","marker":"[31]"},{"why":"Establishes the Sr clock's 8e-19 systematic uncertainty, underpinning the accuracy of the steering reference.","marker":"[32]"},{"why":"Demonstrates the Sr clock's extreme stability, which sets the reference quality for flywheel comparison.","marker":"[33]"}],"fun_headline_variants":["All-optical time scales beat masers to 1e-16 in days","Cryogenic cavities + iodine clock: 1e-16 time scales","Optical flywheels replace masers for ultra-stable time","20-day optical time runs with sub-1e-16 instability","Three optical time scales stay sub-1e-16 for days"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"During gaps when the Sr clock is off, the steering algorithm assumes the silicon cavities' frequency drift stays equal to its average over the past 10 hours (Si3) or zero (Si6); if the true drift changes on shorter timescales, the predicted frequency is wrong and the reported roughly 20 ps per 6-hour gap (hence <100 ps total) is an underestimate.","fun_headline_variants_meta":{"raw":{"variants":["All-optical time scales beat masers to 1e-16 in days","Cryogenic cavities + iodine clock: 1e-16 time scales","Optical flywheels replace masers for ultra-stable time","20-day optical time runs with sub-1e-16 instability","Three optical time scales stay sub-1e-16 for days"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000988,"raw_usage":{"total_tokens":4055,"prompt_tokens":799,"completion_tokens":3256,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":543,"completion_tokens_details":{"reasoning_tokens":3159}},"tokens_in":543,"tokens_out":3256,"duration_ms":26425,"temperature":1.0,"reasoning_tokens":3159,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T00:39:43.550713+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same three-flywheel comparison with an additional continuously operating optical frequency standard that is independent of the Sr clock; if the absolute time error of any all-optical time scale exceeds the pairwise differences reported here, the pairwise comparison is not representative of true timekeeping error.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows the commercial iodine optical clock operating reliably at sea, supporting VA21's use as an optical flywheel."}],"review_version":1}