{"id":"cf73f139-7fd0-41d1-a532-2adf56a99686","arxiv_id":"2411.12706","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":8,"one_line_summary":"A single CME observed by four radially spaced spacecraft is used to show that ensemble hindcast solutions diverge with heliocentric distance, suggesting a maximum useful separation for inner-probe space weather constraints.","lead":"Four spacecraft, roughly lined up between 0.4 and 1 au from the Sun, all measured the same slow coronal mass ejection in September 2021. The authors used an ensemble model to hindcast the CME at each probe and found that the model's predictions spread more widely as distance from the Sun grows, hinting at a limit for using inner probes to predict impacts at Earth.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The inferred inner-to-outer probe separation threshold rests on OSPREI's uniform solar wind; the wrong B_R sign and the required mirroring show the modelled trajectory missed the real CME, so the divergence result is not yet a general CME property.","rationale":"The reader's weakest assumption is exactly the load-bearing concern: OSPREI's uniform, constant solar wind means no interplanetary deflections or rotations, so the ensemble spread is a pure input-parameter propagation result. This matters because the paper's central modelling claim is about when inner-probe constraints lose predictive power at 1 au—a statement about real CMEs, not about OSPREI under a specially simplified background. The internal evidence strengthens the concern: the model predicts the wrong sign of B_R at every spacecraft and only matches observations after all crossings are mirrored to the opposite side of the CME nose. That shows the modelled trajectory is systematically wrong in a way that can affect which ensemble members are ranked as 'best' and how their profiles diverge with distance. The paper is honest about the uniform-wind limitation and hedges its threshold statement, but an acknowledged limitation is still a limitation. I therefore do not change the reader's CONDITIONAL verdict; the observational four-probe identification is solid, while the generalisation to a separation threshold requires the structured-wind test proposed above.","tokens_in":30916,"tokens_out":5114,"duration_ms":56567,"concrete_test":"Re-run the same 200-member Table 4 ensemble with a structured solar wind background for 23 September 2021, for example by coupling the OSPREI perturbation ranges into EUHFORIA with an embedded flux-rope CME or an equivalent heliospheric MHD setup, and compare the spread and B_R sign at PSP and STEREO-A with Figure 9. If the divergence with distance disappears, or B_R still has the wrong sign without the mirroring fix, the inferred separation threshold is an artifact of the uniform-wind model rather than a robust property of CME encounters.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central modelling conclusion—that best-fit solutions diverge with heliocentric distance and that there is a maximum useful inner-to-outer probe separation—is obtained under the Section 4.1/5.2 assumption of a constant, uniform solar wind background. Because OSPREI forbids interplanetary deflections and rotations by construction, the growing ensemble spread is purely the propagation of input-parameter uncertainty through a homogeneous expansion. The paper's own Figure 11 undermines the assumption that this spread is a faithful proxy for real CME orientation uncertainty: the seed run predicted the wrong sign of B_R at all four spacecraft, and agreement was only recovered by artificially mirroring every crossing to the opposite side of the CME nose. That is direct evidence that the model's heliospheric trajectory, and therefore the ranking of ensemble members used to define 'best-fit' divergence, is systematically off. A threshold description of when inner probes cease to constrain 1 au field orientation is thus conditional on the very physics (structured-wind deflection and rotation) that the model omits.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper analyses the 23 September 2021 slow streamer-blowout CME that was observed in situ by BepiColombo, Solar Orbiter, Parker Solar Probe, and STEREO-A at heliocentric distances between 0.4 and 1 au. The authors identify shocks, sheaths, ejecta and core flux ropes at each spacecraft, perform EFF flux rope fits, and run a 200-member OSPREI ensemble in hindcast mode. They report that the spread in predicted in-situ quantities grows with heliocentric distance, and interpret this as evidence for a maximum angular/radial separation beyond which inner-probe constraints on the magnetic field orientation lose power. They also claim priority for the first four-probe radially-separated encounter inside 1 au.","tokens_in":31252,"tokens_out":6447,"duration_ms":56741,"significance":"The manuscript is a valuable contribution in the sense of documenting a rare multi-spacecraft CME encounter inside 1 au with consistent in-situ signatures across four probes. The OSPREI ensemble setup is open and reproducible, the WSA-Enlil comparison provides an independent arrival-time check, and the authors are transparent about many limitations. However, the central interpretive claim about the distance-dependent divergence of magnetic field predictions is not yet established as a general result: it is conditional on the uniform-wind simplification in OSPREI and on a trajectory that appears systematically offset given the wrong B_R sign. If the authors can either temper the claim or demonstrate robustness to structured solar wind, this would become a solid reference case for multi-probe CME modelling.","major_comments":[{"comment":"The central conclusion—that the ensemble spread in predicted in-situ quantities increases with heliocentric distance and that there is a maximum angular/radial separation beyond which inner-probe constraints on magnetic field orientation lose power—is derived from OSPREI runs that assume a constant, uniform solar wind background, as stated in Section 4.1 and Section 5.2. This assumption excludes interplanetary deflections and rotations by construction, so the growing spread reflects only the propagation of input-parameter uncertainties through a homogeneous expansion. The abstract and Section 6 present this result without the uniform-wind caveat; I recommend either softening the claim to a model-dependent result or adding a test with a structured background to show it is not an artifact of the missing physics.","section":"Abstract, Sections 4.1 and 5.2"},{"comment":"The seed run predicts the wrong sign of B_R at all four spacecraft, and the only way agreement is reached is by artificially mirroring each spacecraft crossing to the opposite side of the CME nose (Figure 11). This is direct evidence that the modelled heliospheric trajectory—and hence the CME nose latitude used to define the encounter geometry—is systematically incorrect. Because the 'best-fit' ensemble members are ranked by comparison with the observed profiles, the divergence of best-fit solutions with distance (Figure 9) may be an artifact of this geometric offset rather than a robust property of the CME. The paper should quantify how the mirroring changes the best-fit ranking and the divergence trend, or at minimum present the divergence result as conditional on the assumed trajectory.","section":"Section 5.2, Figure 11"},{"comment":"The ensemble 'best-fit' solutions are selected by comparing synthetic profiles to the same in-situ data used to set the seed parameters and the ensemble ranges (Table 4). This is a hindcast, as the paper states, but the abstract's phrase 'spread in the predicted quantities' and the discussion of using inner-probe observations 'to constrain predictions' could be misread as an out-of-sample forecast result. The divergence of the best-fit members is a measure of model sensitivity within a hindcast setup, not of predictive skill. Please rephrase the abstract and Section 5.2 to make this distinction explicit.","section":"Section 4.2"}],"minor_comments":[{"comment":"The shock parameters in Table 2 are reported without uncertainties, even though the analysis uses averaging windows of 1 to 8 minutes; please provide uncertainty ranges or state that the variations are negligible.","section":"Table 2"},{"comment":"The EFF flux rope fit parameters in Table 3 include goodness-of-fit values but no parameter uncertainties; given the SolO trailing-edge data gap and the STEREO-A double-peak profile, a discussion of fit parameter confidence would strengthen the comparison.","section":"Table 3"},{"comment":"The SolO ejecta trailing edge is defined only by a data gap, and the flux rope fit is truncated at that boundary; the paper notes this, but it should explicitly state how a different choice of the trailing boundary would affect the fitted axis orientation and the multi-spacecraft comparison in Figure 10.","section":"Section 3.2"},{"comment":"The claim to be the 'first report of an event being observed in situ by four well-radially-separated probes inside 1 au' needs a supporting citation or a search statement; as written, it is a strong priority claim that is not documented.","section":"Section 6"},{"comment":"There are a number of minor language issues, for example 'in in Figure 4(b)' in Section 3.2 and 'different than' in Section 3.2; a careful proofread would resolve these.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The paper is well within the scope of the journal, and the event itself is significant. However, I am concerned that the headline claim (the divergence threshold) is presented more strongly than the model's limitations support. Please ask the authors to address the uniform-wind caveat and the B_R sign issue before acceptance. The paper might benefit from a more explicit statement that the result is a model-based inference, not a measured property of the CME."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear [name],\n\nThe thing worth knowing about this paper is that it documents a genuinely rare event: a slow streamer-blowout CME from 23 September 2021 hit four spacecraft between 0.4 and 1 au, and they were close to radial alignment. That alone makes the paper worth reading. The in-situ analysis is careful: shock parameters, flux rope fits, clear identification of ejecta boundaries, and the WSA-Enlil check on arrival times is a nice consistency test.\n\nThe modelling part is a legitimate new application of OSPREI. The seed run gets arrival times within a few hours at all four probes, and the ensemble shows that best-fit members for inner probes become increasingly poor fits at outer probes. That is a real, useful result: it quantifies how quickly inner-probe constraints lose power for 1 au predictions.\n\nNow the soft spots, in proportion. The paper's central claim—that there may be a maximum useful separation between inner and outer probes—is conditional on OSPREI's uniform, constant solar wind background. The model forbids interplanetary deflections and rotations by construction, so the growing spread is purely input-parameter propagation through homogeneous expansion. The paper says this clearly, but it means the threshold language in the abstract overstates what a single event in a simplified model can support.\n\nMore telling is the B_R sign problem. The seed run predicts the wrong sign of B_R at all four spacecraft, and the authors only recover agreement by mirroring every crossing to the opposite side of the CME nose. That is direct evidence that the model's heliospheric trajectory is systematically off, probably because the real CME kept deflecting north after the coronal domain. So the 'divergence' is real as a model output, but it is not yet a general property of CMEs. A structured-wind simulation or a second event would be needed to elevate it.\n\nThere are also smaller technical caveats: no error bars on shock parameters or flux rope fits, a data gap at SolO's trailing edge, and Bepi data not yet public. These are annoying but not fatal in an event study.\n\nWho benefits? Anyone working on multi-spacecraft CME encounters, ensemble forecasting, or inner-heliospheric structure. The event itself will be a reference point for years. I'd send it to review; the referee should push for softening the threshold claim and maybe for an interplanetary-mode run that bypasses ForeCAT, which the authors already mention. But the observational core is solid and the modelling is honest about its limitations.","headline":"A genuinely rare four-probe CME event, carefully analysed; the ensemble-divergence conclusion is a useful hypothesis that still needs a structured-wind test before it becomes a general claim.","tokens_in":31834,"tokens_out":1956,"would_cite":true,"duration_ms":19542,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Ensemble predictions of a CME's magnetic field spread more widely with heliocentric distance, implying a limit to how far an inner probe can anchor an outer probe's forecast.","keywords":["coronal mass ejections","multi-spacecraft observations","ensemble modelling","flux rope magnetic structure","inner heliosphere","space weather forecasting","streamer-blowout CME","OSPREI"],"falsifier":"Run the same 200-member OSPREI ensemble for a CME encountered by outer probes at separations between 0.1 and 0.6 au and measure whether the spread in predicted magnetic field orientation grows monotonically with separation; if an outer probe close in angle but far in radius shows no divergence, or a nearby probe shows large divergence, the proposed threshold picture collapses.","tokens_in":30738,"feed_emoji":"☀️","tokens_out":7483,"duration_ms":65401,"temperature":0.7,"pith_summary":"This paper analyses a slow, streamer-blowout CME that erupted from the Sun on 23 September 2021 and was encountered in situ by four spacecraft spread almost evenly in heliocentric distance between 0.4 and 1 au. It is, to the authors' knowledge, the first reported CME observed by four well-radially-separated probes inside 1 au. Using the OSPREI modelling suite in a 200-member ensemble hindcast, the authors find that the spread of predicted in-situ quantities grows with heliocentric distance: the best-fit solution at an inner probe is not necessarily a good fit at an outer probe. They argue from this that there may be a maximum angular and radial separation between inner and outer probes beyond which inner-probe measurements lose power to constrain the magnetic field orientation at 1 au. The result matters because space weather forecasts increasingly use upstream inner probes to predict what will hit Earth, and this event gives a concrete test of how far that strategy can work.","feed_headline":"CME forecasts spread wider with distance from the Sun","feed_subtitle":"A 2021 CME hit four aligned probes; best-fit ensemble solutions split more the farther the probe lies.","key_machinery":"The argument is carried by the OSPREI modelling suite in ensemble mode. OSPREI chains three analytic modules: ForeCAT, which computes coronal deflections and rotations of the CME's flux rope; ANTEATR, which propagates the CME through interplanetary space and builds the sheath; and FIDO, which generates synthetic in-situ time series along a chosen observer trajectory. The CME body is described by an elliptic-cylindrical flux rope, and 200 ensemble members perturb 24 input parameters (position, tilt, speed, mass, magnetic field, solar wind conditions) around a seed run. A goodness-of-fit score, the sum of fractional mean absolute errors on hourly averaged field and plasma quantities plus a timing error for the shock and ejecta boundaries, selects the best member at each spacecraft and globally.","core_discovery":"The central claim is that for this event, the ensemble spread in predicted magnetic field and plasma quantities increases with heliocentric distance, and that this points to a practical limit on using inner spacecraft to constrain outer-spacecraft forecasts. The four probes, at about 0.44, 0.61, 0.78, and 0.96 au, all encountered the same CME, and OSPREI's seed run reproduced arrival times within the usual few-hour-to-10-hour uncertainty at all four. But in the 200-member ensemble, the single-spacecraft \"best-fit\" members for Bepi and Solar Orbiter become outliers by the time they are propagated to Parker Solar Probe and STEREO-A, while the same member that best fits PSP also best fits STEREO-A. At their CME arrival times, STEREO-A was separated from Bepi by 0.52 au and 12 degrees, from SolO by 0.35 au and 11 degrees, and from PSP by 0.18 au and 5 degrees; the authors propose that beyond some separation like these, inner-probe constraints on the in-situ magnetic field orientation, parameterised through flux rope geometry, increasingly diverge. They also show that mirroring all four encounters to the south of the modelled CME nose fixes a systematic sign error in the radial magnetic field component, suggesting the real CME deflected north of the simulated trajectory.","pith_inferences":["Editorial extension: a robust separation threshold would give a design rule for future heliospheric constellations, placing upstream monitors below roughly 0.2 au and a few degrees of angular separation to keep inner-outer correlation useful.","Editorial extension: because OSPREI assumes a uniform constant solar wind, the growing ensemble spread is a lower bound; realistic stream interaction regions and sector boundaries would add deflections and rotations that make inner-outer correlation fail at even smaller separations.","Editorial extension: the paper's mirroring exercise suggests a cheap test: compare the sign of the radial magnetic field across multiple spacecraft to estimate the CME nose latitude, an observable constraint independent of flux rope fitting.","Editorial extension: the divergence trend could be checked against metric choice; using dynamic time warping or other shape-sensitive scores might identify different \"best\" members, so a robustness study over metrics is a natural next step."],"forward_implications":["The same CME can be consistently identified at four probes from 0.4 to 1 au, and an analytic ensemble model can place all four arrivals within a few hours of observation.","Inner-probe data become a weaker constraint on outer-probe magnetic field orientation as the angular and radial separation grows; beyond some threshold, the best inner solution can mispredict arrival time by about 12 hours and field magnitude by roughly a factor of two at 1 au.","A sub-au probe near Venus's orbit is a plausible sweet spot for 1 au forecasts, close enough to remain correlated and far enough ahead to give lead time.","A systematic sign error in the predicted radial magnetic field can be traced to the assumed CME nose latitude, making the $B_R$ component a useful diagnostic of whether a crossing is north or south of the CME apex."],"supporting_citations":[{"why":"Supplies the complete OSPREI modelling suite used for the ensemble hindcast.","marker":"Kay et al. 2022a"},{"why":"Provides the ANTEATR module that propagates the CME and forms the sheath in interplanetary space.","marker":"Kay et al. 2022b"},{"why":"Supplies the FIDO module that turns model outputs into synthetic in-situ time series.","marker":"Kay & Gopalswamy 2017"},{"why":"Provides the ForeCAT module that computes the coronal deflections and rotations of the CME.","marker":"Kay et al. 2015"},{"why":"Supplies the graduated cylindrical shell method used to estimate the CME's coronal speed, direction, and tilt.","marker":"Thernisien 2011"},{"why":"Provides the expansion-modified force-free model used to fit the flux rope intervals at each spacecraft.","marker":"Farrugia et al. 1993"},{"why":"Establishes the inner-probe-to-1 au prediction strategy that this paper tests and qualifies.","marker":"Laker et al. 2024"},{"why":"Documents the difficulty of reconciling far-separated in-situ flux rope reconstructions under self-similar expansion, supporting the proposed separation threshold.","marker":"Davies et al. 2024"}],"fun_headline_variants":["CME ensemble forecasts diverge with heliocentric distance","Four probes catch same CME, reveal forecast limits","Inner probes may not constrain outer CME field forecasts"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The simulation treats the solar wind as a uniform, unchanging background, so any real solar-wind structures that bend or twist the CME on its way from 0.4 to 1 au are omitted.","fun_headline_variants_meta":{"raw":{"variants":["CME ensemble forecasts diverge with heliocentric distance","Four probes catch same CME, reveal forecast limits","Inner probes may not constrain outer CME field forecasts"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000541,"raw_usage":{"total_tokens":2692,"prompt_tokens":1144,"completion_tokens":1548,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":760,"completion_tokens_details":{"reasoning_tokens":1497}},"tokens_in":760,"tokens_out":1548,"duration_ms":13697,"temperature":1.0,"reasoning_tokens":1497,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T17:13:47.616965+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same 200-member OSPREI ensemble for a CME encountered by outer probes at separations between 0.1 and 0.6 au and measure whether the spread in predicted magnetic field orientation grows monotonically with separation; if an outer probe close in angle but far in radius shows no divergence, or a nearby probe shows large divergence, the proposed threshold picture collapses.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the FIDO module that turns model outputs into synthetic in-situ time series."}],"review_version":1}