{"id":"3302e550-a9c6-4588-9c13-5aaf1696587d","arxiv_id":"2506.03721","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"In bench tests on a VLT-plus-atmosphere simulator, the GRAVITY+ extreme adaptive optics system met its target Strehl ratios (75% at magnitude 9; 50% at magnitude 17), clearing the way for shipment to Paranal.","lead":"The GRAVITY+ adaptive optics system passed its European lab qualification, reproducing its target image sharpness on a simulator of the VLT telescope and atmosphere. The result green-lights shipment to Paranal, where the system will serve the VLTI interferometer for faint-object and high-contrast observations.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Bench performance curves were taken at 0.7\" seeing, not the 0.83\" median seeing used for the GPAO top-level requirement, and the paper does not quantify the resulting Strehl offset.","rationale":"The reader's weakest assumption was that the bench reproduces an UT coudé focus and Paranal turbulence; this concern is a specific instance of that assumption, namely the seeing value used for the headline performance curves. I do not see a reason to move the verdict: the paper is transparent about many limitations and reports an engineering qualification, and the seeing mismatch is a missing quantification rather than an established contradiction. However, it is the most load-bearing correctable issue because the entire claim 'we reproduced the expected performances' rests on Fig. 8, and the environmental condition under which those curves were measured is not aligned with the condition stated in the top-level requirement. If the 0.7\" vs 0.83\" offset is small relative to the margin, the claim stands; if it is large, the bench qualification overstates readiness. One concrete repetition at 0.83\" would settle this. Other concerns (absence of error bars, internal ERIS expectation curves, unsimulable LGS cone effect) are real but secondary: error bars would not change the central argument if the seeing-matched measurement still meets spec, the ERIS curves are a benchmark not a proof, and the cone-effect caveat is already acknowledged by the authors and deferred to on-sky AIV. The paper deserves credit for the iterative baffling result and the extensive template-based calibration workflow, which independently support that the system is functionally sound.","tokens_in":8409,"tokens_out":5737,"duration_ms":59375,"concrete_test":"Repeat the Fig. 8 NGS-VIS Strehl measurement at magR=9 with the phase plate configured to the specification seeing of 0.83\" (and rotation speed corresponding to 9.5 m/s wind), using the same photometric calibration and loop parameters. If the measured SR is below 75% by more than the stated 1%-at-1.31-µm measurement floor (≈20% at 2.2 µm), or below the ERIS expectation curve at 0.83\" by more than that floor, then the bench data as presented do not establish that GPAO meets its top-level NGS requirement. A complementary analytical check is to compute the predicted SR difference between 0.7\" and 0.83\" seeing from the ERIS model and compare it with the reported margin in Fig. 8.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4.1 states that the rotating phase plate on the Nice bench produced a 0.7\" seeing, with additional plates used to stress-test up to 1.4\". Section 5.2 defines the GPAO top-level requirement as SR=0.75 at magR=9 (NGS) and SR=0.50 at magR=17 (LGS) for median seeing conditions of 0.83\" and 9.5 m/s wind, then reports Fig. 8 Strehl curves measured 'with the rotating phase plate to simulate turbulence' without stating that the plate was configured to 0.83\" or that the results were rescaled to specification seeing. Strehl is a strong function of r0, so the difference between 0.7\" and 0.83\" seeing is not negligible: the green curve is described as meeting expectations 'with some margins,' but that margin is never quantified and could be comparable to, or smaller than, the expected degradation from the 0.13\"-seeing offset. The additional robustness tests at 1.4\" are not presented as SR curves, so they do not demonstrate that the headline SR values survive at specification seeing. This is a load-bearing gap because the central claim is that the bench validated the top-level requirements, and the measurement condition for that validation is not matched to the stated requirement.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript reports the European laboratory test campaign of the GRAVITY+ adaptive optics system (GPAO), conducted on a bench in Nice that simulates the VLT UT coudé focus and Paranal-like turbulence. It describes the calibration templates, interaction matrices, rejection-function measurements, NCPA calibration, and Strehl-versus-magnitude performance curves for the four GPAO modes (NGS-VIS, LGS-VIS, NGS-IR, LGS-IR). The authors conclude that the bench measurements reproduce expected performances, including under non-nominal conditions, and that this justified shipping the system to Paranal for AIV. The headline quantitative targets are SR=0.75 at magR=9 for NGS and SR=0.50 at magR=17 for LGS under median seeing (0.83\"), 9.5 m/s wind.","tokens_in":8643,"tokens_out":9153,"duration_ms":90935,"significance":"If the bench results transfer to Paranal, this is an important milestone for VLTI: the 40x40 NGS extreme AO and the LGS capability would be key enablers for high-contrast and faint-science observations. The paper's strengths are the transparency about known limitations (Maréchal breakdown, absent cone effect, parasitic light found and baffled), the detailed description of operational templates and calibration procedures, and the direct measurement of rejection functions and loop delays. The main caveats are that the headline figure-of-merit curves are not tied to the specification seeing, the LGS curve rests on a simple magnitude-shift ansatz, and the quoted Strehl values lack an uncertainty analysis.","major_comments":[{"comment":"The GPAO top-level requirements are quoted for median seeing of 0.83\" and 9.5 m/s wind, but the Strehl-vs-magnitude curves in Fig. 8 were obtained with the rotating phase plate described in Sec. 4.1 as producing 0.7\" seeing. The text does not indicate that the plate was reconfigured to 0.83\" for these measurements, nor does it rescale the results to the specification seeing. Since Strehl is a strong function of r0, the 0.13\" difference is potentially comparable to the unquantified 'some margins' claimed for the green curve; the stress tests at up to 1.4\" are not presented as Strehl-vs-magnitude curves and therefore do not bracket the specification point. Please quantify the expected Strehl degradation at 0.83\" (e.g., from the measured rejection functions or a simulation), report the margin at the specification point, and include error bars or an uncertainty estimate on the Fig. 8 curves.","section":"Section 4.1, 5.2, Fig. 8"},{"comment":"The LGS magnitude calibration is obtained by shifting the NGS-VIS calibration by 5 magnitudes because the 4x4 NGS WFS has 100x fewer subapertures than the 40x40 WFS. This is an assumption, not a measurement of the LGS path: it ignores differences in detector quantum efficiency, read noise, spot sampling, and the separate constant-brightness LGS source used for high-order correction. The yellow curve should be presented as an extrapolated estimate, with an uncertainty, rather than as a measured LGS performance curve. A photon-budget comparison or an end-to-end simulation would make the 5-magnitude shift testable.","section":"Section 4.2 and 5.2 (yellow curve, Fig. 8)"},{"comment":"The paper correctly warns that the Maréchal conversion is not valid at low Strehl and that 1% of Strehl at 1.3 µm maps to about 20% at 2.2 µm. The LGS requirement SR=0.5 at 2.2 µm corresponds to roughly 0.14 Strehl at 1.3 µm under the same formula, i.e., precisely in the regime where the conversion is most uncertain. The yellow curve's compliance with the LGS top-level requirement is therefore not established by the 1.3 µm measurements alone. Please report the 1.3 µm values for the LGS curve, quantify the conversion error at the relevant Strehl levels, or mark the affected points as upper limits.","section":"Section 5.2, Eq. (3) and Fig. 6"},{"comment":"The definition of the Strehl estimator is ambiguous as written. If PSF_perfect,norm and PSF_meas,norm are both normalized to unit total flux, the ratio of their sums is unity and Eq. (2) cannot produce the reported Strehl values; if 'norm' denotes peak normalization or a windowed sum, that must be stated. The omission of the telescope spider from the perfect PSF and the 16% ghost fraction are also potential biases that need to be quantified. Please specify the normalization, the summation window, and the error propagation to Fig. 8.","section":"Section 5.2, Eq. (2)"}],"minor_comments":[{"comment":"The factor S0 is called an 'unknown initial strehl' but is never defined or used; the text should clarify that this equation is a relative flux estimator, not an absolute Strehl measurement.","section":"Section 5.1, Eq. (1)"},{"comment":"The expected curves in Fig. 8 (right) are said to be 'derived from the ERIS design documentation' but no citation is given; please add a public reference or make the model accessible.","section":"Section 5.2"},{"comment":"The spelling 'strehl'/'Strehl' and 'Maréchal'/'Marechal' is inconsistent; please standardize.","section":"Throughout"},{"comment":"The sentence 'GPAO#1, as well as GPAO#3 and 4 have been sent...' is unclear, since GPAO#1 was described as used in Nice and later sent back; please clarify the hardware flow.","section":"Section 2.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is a proceeding-style instrument status report, and the recommendation of major revision is based on the gap between the headline claim ('green light') and the quantitative support in Fig. 8. The reliance on internal ERIS design documentation and on consortium self-citations is typical for this community but makes the 'expected performance' benchmark hard for an outside reader to verify. I would encourage the editor to ask for a numerical table of the Fig. 8 values and a statement of the seeing setting used."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis is the bench qualification report for the GRAVITY+ adaptive optics, the 40x40 XAO headed for all four VLT UTs. What's genuinely new is the end-to-end test of the assembled hardware on the Nice UT+atmosphere simulator: measured rejection functions, the mis-registration tracking algorithm, NCPA calibration by modal flux maximization, and Strehl-versus-magnitude curves for all four modes. That is real engineering data, and the green light to ship to Paranal rests on it.\n\nThe paper deserves credit for being candid. It states that the Marechal approximation breaks down at low Strehl and could make some K-band values optimistic; it says the LGS cone effect cannot be reproduced on the bench; it describes the encoder-LED parasitic light problem, the baffle fix, and the before/after curves. The test workflow---templates, interaction matrices, reference slopes---looks credible and is exactly what you would want from an AIT team.\n\nThe soft spots are in the quantitative chain. The stress-test concern is real: Section 4.1 says the rotating phase plate produced 0.7 arcsec seeing, and Section 5.2 uses that plate for the headline curves without ever stating that it was reset to the 0.83 arcsec median seeing in the top-level requirement, or that results were rescaled. Since Strehl is a steep function of r0, the claimed margins on the green NGS curve may be partly an artifact of testing at better-than-spec seeing. The paper does not quantify that offset. The curves also lack error bars and tabulated values, and the expected curves come from internal ERIS design documentation rather than a citable source. These are not fatal to the qualitative conclusion---the system clearly works---but they weaken the specific claim that SR=0.75 at mag 9 and SR=0.5 at mag 17 were met with margin.\n\nWho is this for? Instrument builders, VLTI users, and anyone planning on GPAO capabilities. It is a proceedings paper, not a methods or science paper; the math is simple Strehl estimation and rejection-function fitting.\n\nI would engage with it, but I would want the final version to state the bench seeing explicitly and give numbers with uncertainties. That is a refereeable request, not a rejection.","headline":"Genuine end-to-end bench qualification of GRAVITY+ AO, but the headline Strehl curves need the bench seeing and uncertainties stated before the claimed margins mean anything.","tokens_in":9661,"tokens_out":2629,"would_cite":false,"duration_ms":25849,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"GRAVITY+ adaptive optics meets all performance targets in European lab tests.","keywords":["GRAVITY+","XAO","adaptive optics","AIT","test","VLTI","laser guide star","Strehl ratio"],"falsifier":"Measure the on-sky K-band Strehl of GRAVITY+ during first commissioning at the VLT under the same conditions as the bench specification (median seeing of 0.83 arcseconds and 9.5 m/s wind speed). If the natural-guide-star Strehl at magnitude 9 falls below 75 percent, or the laser-guide-star Strehl at magnitude 17 falls below 50 percent, after calibration, the bench predictions would be contradicted.","tokens_in":8185,"feed_emoji":"🔭","tokens_out":7269,"duration_ms":72854,"temperature":0.7,"pith_summary":"The paper reports the results of the European test phase of GRAVITY+, the extreme adaptive optics system being added to each Unit Telescope of the VLT interferometer. The authors built a laboratory bench that reproduces the telescope's coude focus and a Paranal-like turbulent atmosphere, and they claim that all four operating modes of the adaptive optics reached their expected Strehl ratios there, including under degraded seeing and misaligned conditions. After optimizing the control loop and shielding a stray-light source, the bench measurements meet the top-level requirements of 75% Strehl at magnitude 9 for the natural guide star and 50% Strehl at magnitude 17 for the laser guide star. The results gave the green light to ship the system to Paranal for assembly, integration, and verification. If the bench predictions transfer to the sky, GRAVITY+ will enable high-dynamic-range observations of exoplanets and faint active galactic nuclei.","feed_headline":"Lab tests clear GRAVITY+ adaptive optics for Paranal","feed_subtitle":"Realistic bench runs hit the Strehl targets for natural and laser guide stars, boosting VLTI sensitivity and contrast.","key_machinery":"The load-bearing object is the test bench built in Nice: it combines an illuminated source whose flux is calibrated in apparent magnitude, a rotating phase plate that reproduces Paranal-like turbulence (about 0.7 arcsecond seeing in the nominal configuration, up to 1.4 arcseconds in stress tests), and optics that emulate the UT coude focus, followed by the adaptive optics system and a camera that measures the PSF at 1.31 micrometers. Strehl is computed from that camera image using a normalized perfect-PSF comparison with a 16-percent ghost-flux correction, and the 1.31-micrometer value is converted to the 2.2-micrometer GRAVITY band with the Marechal approximation. Calibration templates generate interaction matrices, reference slopes, and non-common-path aberration corrections, and a real-time mis-registration algorithm keeps the wavefront sensor aligned to the deformable mirror during operation.","core_discovery":"The central claim is that the GRAVITY+ adaptive optics system performs as designed when tested on a full-scale simulator of the VLT Unit Telescope and Paranal atmosphere. The paper shows measured Strehl-versus-magnitude curves for all four modes: the natural-guide-star visible mode (40x40 sub-apertures), the laser-guide-star visible mode (30x30 plus a 4x4 low-order sensor), and the corresponding infrared 9x9 modes inherited from the CIAO system. After a parasitic infrared LED from a motor encoder was found to contaminate the wavefront sensor and was blocked with baffling, the NGS curve met the 75-percent Strehl requirement at magnitude 9 with margin, and the LGS curve met the 50-percent requirement at magnitude 17. The paper also reports that non-common-path aberrations were calibrated on the bench by mode-by-mode flux maximization, reaching above 75 percent Strehl at 1.31 micrometers, and that the loop remained stable under off-centered pupils, defective actuators, and turbulence up to an estimated 1.4 arcseconds seeing.","pith_inferences":["Because the bench cannot simulate the LGS cone effect, on-sky LGS Strehl at magnitude 17 could come in below the 50 percent target even though the bench shows margin; commissioning will be the real test.","The paper cautions that the Marechal conversion is unreliable below about 20 percent Strehl at 2.2 microns, so the faint-end performance values, especially at high NGS magnitudes, are uncertain and could be lower in practice.","If the phase plate's turbulence statistics match Paranal's median seeing, the measured margins suggest the system has headroom, but real telescope vibrations and thermal flexure were only partially represented, so final margins may shrink on sky.","The stray-light episode shows that bench environments contain their own parasitic sources; similar effects on the telescope could degrade Strehl unless checked with an equivalent health-check template."],"forward_implications":["The four GPAO operating modes are ready for the integration and verification phase at Paranal, with the system shipped in mid-2024.","On-sky NGS performance around 75 percent Strehl in K band at magnitude 9 would give GRAVITY+ the high dynamic range needed to observe faint companions and exoplanets in the VLTI.","Working LGS modes extend the interferometer to fainter targets such as active galactic nuclei, at the cost of a cone effect that cannot be tested in the lab.","The automated calibration templates should make daytime setup and night operations faster and more repeatable than earlier AO systems.","Robustness tests against vignetting, actuator failure, and target wandering suggest GPAO can maintain performance during real observing conditions."],"supporting_citations":[{"why":"Describes the test bench that emulates the UT coude focus and Paranal turbulence; it is the central apparatus of all performance measurements.","marker":"[2]"},{"why":"Presents the GRAVITY+ wavefront sensors, defining the four modes whose performance is reported here.","marker":"[1]"},{"why":"Sets the GRAVITY+ science goals and faint-object requirements that the top-level Strehl specifications come from.","marker":"[4]"},{"why":"Specifies the ALPAO deformable mirror used in the corrective optics, whose 40x40-class actuator count underlies the extreme AO capability.","marker":"[9]"},{"why":"Provides the mis-registration algorithm used to track wavefront-sensor/deformable-mirror alignment during the robustness tests.","marker":"[10]"},{"why":"Characterizes the OCAM2K camera used in the wavefront sensors, supporting the noise and sensitivity assumptions in the bench measurements.","marker":"[11]"}],"fun_headline_variants":["GRAVITY+ AO bench tests meet Strehl specs, green light for Paranal","Lab simulator confirms GRAVITY+ AO hits performance goals","GRAVITY+ AO passes simulated Paranal tests, ready for AIV","GRAVITY+ AO lab runs match specs, clearing way for Paranal"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The bench faithfully reproduces the VLT telescope focus and Paranal turbulence, so that the Strehl ratios measured in the laboratory will also be reached on the sky at Paranal; the paper itself notes that the laser-guide-star cone effect cannot be simulated in the bench.","fun_headline_variants_meta":{"raw":{"variants":["GRAVITY+ AO bench tests meet Strehl specs, green light for Paranal","Lab simulator confirms GRAVITY+ AO hits performance goals","GRAVITY+ AO passes simulated Paranal tests, ready for AIV","GRAVITY+ AO lab runs match specs, clearing way for Paranal"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000434,"raw_usage":{"total_tokens":2198,"prompt_tokens":917,"completion_tokens":1281,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":533,"completion_tokens_details":{"reasoning_tokens":1196}},"tokens_in":533,"tokens_out":1281,"duration_ms":10422,"temperature":1.0,"reasoning_tokens":1196,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T10:56:25.641015+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure the on-sky K-band Strehl of GRAVITY+ during first commissioning at the VLT under the same conditions as the bench specification (median seeing of 0.83 arcseconds and 9.5 m/s wind speed). If the natural-guide-star Strehl at magnitude 9 falls below 75 percent, or the laser-guide-star Strehl at magnitude 17 falls below 50 percent, after calibration, the bench predictions would be contradicted.","supporting_citations":[{"cited_title":"Building a ...,","cited_arxiv_id":null,"evidence_quote":"Describes the test bench that emulates the UT coude focus and Paranal turbulence; it is the central apparatus of all performance measurements."},{"cited_title":"GRA VITY+ Wavefront Sensors: fully AO assisted interferometry on 8m class telescope at VLTI ,","cited_arxiv_id":null,"evidence_quote":"Presents the GRAVITY+ wavefront sensors, defining the four modes whose performance is reported here."},{"cited_title":"GRA VITY+: Towards faint science,","cited_arxiv_id":null,"evidence_quote":"Sets the GRAVITY+ science goals and faint-object requirements that the top-level Strehl specifications come from."},{"cited_title":"Recent improvements of high density magnetic deformable mirrors: faster, larger and stronger,","cited_arxiv_id":null,"evidence_quote":"Specifies the ALPAO deformable mirror used in the corrective optics, whose 40x40-class actuator count underlies the extreme AO capability."},{"cited_title":"Estimation of the lateral mis-registrations of the GRA VITY + adaptive optics system,","cited_arxiv_id":null,"evidence_quote":"Provides the mis-registration algorithm used to track wavefront-sensor/deformable-mirror alignment during the robustness tests."},{"cited_title":"Characterization of OCam and CCD220: the fastest and most sensitive camera to date for AO wavefront sensing,","cited_arxiv_id":null,"evidence_quote":"Characterizes the OCAM2K camera used in the wavefront sensors, supporting the noise and sensitivity assumptions in the bench measurements."}],"review_version":1}