{"id":"1e6cc44b-83e8-4a5f-be1e-6e1163bbbbab","arxiv_id":"2509.03325","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"For SPHERE/IRDIS RDI reductions, mixed reference libraries beat single-criterion libraries in disc S/N and consistency, while Pearson-correlation-selected libraries give the best mean contrast.","lead":"This paper tests different ways of choosing reference-star images for cleaning starlight noise when imaging faint dusty discs around other stars. It finds that mixing several selection criteria gives the most reliable results, while pure correlation-based selection gives the best average contrast.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Mixed-library S/N advantage rests on visually chosen S/N apertures for 7 discs; the paper itself admits region shifts reorder the metrics.","rationale":"The reader's weakest assumption identifies the same load-bearing issue: the S/N ranking depends on manually drawn regions, and slight changes can reorder selection metrics. This is the single most important risk to the paper's central claim because the preference for mixed libraries over PCC and epoch in the conclusions rests primarily on the real-disc S/N analysis, which uses only seven targets, no significance testing, and a measurement region that the authors acknowledge is not unique. The contrast analysis is more extensive (20 disc-free targets, four radii, two widths) and already shows PCC and mixed within 4% of each other, so the main contested point is the S/N advantage. The suggested test directly targets the stated caveat and would tell whether the headline 'best disc S/N' is a real effect or a region-selection artifact. This does not change the reader's CONDITIONAL verdict: the concern is addressable but currently unresolved, so the paper should not be unconditionally accepted without the robustness check.","tokens_in":22331,"tokens_out":2049,"duration_ms":24436,"concrete_test":"Recompute the disc S/N ranking with a reproducible, algorithmically defined measurement region (e.g., an annular sector from literature-based disc position angle and radius, or a threshold-based mask on the coadded best-S/N reduction), then apply a paired bootstrap or permutation test over the seven targets and also perturb the region boundaries by ±2 px and ±20 degrees. If mixed no longer significantly outperforms epoch and PCC under these region variations, the abstract's 'best disc S/N' claim should be weakened to 'comparable S/N with better robustness.'","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that mixed libraries achieve the best disc S/N (abstract, Sect. 5, conclusions) is supported by only seven real-disc targets and by S/N measured in a region defined via visual inspection (Sect. 3.4, Appendix B). The quoted advantage is numerically tiny: mixed has mean normalized S/N 0.84 versus 0.83 for epoch and PCC, with minimum 0.55 versus 0.52 and 0.66. The paper states in Sect. 5 that 'slight variations of the measurement regions can change the ordering of selection metrics with similar mean S/N values.' Since the region is chosen to include brightest disc flux while excluding negative oversubtraction artifacts, and the same region is used across libraries, the small mean differences could be an artifact of one subjective aperture choice. No bootstrap, permutation, or leave-one-target-out test is provided, so there is no estimate of whether the mixed-vs-epoch/PCC difference is statistically distinguishable from region-placement noise. The contrast prong of the analysis (Sect. 4) is more robust and does support mixed as a consistently good choice, but the 'best S/N' prong is load-bearing for the recommendation of mixed over PCC/epoch in a blind archival search.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper addresses reference-library selection for PCA-based reference-star differential imaging (RDI) of circumstellar discs, with a focus on pole-on discs observed with SPHERE/IRDIS. The authors construct 13 library types: ten metadata-matched libraries, a mixed library combining subsets matched on different parameters, a pure Pearson-correlation (PCC) library, and a random library. These are tested on 20 disc-free targets with synthetic disc throughput corrections at four radii and two disc widths, and on seven known disc targets whose disc S/N is measured in visually defined regions. The main claims are that mixed libraries achieve the best disc S/N and the smallest deviation from the best contrast per target, while PCC-only libraries achieve the best mean throughput-corrected contrast; epoch-based libraries perform well for small-separation discs. The authors recommend mixed libraries for large-scale archival reductions.","tokens_in":22641,"tokens_out":4746,"duration_ms":48918,"significance":"If the conclusions are robust, this is a practical and timely contribution: it directly informs the planned large-scale reduction of archival SPHERE/IRDIS data for disc searches, and it compares physically motivated selection metrics under realistic observing conditions. The study is carefully designed, uses a moderate but well-characterized sample (20 disc-free + 7 disc targets), and is transparent about many limitations. The contrast analysis with synthetic disc throughput is a useful empirical benchmark, and the finding that the mixed library consistently stays within a factor of two of the best contrast is a concrete, falsifiable claim. However, the headline S/N result is supported by only seven targets and by mean differences of 0.01 in normalized S/N, with no significance testing or region-robustness analysis. The paper's own caveats indicate that the ranking is fragile, yet the abstract and conclusions present the mixed library as the unequivocally best S/N choice. Strengthening or softening this claim is essential before the recommendation is adopted.","major_comments":[{"comment":"The headline result that the mixed library gives the best disc S/N is not statistically supported. The mean normalized S/N is 0.84 for mixed versus 0.83 for epoch and PCC over only seven targets, and the paper itself states that slight variations of the manually defined S/N regions (Sect. 3.4, Appendix B) can change the ordering. No bootstrap, permutation, or leave-one-target-out analysis is provided, and no test of sensitivity to the chosen S/N aperture is reported. Furthermore, the mixed library has a minimum normalized S/N of 0.55, worse than PCC's 0.66, so 'best' depends entirely on a small mean difference. Please add uncertainty estimates or a region-robustness test, and if the difference is not significant, revise the abstract and conclusion to say mixed is among the best rather than the best.","section":"Section 5, Fig. 5"},{"comment":"For non-wind-effect observing conditions the paper reports that the five best selection metrics fall within 10% of each other and that PCC, seeing, and mixed are within ~5%, then states that no meaningful ranking can be drawn. This caveat is not carried into the Conclusions, where mixed and PCC are described as giving the best mean contrast for LWE/WDH and PCC as best for the non-wind groups. With only about four science targets per observing-condition group (Table 2) and no significance tests, these subgroup rankings are not established. Please add confidence intervals or a by-target bootstrap, or explicitly label the condition-dependent analysis as exploratory.","section":"Section 4.2, Fig. 2"},{"comment":"The contrast comparison reports means such as PCC 1.16 and mixed 1.21 with no associated uncertainties. The 160 measurements (20 targets x 4 radii x 2 widths) are treated as independent, but they are not: the same target appears eight times, and measurements at different radii and widths for a given target are correlated. A difference of 0.05 in normalized contrast may or may not be meaningful, and the claim that PCC achieves the best mean contrast is therefore not quantitatively supported. A hierarchical or target-bootstrap resampling scheme should be used to attach uncertainties to the means and to the statement that PCC is best on average.","section":"Section 4.1, Fig. 1"}],"minor_comments":[{"comment":"The phrase 'single criteria' should be 'single criterion'.","section":"Abstract"},{"comment":"Typo: 'HD 1000453' should be 'HD 100453'.","section":"Section 5"},{"comment":"The caption states the metrics are 'ordered from top to bottom by ascending mean', but the figure shows the best-performing mixed library at the top and the worst-performing random library at the bottom, which is descending mean. This is inconsistent with the other figure captions.","section":"Fig. 5 caption"},{"comment":"The sentence 'For this study, we chose to build unique reference libraries for each science frame and wavelength channel, consisting of 1000 reference frames each' could be clarified to state explicitly whether 1000 frames are selected per wavelength channel or per cube.","section":"Section 3.2"},{"comment":"The small-sample penalty term from Mawet et al. (2014) is mentioned, but the reader must consult the reference to see the exact formula. Including the formula or a reference to the equation in that paper would improve reproducibility.","section":"Section 3.4"}],"recommendation":"major_revision","confidential_remarks":"The paper is commendably honest about its limitations, including the region-sensitivity of the S/N rankings and the 10% proximity of many contrast metrics. The editor may wish to ask the authors to reconcile the abstract/conclusions with those caveats. The central claim needing work is not a methodological error but a missing statistical treatment; this is within the manuscript's scope and should be fixable with a moderate revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The one-line take: this is a useful, honest empirical paper on how to pick reference libraries for RDI disk work, and the contrast half of it is solid. The S/N half that favors mixed libraries is more fragile than the abstract implies, and the authors largely know it.\n\nWhat's new: nobody had systematically tested the ten metadata preselection criteria side by side for disk RDI, and the mixed-library idea—drawing frames from parameter-matched subsets—is genuinely new. The epoch result for small-separation disks is also a practical finding worth having. The paper does the work properly: twenty disk-free targets spanning LWE, WDH, and good/average/bad seeing, throughput-corrected synthetic disk injection for contrast, seven known disks for S/N.\n\nThe contrast analysis is the sturdy part. PCC libraries give the best mean contrast, mixed libraries have the smallest worst-case deviation and always stay within a factor of two of the best. SpT and DIT are useless, random is worst. For a blind archival search, that is a real, actionable recommendation, and the computational argument for mixed libraries is sensible.\n\nNow the soft spots. The S/N ranking that pushes mixed to the top is based on seven real disks, with S/N measured in a region defined by eye to include the brightest disk flux and exclude negative oversubtraction. The paper itself says slight variations of the region can reorder the metrics. The numbers are close: mixed mean normalized S/N 0.84 versus 0.83 for epoch and PCC. No bootstrap, permutation, or leave-one-out test. So \"mixed gives the best S/N\" should not be treated as established; it is at best a hint. The stress-test concern lands here. That said, the authors are transparent about it, and the contrast prong still supports mixed as the safe all-around choice. The lack of significance testing is a genuine weakness, but it is addressable and not the kind of flaw that sinks the paper.\n\nThe paper keeps to its scope. The central claim isn't circular, there are no fitted parameters driving the result, and the self-citations are context, not load-bearing. It is a methods paper, not a physics breakthrough; the audience is people doing RDI disk reductions on SPHERE/IRDIS archival data.\n\nI would send this to peer review. The referee should ask for a robustness check on the S/N ranking—reproducible region definitions, bootstrapping—and a softened abstract on \"best S/N.\" But the core empirical work is worth publishing.","headline":"Solid, honest empirical comparison of RDI reference-selection metrics; the contrast analysis is the robust part, but the S/N ranking that favors mixed libraries is more fragile than the abstract lets on.","tokens_in":23090,"tokens_out":3314,"would_cite":true,"duration_ms":33405,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"In reference-star differential imaging, blended reference libraries give the best disc signal and most consistent contrast, while pure correlation picks give the best average contrast.","keywords":["reference-star differential imaging","RDI","reference library selection","SPHERE/IRDIS","circumstellar discs","pole-on discs","principal component analysis","Pearson correlation coefficient"],"falsifier":"Run the same 20-target synthetic-disc experiment with the S/N measured by an automated aperture centered on the injected disc model, with no hand-drawn regions, and check whether the mixed library still ranks first; if the ordering changes, the reported S/N advantage is a measurement artifact.","tokens_in":22280,"feed_emoji":"🔭","tokens_out":7963,"duration_ms":77982,"temperature":0.7,"pith_summary":"This paper tests how the choice of reference frames affects reference-star differential imaging (RDI) of faint, face-on circumstellar discs with SPHERE/IRDIS. The authors compare libraries built by matching ten different observational, atmospheric, and stellar parameters, by frame-to-frame Pearson correlation alone, by mixing all ten parameter-matched subsets, and by random selection. Their central result is that diverse 'mixed' libraries give the best disc signal-to-noise and the most consistent contrast across targets, while pure correlation-based libraries give the best average contrast. They conclude that mixed libraries are the practical choice for the upcoming large-scale RDI reduction of archival SPHERE/IRDIS data in the search for new discs.","feed_headline":"Mixed reference libraries beat single-criterion picks for disc imaging","feed_subtitle":"For RDI reductions of SPHERE/IRDIS data, ten blended selection metrics beat any single one on disc signal.","key_machinery":"The central object is the reference library construction scheme. For each science frame, the authors preselect the 10,000 master-library frames whose metadata—seeing, coherence time, 30m wind, 200mbar wind, elevation, epoch, detector integration time, G/H magnitudes, and spectral type—most closely match the science frame; the final 1000-frame library keeps the highest frame-to-frame Pearson correlation coefficient computed in the speckle-dominated annulus between 0.18 and 0.43 arcseconds. A 'mixed' library combines 100 such frames from each of the ten parameter preselects. These libraries feed a PCA-based RDI subtraction, and performance is judged by throughput-corrected contrast on syntheti","core_discovery":"The paper's claim is that for PCA-based RDI reductions of pole-on discs, no single selection criterion is enough: a reference library assembled by taking equal numbers of high-correlation frames from sub-libraries preselected to match different noise-determining parameters outperforms libraries built from any one criterion. On 20 disc-free sequences with synthetic pole-on discs, mixed libraries always stayed within a factor of two of the best contrast, and on seven real disc targets they produced the highest mean disc S/N. Pure Pearson-correlation libraries produced the best mean contrast overall, but with a larger worst-case spread, while libraries selected by frames close in time performed","pith_inferences":["A likely reason mixed libraries win is that real speckle patterns combine several decorrelated noise sources—wind-driven halos, low-wind effect, seeing residuals—and no single parameter or correlation score captures all of them; a testable extension would be adding more noise-descriptive parameters, such as scintillation strength, and checking whether the mixed-library margin grows.","The manual S/N region selection is the fragile point: an automated, model-matched measurement could decide whether the mixed library's S/N advantage is real or a region-choice artifact.","The epoch result hints that for observing strategies with dense temporal sampling, such as star-hopping, time-nearest reference selection may be competitive with image-similarity metrics, a hypothesis that could be tested directly on star-hopping sequences."],"forward_implications":["If this preference is right, the coming large-scale reduction of archival SPHERE/IRDIS data should use mixed reference libraries as the default selection, gaining the most consistent contrast and best disc S/N at lower computational cost than pure correlation libraries.","Discs at small separations, around 20px, should additionally or alternatively be reduced with epoch-matched libraries; these were best for about 30% of configurations and reached normalized S/N of at least 0.95 for the real small discs in the sample.","Pure Pearson-correlation libraries remain the best choice when the goal is lowest average contrast, such as for individual targeted reductions where computational cost is not a concern.","Random selection is strongly suboptimal, and selection on spectral type or detector integration time performs little better than random, so these should not be used as primary selection metrics.","The same selection scheme should transfer to other ground-based high-contrast instruments with large archives and atmospheric metadata, such as the Gemini Planet Imager."],"supporting_citations":[{"why":"Supplies the principal component analysis (Karhunen-Loève) algorithm used for RDI speckle subtraction.","marker":"Soummer et al. 2012"},{"why":"Jointly cited PCA approach for reference-star differential imaging.","marker":"Amara & Quanz 2012"},{"why":"Prior comparison of reference-selection metrics (SSIM vs Pearson correlation) for RDI; the baseline this study extends.","marker":"Ruane et al. 2019"},{"why":"Found that Pearson-correlation-selected reference frames give best S/N; the pure-PCC comparison point.","marker":"Romero et al. 2024"},{"why":"Showed ARDI improves most when reference frames are poor and that highest-PCC frames are not always best, motivating mixed libraries.","marker":"Juillard et al. 2024"},{"why":"Introduced the master-reference-library concept and showed contrast improves with library size; defines the frame pool.","marker":"Xie et al. 2022"},{"why":"Provides the small-sample-statistics penalty used in the disc S/N measurement.","marker":"Mawet et al. 2014"},{"why":"Defines the disc-free sample by infrared-excess and fractional-luminosity thresholds.","marker":"McDonald et al. 2017"}],"fun_headline_variants":["Blend reference criteria to lift disc signal in RDI","Mixed reference libraries improve disc S/N over single picks","Combine selection metrics for best disc signal in RDI","Diverse reference libraries win for pole-on disc imaging","No single criterion: blend RDI references for better discs"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The comparisons assume the hand-drawn regions used to measure disc S/N are unbiased; the paper itself reports that slightly changing these regions can reorder the selection metrics, so a biased region choice could make mixed libraries look best when they are not.","fun_headline_variants_meta":{"raw":{"variants":["Blend reference criteria to lift disc signal in RDI","Mixed reference libraries improve disc S/N over single picks","Combine selection metrics for best disc signal in RDI","Diverse reference libraries win for pole-on disc imaging","No single criterion: blend RDI references for better discs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000612,"raw_usage":{"total_tokens":2723,"prompt_tokens":827,"completion_tokens":1896,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":571,"completion_tokens_details":{"reasoning_tokens":1832}},"tokens_in":571,"tokens_out":1896,"duration_ms":13823,"temperature":1.0,"reasoning_tokens":1832,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T10:59:12.286025+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same 20-target synthetic-disc experiment with the S/N measured by an automated aperture centered on the injected disc model, with no hand-drawn regions, and check whether the mixed library still ranks first; if the ordering changes, the reported S/N advantage is a measurement artifact.","supporting_citations":[{"cited_title":"2012, ApJ, 755, L28","cited_arxiv_id":null,"evidence_quote":"Supplies the principal component analysis (Karhunen-Loève) algorithm used for RDI speckle subtraction."},{"cited_title":"& Quanz, S","cited_arxiv_id":null,"evidence_quote":"Jointly cited PCA approach for reference-star differential imaging."},{"cited_title":"2019, AJ, 157, 118","cited_arxiv_id":null,"evidence_quote":"Prior comparison of reference-selection metrics (SSIM vs Pearson correlation) for RDI; the baseline this study extends."},{"cited_title":"M., et al","cited_arxiv_id":null,"evidence_quote":"Found that Pearson-correlation-selected reference frames give best S/N; the pure-PCC comparison point."},{"cited_title":"2024, A&A, 688, A185","cited_arxiv_id":null,"evidence_quote":"Showed ARDI improves most when reference frames are poor and that highest-PCC frames are not always best, motivating mixed libraries."},{"cited_title":"2022, A&A, 666, A32 Article number, page 11 of 16 A&A proofs: manuscript no","cited_arxiv_id":null,"evidence_quote":"Introduced the master-reference-library concept and showed contrast improves with library size; defines the frame pool."},{"cited_title":"A., & Watson, R","cited_arxiv_id":null,"evidence_quote":"Defines the disc-free sample by infrared-excess and fractional-luminosity thresholds."}],"review_version":1}