{"id":"b747a52b-8d8d-4c8f-9758-15c7efa24a76","arxiv_id":"1908.02815","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A U-net trained on noisy image pairs alone, without clean targets, denoises solar Stokes images to about 6e-4 continuum residual, matching clean-target training on synthetic data.","lead":"Solar polarization signals that trace magnetic fields are often buried in noise, and clean images or noise models are usually unavailable; this paper adapts the Noise2Noise deep-learning trick by training a network on pairs of noisy solar frames of the same scene, with no clean data. Tests on simulations match supervised training, and examples on Swedish 1-meter Solar Telescope data show cleaner Stokes maps and profiles.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Real-data validation rests on an unverified assumption that time-separated CRISP frames are independent zero-mean noise realizations of the same signal; if temporal evolution is not negligible, the network may suppress real signal, and no quantitative ground-truth test distinguishes this.","rationale":"The reader identifies the same load-bearing assumption: Noise2Noise requires independent noise realizations of the same underlying signal with zero-mean noise, and real observations approximate this by temporal redundancy. My reading of the paper confirms this is the weakest point of the central claim, because the real-data demonstration has no ground truth and the only quantitative evidence of success is agreement with time-averaged profiles, which is also the expected behavior of a network that suppresses temporal evolution. The synthetic demonstration is solid for its controlled setting, and the authors are transparent about the need for stable training data and about the residual correlation introduced by MOMFBD. However, none of that directly tests whether the network preserves real small-scale signal when the two training frames are not identical. This is not a reason to reject the paper; the method and synthetic validation are valuable, and the code and weights are provided. But the real-data claim should remain conditional until a quantitative test with known signal (injected or simulated) is performed. I therefore keep the reader's CONDITIONAL verdict rather than moving to ACCEPT or REJECT. The abstract's broad generality claim is also overreaching relative to the tested spectral lines and samplings, but that is secondary to the core validation gap.","tokens_in":18302,"tokens_out":4463,"duration_ms":52495,"concrete_test":"Use the MURaM time series already used in Section 3.1 to build training pairs from temporally consecutive snapshots rather than two independent noise draws of the same snapshot, with known clean signals for both frames and with noise added to mimic CRISP/MOMFBD correlated residuals. Train the same U-net with the Noise2Noise loss and compare its output to the true first-frame signal, quantifying residual error and specifically the amplitude of real inter-frame changes that are removed. If the network suppresses known transient features or the residual exceeds the 6x10^-4 Ic level claimed in Section 3.1, the real-data conclusions in Section 3.2 are not supported. A complementary check is to inject a synthetic Stokes-V blob of known amplitude into a quiet CRISP frame and measure the recovered amplitude after applying the trained network.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that Noise2Noise training denoises real CRISP/SST Stokes data without clean targets depends on the assumption, stated in Section 2.1 and used in Section 3.2, that the two images in each training pair differ mainly by independent zero-mean noise. For real data this is enforced only by selecting 'stable over time' periods and by data augmentation (Section 3.2). The paper provides no quantitative test on real data that the network is recovering the true signal rather than suppressing real temporal evolution or signal-correlated MOMFBD residuals. The only real-data validations are qualitative: residual maps, power spectra, and comparison of reconstructed profiles to time-averaged profiles (Section 3.2.2). Agreement with a time average is exactly what a smoothing or evolution-suppressing network would produce, so it cannot establish that temporal evolution is preserved. The synthetic experiment in Section 3.1 validates the method only for identical-clean-image pairs with zero-mean Gaussian noise; it does not validate the temporal-drift case. Figure 9 shows sensitivity to noise character, but no analogous test is provided for nonzero-mean or signal-correlated corruption or scene evolution.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper applies the Noise2Noise paradigm to denoising solar spectropolarimetric images. The authors train a U-Net-like convolutional network on pairs of independent noisy realizations of the same scene, without any clean targets or explicit noise model, and compare it to a network trained on noisy-clean pairs. In a synthetic test based on MURaM MHD simulations, both training strategies produce similar residuals (about 6e-4 Ic against an injected noise level of 3e-3 Ic). The same architecture is then trained on CRISP/SST observations of Ca II 8542 Å using temporally separated frames as noisy pairs, and the denoised maps and spectral profiles are compared qualitatively with the originals and with time-averaged profiles. The paper also presents an uncertainty-estimation extension using heteroscedastic aleatoric uncertainty and MC-dropout, and explores a spatial-plus-spectral variant. The central claim is that the method recovers weak chromospheric polarization signals without needing clean data or a noise model.","tokens_in":18540,"tokens_out":3801,"duration_ms":41465,"significance":"If the central claim holds, the method is genuinely useful for chromospheric spectropolarimetry, where polarization signals are often at the detection limit and where post-processing such as MOMFBD produces correlated, non-Gaussian noise that is difficult to model. A notable strength is that the synthetic proof of concept is independently grounded: the noisy-noisy network is benchmarked against external clean synthetic images, so the equivalence of the two training strategies is demonstrated rather than assumed. The authors also make the code and trained weights publicly available, which supports reproducibility. The limitations of the work are openly discussed, including the possibility that the uncertainty estimates are too low and that the spectral variant may introduce unphysical correlations. However, the real-data validation is qualitative and rests on an unverified assumption about the independence of temporal frames, and the uncertainty calibration is partly circular. These issues need to be addressed before the method can be recommended for routine use on real observations.","major_comments":[{"comment":"The real-data validation does not establish that the network preserves real signal, including temporal evolution, rather than merely smoothing it. The comparison of reconstructed profiles to time-averaged profiles is exactly what a smoothing or evolution-suppressing network would produce, so agreement with the average cannot discriminate between noise suppression and signal suppression. The authors should provide a quantitative test on real data, for example by injecting synthetic (known) signals into real frames and measuring recovery, or by comparing the network output against an independent higher-S/N observation, to support the claim that temporal evolution is preserved.","section":"Section 3.2.2, Fig. 8"},{"comment":"The uncertainty calibration is circular as presented. The manuscript states that the MC-dropout rate is 'chosen to produces the expected error', and the dropout rate is the free parameter controlling the epistemic uncertainty. Therefore the agreement between the uncertainty estimate and the residual error shown in Fig. A.1 is not an independent confirmation that the network's uncertainties are well calibrated; it is a consequence of tuning that parameter. The authors should either fix the dropout rate a priori on a validation set independent of the residual-error comparison, or present the calibration as a demonstration of the tuning procedure rather than as evidence that the uncertainty estimates are trustworthy.","section":"Section 3.1.1, Appendix A"},{"comment":"The synthetic proof of concept covers one MHD snapshot, one noise level (3e-3 Ic), and reports a single residual statistic (6e-4 Ic) with no repeated runs, multiple noise levels, or error bars. In addition, Fig. 4 shows that in the critically sampled case the network suppresses the highest spatial frequencies, so the recovery is not complete. The claim that the method 'succeeds in correcting up to a level of 6e-4 Ic' should be framed as a single realization rather than a general performance estimate, and the power-spectrum suppression should be discussed as a known limitation of the method.","section":"Section 3.1, Fig. 4"},{"comment":"The load-bearing assumption for the real-data application is that two time-separated frames are independent zero-mean noise realizations of the same underlying signal, as stated in Section 2.1 ('the noise is the main difference between the two images'). For the CRISP data, this is mitigated only by selecting periods that are 'stable over time' and by data augmentation, but no quantitative test on real data is provided to show that the network is not suppressing real temporal evolution or signal-correlated MOMFBD residuals. Given that the entire real-data claim depends on this assumption, the authors should provide an explicit diagnostic, such as measuring the residual statistics against a known injected signal on real frames, or comparing network outputs with an independent data product.","section":"Sections 2.1 and 3.2"}],"minor_comments":[{"comment":"The statement that the method 'can recover weak signals equally well no matter on what spectral line or spectral sampling is used' overreaches the presented evidence; the experiments cover one spectral line pair (Fe I 6302/6301 Å at 50 mÅ sampling) and one real line (Ca II 8542 Å). Please soften the claim to reflect the tested configurations.","section":"Abstract"},{"comment":"In the sentence about optical flow, 'wrap one frame to the other' should be 'warp one frame to the other'.","section":"Section 3.2"},{"comment":"The phrase 'around 1 .10−4Ic' should read 'around 1x10^-4 Ic'.","section":"Section 3.2.3"},{"comment":"The 'Variance Output' row is unclear: '−Abs(x)−12' does not specify the operation unambiguously. Please state that the final variance is computed as exp of the network output or otherwise clarify the parameterization.","section":"Appendix B, Table B.2"},{"comment":"The labels in the figure captions such as 'Difference (3e-03)/(6e-04)' are ambiguous. Please clarify that the difference panels are scaled to the stated levels or indicate the color scale explicitly.","section":"Figures 3 and 6"},{"comment":"The claim that the chosen topology is 'the perfect balance between accuracy and speed of execution' is not quantified; the comparison with other architectures in Section 3.2.3 is qualitative. Please provide quantitative comparisons (e.g., network size, runtime, and residuals) or remove the word 'perfect'.","section":"Section 2.2"}],"recommendation":"major_revision","confidential_remarks":"The paper is a reasonable application of Noise2Noise to solar spectropolarimetry, and the synthetic proof of concept is clean and independently grounded. The main gap is the lack of quantitative real-data validation: the method's usefulness for real observations rests on an assumption of temporal-frame independence that is not tested. The circularity in the dropout-rate tuning for uncertainty calibration is also a concern that should be fixed, not merely acknowledged. I would encourage the editor to request a revision that addresses these two points explicitly."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague—\n\nQuick take: this is a solid, honest application paper. Díaz Baso et al. bring Noise2Noise to spectropolarimetric Stokes images, which is new in the solar literature. They show on synthetic MURaM data that training on noisy–noisy pairs matches training on noisy–clean pairs to about 6e-4 Ic residual, and they make code and weights available. The synthetic setup uses an external clean simulation as truth, so the central denoising claim is not circular. The uncertainty work is a real addition: MC-dropout plus heteroscedastic aleatoric loss, with a calibration check against residuals. Also to their credit, they test a spatio-spectral variant, find it gives marginal gains, and recommend PCA for spectral compression—that is the kind of judgment you want.\n\nThe soft spots are real but proportionate. First, the abstract claims the method recovers weak signals 'equally well no matter on what spectral line or spectral sampling is used.' The evidence covers one synthetic Fe i line pair and one real Ca ii line with a single sampling. That is an overgeneralization. Second, the real-data validation is qualitative. The comparison of reconstructed profiles to time-averaged profiles is weak evidence: a network that smooths or suppresses temporal evolution would also agree with a time average, so it cannot show that real evolution is preserved. The paper explicitly says separating noise from evolution is 'another problem,' but it does not quantify it. This is the load-bearing assumption of Noise2Noise—pairs must be independent zero-mean noise realizations of the same signal. They mitigate by choosing stable periods and augmenting, but no test distinguishes signal preservation from smoothing. Third, the uncertainty calibration is partly tuned: the dropout rate is chosen to produce the expected error, so Fig. A.1's agreement is not an independent check. Minor: no repeated synthetic runs or error bars, and no quantitative comparison to PCA or time-averaging baselines.\n\nThe central claim—that Noise2Noise can match clean-target training when its assumption is satisfied—holds up on the synthetic test. What is not proven is how much of the real-data gain is signal rather than smoothing. That is the gap a revision should close.\n\nWho this is for: solar physicists working on weak chromospheric polarization in CRISP, DKIST, or EST data, and method people interested in self-supervised denoising. It deserves serious peer review; I would send it out with the expectation of revision. The paper is not a desk reject.","headline":"Solid first application of Noise2Noise to solar Stokes polarimetry; abstract overclaims generality and real-data validation is qualitative, but the synthetic proof-of-concept and honest limitations make it worthy of review.","tokens_in":19080,"tokens_out":2583,"would_cite":true,"duration_ms":29453,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A convolutional neural network trained only on paired noisy solar spectropolarimetric images, with no clean targets and no noise model, denoises the data about as well as a network trained with clean targets.","keywords":["convolutional neural networks","image denoising","spectropolarimetry","solar chromosphere","Stokes parameters","Noise2Noise","U-network","temporal redundancy"],"falsifier":"Use a synthetic dataset with known ground truth where each paired frame carries the same scene plus independent noise, but inject a small, known, slowly varying shift in one frame. If the trained network output attenuates or removes that injected change, the zero-mean assumption has been violated and the method cannot separate real temporal evolution from noise. A realistic version applies the network to two real frames separated by a known flow and compares the output in the moving region with a third, temporally averaged reference.","tokens_in":18104,"feed_emoji":"☀️","tokens_out":7068,"duration_ms":76725,"temperature":0.7,"pith_summary":"This paper claims that a convolutional neural network can denoise solar spectropolarimetric images using only pairs of independent noisy observations of the same scene, with no clean images and no explicit noise model. On synthetic data, the noisy-pair network leaves residuals of about $6\\times10^{-4}I_c$, the same level as a clean-target network and roughly what averaging 25 frames would achieve. On real CRISP/SST observations of the Ca ii 8542 Å line, the procedure removes noise and fixed-pattern artifacts while yielding spectrally coherent Stokes profiles. If the claim holds, weak chromospheric magnetic-field polarization signals, normally buried near the detection limit, become recoverable from existing time series without clean calibrations.","feed_headline":"Solar Stokes denoising needs no clean images, only noisy pairs","feed_subtitle":"Trained on paired noisy frames alone, it matches clean-reference results and reveals weak chromospheric signals.","key_machinery":"The central mechanism is Noise2Noise training: the network minimizes the mean squared difference between its output and a second, independent noisy image of the same underlying scene rather than a clean image, so the expected loss is minimized by the clean scene when the noise has zero mean. The architecture is a U-network with encoder-decoder convolutional blocks and skip connections; it predicts each pixel from its spatial neighbors, using spatial coherence as the signal. Independent pairs come from temporal redundancy in time series, with data augmentation (rotations, sign flips, reversed wavelength order) used to balance any solar evolution between frames. The paper also pairs this with a Bayesian uncertainty estimate using dropout and a heteroscedastic loss to return per-pixel error bars.","core_discovery":"A U-network convolutional encoder-decoder trained with the Noise2Noise objective recovers weak Stokes signals under complex, signal-correlated corruption. The paper demonstrates this in two settings: on synthetic profiles from an MHD simulation, where training on noisy input/noisy target pairs leaves a residual standard deviation of about $6\\times10^{-4}I_c$, matching the clean-target baseline; and on real full-Stokes observations from the Swedish 1-meter Solar Telescope, where the same architecture suppresses photon noise and post-processing artifacts. The recovered Stokes Q and V profiles are spectrally coherent even though the network is trained monochromatically, and the network can be extended to take wavelength cubes as input for slightly smoother profiles. The conclusion is that clean targets are unnecessary for this denoising task: temporal redundancy supplies the independent noise realizations the loss function needs.","pith_inferences":["Any instrument that repeatedly images the same target, not only Fabry-Perot filtergraphs, could use the same pairing trick; the practical limit is how stationary the scene remains between paired exposures.","The cleanest stress test is to inject a small, known temporal evolution into one frame of a synthetic pair: if the network suppresses or attenuates it, the zero-mean-noise condition has failed in exactly the regime where real chromospheric dynamics vary.","For real data, cross-checking the denoised output against temporally averaged profiles at the same pixels, as the paper does qualitatively, is a cheap way to detect whether the network has learned to erase genuine evolution rather than noise."],"forward_implications":["Weak chromospheric polarization signals can be recovered from existing observations with no clean reference data, improving the empirical basis for chromospheric magnetic-field studies.","Because the pairing relies only on repeated exposures of the same scene, the method transfers across spectral lines and wavelength samplings, including cases where PCA degrades because wavelength sampling is scarce.","On real data, the network removes not just random noise but also fixed-pattern post-processing artifacts, such as vertical stripes, that are hard to model analytically.","Adding the spectral dimension as input yields smoother Stokes profiles, with average differences around $1\\times10^{-4}I_c$ relative to the spatial-only version.","The Bayesian extension produces uncertainty maps of order $6\\times10^{-4}I_c$, giving downstream magnetic-field inversions a quantitative error estimate."],"supporting_citations":[{"why":"Supplies the Noise2Noise principle that independent noisy realizations can replace clean targets in the loss.","marker":"Lehtinen et al. (2018)"},{"why":"Supplies the U-network encoder-decoder topology with skip connections used for denoising.","marker":"Ronneberger et al. (2015)"},{"why":"Defines the CRISPRED pipeline used to reduce the real CRISP/SST observations.","marker":"de la Cruz Rodríguez et al. (2015)"},{"why":"Describes the multi-frame blind deconvolution reconstruction whose correlated residuals the network must handle.","marker":"Löfdahl (2002)"},{"why":"Extends the MOMFBD method used for the real data and its residual artifacts.","marker":"van Noort et al. (2005)"},{"why":"Provides the earlier network topology and training details that the present architecture builds on.","marker":"Díaz Baso & Asensio Ramos (2018)"},{"why":"Supplies the MHD simulation used to synthesize the training and test spectropolarimetric data.","marker":"Vögler et al. (2005)"},{"why":"Supplies the flux-emergence simulation snapshot used in the synthetic experiment.","marker":"Rempel (2017)"}],"fun_headline_variants":["Solar denoising CNNs learn without clean images","Noise2Noise cleans solar images, no clean data needed","Deep learning removes solar noise using noisy pairs only","CNN denoising for the Sun: clean images not required"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that each training pair consists of two independent noise realizations of the same underlying scene with zero-mean noise; if the solar scene changes between frames or post-processing leaves signal-correlated, non-zero-mean residuals, the network will learn to suppress real changes or true signals.","fun_headline_variants_meta":{"raw":{"variants":["Solar denoising CNNs learn without clean images","Noise2Noise cleans solar images, no clean data needed","Deep learning removes solar noise using noisy pairs only","CNN denoising for the Sun: clean images not required"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000133,"raw_usage":{"total_tokens":1159,"prompt_tokens":991,"completion_tokens":168,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":607,"completion_tokens_details":{"reasoning_tokens":101}},"tokens_in":607,"tokens_out":168,"duration_ms":2891,"temperature":1.0,"reasoning_tokens":101,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T14:33:14.895557+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Use a synthetic dataset with known ground truth where each paired frame carries the same scene plus independent noise, but inject a small, known, slowly varying shift in one frame. If the trained network output attenuates or removes that injected change, the zero-mean assumption has been violated and the method cannot separate real temporal evolution from noise. A realistic version applies the network to two real frames separated by a known flow and compares the output in the moving region with a third, temporally averaged reference.","supporting_citations":[],"review_version":1}