{"id":"1f5f47cf-83cb-45e8-afa2-da6ea8681eb0","arxiv_id":"2510.12173","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A CNN trained on mock CANDELS images from TNG50 identifies major and minor galaxy mergers at z~1 down to 10^8 solar masses with ~73% accuracy on a held-out simulated test set.","lead":"Astronomers trained a neural network on simulated Hubble images of galaxies at redshift ~1, including the small, faint systems and unequal-mass pairs that earlier merger finders mostly ignored. The network spots about 73% of simulated mergers correctly and could help build more complete merger catalogs for CANDELS and future surveys.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Mock-to-real transfer is the load-bearing gap: the 73% accuracy may partly reflect SFR/color differences in TNG50, and has not been validated on real CANDELS data.","rationale":"The paper's strongest measurable claim is the ~73% accuracy on held-out mock images, and that claim is reasonably supported by the per-galaxy split and three seeds. However, the headline scientific promise is applicability to real CANDELS. That step is untested and the paper itself flags the domain gap. The SFR analysis strengthens this concern by showing a specific confound: the network partly uses star-formation level as a proxy, and the training nonmerger sample is not SFR-matched. Therefore the reader's conditional verdict is appropriate: the mock results are credible but do not yet justify the CANDELS exploration claim. No internal inconsistency in the mock evaluation was found that would invalidate the 73% on the mock distribution, so the verdict should not move to reject; it remains conditional pending real-data validation. My concrete test is the direct check of transfer performance.","tokens_in":25829,"tokens_out":7224,"duration_ms":63444,"concrete_test":"Apply the trained CNN to real CANDELS galaxies that have independent merger classifications (e.g., the CANDELS visual classification catalog of Kartaltepe et al. 2015, or spectroscopic close-pair samples), and compute the confusion matrix at the default 0.5 threshold. If the accuracy/purity on real data is substantially lower than the mock-image 73% (e.g., <60%), the mock-to-real transfer concern is confirmed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that the CNN 'enables the exploration of mergers in CANDELS' rests entirely on performance measured on mock images built from TNG50. The paper's own analysis shows the network partially bases decisions on star formation: misclassified nonmergers have higher sSFR, misclassified mergers have lower sSFR, and the latent space tracks sSFR even in greyscale CNNs (Figures 10-11, §4.3, §5.2.3). Because the nonmerger sample is only mass-matched, not SFR-matched, the 73% accuracy/purity/completeness may partly measure the known sSFR offset between mergers and nonmergers in TNG50, rather than robust morphological merger identification. Real CANDELS contains many clumpy, high-sSFR nonmergers at z~1, so this confound could substantially reduce performance on real data. Additionally, the balanced test set yields a purity of ~74% at a 50% prior; at realistic CANDELS merger fractions (perhaps 10-20%), the same false-positive rate would drop precision to roughly 25-40% unless the decision threshold is recalibrated. The paper explicitly acknowledges in §5.3 that 'there will always be differences between mock images and real data' and defers domain adaptation to future work, so the externally useful claim remains unverified.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper trains a ResNet18 CNN, initialized with Zoobot weights, on mock HST/CANDELS three-band images of TNG50 galaxies at 1 ≤ z ≤ 1.5, spanning stellar masses 10^8–10^12.5 M_sun and merger mass ratios μ > 1:10. Merger labels are taken from SubLink merger trees, and mock images are produced with SKIRT radiative transfer, TinyTim PSFs, and real CANDELS background cutouts. The reported test-set performance on a balanced, per-galaxy held-out set is ~73% accuracy, purity, and completeness, with early-stage mergers identified at ~80% and major mergers at ~76%. The paper also analyzes the effect of orientation angle, presents Grad-CAM and UMAP interpretability results, and identifies star formation rate as a confound. The authors conclude that the network is ready to enable merger identification in CANDELS and future surveys.","tokens_in":26135,"tokens_out":6264,"duration_ms":59743,"significance":"If the reported performance transfers to real observations, this would be a valuable extension of CNN-based merger identification to lower stellar masses and higher mass ratios at z ~ 1, a regime that previous work has largely excluded. The experimental design is in several respects careful: all six viewpoints of a galaxy are kept in the same split, results are averaged over three random seeds, and standard metrics plus calibration statistics are reported. The orientation analysis and the explicit sSFR misclassification analysis are honest and informative. However, the central external-validity claim is not yet supported: all evaluation is performed on mock images from the same simulation/pipeline used for training, and the paper's own sSFR analysis indicates a potential confound that may not transfer to real CANDELS data.","major_comments":[{"comment":"The headline claim that the network 'enables the exploration of ... mergers ... in CANDELS' is supported only by metrics on mock images built from TNG50 with the same SKIRT/CANDELS-background pipeline. The authors concede in §5.3 that 'there will always be differences between mock images and real data' and defer domain adaptation to future work. The 73% accuracy/purity/completeness (Table 1) is therefore a simulation-validated figure, not a demonstrated real-data performance. Please either add a sanity check against existing CANDELS merger catalogs (e.g., visual or non-parametric classifications) or reframe the abstract/conclusion to say the classifier is 'prepared for' rather than 'enables' real-data science.","section":"Abstract; §5.3"},{"comment":"The nonmerger sample is mass-matched but not SFR-matched, and the paper's own UMAP and misclassification analysis show that the network's decisions track sSFR: misclassified nonmergers have higher sSFR and misclassified mergers lower sSFR. Because TNG50 mass-matched mergers have higher sSFR on average (Figure 1), some of the reported accuracy may reflect the network using blue/star-forming structure as a proxy for merging rather than a robust morphological merger signature. This is not circular, but it weakens the morphological claim. Please report performance stratified by sSFR (e.g., accuracy on high-sSFR nonmergers versus low-sSFR nonmergers) or evaluate on an SFR-matched test set, and discuss how much separation remains after controlling for sSFR.","section":"§4.3, Figures 10–11; §2.1"},{"comment":"All headline metrics are quoted at the default decision threshold of 0.5 on a balanced test set. At realistic CANDELS merger fractions, the same per-class error rates (completeness ≈ 0.73, false-positive rate ≈ 0.26 from Figure 4) imply purity of roughly 25–40% (e.g., ≈34% at a 15% merger fraction). The paper does not discuss prior correction, threshold selection, or the resulting purity/completeness at realistic operating points. To support the claimed applicability to CANDELS, please include precision–recall curves or recalibrate the threshold for representative merger fractions and state the expected purity at those operating points.","section":"Table 1, §4.1"},{"comment":"The abstract in the arXiv header reports overall accuracy of ~65%, while the abstract in the full text and §4/Table 1 report 73.02 ± 0.41%. This is a direct discrepancy in the central quantitative result. The inconsistency must be reconciled, and if the 65% figure refers to a different configuration it should be clearly explained.","section":"Abstract vs. Table 1"}],"minor_comments":[{"comment":"The Brier score is defined with o_t as the network's own final class ('cutoff score of 0.5'), not the true binary label. The standard Brier score compares the predicted probability to the true outcome, (p_t − y_t)^2. As written, the reported Brier score is not a proper scoring rule and its value may be misinterpreted.","section":"Eq. (4)"},{"comment":"The text says the best model is chosen by the epoch with the lowest validation loss, but the Figure 3 caption says 'lowest validation accuracy.' Please make the selection criterion consistent.","section":"§3.1 and Figure 3 caption"},{"comment":"The description of 'one mini snapshot on each side of the central full snapshot' within 250 Myr may be asymmetric in time at z = 1.5; a short clarification of the actual snapshot spacing would help reproducibility.","section":"§2.1"},{"comment":"The abstract uses q ≥ 1:10 while the text uses μ > 1:10. Please use one notation consistently and define q/μ at first use.","section":"Abstract and §2.1"},{"comment":"Typo: 'classifiy' should be 'classify'.","section":"Figure 2 caption"}],"recommendation":"major_revision","confidential_remarks":"This is a borderline major-revision case. The experiment is well-designed and the mock pipeline is a strength, but the abstract overclaims real-data applicability, and the sSFR confound and balanced-test operating point are substantive gaps that need explicit treatment. I would be open to accepting a revision that either adds a real-data sanity check or carefully scopes the claims, and that fixes the headline-number inconsistency."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things. First, the central claim — a CNN trained on TNG50 mock CANDELS images at z∼1 reaching ~73% accuracy on a held-out test set with stellar masses from 10^8 to 10^12.5 M_sun and mass ratios down to 1:10 — is solidly supported. Per-galaxy splits, three random seeds, calibration metrics, and stage/mass breakdowns are all there, and the minor-merger performance is a genuine addition. Second, the abstract overpromises. It says the network 'enables the exploration' of such mergers in CANDELS, but every evaluation is on mock images. The authors admit in §5.3 that mock and real data will always differ and defer domain adaptation to future work. That honesty is welcome, but it should be in the abstract.\n\nWhat's new is the parameter range: previous CNN merger work mostly started at 10^9.5 or 10^10 M_sun, and this pushes to 10^8 and 1:10 mass ratios at z∼1. The orientation-angle analysis — 98% of mergers correctly identified from at least one angle, 61% from most angles — and the star-formation confounding study are well done. The greyscale CNN control experiments are exactly the kind of check that should be standard.\n\nThe soft spots are real but not fatal. The sSFR issue is the biggest: the nonmerger sample is mass-matched but not SFR-matched, and Figures 10–11 show the network is partly separating mergers from nonmergers on star formation. Since real CANDELS is full of clumpy, high-sSFR nonmergers, the 73% will not transfer at face value. The balanced test set is also optimistic in a different way: at a realistic merger fraction of 10-20%, the same false-positive rate would drop precision to roughly 25-40% unless the decision threshold is recalibrated. Both are addressable. Also annoying: the abstract says 65% while the body and conclusion say 73%, and no code, data, or trained model is released, which hurts reproducibility.\n\nThe citation pattern is fine. The paper engages with the prior CNN and nonparametric work, including the numbers it compares against. This is a serious piece of work with a clear head on it.\n\nWho it's for: anyone building merger classifiers for high-z surveys, and anyone who wants a careful XAI example on a messy astronomy problem. I'd send it to review. The referee should ask that the abstract match the body, that the CANDELS-enabling claim be softened or backed by a domain-adaptation test, and that code and model weights be released. The science is worth the referee time.","headline":"Solid mock-image merger classifier with honest failure analysis, but the abstract overclaims real-CANDELS applicability and contradicts the body on the headline accuracy.","tokens_in":26706,"tokens_out":3925,"would_cite":true,"duration_ms":55713,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A neural network trained on simulated Hubble images can identify faint, minor galaxy mergers at z~1 with about 73% accuracy, a level comparable to networks trained on brighter, lower-redshift samples.","keywords":["galaxy mergers","convolutional neural networks","low-mass galaxies","minor mergers","cosmic noon","CANDELS","IllustrisTNG","mock observations"],"falsifier":"Run the trained network on a sample of real CANDELS galaxies with independent, reliable merger labels (e.g., from spectroscopic close pairs or expert visual inspections in fields not used for background cutouts) and measure whether accuracy, purity, and completeness remain near 73% for low-mass and minor mergers. Another concrete test: apply the network to mock images from a different simulation (e.g., with different subgrid physics) and see if accuracy drops; if it does, the network has learned simulation-specific morphologies rather than universal merger features.","tokens_in":25705,"feed_emoji":"🔭","tokens_out":5651,"duration_ms":51165,"temperature":0.7,"pith_summary":"The paper argues that a convolutional neural network, trained on realistic mock Hubble/CANDELS images built from a high-resolution cosmological simulation, can identify galaxy mergers at redshift ~1 that are much fainter and at lower mass ratios than previous methods could find. The network achieves about 73% accuracy, purity, and completeness, and recovers early-stage major mergers about 74-80% of the time. This matters because low-mass and minor mergers are common and likely drive much of galaxy growth at cosmic noon, yet they are easily confused with clumpy star-forming galaxies. If the network transfers to real CANDELS data, it would enable statistical studies of these previously overlooked mergers.","feed_headline":"Neural network finds minor galaxy mergers at z~1 with 73% accuracy","feed_subtitle":"Trained on simulated Hubble images, it catches low-mass and early-stage mergers that human and statistical methods miss.","key_machinery":"The essential mechanism is the mock-image pipeline: simulated galaxies are post-processed with full dust radiative transfer, filtered into three HST bands, convolved with a simulated point-spread function, and embedded in real CANDELS background cutouts, so each image has a known ground-truth label, realistic noise, and contamination by background objects. This gives the network labeled training data close to actual HST observations, letting it learn morphology-based merger signatures rather than artifacts. The CNN itself is a residual convolutional network initialized with weights from a model trained on millions of citizen-scientist galaxy classifications, then fine-tuned on the mock image","core_discovery":"The central discovery is that a CNN fed three HST bands with realistic PSF, noise, and background sources classifies mergers and nonmergers among galaxies with stellar masses from 10^8 to 10^12.5 solar masses and mass ratios down to 1:10 with ~73% balanced accuracy. The authors claim this is the first demonstration that a CNN trained on such a diverse mass and mass-ratio range performs comparably to networks trained on more massive or lower-redshift samples, with ~76% accuracy on major mergers and ~68% on minor mergers. They also find that orientation angle is a fundamental limit: 98% of mergers are identified from at least one of six viewpoints, but only 61% from the majority, meaning some","pith_inferences":["A testable extension is to apply this network (after domain adaptation) to real CANDELS galaxies and compare its predictions with spectroscopic close-pair catalogs; agreement would validate the mock-to-real transfer, while disagreement would localize where the simulation assumptions break.","The per-galaxy detectability across six viewpoints could be used as a prior to derive an orientation-unbiased merger rate from single-view observations, a step the paper leaves implicit.","The sSFR-sensitivity suggests that a real CANDELS catalog built this way will preferentially include star-forming mergers; matching nonmergers in both mass and SFR would reduce false positives but might also suppress the very merger-induced starbursts of scientific interest.","The same pipeline can be re-run with JWST, Rubin, Roman, or Euclid PSFs and filters; the paper's principle that realistic environments matter more than radiative transfer predicts that retraining on appropriate mock backgrounds is the main cost, not redoing the physics."],"forward_implications":["The method can produce a merger catalog from existing CANDELS imaging that extends to stellar masses about 100 times lower than previous CNN-based catalogs at z~1.","Such a catalog would let astronomers measure merger rates for minor mergers at cosmic noon and test whether minor mergers drive the growth of low-mass galaxies.","The orientation-angle statistics imply that no single-angle image can ever find every merger; catalogs must state incompleteness as a function of viewing angle, and multi-angle or multi-epoch data would be needed to push above the ~90% ceiling.","The network's reliance on star formation as a discriminating feature means real-data merger samples will inherit an sSFR-dependent selection function; the authors show that removing color information lowers accuracy by ~10%, so color is a genuine aid, not just a confounder."],"fun_headline_variants":["Neural net finds minor galaxy mergers at z=1","Deep learning uncovers low-mass galaxy mergers","CNN excels at early-stage galaxy mergers","First CNN to catch minor mergers at z~1","AI goes beyond brightest to find faint mergers"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing assumption is that mock CANDELS images built from TNG50 with radiative transfer are similar enough to real CANDELS data—in morphology, noise, and normalization—that the CNN's 73% accuracy on mock images will transfer to real galaxies; the paper itself notes there will always be differences between mock images and real data.","fun_headline_variants_meta":{"raw":{"variants":["Neural net finds minor galaxy mergers at z=1","Deep learning uncovers low-mass galaxy mergers","CNN excels at early-stage galaxy mergers","First CNN to catch minor mergers at z~1","AI goes beyond brightest to find faint mergers"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000625,"raw_usage":{"total_tokens":2801,"prompt_tokens":886,"completion_tokens":1915,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":630,"completion_tokens_details":{"reasoning_tokens":1845}},"tokens_in":630,"tokens_out":1915,"duration_ms":14063,"temperature":1.0,"reasoning_tokens":1845,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T09:59:02.692233+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the trained network on a sample of real CANDELS galaxies with independent, reliable merger labels (e.g., from spectroscopic close pairs or expert visual inspections in fields not used for background cutouts) and measure whether accuracy, purity, and completeness remain near 73% for low-mass and minor mergers. Another concrete test: apply the network to mock images from a different simulation (e.g., with different subgrid physics) and see if accuracy drops; if it does, the network has learned simulation-specific morphologies rather than universal merger features.","supporting_citations":[],"review_version":1}