{"id":"a65e7a9d-e01e-4122-86e8-856155cfb0a2","arxiv_id":"2507.01943","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A ResNet and U-Net pipeline trained on simulated JWST images can detect and locate strong lenses with Einstein radii down to 0.03 arcseconds, forecasting about 17 detectable low-halo-mass lenses per square degree.","lead":"This paper trains AI on simulated telescope images to find tiny gravitational lenses, down to a few hundredths of an arcsecond, that human eyes would miss. It forecasts that JWST can find about 17 low-mass galaxy lenses per square degree, a new regime that could test dark matter models.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claimed 100% precision for 1.1/deg² low-mass lenses is an in-sample validation artifact: thresholds are chosen on the same ~2000 non-lenses, and zero false positives there cannot bound the deployment false-positive rate below the paper's own ~1/3800 lens prior.","rationale":"I took the central claim to be the forecast of 1.1 (and 7.0) low-mass lenses per square degree at 100% (99%) precision. The reader's weakest assumption is simulation-to-real transfer and the Z18 mass-velocity dispersion relation; I agree those are real uncertainties, but the tighter problem is the statistical basis of the purity claim. It is checkable without new observations, and the paper itself contains the ingredients to show the problem: at the 0.95 threshold the authors estimate 10 false positives per 1.8 true positives in a field prior, so 'zero false positives' cannot be inferred from a small validation set. A held-out test set and a binomial upper-bound calculation would settle whether the precision claim survives. The numerical inconsistency between Table 8 and Table 9 (240 vs 490 forecast lenses/deg^2 in the same theta_E bin) reinforces the need to recompute the forecast tables before accepting the headline rates. None of this invalidates the interesting model-transfer results for conventional lenses or the two HST candidates; it says the low-mass purity/rate headline needs revision, which is what conditional acceptance is for. I therefore keep the reader's conditional verdict.","tokens_in":43252,"tokens_out":10906,"duration_ms":118138,"concrete_test":"Fix thresholds p=0.96 and q=0.72, then evaluate the RUN pipeline on a freshly generated simulated JWST test set that was not used in any threshold or model selection (or use 10-fold cross-validation repeated over threshold choices). Count false positives among the non-lens cutouts; if zero in N=2000, compute the 95% binomial upper bound on FPR and fold it into the paper's 1/3800 lens prior to obtain an expected precision for a real 10,000-source field. As part of the same reanalysis, independently recompute Table 9 from Table 2, Figure 19, and Figure 21 to resolve the 240 vs 490 discrepancy. Report expected precision rather than validation precision; if it is below 99%, revise the abstract and Table 9 claims.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing part of the headline claim is the '100% precision' attached to the 1.1 and 7.0 deg^-2 forecasts. This precision is not an expected field precision; it is the validation-set precision at thresholds (ResNet 0.96, U-Net 0.72) that were chosen after looking at that same validation set. With ~2000 Type-2 non-lenses, zero false positives gives a 95% Poisson/binomial upper limit of about 3.7 FPs, i.e. FPR ~0.0019. Using the paper's own 1 lens per 3800 sources prior, expected FPs per detected lens are of order 10, not 0. Indeed Section 3.2.2.4 shows a 5.56:1 FP:TP ratio at threshold 0.95, yet the abstract and Table 9 quote zero false positives. The 1.1/deg^2 number also inherits an internal inconsistency: Table 9 lists 490 forecast lenses/deg^2 for 0.02''<theta_E<0.05'', while Table 8 (and the difference of the first two rows of Table 2) gives 240 total lenses/deg^2 for that bin. These two issues mean the headline purity and rate are not yet demonstrated.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper forecasts the number of strong gravitational lenses detectable with JWST, simulates lensed and unlensed images using CosmoDC2, JAGUAR/W18, and VELA sources, trains a shielded ResNet for three Einstein-radius ranges, and trains a U-Net to localize small-Einstein-radius lenses. The central new claims are that JWST can find ~17 deg^-2 lenses with 0.02''<θ_E<0.05'' and M_halo<10^11 M_sun using a ResNet, and that the combined ResNet+U-Net pipeline can localize ~1.1 deg^-2 of them at ~100% pixel-level precision (or ~7.0 deg^-2 at 99% precision). The paper also reports two new HST lens candidates missed by the Garvin et al. (2022) crowdsourcing search.","tokens_in":43508,"tokens_out":6252,"duration_ms":68060,"significance":"If the small-lens forecasts hold, this work opens a genuinely new observational regime: galaxy-scale strong lensing below θ_E≈0.03'' and M_halo≈10^11 M_sun would provide a new probe of low-mass dark matter halos. The simulation effort is unusually thorough for this type of forecast, and the k-fold VELA split, truncated-SIS test, environmental-galaxy variation, and real-HST validation for conventional lenses are genuine strengths. The two new HST candidates and the demonstration that Model 1a separates real lenses from non-lenses are concrete, falsifiable results that support the conventional-lens part of the paper. The small-lens part is scientifically important but currently rests on in-sample validation and unpropagated systematics.","major_comments":[{"comment":"The forecast base for the headline bin is internally inconsistent. Table 8 states that the 0.02''<θ_E<0.05'' bin contains 240 lenses/deg^2 with M_halo<10^11 M_sun (the difference of the first two rows of Table 2), while Table 9 lists 490 lenses/deg^2 for the same bin and the same mass cut. Because the 17 and 1.1 deg^-2 numbers are derived from this bin, the authors must determine which base is correct, correct the affected tables/figures, and recompute the predicted yields.","section":"§3.3.3, Table 9 vs. Table 8"},{"comment":"The claim of ~100% precision for the 1.1/deg^2 (and 7.0/deg^2 at 99%) forecast is a validation-set zero-count statement, not a statistically supported field-deployment bound. With ~2000 Type-2 non-lenses, zero false positives gives a 95% upper limit of order 3.7 false positives (FPR≈0.0019); combined with the paper's own 1/3800 lens prior and a recall of ~0.67, the expected precision at the chosen ResNet threshold is not 100%. The paper itself computes a 5.56:1 FP:TP ratio at threshold 0.95 in §3.2.2.4, and the 0.96 and 0.72 thresholds were selected after inspecting the same validation set. The abstract and Table 9 should report a confidence interval for precision or an expected field precision under the stated prior, not a point value of 100% from zero counts.","section":"§3.2.2.4, §4.2.2, Table 9, Abstract"},{"comment":"The headline yields (17, 7.0, and 1.1 deg^-2) are point estimates with no propagated uncertainty. The M_halo–σ relation of Zahid et al. (2018) is unconstrained below σ≈100 km/s, as the paper acknowledges in §2.3, and θ_E scales as σ^2; plausible changes in the normalization, slope, scatter, or low-σ behavior can shift the small-θ_E counts by an order of magnitude. The authors should provide credible intervals by varying the Z18 relation parameters, its scatter, and the VDF assumptions, and propagate these through Figure 21 and the U-Net recall factors.","section":"§2.1.3, §2.3, §3.3.3"},{"comment":"The small-lens completeness and purity used in the forecast are measured on simulations generated with the same source catalog (VELA), lens profile (SIE), PSF model, and noise model used to construct the forecast, and no real JWST small-θ_E validation exists. The k-fold VELA split and truncated-SIS tests are useful robustness checks, but they do not constrain performance on real JWST images with realistic PSF structure, blending, and source morphologies at θ_E≈0.03''. The deployment rates should be explicitly framed as simulation-based predictions, and a concrete validation path—such as a blinded injection test on real JWST imaging or a pilot search with spectroscopic follow-up—should be identified.","section":"§3.1, §3.2.2, §4.2, §3.3.3"}],"minor_comments":[{"comment":"The exposure times for Model 2a and Model 2b appear swapped in the text: Table 3 assigns 10,000 s to Model 2a (JWST-long) and 1,000 s to Model 2b (JWST-short), but §3.1.5 states texp=1000 s for Model 2a and texp=10,000 s for Model 2b.","section":"§3.1.5"},{"comment":"Table 6 lists the recall at probability threshold 0.5 as 0.9034, which is an outlier relative to the neighboring thresholds (0.9824 and 0.9794) and inconsistent with the text's statement of 0.980 recall at a 0.5 threshold; this is likely a typographical error and should be corrected.","section":"§3.2.1.1, Table 6"},{"comment":"The deployment calculation contains a wording error: after estimating 1.8 true positives and 10 false positives, the text says '1.8 are false positives,' but the context indicates it should read '1.8 are true positives.'","section":"§3.2.2.4"},{"comment":"The text claims zero false positives for the U-Net above threshold 0.72 in one place, while §4.2.2 reports one false positive at that threshold and Figure 25 shows a red point above the 0.72 line; the 'zero' claim should be restricted to the ResNet-filtered subset used for the final pipeline, or the count should be stated consistently.","section":"§4.2.2 and §4.3"},{"comment":"The abstract describes going down to θ_E≲0.03'', but Figure 14 shows that recall degrades sharply toward θ_E=0.02'' (roughly 40% at threshold 0.5); the completeness caveat at the lowest Einstein radii should appear near the headline claim.","section":"Abstract and §3.2.2.2"}],"recommendation":"major_revision","confidential_remarks":"The conventional-lens part of this paper is solid and the two HST candidates are a nice result. My recommendation hinges on the small-lens forecast: the headline 100% precision and the 1.1/deg^2 rate are not yet demonstrated because of the validation-selection issue, the Table 8/Table 9 inconsistency, and the lack of propagated uncertainties in the M_halo–σ relation. These are fixable in revision, so I do not recommend rejection, but the paper should not be accepted with the current headline numbers."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things about arXiv:2507.01943. First, it is genuinely new: the first ML search designed for strong lenses with Einstein radii below 0.1 arcsec on simulated JWST images, forecasting about 17 detectable low-halo-mass lenses per square degree and 1.1 per square degree localizable by a U-Net at claimed 100% precision. It also reports two new HST lens candidates missed by the Garvin et al. crowdsourcing project. Second, the headline purity numbers are not as solid as the abstract suggests, and there is an internal inconsistency in the forecast tables that needs fixing.\n\nWhat the paper does well: the simulation pipeline is careful, using VELA sources, CosmoDC2 lenses, realistic noise, and a sensible definition of non-lenses. The robustness tests are above average for this literature: k-fold splits on VELA, truncated SIS profiles, environmental galaxy variations, and a real HST test for conventional lenses. The conventional-lens results, including the two new candidates, look credible. The authors also acknowledge some obvious limitations, such as the uncertain velocity-dispersion function below 100 km/s and the small validation set.\n\nThe soft spots are real. The claimed 100% precision is an in-sample artifact: the ResNet and U-Net thresholds were chosen on the same validation set where zero false positives were seen among about 2000 non-lenses. That zero cannot bound the deployment false-positive rate; the paper's own calculation at a 0.95 threshold gives a 5.6:1 false-positive-to-true-positive ratio, and the 0.96 threshold is only one step away. A proper binomial upper limit on the validation FPR would put the expected contamination at order ten per detected lens, not zero. The abstract and conclusions should be tempered, or the claims presented with uncertainty.\n\nThere is also a numerical inconsistency: Table 9 lists 490 forecast lenses per square degree in the 0.02-0.05 arcsec bin, while Table 8 (and the difference of the first two rows of Table 2) gives 240. This is a factor of two in the forecast itself, so the 17/deg^2 and 1.1/deg^2 numbers inherit that ambiguity. The authors need to reconcile this before the paper is taken at face value.\n\nFinally, the Z18 mass-velocity-dispersion relation is unconstrained below sigma~100 km/s, and the forecast has no propagated error bars; an order-of-magnitude shift in the 17/deg^2 number is not implausible. That does not invalidate the approach, but it means the paper should be read as a proof of concept with illustrative numbers, not a precise prediction.\n\nBottom line: this deserves a serious referee, but it needs major revision — fix the tables, add realistic uncertainties to the purity and rate claims, and either release code/weights or show more out-of-sample validation. I would send it to review, and I would want to see the revised version before citing it.","headline":"A careful simulation-driven forecast of low-mass strong lenses with JWST, but the headline 100% purity is not demonstrated and one table contradicts another.","tokens_in":44109,"tokens_out":3782,"would_cite":false,"duration_ms":38999,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that machine-learned classifiers trained on realistic simulations can find gravitational lenses with Einstein radii down to about 0.03 arcseconds and halo masses below 10^11 solar masses, a regime that is effectively…","keywords":["strong gravitational lensing","JWST","machine learning","ResNet","U-Net","low-mass dark matter halos","Einstein radius","CDM tests"],"falsifier":"Run the trained RUN pipeline on a few square degrees of deep JWST imaging; if the yield of localized lenses with halo mass below $10^{11}M_\\odot$ and Einstein radii between $0.02''$ and $0.05''$ is far below 1.1 per square degree at the claimed precision threshold, the forecast and transferability assumption fail.","tokens_in":43026,"feed_emoji":"🔭","tokens_out":7057,"duration_ms":73266,"temperature":0.7,"pith_summary":"The paper argues that machine learning, trained on simulations, can extend strong gravitational lens searches into a new regime: lenses with Einstein radii as small as 0.03 arcseconds, produced by halos below $10^{11}M_\\odot$, which human inspectors cannot see. It forecasts that JWST can find about 17 such low-halo-mass lenses per square degree, and that a combined ResNet-plus-U-Net pipeline can localize about 1.1 per square degree at 100 percent precision, or about 7 per square degree at 99 percent precision. For conventional lenses ($\\theta_E>0.5''$), the same ResNet reaches near-100 percent completeness and purity on simulated and real HST images, and it found two HST lens candidates that a crowdsourced human search missed. The payoff is a new way to test cold dark matter, because halo abundance below the typical galactic scale is where CDM predictions are largely untested.","feed_headline":"Machine learning finds gravitational lenses down to 0.03 arcseconds","feed_subtitle":"A two-stage ResNet and U-Net pipeline localizes 1.1 low-mass lenses per square degree at 100 percent precision.","key_machinery":"The machinery is a two-stage pipeline. First, a 'shielded' ResNet classifier scans image cutouts and assigns each a lens probability; it was trained on thousands of simulated images built from CosmoDC2 lens halos, VELA hydrodynamic source galaxies, and Sersic/environmental galaxies, with HST or JWST noise added during training. Second, a U-Net segmentation model takes ResNet-positive cutouts and predicts a per-pixel lens-location probability, so the lens position can be pinpointed. The forecasts that set the discovery rates use the CosmoDC2 halo catalog plus the Zahid et al. (2018) halo-mass-to-velocity-dispersion relation to convert halo masses into Einstein radii, and count lens-source overlap with three geometric methods: source in lens area, lens in source area, or either.","core_discovery":"The central claim is that strong lens detection is no longer limited by human visual inspection or by Einstein radii large enough to show visible arcs: a shielded ResNet classifier trained on simulations with realistic hydrodynamic sources can identify lenses down to JWST's diffraction limit, and a U-Net can pinpoint their locations. On simulated JWST data with 10,000-second exposures, the ResNet achieves high classification accuracy for lenses with $0.02''<\\theta_E<0.15''$, and the combined RUN pipeline reaches 100 percent pixel-level precision at a 0.96 ResNet threshold and a 0.72 U-Net threshold. Applied to real HST data, the conventional-lens model classified every test lens as a lens with high probability and identified two candidates missed by a crowdsourced citizen-science search. The authors conclude that the bottleneck is no longer finding candidate lenses but confirming them spectroscopically.","pith_inferences":["If real JWST images match the simulation inputs, then existing deep JWST surveys already contain these small lenses, and rerunning the trained pipeline on archival data is a direct test of the forecast.","Because the halo-mass-to-velocity-dispersion relation is unconstrained below about 100 km/s, the forecast rates could shift by an order of magnitude; measuring that relation from dwarf-galaxy kinematics would firm up or revise the numbers.","The same approach may extend to even smaller, effectively dark halos below $10^{10}M_\\odot$ with higher-resolution instruments, since the lensing signal is not obscured by lens light.","Multi-band imaging, by separating lens and source light, could help confirm small-Einstein-radius candidates and improve precision beyond single-band detection."],"forward_implications":["Near-100 percent completeness and purity for conventional lenses means space-based surveys with HST, JWST, Roman, and Euclid can automate lens discovery instead of relying on human scanning.","JWST should reveal about 17 lenses per square degree with halo mass below $10^{11}M_\\odot$ and Einstein radii $0.02''$ to $0.05''$, with about 1.1 per square degree localized by the RUN pipeline at zero false positives.","The two HST candidates show that even comprehensive crowdsourced searches miss discoverable lenses; machine learning can recover such systems.","The low-mass lens population provides a new observational handle on CDM predictions for halo abundance below about $10^{11}M_\\odot$."],"supporting_citations":[{"why":"Supplies the halo-mass to stellar velocity dispersion relation used to convert CosmoDC2 halo masses into Einstein radii for the forecasts and simulations.","marker":"Zahid et al. (2018)"},{"why":"Provides the CosmoDC2 lens catalog of halo masses and redshifts down to $10^{10}M_\\odot$.","marker":"Korytov et al. (2019)"},{"why":"Defines the JAGUAR source catalog and JWST depth used for the lens-count forecast and for source light in simulations.","marker":"Williams et al. (2018)"},{"why":"Describes the VELA hydrodynamic simulations that supply realistic high-redshift source galaxies for the lensed images.","marker":"Simons et al. (2019)"},{"why":"Provides the VELA simulation mock images used as lensed source galaxies.","marker":"Snyder (2018)"},{"why":"Supplies the shielded ResNet architecture and training scheme that the paper adapts for each Einstein-radius model.","marker":"Huang et al. (2021)"},{"why":"Establishes the ResNet approach for strong lens finding on simulated Euclid data, the baseline this work extends to lower radii.","marker":"Lanusse et al. (2018)"},{"why":"LensPop forecast code and velocity-dispersion comparison set the baseline for the JWST lens-count forecast.","marker":"Collett (2015)"},{"why":"The crowdsourced HST lens catalog whose missed systems provide the two new candidate discoveries.","marker":"Garvin et al. (2022)"},{"why":"Introduces the U-Net architecture used to localize small lenses in cutout images.","marker":"Ronneberger et al. (2015)"}],"fun_headline_variants":["AI spots gravitational lenses down to 0.03 arcseconds","Machine learning finds tiny lenses missed by humans","JWST plus AI detects lenses at 0.03 arcseconds","Machine learning discovers tiny gravitational lenses with JWST"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The forecast numbers rest on the assumption that the simulated images, built from VELA sources, CosmoDC2 lenses, SIE mass profiles, and the JWST noise model, capture the appearance of real small-Einstein-radius lenses, and that the unmeasured halo-mass-to-velocity-dispersion relation below about 100 km/s used to convert halo mass into Einstein radius is roughly right.","fun_headline_variants_meta":{"raw":{"variants":["AI spots gravitational lenses down to 0.03 arcseconds","Machine learning finds tiny lenses missed by humans","JWST plus AI detects lenses at 0.03 arcseconds","Machine learning discovers tiny gravitational lenses with JWST"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000906,"raw_usage":{"total_tokens":4005,"prompt_tokens":1165,"completion_tokens":2840,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":781,"completion_tokens_details":{"reasoning_tokens":2775}},"tokens_in":781,"tokens_out":2840,"duration_ms":26315,"temperature":1.0,"reasoning_tokens":2775,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T20:40:15.931397+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the trained RUN pipeline on a few square degrees of deep JWST imaging; if the yield of localized lenses with halo mass below $10^{11}M_\\odot$ and Einstein radii between $0.02''$ and $0.05''$ is far below 1.1 per square degree at the claimed precision threshold, the forecast and transferability assumption fail.","supporting_citations":[{"cited_title":"J., Sohn , J., & Geller , M","cited_arxiv_id":null,"evidence_quote":"Supplies the halo-mass to stellar velocity dispersion relation used to convert CosmoDC2 halo masses into Einstein radii for the forecasts and simulations."},{"cited_title":"C., Kassin , S","cited_arxiv_id":null,"evidence_quote":"Describes the VELA hydrodynamic simulations that supply realistic high-redshift source galaxies for the lensed images."},{"cited_title":"2018, Vela- Sunrise Mock Observations , doi:10.17909/t9-ge0b-jm58","cited_arxiv_id":null,"evidence_quote":"Provides the VELA simulation mock images used as lensed source galaxies."}],"review_version":1}