{"id":"8132122f-b880-453b-9ba8-0852ee554b96","arxiv_id":"1908.06344","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"Narrow-line flux measurements for eight quadruply lensed quasars double the compact-source lens sample and show flux ratios that smooth lens models cannot reproduce, indicating small-scale dark matter structure.","lead":"This paper measures the narrow-line light of eight gravitationally lensed quasars and shows that ordinary smooth lens models cannot explain the relative brightness of the lensed images. The mismatch suggests small clumps of dark matter are bending the light, and the new measurements double the number of systems available for this kind of dark matter test.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Figure 4's p<0.005 is not a calibrated test: best-fit-only position draws, ignored flux-ratio covariances, and non-Gaussian marginals mean the abstract's 'demonstrate' claim needs a proper joint posterior predictive check.","rationale":"The paper's measurement program is valuable: it presents new narrow-line flux ratios for eight lenses, uses a documented forward-modeling pipeline (grizli), evaluates source-size systematics for HS 0810 in Section 7, and explicitly defers dark-matter inference to a companion paper. I read the central claim as the abstract's assertion that smooth models fail and that the discrepancy indicates small-scale dark matter structure, with Figure 4's p<0.005 as its in-paper support. That support is not statistically calibrated, for the reasons in the attack. The reader's identified weakest assumption (NLR size and centroid) is a real secondary concern, but it is not the most load-bearing issue: even granting that assumption, the p-value as computed does not establish the claim. I therefore keep the verdict conditional, but the condition should be stronger than the reader's proposed abstract correction: either provide a calibrated joint posterior predictive test with a defensible p-value, or soften the abstract to say the data are inconsistent with the smooth-model flux-ratio distribution under a simplified comparison and that quantitative dark-matter inference is deferred. The source-size/position-offset assumption should also be acknowledged as a caveat in the revised abstract.","tokens_in":26838,"tokens_out":10866,"duration_ms":119991,"concrete_test":"Recompute the Section 6.2/Figure 4 comparison from the existing lenstronomy setup: for each lens, run an MCMC (or importance sampling) over the macromodel parameters using the same priors and the same 0.005-arcsec position likelihood, generate the joint posterior predictive distribution of all flux ratios for that lens, and compute a posterior predictive p-value with a joint test statistic (e.g., likelihood-ratio or Mahalanobis distance that includes flux measurement covariance). If the calibrated p-value is no longer below 0.05, the headline dark-matter-discrepancy claim is unsupported by this paper; if it remains below 0.005, the central conclusion survives this objection.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The quantitative basis for the abstract's central claim is the p<0.005 from Section 6.2 and Figure 4. That p-value is obtained by, for each lens, drawing 1.4e4 sets of image positions, fitting each with the best-fit macromodel parameters, and then comparing the 1D marginalized flux-ratio posteriors with the measured flux ratios as if they were independent chi-square variates with one degree of freedom. This is not a calibrated posterior predictive test. Taking only the best-fit parameters for each position realization does not explore the full macromodel posterior given the positions, so the model-predicted flux-ratio distribution is likely too narrow. The several flux ratios within a lens are correlated through the shared macromodel and source, so treating them as independent observations overstates the effective sample size. The appendix contours show strongly non-Gaussian and asymmetric model marginals, for which the 1-dof chi-square distribution is not the correct reference. The authors explicitly note that the comparison ignores covariance and non-Gaussianity. Even if the narrow-line source assumption is correct, this uncalibrated p-value cannot support the abstract's wording that smooth models 'fail' and that the discrepancy 'indicates' dark matter substructure; that requires a calibrated joint posterior predictive test (or weakening the claim).","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents WFC3/IR grism observations of eight quadruply imaged quasar lenses and uses a forward-modeling spectral extraction pipeline to measure narrow-line ([OIII] or [NeIII]) flux ratios with reported uncertainties of 2–10%. The narrow-line fluxes are intended to provide compact-source flux ratios that are free of stellar microlensing. The paper then fits smooth power-law ellipsoid plus external shear lens models to the image positions only, derives model-predicted flux-ratio distributions, and compares them with the measured narrow-line flux ratios. The authors report that the smooth models fail to reproduce the observed flux ratios with p<0.005 and interpret this as evidence for small-scale dark matter structure, with the detailed dark-matter interpretation deferred to a companion paper. HS 0810 and SDSS J1330 are excluded from the comparison, so Figure 4 is based on six lenses. Section 7 contains a resolved-source analysis for HS 0810 indicating a narrow-line source size of order tens of parsecs.","tokens_in":27167,"tokens_out":5735,"duration_ms":63492,"significance":"If the statistical claim is supported, the paper would meaningfully expand the compact-source lens sample for dark-matter substructure studies, doubling the number of systems relative to the radio-loud sample and demonstrating a viable optical path for the technique. The spectral extraction is a genuine advance: it forward-models the 2D grism frame in the native FLT frame, accounts for blending, tests multiple FeII and H-beta templates, and validates the pipeline against alternative models. The use of public tools (grizli, lenstronomy) and the presentation of per-lens model and data flux-ratio contours are strengths. The central interpretive claim, however, rests on a simplified chi-square comparison that is not a calibrated posterior predictive test, and the paper's own caveats are in tension with the abstract's wording. The measurement campaign and the individual flux-ratio measurements are valuable regardless, but the paper in its current form overstates what the statistical test demonstrates.","major_comments":[{"comment":"The p<0.005 result is not a calibrated posterior predictive test. For each of the 1.4×10^4 position draws the procedure uses only the best-fit macromodel parameters, so the model-predicted flux-ratio distribution does not marginalize over the macromodel posterior and is likely too narrow; several flux ratios within one lens are correlated through the shared macromodel and source parameters, yet they are combined as independent one-degree-of-freedom chi-square variates, inflating the effective sample size; and the appendix contours show strongly asymmetric, non-Gaussian model marginals for which the 1-dof chi-square reference is not appropriate. The text acknowledges that covariance and non-Gaussianity are ignored and asserts that this 'under-represents' the discrepancy, but that directional claim is not demonstrated and the opposite can hold. The abstract's wording that smooth models 'fail' and that the discrepancy 'indicates' dark matter substructure therefore goes beyond what the statistic supports. A joint posterior predictive test, or a substantially weakened claim, is needed before publication.","section":"§6.2"},{"comment":"The abstract states that the smooth models fail to produce the observed flux distribution 'over the entire sample of lenses,' but the comparison in Figure 4 and §6.2 excludes HS 0810 and SDSS J1330 and therefore uses six of the eight lenses. The individual exclusions are motivated (blended high-magnification fold images for HS 0810; a disk galaxy for SDSS J1330), but the sample-wide wording is inaccurate. The abstract and summary should state explicitly that the test is based on six systems, and the reported p-value should be presented together with that sample definition.","section":"Abstract"},{"comment":"The dark-matter interpretation assumes that the narrow-line emission is unresolved at grism resolution, free of microlensing, and centered on the quasar continuum position used in the lens model. The paper itself notes that the Müller-Sánchez et al. (2011) sample is small and that high-redshift, luminous quasars may have different narrow-line region sizes or centroid offsets. The resolved-source test in §7 is performed only for HS 0810, which is excluded from the main comparison, and no analogous test is presented for the six lenses that drive the p-value. A population of somewhat extended or offset narrow-line regions could produce flux-ratio anomalies that mimic the dark-matter signal, so this systematic should be quantified, or at minimum explicitly budgeted, before the abstract's 'indicates' claim is made.","section":"§4"}],"minor_comments":[{"comment":"The G141 grism wavelength range is given as 0.8–1.15 µm, which is the same range listed for G102; WFC3 G141 covers approximately 1.1–1.7 µm. Please correct this typo.","section":"§3"},{"comment":"The sentence 'the chi2 values should be Gaussian with one degree of freedom' should read that the chi-square values should follow a chi-square distribution with one degree of freedom.","section":"§6.2"},{"comment":"The statistic used to obtain p<0.005 is not identified; if it is a Kolmogorov-Smirnov or Anderson-Darling test, name it and report the test statistic so the reader can assess the comparison.","section":"Figure 4"},{"comment":"The resolved-source comparison for HS 0810 reports log-likelihood improvements without a formal correction for the three fewer degrees of freedom or for the noise introduced by the drizzling/blotting procedure; the text discusses the latter qualitatively, but an information-criterion-style comparison would make the source-size constraint easier to evaluate.","section":"§7"}],"recommendation":"major_revision","confidential_remarks":"The paper's data product is valuable and the measurement pipeline appears carefully constructed, but the headline statistical claim ('demonstrate that smooth models fail') is currently supported only by an uncalibrated chi-square comparison. Because the companion paper (Gilman et al. 2019a) is framed as using these data to constrain dark matter models, the statistical calibration issue matters beyond this paper's abstract. A major revision that either supplies a proper joint posterior predictive check or clearly downgrades the claim to 'inconsistent under a simplified test' seems appropriate; I would not reject, since the narrow-line flux measurements themselves are the main contribution."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The real product here is the data: narrow-line flux ratios for eight lenses, doubling the compact-source sample available for dark matter work. That alone makes the paper worth serious attention. The spectral extraction is careful, forward-modeled in the native FLT frames, tested against alternative FeII and H-beta templates, and the lens model is fitted to image positions only, so the flux ratios are not circularly fitted. The resolved-source test for HS 0810 is a nice extra. The body of the paper is honest about methodology.\n\nThe soft spots are mostly in the abstract. The Section 6.2 comparison that yields p<0.005 is explicitly simplified: best-fit macromodel parameters per position draw, one-dimensional marginal flux-ratio posteriors treated as independent chi-square variates, with the authors themselves noting the ignored covariance and non-Gaussianity. That is not a calibrated posterior predictive check, and the abstract's word 'demonstrate' is stronger than the method supports. Also, the abstract says 'entire sample' but two of the eight lenses (HS 0810 and SDSS J1330) are excluded from that comparison for good reasons; that phrasing should be fixed. The narrow-line region size assumption is a real systematic, but the paper discusses it honestly, and the HS 0810 test suggests sizes of 10-100 pc; for the other lenses the magnifications are lower, so it is probably not a fatal issue.\n\nThis is primarily a data paper. Its value is the new measurements and the pipeline improvements. The rough lens model comparison is a motivation for the companion paper, not a standalone detection. A reader working on flux-ratio lensing will want this. The technical content deserves peer review, but the abstract needs to be aligned with the caveats.\n\nI would send it to review, asking the authors to correct the 'entire sample' wording, soften 'demonstrate' to something like 'are inconsistent with' or 'disfavor', and ideally add a calibrated posterior predictive check or at least label the existing test as a rough diagnostic. The measurements themselves are solid and the field needs them.","headline":"Solid new flux-ratio measurements double the compact-source lens sample, but the abstract's smooth-model rejection claim is not supported by the paper's own simplified p-value.","tokens_in":27793,"tokens_out":2747,"would_cite":true,"duration_ms":27455,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Narrow emission lines from lensed quasar nuclei provide a microlensing-free way to measure image fluxes, doubling the compact-source lens sample and showing that smooth lens models fail to explain the observed flux ratios, pointing to…","keywords":["gravitational lensing","dark matter substructure","quasar narrow-line region","flux ratios","WFC3 grism spectroscopy","microlensing","quadruply imaged quasars","halo mass function"],"falsifier":"A decisive check would be diffraction-limited integral-field spectroscopy of one of the six comparison lenses, resolving the narrow-line region and measuring its centroid relative to the quasar continuum. If the region is found to be larger than roughly 100 pc, or offset from the continuum by more than about 10 pc, the measured flux ratios could be diluted or mis-centred, and the $p<0.005$ smooth-model discrepancy would no longer uniquely indicate dark matter.","tokens_in":26650,"feed_emoji":"🔭","tokens_out":8675,"duration_ms":85126,"temperature":0.7,"pith_summary":"The paper argues that the narrow forbidden-line emission of quasars can replace radio jets as a microlensing-free probe of dark matter substructure in strong gravitational lenses, and it delivers the measurements that make the case. For eight quadruply imaged quasars, narrow-line fluxes are measured with 2–10% uncertainties, doubling the number of compact-source lenses suitable for this analysis. Fitting the image positions with smooth mass models and comparing the predicted flux ratios to the measured ones rejects the smooth models at $p<0.005$, with deviations larger than macromodel uncertainties. The paper reads this as evidence for perturbations by low-mass dark matter halos along the line of sight, with the quantitative dark matter interpretation left to a companion paper.","feed_headline":"Narrow quasar lines double dark matter lens probes","feed_subtitle":"These flux ratios, immune to microlensing, rule out smooth lens models, pointing to dark matter substructure.","key_machinery":"The central object is the quasar narrow-line region, traced by forbidden lines such as [OIII] and [NeIII]: at milliarcsecond scales it is too large to be significantly microlensed by stars, which act on microarcsecond scales, yet compact enough to be treated as a point source at grism resolution and centred on the continuum image position used for lens modelling. The measurement machinery is a forward-modelling spectral extraction that builds a full model of the two-dimensional grism image, including quasar point sources, the lens galaxy, the lensed quasar host, continuum, broad FeII and Balmer emission, and the narrow lines, and fits it in the native detector frames. The statistical machinery is a flux-ratio posterior: image positions are drawn from their measured uncertainties, a smooth power-law ellipsoid plus external shear macromodel is solved for each draw, and the predicted flux ratios are compared with the measured narrow-line flux ratios through a chi-square test.","core_discovery":"The central claim is that smooth lens models fail to describe the narrow-line flux ratios of a sample of quadruply imaged quasars, and that this failure is a signal of small-scale dark matter structure. After measuring [OIII] 4959/5007 Å and [NeIII] 3869/3969 Å narrow-line fluxes in eight systems with WFC3/IR grism spectroscopy, the authors fit the quasar image positions with flexible power-law ellipsoid mass models plus external shear and found that the flux ratios predicted by those models disagree with the measured ratios. The statistical comparison, made on six lenses (excluding HS 0810, whose fold images blend at magnifications near 120, and SDSS J1330, whose disk requires extra macromodel complexity), rejects the smooth-model flux ratio distribution at $p<0.005$, with typical deviations larger than expected from macromodel uncertainties. The authors interpret this as evidence for perturbations from low-mass dark matter halos along the entire line of sight.","pith_inferences":["Beyond the paper, comparing these narrow-line flux ratios with mid-infrared or radio continuum flux ratios for the same lenses would isolate any residual source-size or dust-extinction effects and independently test the dark matter interpretation.","A testable extension: re-observing a few lenses at a later epoch should leave narrow-line flux ratios unchanged even while continuum ratios vary; any epoch-dependent narrow-line variation would point to contamination rather than dark matter.","The resolved-source analysis of HS 0810 hints that high-magnification fold pairs can serve as physical-size measurements of high-redshift narrow-line regions, turning a discarded system into a useful probe.","Coupling the narrow-line flux ratios with lensed host-galaxy arc constraints in the macromodel fit would tighten the predicted flux-ratio posterior and, in principle, lower the halo mass scale the method can detect."],"forward_implications":["The usable sample of compact-source lenses for dark matter studies roughly doubles, from about seven radio-loud systems to around fifteen including the eight narrow-line lenses presented here.","Because narrow-line ratios are insensitive to stellar microlensing, the discrepancy with smooth models is attributed to low-mass halos rather than to stars in the lens galaxy.","Five of the eight lenses show large differential magnification between broad and narrow emission, directly confirming that the narrow lines are the microlensing-free component and that the continuum/broad lines are microlensed.","With 2–10% flux measurement precision, the sample approaches the ~4% precision level at which simulations indicate that roughly 10–40 lenses can rule out a 3.3 keV warm dark matter particle, materially strengthening the statistical reach.","The same pipeline can be applied to the growing number of quasar lenses from wide-field surveys, which are forecast to contain thousands of such systems in the coming decade."],"supporting_citations":[{"why":"Establishes the theoretical basis for using narrow-line emission as a microlensing-free compact source in gravitational lens flux-ratio studies.","marker":"Moustakas & Metcalf 2003"},{"why":"Provides measured local narrow-line region sizes (FWHM roughly 10-60 pc), justifying the milliarcsecond source-size assumption and the 20-50 pc prior.","marker":"Müller-Sánchez et al. (2011)"},{"why":"Demonstrates the WFC3 grism narrow-line flux-ratio method on the lens HE 0435 and constrains the narrow-line source size to be point-like; the direct precursor of this pipeline.","marker":"Nierenberg et al. (2017)"},{"why":"Foundational flux-ratio analysis showing compact-source lenses are sensitive to low-mass dark matter halos, motivating the sample.","marker":"Dalal & Kochanek (2002)"},{"why":"Simulated flux-ratio lens samples to show that roughly 10-40 lenses at ~4% precision can rule out 3.3 keV warm dark matter, quantifying why doubling the sample matters.","marker":"Gilman et al. (2019b)"},{"why":"The radio-loud compact-source lens sample of seven systems whose size this paper doubles; supplies the warm dark matter constraint this method aims to improve.","marker":"Hsueh et al. (2019)"},{"why":"Companion paper that interprets the flux-ratio discrepancy in terms of dark matter halo population models, completing the argument.","marker":"Gilman et al. (2019a)"},{"why":"Supplies the lens modelling software used to fit image positions with smooth mass models and compute predicted flux ratios.","marker":"Birrer et al. (2015); Birrer & Amara (2018)"}],"fun_headline_variants":["Narrow-line fluxes double dark matter lens probes","Smooth models fail narrow-line test, dark matter substructure seen","Eight lenses, no microlensing: dark matter signal from narrow lines","Narrow-line lensing doubles dark matter sightlines, rules out smooth models","Grism narrow lines expose dark matter halos in quasar lenses"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire signal rests on the assumption that each quasar's narrow-line region is compact enough (milliarcsecond scale) to be treated as point-like at grism resolution, extended enough to be free of stellar microlensing, and centred on the quasar continuum position used in the lens models.","fun_headline_variants_meta":{"raw":{"variants":["Narrow-line fluxes double dark matter lens probes","Smooth models fail narrow-line test, dark matter substructure seen","Eight lenses, no microlensing: dark matter signal from narrow lines","Narrow-line lensing doubles dark matter sightlines, rules out smooth models","Grism narrow lines expose dark matter halos in quasar lenses"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000684,"raw_usage":{"total_tokens":3159,"prompt_tokens":1055,"completion_tokens":2104,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":671,"completion_tokens_details":{"reasoning_tokens":2014}},"tokens_in":671,"tokens_out":2104,"duration_ms":16141,"temperature":1.0,"reasoning_tokens":2014,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T12:48:41.656463+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A decisive check would be diffraction-limited integral-field spectroscopy of one of the six comparison lenses, resolving the narrow-line region and measuring its centroid relative to the quasar continuum. If the region is found to be larger than roughly 100 pc, or offset from the continuum by more than about 10 pc, the measured flux ratios could be diluted or mis-centred, and the $p<0.005$ smooth-model discrepancy would no longer uniquely indicate dark matter.","supporting_citations":[{"cited_title":"A., & Metcalf, R","cited_arxiv_id":null,"evidence_quote":"Establishes the theoretical basis for using narrow-line emission as a microlensing-free compact source in gravitational lens flux-ratio studies."},{"cited_title":"M., Treu, T., Brammer, G., et al","cited_arxiv_id":null,"evidence_quote":"Demonstrates the WFC3 grism narrow-line flux-ratio method on the lens HE 0435 and constrains the narrow-line source size to be point-like; the direct precursor of this pipeline."}],"review_version":1}