{"id":"86ea842a-d141-4075-a272-5f45fd479059","arxiv_id":"2604.24908","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Uniform modeling of eight double quasar lenses shows that host arc surface brightness determines Fermat potential precision.","lead":"This paper performs the first uniform gravitational lens modeling of eight doubly imaged quasars from Hubble Space Telescope multi-band data using the Lenstronomy framework. It identifies arc surface brightness as the primary driver of mass model precision, supporting future expansion of time-delay cosmography samples.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"No significant objection identified","rationale":"Reader correctly flagged the tailored pipeline as the key assumption, but the manuscript's multiple independent diagnostics (conjugate analysis, hypervolume anti-correlation, external consistency checks) address it directly. No internal inconsistency or missing control appears in the described results, so the UNVERDICTED verdict and low confidence (due to abstract-only read) require no adjustment.","tokens_in":1803,"tokens_out":289,"duration_ms":33931,"concrete_test":"Extract the 8 Fermat-potential uncertainty values and arc surface-brightness measurements from the posterior tables; recompute the Spearman rank correlation and its p-value after controlling for Einstein radius via partial correlation; if the coefficient falls below 0.6 or p-value rises above 0.05 the primary-driver conclusion weakens.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim rests on a reported strong correlation between Fermat potential precision and arc surface brightness, backed by conjugate-point analysis (showing broader posteriors from quasar positions alone) and an anti-correlation with mass-parameter hypervolume. Literature comparisons yield 1.5σ agreement on Einstein radii and 3.6 mas rms on image separations, with an open-source Lenstronomy pipeline applied uniformly. These elements provide internal cross-checks that the brightness trend is not an artifact of modeling heterogeneity or unaccounted systematics.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript presents the first uniform gravitational lens modeling of eight doubly imaged quasars from HST multi-band data using a tailored open-source Lenstronomy pipeline. It reports 1.5σ average agreement on Einstein radii with literature values, 3.6 mas rms difference on image separations versus Gaia DR2, a strong correlation between Fermat potential precision and the surface brightness of the lensed host arcs, and an anti-correlation with mass-parameter hypervolume. A conjugate-point analysis using only quasar image positions yields substantially broader posteriors, confirming that the arcs provide the dominant constraints on the lens mass profiles. The work is positioned as preparation for a follow-up hierarchical cosmographic analysis to constrain H0.","tokens_in":1902,"tokens_out":602,"duration_ms":106415,"significance":"If the reported correlation holds, the result is significant for time-delay cosmography because doubly imaged systems are far more abundant than quads yet have been underutilized due to fewer constraints. Demonstrating that arc surface brightness is the primary driver of Fermat-potential precision, backed by internal cross-checks (conjugate-point posteriors and hypervolume anti-correlation) and external consistency (Einstein radii and image positions), supplies a practical selection criterion for the large samples expected from LSST, Roman, and Euclid. The uniform, open-source pipeline is a clear strength that improves reproducibility across the TDCOSMO series.","major_comments":[{"comment":"§4 (correlation analysis): The central claim that arc surface brightness is the 'primary driver' of mass-model precision rests on a reported 'strong correlation,' yet no quantitative statistic (Spearman coefficient, Pearson r with uncertainty, or p-value) or explicit fit to the eight data points is provided; without this, it is impossible to judge whether the trend is statistically robust or dominated by outliers.","section":"§4"},{"comment":"Conjugate-point analysis section: The statement that posteriors are 'substantially broader' when only quasar positions are used is presented without a quantitative metric (e.g., ratio of credible-interval widths or hypervolume ratio) or the exact number and selection of conjugate points; this leaves the confirmation that arcs dominate the constraints qualitative rather than rigorous.","section":"Conjugate-point analysis"}],"minor_comments":[{"comment":"Abstract: The phrase 'a subsequent publication' for the H0 analysis should include the planned paper number or title to aid readers following the TDCOSMO series.","section":"Abstract"},{"comment":"Figure captions (e.g., those showing posterior contours): Captions should explicitly state the number of MCMC samples, burn-in length, and whether the displayed contours are 68 % and 95 % credible regions.","section":"Figure captions"}],"recommendation":"minor_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the positive assessment of our manuscript and the recommendation for minor revision. We address each major comment below and will incorporate the suggested quantitative improvements.","responses":[{"response":"We agree that quantitative statistics are needed to substantiate the correlation claim. In the revised manuscript we will report the Spearman rank correlation coefficient with its p-value for Fermat potential precision versus arc surface brightness. We will also add a linear fit to the eight data points, including uncertainties on the slope and intercept, to allow evaluation of robustness and outlier influence. These results will be included in §4.","revision_made":"yes","referee_comment":"[§4] §4 (correlation analysis): The central claim that arc surface brightness is the 'primary driver' of mass-model precision rests on a reported 'strong correlation,' yet no quantitative statistic (Spearman coefficient, Pearson r with uncertainty, or p-value) or explicit fit to the eight data points is provided; without this, it is impossible to judge whether the trend is statistically robust or dominated by outliers."},{"response":"We agree that quantitative metrics would make this section more rigorous. In the revision we will state the exact number of conjugate points (the four quasar image positions) and their selection criteria. We will report the ratios of 68% credible-interval widths for key parameters (Einstein radius, power-law slope) between the full and conjugate-point-only models, and the posterior hypervolume ratio where relevant. These details will be added to the conjugate-point analysis section.","revision_made":"yes","referee_comment":"[Conjugate-point analysis] Conjugate-point analysis section: The statement that posteriors are 'substantially broader' when only quasar positions are used is presented without a quantitative metric (e.g., ratio of credible-interval widths or hypervolume ratio) or the exact number and selection of conjugate points; this leaves the confirmation that arcs dominate the constraints qualitative rather than rigorous."}],"tokens_in":1561,"tokens_out":429,"duration_ms":19034,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main takeaway is that brighter lensed host arcs give tighter constraints on the mass model and Fermat potential in doubly imaged quasars. This is the first uniform Lenstronomy analysis of these eight specific doubles, and the reported correlation between arc surface brightness and model precision is new relative to prior TDCOSMO work on quads. They also run a conjugate-point test that drops the extended arcs and shows much broader posteriors, plus an anti-correlation with mass-parameter hypervolume. That internal check is useful and helps rule out obvious modeling artifacts. Literature comparisons are reasonable: Einstein radii agree at 1.5 sigma on average, and image separations match Gaia to 3.6 mas rms. The open-source pipeline is a practical plus for anyone who wants to scale this to larger samples from LSST or Euclid. The central claim holds up on the checks described. Soft spots are minor and mostly about missing detail in the abstract: no numbers on the actual correlation strength, scatter, or exact data cuts, and limited discussion of how data heterogeneity across the eight systems was handled. The tailored pipeline for doubles is presented as accurate, but full posterior validation and any residual systematics would need the full text to confirm. This is squarely for the time-delay cosmography community that wants to move beyond quads. It deserves a serious referee because the uniform analysis and the practical implication for sample size are worth verifying in detail, even if the H0 results come in the follow-up paper.","headline":"The paper shows arc brightness sets Fermat potential precision in doubles via a uniform Lenstronomy pipeline on eight systems, with solid literature checks and a conjugate-point test.","tokens_in":2411,"tokens_out":373,"would_cite":true,"duration_ms":34548,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Surface brightness of lensed host arcs sets the precision of mass models in doubly imaged quasars.","keywords":["gravitational lensing","time-delay cosmography","doubly imaged quasars","lens mass modeling","Fermat potential","surface brightness","Hubble Space Telescope","Hubble constant"],"falsifier":"Finding no correlation between measured arc surface brightness and Fermat-potential uncertainty when the same eight systems are re-modeled with an independent code or additional high-resolution data.","tokens_in":2719,"feed_emoji":"🔭","tokens_out":707,"duration_ms":29899,"temperature":0.7,"pith_summary":"This paper applies a uniform Lenstronomy pipeline to eight doubly imaged quasars observed with the Hubble Space Telescope. The analysis reconstructs each system's lensing geometry and mass profile, then quantifies the uncertainty in the Fermat potential that enters time-delay cosmography. It finds a strong correlation showing that brighter extended host-galaxy arcs produce tighter constraints on the lens mass, while fainter arcs leave the models underconstrained. A separate conjugate-point test that ignores the arcs entirely reproduces broader posteriors and confirms the same brightness trend. The result identifies arc surface brightness as the dominant factor controlling model quality in doubles, which are far more numerous than quadruples.","feed_headline":"Arc brightness sets mass-model precision for double quasar lenses","feed_subtitle":"Uniform analysis of eight systems shows host-galaxy arc surface brightness drives Fermat-potential accuracy in time-delay cosmography.","key_machinery":"Correlation between Fermat potential uncertainty and host-arc surface brightness, measured through full extended-image modeling versus a conjugate-point analysis restricted to quasar positions.","core_discovery":"Uniform modeling of the eight systems yields Einstein radii consistent with literature values at 1.5 sigma and image separations matching Gaia DR2 to 3.6 mas rms. Full image reconstruction and a conjugate-point analysis that uses only the quasar positions demonstrate that Fermat-potential precision improves directly with the surface brightness of the spatially extended host arcs. An anti-correlation between mass-parameter hypervolume and arc magnitude further isolates arc brightness as the primary driver of how well the lens mass profile can be recovered in doubly imaged systems.","pith_inferences":["Survey strategies could prioritize follow-up of doubles that already show bright arcs in discovery imaging to maximize cosmographic return.","The same brightness trend may allow statistical marginalization over lens-model uncertainty when building large H0 samples.","Applying the conjugate-point test to triples or other configurations could reveal whether arc brightness remains the dominant constraint across image multiplicities."],"forward_implications":["Doubly imaged systems can now be ranked for cosmographic usefulness by a directly observable quantity: arc surface brightness.","The larger population of doubles can be incorporated into hierarchical H0 analyses once arc brightness is used to select or weight targets.","Uniform pipelines become feasible for the thousands of new lenses expected from LSST, Roman, and Euclid.","Model precision in doubles is shown to be limited by extended emission rather than by the mere number of point images."],"fun_headline_variants":["Arc brightness drives Fermat potential precision in double quasars","Host arc brightness controls lens model precision for doubles","Uniform double quasar models tie accuracy to arc surface brightness","Arc brightness determines Fermat potential constraints in eight doubles"],"cache_read_input_tokens":64,"weakest_assumption_plain":"The tailored Lenstronomy pipeline accurately recovers the true lensing geometry and mass profiles of doubles without large biases from data heterogeneity or unmodeled systematics.","fun_headline_variants_meta":{"raw":{"variants":["Arc brightness drives Fermat potential precision in double quasars","Host arc brightness controls lens model precision for doubles","Uniform double quasar models tie accuracy to arc surface brightness","Arc brightness determines Fermat potential constraints in eight doubles"]},"model":"grok-4.3","cost_usd":0.009074,"raw_usage":{"total_tokens":4051,"prompt_tokens":789,"num_sources_used":0,"completion_tokens":62,"cost_in_usd_ticks":90740500,"prompt_tokens_details":{"text_tokens":789,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":3200,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":789,"tokens_out":62,"duration_ms":38661,"temperature":1.0,"reasoning_tokens":3200,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-08T01:21:29.635477+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Finding no correlation between measured arc surface brightness and Fermat-potential uncertainty when the same eight systems are re-modeled with an independent code or additional high-resolution data.","supporting_citations":[],"review_version":1}