{"id":"dc279e78-f91b-48d7-b55c-c9d82cb96252","arxiv_id":"2501.09883","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A new catalog of 4,876 bent-tail radio galaxies, of which 3,871 are new, built from FIRST survey images using a deep learning source finder followed by visual inspection.","lead":"Astronomers used a deep learning model plus human review to find bent-tail radio galaxies in the FIRST survey, building a catalog of 4,876 such galaxies, 3,871 of which are new. Because bent-tail galaxies trace motion through galaxy clusters, this larger sample enables better statistical studies of cluster environments and galaxy evolution.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No survey-wide recall check supports the 11,473-candidate list; the reported mAP is measured on a labeled evaluation set (Section 2.2), so true BTRGs missed by RGCMT are irrecoverably absent from the 4876-entry catalog.","rationale":"The reader's weakest_assumption is the recall/completeness of the RGCMT candidate list. I agree that this is the load-bearing point. The paper's strongest claim is the size and comprehensiveness of the catalog; that claim requires that the 11,473 candidates contain most true BTRGs in FIRST before human filtering. The reported mAP is a standard detection metric on a validation set, not a survey-wide completeness estimate; it cannot rule out systematic misses of faint, diffuse, or atypical BTRGs. The overlap with nine published samples is supportive (1005 recovered), but the paper gives only overlap counts, not missing-source audits, and even acknowledges some known sources are absent without a full accounting. Secondary issues (WAT count inconsistency 4424 vs 4224; note added in proof removing one source) are sloppiness that undercuts trust but do not by themselves overturn the catalog's existence. These are addressable, so conditional acceptance is the right level: the catalog should be released with a clear completeness caveat and the missing-source audit or a random human-labeled recall estimate should be provided. Therefore I do not change the reader's verdict.","tokens_in":30991,"tokens_out":7705,"duration_ms":83718,"concrete_test":"Cross-match BTRGcat against all nine published samples used in Section 4.1, using the full published catalogs rather than the overlap counts; compute for each sample how many sources in the FIRST footprint are absent from BTRGcat and audit every absence against the four exclusion criteria in Section 2.4. If any appreciable number (e.g. >2%) of absences are genuine BTRGs by the paper's own definition, the 11,473-candidate list is incomplete and the 'most comprehensive' claim is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central catalog claim depends on the candidate list in Section 2.3 being essentially complete: after RGCMT processes all 946,432 FIRST cutouts, the 11,473 candidates are reduced by visual inspection only to 4876. The paper's evidence for detection quality is the 98.4% mAP (AP for BT = 98.4%) on a 1946-image evaluation set, but that is not a survey-wide recall estimate at the deployed confidence threshold of 0.5. The evaluation set was labeled under the same scheme and likely carries the same selection biases as the training set; it does not sample the full range of FIRST images, including faint, confused, or atypical bent morphologies. The overlap checks in Section 4.1 show that 1005 sources from nine published samples are recovered, but they do not report how many known BTRGs in the FIRST footprint are absent; the paper explicitly acknowledges missing sources from Sasmal et al. (2022) and attributes them to selection criteria (OA > 170, straight one-sided, S-shaped) without a quantitative audit. Because Section 2.4 only removes false positives, any true BTRG that RGCMT never proposes is permanently lost from the catalog. This affects the 'largest and most comprehensive' claim and the 3871 'newly discovered' count. Secondary internal inconsistencies (4424 vs 4224 WATs between the abstract and Section 4.4; the note added in proof removing one source) reinforce that the catalog numbers need external verification before the headline counts are taken at face value.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript presents a catalog of 4876 bent-tail radio galaxies (BTRGs) from the FIRST survey, built by (i) running the deep-learning detector RGCMT over 946,432 FIRST cutouts, (ii) retaining 11,473 candidates with score >= 0.5 and total flux >= 1.64 mJy, and (iii) visually inspecting each candidate to remove false positives, leaving 4876 sources, of which 3871 are claimed to be new discoveries. The catalog includes host-galaxy identifications from DESI LS (4193 hosts), spectroscopic and photometric redshifts (4171 sources), 1.4-3 GHz spectral indices from NVSS/VLASS, opening angles and radii of curvature derived from the detector's predicted masks, and cluster associations from Wen & Han (2024) and NED (3286 sources, 1825 within r500). The overlap of the catalog with nine published BTRG samples is quantified (1005 sources recovered after de-duplication), and physical statistics (spectral index, luminosity, host colors, black-hole masses, cluster properties) are presented. The paper claims this is the largest and most comprehensive BTRG catalog to date.","tokens_in":31339,"tokens_out":10855,"duration_ms":98840,"significance":"If validated, this catalog would roughly triple the number of known BTRGs and provide a statistically powerful sample for studying jet-ICM interactions, cluster environments, and AGN feedback; the luminosity range (10^20-10^28 W Hz^-1) and the cluster association statistics (1825 BTRGs within r500) are genuinely useful products. The paper's strengths are real: the pipeline is documented step by step, the candidate thresholds are explicitly stated, the overlap with nine published samples is quantified (1005 recovered sources), and the catalog is publicly deposited with a DOI (10.5281/zenodo.14271760). The RGCMT model is prior published work (Lao et al. 2023) with a stated mAP of 98.4%, but the catalog paper's added value is the survey-wide application, which is exactly where the recall question arises. The internal count inconsistencies and the unverified survey recall are the main risks to the headline claims.","major_comments":[{"comment":"The central completeness assumption is not verified. The catalog pipeline is one-directional: RGCMT proposes 11,473 candidates (Section 2.3) and visual inspection only removes candidates (Section 2.4), so any true BTRG that RGCMT fails to propose is permanently absent from the final 4876-source catalog. The only quantitative evidence for detection quality is the 98.4% mAP (BT AP 98.4%) on the 1946-image evaluation set (Section 2.2), which is not a survey-wide recall measurement at the deployed operating point (score >= 0.5, total flux >= 1.64 mJy), and the evaluation set was labeled by the same team under the same scheme as the training set. The overlap analysis in Section 4.1 confirms that 1005 sources from nine published samples are recovered, but for the largest comparison sample (Sasmal et al. 2022) only 506 of 717 sources are recovered, and the 211 missing sources are explained only by qualitative examples (J0044+1026, J1321-0637, J1521+5104, J1138+2039) rather than a complete audit. Because the headline claims (4876 BTRGs, 3871 new discoveries, and the 'largest and most comprehensive' statement in Section 5) inherit this unverified recall assumption, I request a survey-level recall test (for example, injecting synthetic BTRGs with a range of bending angles, sizes, and fluxes into FIRST images and measuring the recovery rate) and a full quantitative accounting of the non-recovered Sasmal et al. (2022) sources.","section":"Sections 2.2-2.4"},{"comment":"The WAT/NAT counts are internally inconsistent. The abstract and Section 5 state 4424 WATs and 652 NATs, which sum to 5076 and conflict with the stated total of 4876; Section 4.4 gives 4224 WATs and 652 NATs, which sum correctly to 4876. The abstract, Section 4.4, Section 5, and the deposited table must be reconciled, since the WAT/NAT split is one of the paper's headline results and any reader using the catalog needs a single authoritative count.","section":"Abstract vs. Section 4.4"},{"comment":"The note added in proof contradicts the body text and the catalog as presented. It asserts that the source J125648.57+481749.8, which appears in Table 1 and is included in the 4876 total, 'should be removed from the sample of BTRGs'; it also states that 12 of the 17 'blue' hosts are spectroscopic QSOs and a further three are blazars, which directly contradicts Section 4.5's conclusion that no common properties were found among these hosts. The catalog counts, host statistics, and the g-r analysis of Section 4.5 must be revised in the main text and the deposited table (Zenodo DOI 10.5281/zenodo.14271760) re-issued, rather than leaving the correction in a note.","section":"Note added in proof"}],"minor_comments":[{"comment":"The sentence reporting the mean and median of 'the logarithmic ratio between log10(M500) and log10(r500)' as 2.43 x 10^14 M_sun/Mpc is dimensionally confused; if the intended quantity is the ratio M500/r500, it should be stated without the 'log' terminology.","section":"Section 4.6"},{"comment":"The notation in 'Almost all of them (99.3%) have -21 gm g Mr g g -25' (and the analogous MBH statement) is mathematically equivalent to -25 <= Mr <= -21 but is written in an order that is easy to misread; please reverse the inequalities for clarity.","section":"Section 4.5"},{"comment":"The caption says the histogram shows the distribution of OA for 'all BTRG candidates', but the text of Section 4.4 describes the final 4876 sources after the OA < 170 deg cut; please make this unambiguous.","section":"Figure 8 caption"},{"comment":"The reported mAP is evaluated at IoU = 0.5; since the predicted masks are subsequently used to measure OA, Rc, and LAS in Section 4.4, reporting a stricter IoU threshold (0.75) or a boundary-distance metric would better substantiate the geometric measurements.","section":"Section 2.2"},{"comment":"The choice of the 1.64 mJy total-flux threshold is not motivated (for example, relative to the typical FIRST rms of 0.15 mJy); a sentence explaining its basis would help readers assess the faint-end cutoff of the candidate list.","section":"Section 2.3"}],"recommendation":"major_revision","confidential_remarks":"The main risk is completeness: the paper's own note added in proof and the 4424/4224 WAT discrepancy show that the catalog numbers are still shifting, and the deposited table should be re-issued after the authors correct the counts. The recall question (no survey-wide completeness test) is the central review concern; I would recommend asking for the synthetic-injection test or an equivalent before the catalog claims are treated as final. The paper is within scope for a radio-astronomy journal and the resource is potentially valuable, so revision rather than rejection is appropriate."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear [Colleague],\n\nThe paper is essentially a catalog paper: 4876 bent-tail radio galaxies from FIRST, 3871 not in nine earlier samples, with host galaxies, redshifts, 1.4/3 GHz fluxes, spectral indices, opening angles, and cluster associations. The catalog is real and useful—this is the largest BTRG sample by a factor of several, and the machine-readable table is on Zenodo.\n\nWhat the paper does well: the pipeline is clearly described, the overlap with previous published samples is quantified (1005 known BTRGs recovered after de-duplication), hosts and redshifts are matched through a documented visual-inspection process, and the derived statistics (luminosity distributions, FR-I/II breakdown, cluster membership) are presented with enough detail to be checked. The authors also explain why some Sasmal et al. sources are absent—selection criteria—rather than pretending completeness.\n\nThe soft spots are real but addressable. The main one is that the \"newly discovered\" count depends on RGCMT's candidate list being essentially complete, and there is no survey-wide recall measurement. The 98.4% mAP is on a 1946-image evaluation set labeled under the same scheme as the training data; it does not tell you how many true BTRGs the model misses across the full FIRST footprint at the deployed score threshold. Since visual inspection only removes false positives, any missed source is permanently absent. The overlap checks show that 1005 known BTRGs are recovered, but the paper does not audit how many known FIRST BTRGs are absent—and it explicitly acknowledges some are missing without a quantitative count. So the 3871 \"newly discovered\" figure is better read as \"not in these nine samples\" than as a completeness-corrected census.\n\nThere is also an internal inconsistency: the abstract and conclusions say 4424 WATs, but Section 4.4 says 4224 WATs with 652 NATs (4876 - 652 = 4224). The note added in proof says one source should be removed but the catalog numbers are not updated. These are minor fixable issues, but they need to be corrected.\n\nThe circularity concern from the stress test does not really land: the thresholds are selection cuts, not fitted parameters, so there is no fitted-law circularity here. The real weakness is the missing recall audit, not circular reasoning.\n\nVerdict: this deserves a serious referee. The catalog is a contribution that people in the field will use, and the flaws are patchable. I would send it to review with a request for a completeness check and a fix to the WAT/NAT counting.\n\nBest.","headline":"A large, genuinely useful BTRG catalog that deserves refereeing, but the headline counts need a completeness audit and a fix to an internal inconsistency.","tokens_in":31938,"tokens_out":2432,"would_cite":true,"duration_ms":23434,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A deep-learning search of the entire FIRST survey plus visual inspection produces a catalog of 4,876 bent-tail radio galaxies, 3,871 of them newly discovered.","keywords":["bent-tail radio galaxies","wide-angle-tail sources","narrow-angle-tail sources","FIRST survey","deep learning source detection","radio morphology","active galactic nuclei","galaxy clusters"],"falsifier":"Have independent inspectors visually scan a random sample of FIRST images with no knowledge of RGCMT's candidates and tally bent-tail sources missing from the 4,876-entry catalog; if the missed fraction is comparable to or larger than the 3,871 new discoveries, the catalog's completeness claim fails. A cheaper proxy is to run a second, independently trained detector over the same survey images and count bent sources found by only one method.","tokens_in":30803,"feed_emoji":"📡","tokens_out":12014,"duration_ms":106286,"temperature":0.7,"pith_summary":"Bent-tail radio galaxies are active galaxies whose radio jets are swept backward by motion through the hot gas of a galaxy cluster, producing C-, V-, or U-shaped radio lobes. This paper reports a systematic search of the entire FIRST radio survey for these objects, running a deep-learning detector over all 946,432 survey images and then visually inspecting every candidate. The result is a catalog of 4,876 bent-tail radio galaxies, of which 3,871 are newly discovered, more than quadrupling the number known from previous work. The catalog also supplies optical host identifications for 4,193 sources, redshifts from 0.0023 to 3.43, and derived radio luminosities, so it turns a rare morphological class into a large enough sample for statistical studies of cluster environments and jet physics.","feed_headline":"Deep-learning scan finds 4,876 bent-tail radio galaxies","feed_subtitle":"Almost 4,000 of them are newly discovered, giving astronomers a large sample to probe how cluster gas bends radio jets","key_machinery":"The engine of the search is RGCMT, a deep-learning detector built from a convolutional mask-prediction network with a quadtree ambiguity-area detector and a transformer refinement block, trained on 3,172 labeled FIRST images to recognize five radio-morphology classes (compact sources, straight FR-I, straight FR-II, bent tails, and one-sided extended sources). For each detected source it produces a predicted mask, bounding box, and confidence score, and the mask is what allows positions, flux densities, opening angles, and radii of curvature to be computed automatically. The candidate list is then filtered by human visual inspection of FIRST contour overlays, removing false positives such as S/Z-shaped sources, lobes of larger galaxies, and artifacts. The morphological parameter that finalizes membership is the opening angle between the two jets, measured from the predicted mask's skeleton via a Voronoi diagram, with an upper cutoff of 170 degrees.","core_discovery":"The authors claim that a combination of the RGCMT deep-learning source finder and human visual inspection can find bent-tail radio galaxies in the FIRST survey at scale and with high reliability. Applying RGCMT to all 946,432 FIRST catalog components yielded 11,473 candidate detections after removing duplicates and applying flux and confidence cuts; visual inspection of FIRST contour images rejected sources that were not bent in a common direction, lobes of larger galaxies, artifacts, or had opening angles above 170 degrees, leaving 4,876 BTRGs. Of these, 4,424 are wide-angle-tail sources and 652 are narrow-angle-tail sources; 4,193 have optical counterparts in DESI Legacy Surveys DR10, 4,171 have redshift measurements, and 1,825 lie within known galaxy clusters by the nearest-neighbor criterion. The paper presents this as the largest and most comprehensive BTRG catalog to date, with 3,871 newly discovered sources.","pith_inferences":["If the candidate list is nearly complete, the catalog can be used to estimate the space density of bent-tail galaxies in the FIRST footprint; the fact that only 43.5% of hosts match known clusters then suggests either many bent-tail galaxies live in poorer or unidentified environments or the 3-arcmin match radius limits the association rate.","The same mask-based measurement pipeline could be applied to other radio surveys with similar resolution, yielding directly comparable opening angles and curvature radii and testing whether the WAT/NAT split changes with resolution.","The note added in proof, which removes one source and reclassifies several blue hosts as QSOs or blazars, implies that host-type classification is sensitive to spectroscopic follow-up; a systematic spectroscopic campaign on the 1,814 photometric-redshift hosts would sharpen the luminosity and cluster-association statistics.","Because the training and evaluation sets were labeled by human visual inspection, the detector's notion of 'bent' is anchored to human judgment; transferring it to another survey would require re-checking that its bending criterion matches what observers count at other resolutions."],"forward_implications":["The known population of bent-tail radio galaxies grows from roughly 1,005 previously cataloged sources to 4,876, giving a sample large enough for statistical studies of the class.","The catalog's cluster matches place 1,825 of the sources inside known galaxy clusters, strengthening the picture that bent morphology is produced by motion through dense intra-cluster gas and offering a large sample for environment studies.","The derived FR-I and FR-II luminosities show that many FR-I-classified bent sources lie above the classical log L = 25 dividing line and many FR-IIs below it, reinforcing recent evidence that the luminosity break is not a clean morphological divider.","The automated masks let physical parameters such as opening angle, radius of curvature, and largest linear size be measured consistently for thousands of sources at once, which previously required labor-intensive individual measurement.","The median spectral index of 0.85 between 1.4 and 3 GHz, noted by the authors as likely an upper limit because VLASS misses extended emission, gives a first statistical view of the radio spectra of bent-tail galaxies."],"supporting_citations":[{"why":"Introduces the RGCMT deep-learning detector whose masks and confidence scores generate the BTRG candidate list.","marker":"Lao et al. 2023"},{"why":"Defines the FIRST survey whose 946,432 catalog components and images are searched.","marker":"Becker et al. 1995"},{"why":"Provides the FIRST-14dec17 catalog release used as the input source list.","marker":"Helfand et al. 2015"},{"why":"Describes the HAPPY source extraction that produced the FIRST component catalog.","marker":"White et al. 1997"},{"why":"Previous BTRG catalog of 717 sources used to measure how many catalog entries are new and to explain selection differences.","marker":"Sasmal et al. 2022"},{"why":"Describes the DESI Legacy Surveys DR10 used to find optical host candidates for the BTRGs.","marker":"Dey et al. 2019"},{"why":"Supplies the galaxy-cluster catalog cross-matched to identify cluster memberships.","marker":"Wen & Han 2024"}],"fun_headline_variants":["AI + human eyes spot 4,876 bent-tail radio galaxies","Largest bent-tail catalog yet: 4,876 galaxies found","Deep learning finds 3,871 new bent-tail radio galaxies","Bent-tail galaxy haul: 4,876 from FIRST survey"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The deep-learning model's detection accuracy, measured on its own labeled test set, transfers to the whole FIRST survey, so the human visual-inspection step removes false positives without missing a large share of true bent-tail galaxies.","fun_headline_variants_meta":{"raw":{"variants":["AI + human eyes spot 4,876 bent-tail radio galaxies","Largest bent-tail catalog yet: 4,876 galaxies found","Deep learning finds 3,871 new bent-tail radio galaxies","Bent-tail galaxy haul: 4,876 from FIRST survey"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000222,"raw_usage":{"total_tokens":1493,"prompt_tokens":1023,"completion_tokens":470,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":639,"completion_tokens_details":{"reasoning_tokens":396}},"tokens_in":639,"tokens_out":470,"duration_ms":4768,"temperature":1.0,"reasoning_tokens":396,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T19:34:37.449622+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have independent inspectors visually scan a random sample of FIRST images with no knowledge of RGCMT's candidates and tally bent-tail sources missing from the 4,876-entry catalog; if the missed fraction is comparable to or larger than the 3,871 new discoveries, the catalog's completeness claim fails. A cheaper proxy is to run a second, independently trained detector over the same survey images and count bent sources found by only one method.","supporting_citations":[],"review_version":1}