{"id":"4d881fe1-4248-4a03-8ce9-bb9ec59ae03c","arxiv_id":"2604.04810","paper_version":1,"verdict":"CONDITIONAL","confidence":"LOW","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Independent pipeline recovers 63.9% of a prior 5,399-source vanished catalog from POSS-I plates, produces 2.85 million new candidates, and finds no significant temporal association in candidate counts.","lead":"The paper describes an automated pipeline that detects candidate vanished sources on 1950s photographic sky survey plates by finding sources in one epoch and checking if they are absent in others while rejecting modern detections. It recovers most previously claimed vanished sources but finds no statistically significant evidence that disappearances cluster on particular days.","discovery_kind":"replication","skeptic_critique":{"model":"grok-4.3","headline":"Benchmark limited to 20 cases on two dates and FP rate measured only on non-crowded controls; extrapolation to full-survey 2.85M catalog purity untested across plate heterogeneity","rationale":"The reader's weakest assumption (systematic bias in PSF filtering, astrometric refinement, and Pan-STARRS rejection across varying conditions) is the precise point where the small benchmark and non-crowded controls fail to constrain the full-survey output. The paper itself presents the temporal result as inconclusive and the catalog as candidates, so the internal numbers are consistent; the load-bearing uncertainty is simply the untested generalization of the validation metrics.","tokens_in":1977,"tokens_out":436,"duration_ms":56201,"concrete_test":"Select 200 random 10-arcmin fields stratified by local source density (low/medium/high) and plate quality metrics from the POSS-I metadata; run the full pipeline (PSF filter + local astrometry + Pan-STARRS rejection) on each and count surviving candidates; if the median FP rate exceeds 0.5 per field in any stratum, recompute the expected contamination fraction in the 2.85M catalog.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The pipeline recovers 8/9 and 3/3 known sources on the 20-case harness for the April 1950 and July 1952 fields and reports ~0.2 FP per 10 arcmin on random non-crowded controls. The full-footprint run applies the same PSF filtering, local astrometric refinement, and Pan-STARRS DR1 rejection over 30 arcmin patches spanning the entire POSS-I coverage. No larger control sample, injection-recovery test, or crowding-stratified FP measurement is described, so it is unclear whether the reported purity holds when plate quality, stellar density, and artifact rates vary. If the effective FP rate rises even modestly in typical survey regions, the 2.85M catalog could be dominated by artifacts while still producing the observed 63.9% overlap with the prior 5,399-source catalog.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper presents an independent automated pipeline for detecting candidate vanished sources on digitized POSS-I Red photographic plates. The method performs source detection and PSF filtering on DSS cutouts, applies local astrometric registration refinement, and identifies candidates via cross-epoch matching to POSS-I Blue and POSS-II Red plates followed by Pan-STARRS DR1 rejection. On a 20-case benchmark, it recovers 8/9 sources in the April 1950 field and 3/3 in the July 1952 field with a reported false-positive rate of ~0.2 per 10 arcmin on non-crowded controls. A full-footprint run over 30-arcmin patches produces a filtered catalog of 2.85 million candidates. This catalog overlaps 63.9% (3,450 matches, median separation 0.94 arcsec) with the Solano et al. (2022) catalog of 5,399 sources. Temporal association tests using Bruehl & Villarroel-style windows over 368 nights yield a non-significant relative risk of 1.35 (p=0.17) and a null negative binomial result (IRR=1.03, p=0.71).","tokens_in":2182,"tokens_out":869,"duration_ms":32360,"significance":"If the pipeline's purity and completeness hold across the full survey, the 2.85M-candidate catalog would constitute a substantial independent resource for archival transient studies, enabling statistical analyses of source disappearance rates in the 1949-1957 era. The 63.9% replication of the prior Solano catalog is a clear strength, demonstrating consistency between independent detection methods. However, the temporal clustering test remains inconclusive, limiting immediate astrophysical conclusions. The work advances automated archival searches but its impact depends on robust validation of the large catalog.","major_comments":[{"comment":"Benchmark harness and false-positive controls: The recovery rates (8/9 and 3/3) and FP rate (~0.2 per 10 arcmin) are derived from only 20 cases on two specific dates using non-crowded random controls. The full 2.85M catalog is generated across the heterogeneous POSS-I footprint (varying plate quality, stellar density, and artifact rates) without reported injection-recovery tests, crowding-stratified FP measurements, or plate-quality stratification. This gap directly affects whether the reported purity can be extrapolated to the full catalog.","section":"Benchmark harness and full-footprint sweep"},{"comment":"Pan-STARRS DR1 rejection and post-processing: The pipeline relies on Pan-STARRS DR1 rejection plus PSF cuts and deduplication to produce the 2.85M catalog, yet no quantitative assessment of completeness, false-negative rate, or systematic bias in rejection across the survey is provided. The note that unrecovered Solano entries lack Pan-STARRS counterparts within 3 arcsec is useful but insufficient without an error budget or simulation to confirm the rejection step does not preferentially remove real transients.","section":"Catalog generation and cross-matching"}],"minor_comments":[{"comment":"The abstract and methods would benefit from explicit listing of the exact PSF cut thresholds, local astrometric refinement tolerances, and deduplication criteria rather than referring to them generically as 'post-processing cuts'.","section":"Methods and abstract"},{"comment":"Consider adding a table or figure summarizing the 20-case benchmark recoveries, including individual source properties, separations, and reasons for the single non-recovery.","section":"Benchmark results"},{"comment":"The temporal analysis correctly reports non-significance and sensitivity to day coding, but a brief power calculation for the negative binomial model would help readers interpret the null result.","section":"Temporal association test"}],"recommendation":"major_revision","confidential_remarks":"The manuscript fits well within astro-ph.IM scope as a methods paper. The citation to Solano et al. (2022) is appropriate and the overlap analysis is a strength; no obvious citation gaps or scope mismatch noted. The limited benchmark scope is the primary concern for acceptance."},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their constructive feedback, which highlights both the potential utility of our 2.85M-candidate catalog and the need for clearer discussion of validation limits. We address each major comment below with clarifications on our approach and scope. Revisions to the manuscript will incorporate expanded caveats and a dedicated limitations section to better contextualize the benchmark and rejection steps.","responses":[{"response":"We acknowledge that the 20-case benchmark on two fields with non-crowded controls provides only a baseline demonstration rather than a comprehensive purity assessment across the full heterogeneous footprint. Our primary validation metric remains the 63.9% overlap with the independent Solano et al. (2022) catalog, which offers external consistency without relying solely on our internal FP estimates. The reported FP rate serves as a conservative lower bound for non-crowded regions. In the revised manuscript we will add a 'Limitations and Future Work' section explicitly discussing the absence of injection-recovery tests and the potential impact of crowding and plate quality variations, while noting that the catalog is presented as a resource for further statistical studies rather than a purity-guaranteed list.","revision_made":"partial","referee_comment":"Benchmark harness and false-positive controls: The recovery rates (8/9 and 3/3) and FP rate (~0.2 per 10 arcmin) are derived from only 20 cases on two specific dates using non-crowded random controls. The full 2.85M catalog is generated across the heterogeneous POSS-I footprint (varying plate quality, stellar density, and artifact rates) without reported injection-recovery tests, crowding-stratified FP measurements, or plate-quality stratification. This gap directly affects whether the reported purity can be extrapolated to the full catalog."},{"response":"The 3-arcsec Pan-STARRS DR1 rejection radius is chosen to accommodate known astrometric uncertainties in the digitized POSS-I plates while removing sources that remain detectable in modern surveys. The observation that unrecovered Solano entries also lack PS counterparts within this radius supports that the step is not selectively discarding real transients from the prior catalog. We agree that without dedicated injection simulations we cannot provide a full error budget or false-negative quantification. The revised manuscript will expand the methods and discussion sections to justify the radius choice, report the fraction of Solano sources rejected by this criterion, and include a paragraph on possible biases, while emphasizing that the approach is intentionally conservative to minimize contamination from persistent sources.","revision_made":"partial","referee_comment":"Pan-STARRS DR1 rejection and post-processing: The pipeline relies on Pan-STARRS DR1 rejection plus PSF cuts and deduplication to produce the 2.85M catalog, yet no quantitative assessment of completeness, false-negative rate, or systematic bias in rejection across the survey is provided. The note that unrecovered Solano entries lack Pan-STARRS counterparts within 3 arcsec is useful but insufficient without an error budget or simulation to confirm the rejection step does not preferentially remove real transients."}],"tokens_in":1825,"tokens_out":647,"duration_ms":62812,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main takeaway is a new automated pipeline that processes the entire POSS-I coverage and produces a catalog of 2.85 million candidate vanished sources, along with a null result on temporal clustering using the same style of test as earlier papers. The catalog overlaps 63.9% with the Solano et al. list at sub-arcsecond separations, which is a concrete replication point. Recovery on the 20-case benchmark hits 8/9 and 3/3 on the two test fields, and the false-positive rate is quantified at roughly 0.2 per 10 arcmin on non-crowded controls. The statistical checks include confidence intervals and p-values on both the calendar-day risk and the negative binomial model of nightly counts, so the null findings are reported with some rigor. That is the useful output: a large independent list plus explicit numbers on how it lines up with previous efforts. The methods are described at a level that lets you see the steps—PSF filtering, local astrometric refinement, Pan-STARRS rejection, and deduplication. The work is therefore straightforward to reproduce or extend if someone wants to apply similar cuts elsewhere. The soft spot is the validation scope. The benchmark and false-positive measurements stay limited to non-crowded random fields and only 20 known cases across two dates. When the same cuts run over the full survey, plate quality, crowding, and artifact rates vary, yet no injection-recovery test or density-stratified control is shown. If the effective false-positive rate climbs even a little in typical regions, the 2.85 million entries could contain a much higher fraction of artifacts while still producing the observed overlap with the smaller prior catalog. Post-processing cut details and the full error budget are also not fully visible from the abstract. This paper is for researchers who mine historical plates or build cross-epoch catalogs for transient searches. Readers who need a large, independently generated list with basic replication stats will find it usable as a starting point or comparison set. It deserves a serious referee because the quantitative replication and statistical tests give something concrete to evaluate, even if the purity across the whole footprint needs more work. I would send it to peer review and ask the authors to add crowding-stratified controls or injection tests before final acceptance.","headline":"The paper delivers a new 2.85-million-candidate catalog from a full-footprint POSS-I pipeline with decent replication of prior work, but the purity claim rests on thin validation.","tokens_in":2659,"tokens_out":541,"would_cite":false,"duration_ms":37394,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":{"model":"grok-4.3","evidence":[{"relation":"unclear","rs_module":"IndisputableMonolith/Foundation/Cost/FunctionalEquation.lean","rs_theorem":null,"paper_passage":"The pipeline detects and PSF-filters sources on POSS-I Red DSS cutouts, applies local astrometric registration refinement, and identifies candidates by cross-epoch matching against POSS-I Blue and POSS-II Red with Pan-STARRS DR1 rejection."}],"headline":"Standard astronomical transient-detection pipeline with PSF filtering, cross-epoch matching and GLM statistics; no J-cost, phi-ladder or 8-tick structure","alignment":"orthogonal","rationale":"The paper's central machinery (sigma-clipped detection, FWHM/ellipticity PSF cuts, local affine registration, 8-arcsec cross-epoch absence test, Pan-STARRS rejection, negative-binomial GLM on nightly counts) is conventional astro-image processing and catalog validation. It contains none of the RS forcing elements (J(x) = ½(x + x⁻¹) − 1, cosh-cost identities, golden-ratio fixed points, 8-tick periodicity, parameter-free constant derivations). The domain is therefore one on which the RS framework has no opinion.","tokens_in":46212,"confidence":"high","tokens_out":275,"duration_ms":18628,"cache_read_input_tokens":32896,"cache_creation_input_tokens":0},"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"An independent pipeline recovers most reported vanished sources from POSS-I plates but detects no significant temporal clustering.","keywords":["vanished sources","POSS-I plates","source detection","cross-epoch matching","photographic plates","astronomical transients","sky surveys","replication"],"falsifier":"Detection of faint counterparts to most of the 2.85 million candidates in deeper modern surveys such as Gaia or LSST within 3 arcsec would indicate they are persistent rather than vanished.","tokens_in":2858,"feed_emoji":"🔭","tokens_out":785,"duration_ms":27410,"temperature":0.7,"pith_summary":"The paper builds an automated pipeline to detect candidate vanished sources on digitized POSS-I Red photographic plates from the early 1950s. It combines PSF filtering of detected sources, local astrometric refinement, and cross-epoch matching against later blue plates and modern catalogs to filter out persistent objects. Benchmark tests on known cases recover eight of nine sources in one field and all three in another, with a low false-positive rate on control fields. A full scan across the survey footprint produces a catalog of 2.85 million candidates after additional cuts. Statistical tests for day-specific clustering in the 1949-1957 observation window return null results, leaving the temporal pattern inconclusive while the catalog-level replication holds.","feed_headline":"Pipeline recovers 64% of vanished sources from 1950s plates","feed_subtitle":"2.85 million candidates identified but day-specific clustering test returns null result","key_machinery":"The automated detection pipeline that performs PSF-filtered source detection on POSS-I Red DSS cutouts, applies local astrometric registration refinement, performs cross-epoch matching to POSS-I Blue and POSS-II Red plates, and rejects candidates using Pan-STARRS DR1.","core_discovery":"The pipeline recovers 8/9 sources in the April 1950 benchmark field and 3/3 in the July 1952 field. It matches 3450 of the 5399 entries in the published Solano et al. catalog (63.9 percent) at a median separation of 0.94 arcsec. A full-footprint search over POSS-I coverage yields 2.85 million candidates after PSF cuts, deduplication, and Pan-STARRS DR1 rejection. Application of day-window tests across the 368 observation nights produces a post-test relative risk of 1.35 that is not statistically significant, and a negative binomial model of nightly counts likewise shows no effect.","pith_inferences":["The recovered candidates could be followed up with targeted high-resolution imaging to check for reappearance or other properties.","Larger samples or alternative statistical models might be needed to detect any subtle day-of-year pattern if one exists.","The same cross-epoch rejection approach could be extended to additional historical plate archives to expand the search for vanished sources."],"forward_implications":["The pipeline reproduces 63.9 percent of the prior catalog entries with sub-arcsecond positional agreement.","A large, filtered catalog of 2.85 million candidates is generated for further investigation after post-processing.","No statistically significant excess of candidates appears on any particular calendar day within the 1949-1957 interval.","The false-positive rate remains low at approximately 0.2 per 10-arcmin field in non-crowded control regions."],"fun_headline_variants":["Independent pipeline matches 64% of 1950s vanished sources","POSS-I scan identifies 2.85 million vanished source candidates","Pipeline recovers 8/9 sources in 1950 benchmark field","3.45 thousand matches to published vanished sources catalog"],"cache_read_input_tokens":64,"weakest_assumption_plain":"That PSF filtering, local astrometric refinement, and Pan-STARRS DR1 rejection correctly separate true vanished sources from artifacts or misdetections without systematic bias across the full survey footprint and varying plate conditions.","fun_headline_variants_meta":{"raw":{"variants":["Independent pipeline matches 64% of 1950s vanished sources","POSS-I scan identifies 2.85 million vanished source candidates","Pipeline recovers 8/9 sources in 1950 benchmark field","3.45 thousand matches to published vanished sources catalog"]},"model":"grok-4.3","cost_usd":0.005601,"raw_usage":{"total_tokens":2798,"prompt_tokens":900,"num_sources_used":0,"completion_tokens":70,"cost_in_usd_ticks":56012000,"prompt_tokens_details":{"text_tokens":900,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1828,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":900,"tokens_out":70,"duration_ms":27038,"temperature":1.0,"reasoning_tokens":1828,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-10T18:55:53.951099+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Detection of faint counterparts to most of the 2.85 million candidates in deeper modern surveys such as Gaia or LSST within 3 arcsec would indicate they are persistent rather than vanished.","supporting_citations":[],"review_version":1}