{"id":"2979b13c-8986-4dea-a88b-d02e171bbedb","arxiv_id":"2607.14928","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"The EHT's iconic black hole image is one of countless plausible renderings; the collaboration's most valuable evidence is the limited variability across choices, not the single image.","lead":"This paper analyzes how the Event Horizon Telescope turned noisy data into one iconic image, showing that many plausible alternative images exist and that choices about algorithms, parameters, and colors could have changed the result. It argues that the EHT's most valuable evidence is not the single advertised image but the range of variability across reasonable processing choices.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central claim rests on 'limited variability' that is asserted but never quantified; an expanded, more independent parameter survey is needed to confirm the ring stability is not an artifact of shared imaging priors.","rationale":"The reader's weakest assumption concerned whether the authors' open-source software and parameter choices faithfully capture the EHT's actual decision space. My concern is related but broader: even taking the EHT's own survey at face value, the central claim requires that the demonstrated variability be 'limited' in a way that is currently only asserted, not quantified. The EHT's Top Set analysis is qualitative, and the paper's own demonstrations cover a narrow, shared-prior subspace. If the ring is not stable under a wider, more independent exploration, the claim that the variability demonstration is the most valuable evidence loses its empirical foundation. This is not an objection to the paper's historical or philosophical analysis, which is careful and well-supported; it is a request to make the central empirical premise explicit and testable. A quantitative stability analysis would settle the question and either strengthen or qualify the conclusion. Therefore a conditional acceptance—pending such a demonstration—is appropriate.","tokens_in":34692,"tokens_out":9915,"duration_ms":98426,"concrete_test":"Using the public EHT 2017 M87* data and open-source eht-imaging/SMILI, reproduce the fiducial imaging and extend the parameter survey: sample regularizer weights on a continuous grid (not only powers of 10), include configurations without the circular Gaussian prior, vary the systematic error and field-of-view settings, and add an independent imaging package (e.g., THEMIS) that does not share the same priors. For every reconstruction, measure ring presence, ring diameter, and asymmetry via image moments or geometric modeling. Report the distributions. If a substantial fraction (>5%) of plausible reconstructions lack a ring, or if ring diameter varies by more than the EHT's reported statistical uncertainty, the 'limited variability' claim is not supported; if the ring is stable across the expanded, independent set, the central claim holds.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's central claim (Conclusion, §8) is that 'the most valuable scientific evidence produced by the EHT comes not from the single image it ultimately advertised, but from the demonstration of the limited variability that emerged from the specific choices made.' This claim depends on the premise that the variability is genuinely limited in a meaningful, robust sense. Yet the paper never quantifies 'limited': it provides no measure of ring-diameter spread, no fraction of Top Set reconstructions that retain a ring, and no comparison of variability within vs. across pipelines. The EHT's own parameter survey (described in §6) is cited qualitatively: 37,500 combinations for eht-imaging, a Top Set of 1,572, and 'stability of image features (particularly a ring).' The paper's interactive demonstration samples only one algorithm with seven parameters on a coarse, powers-of-10 grid. More importantly, the explored space is constrained by shared assumptions: all RML pipelines use a circular Gaussian prior (§4) that was partly justified by CLEAN results, creating a potential circularity; all pipelines use the same calibrated visibility data. If a broader or more independent set of plausible choices—without the circular Gaussian prior, with continuously varying regularizer weights, or using a genuinely independent algorithm—produced images without a consistent ring, the 'limited variability' would be an artifact of the methods, not a property of the evidence. The paper acknowledges 'different choices, not tried or reported, could lead to greater variability,' but this hedge directly undercuts the strength of the 'most valuable evidence' claim: if the variability is not robustly bounded, the demonstration is less convincing than asserted.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reconstructs the full pipeline that produced the Event Horizon Telescope Collaboration's 2019 image of M87*, from observing campaign through correlation, calibration, algorithm design, blind and fiducial imaging, and finally averaging, coloring, and public presentation. Drawing on EHT papers, technical reports, open-source code, and the authors' own re-runs of the imaging software, it argues that at every stage the data underdetermine the image: many plausible choices of algorithms, parameters, regularizers, and color maps were available, and the published image is the average of three independently produced pipeline images. The paper's central normative claim, stated in the Conclusion (§8), is that the most valuable scientific evidence from the EHT is not the single advertised image but the demonstration of limited variability across the specific choices the collaboration actually made. The authors explicitly acknowledge that untried choices could yield greater variability.","tokens_in":35061,"tokens_out":7276,"duration_ms":76583,"significance":"If the central claim holds, the paper provides a historically well-grounded and epistemically substantive case study of model-data symbiosis in computational imaging, going beyond earlier philosophical discussions by actually running the open-source tools and exhibiting alternative images. The interactive parameter demonstration and animation (archived on Zenodo) are concrete, reproducible contributions; the authors also carefully document the EHT's own language ('blind', 'fiducial', 'ground truth') and contrast it with practices at LHC and LIGO. The paper is appropriately cautious: it does not argue that the image is not a photograph or observation, and it flags the circular-Gaussian-prior concern and the possibility of greater variability outside the surveyed space. This should be a valuable contribution to history and philosophy of astronomy and of data-intensive science.","major_comments":[],"minor_comments":[{"comment":"The central phrase 'limited variability' is used as an unquantified, qualitative assessment. Since the authors state they precomputed images for all 37,500 combinations of the eht-imaging parameter survey, a brief quantitative summary (e.g., distribution of ring diameters, fraction of Top Set images retaining a ring, or inter-image pixel variance) would make the claim more transparent and would preempt the skeptical reading that 'limited' is doing unearned work. If such metrics are not easily reported, the conclusion could be hedged to 'limited variability within the surveyed combinations'.","section":"§8, Conclusion"},{"comment":"The potential circularity of the circular Gaussian prior is acknowledged in §4, but the paper does not report what happens to the reconstructed images when that regularizer is removed or its FWHM is pushed outside the EHT survey range. Since the interactive demonstration appears to include the Gaussian prior as one of its seven parameters, a short sensitivity statement (e.g., 'the ring persists when the Gaussian prior is disabled' or 'the ring disappears') would materially strengthen the robustness discussion and address the shared-assumption concern.","section":"§4 and §6"},{"comment":"The caption says the demonstration 'samples seven parameters, totaling 37,500 possible combinations,' but it does not list how many values are sampled for each parameter. Adding this breakdown (e.g., 5 values for each regularizer weight, 4 values for total flux, etc.) would let readers verify the total and understand the coarseness of the grid.","section":"Figure 11 caption"},{"comment":"The statement that the published image is 'simply the pixel-by-pixel average of the results from the three pipelines' is load-bearing for the paper's narrative but is not tied to a specific EHT Paper IV section. A citation to the relevant part of EHT (2019e) would make the historical claim easier to check.","section":"§7, Averaging"}],"recommendation":"accept","confidential_remarks":"This manuscript is a well-documented historical-epistemological study with a genuinely reproducible element: the archived interactive demonstrations and the authors' use of open-source EHT software are strengths. The main skeptical concern about unquantified 'limited variability' is already partly defused by the paper's explicit limitation statement that untried choices could produce greater variability, but the authors could further immunize the central claim with a short quantitative appendix. This is not, in my judgment, a blocking issue. The paper is suitable for publication in a history/philosophy of science journal and will likely generate useful discussion."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The thing to know: this is the first systematic stage-by-stage reconstruction of the EHT imaging pipeline from open-source code, and it makes a genuinely useful reframing—the single advertised image is not the primary evidence; the demonstrated variability across defensible processing choices is. That claim is well-documented and carefully argued, and the interactive demonstrations (animation of CLEAN, parameter survey explorer) are a real plus. The historical material on the raw-data controversy and on Pannekoek's averaging is well chosen and not decorative. The paper earns its place in the HPS literature on the EHT, next to Galison and Doboszewski & Elder, but it goes beyond them by actually re-running the software and showing alternatives.\n\nSoft spots, in proportion. The stress-test worry about quantification is legitimate but not fatal. The paper never measures 'limited variability': no ring-diameter spread, no fraction of the Top Set retaining a ring, no within- vs across-pipeline comparison. Given the authors had the code and the survey results, they could have given numbers instead of the qualitative 'stability of a ring.' That would have strengthened the central claim. Second, the parameter survey they present is coarse (powers of ten, one RML package) and shares the circular Gaussian prior across RML pipelines, so the explored space is narrower than the phrase 'countless plausible ways' suggests. The paper does hedge in the conclusion—'different choices, not tried or reported, could lead to greater variability'—but that hedge sits in tension with the strong wording of the 'most valuable evidence' claim. If the variability can't be robustly bounded, the demonstration is suggestive rather than conclusive. I don't think this breaks the paper; the central epistemic point—that the EHT's evidence is the ensemble, not the icon—holds up. But the authors should either quantify the variability or soften the claim.\n\nThe citation pattern looks honest, with self-citations peripheral and prior robustness work squarely acknowledged. The oral-history recollections are used lightly. No red flags in the math or data handling.\n\nWho this is for: historians and philosophers of science working on imaging, data assimilation, and robustness; also anyone teaching the EHT as a case study. It deserves a serious referee, and I'd take the review if asked. Recommend: engage, with the request that the authors quantify what they mean by 'limited variability' before publication.","headline":"A careful, credit-worthy HPS paper that reframes the EHT's evidence as the demonstrated variability across plausible imaging choices, though 'limited variability' remains more asserted than quantified.","tokens_in":35553,"tokens_out":1126,"would_cite":true,"duration_ms":17267,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["00A30","85-03"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that the 2019 M87* black hole image is not a single observation but the product of many defensible choices, and that the strongest evidence is the stability of the ring across those choices.","keywords":["Event Horizon Telescope","black hole imaging","M87*","VLBI","image reconstruction","scientific evidence","epistemic choices","model-laden data"],"falsifier":"Run the public 2017 M87* data through the EHT's own imaging pipelines with a denser parameter sweep—including regularizer weights between the powers of ten the collaboration tested—and check whether any plausible combination removes the ring or changes its diameter by more than the reported uncertainty; if so, the limited-variability claim collapses. Alternatively, if the EHT were to release fully unprocessed correlator data and a unique image emerged from independent re-analysis, the underdetermination claim would be undercut.","tokens_in":34637,"feed_emoji":"🕳️","tokens_out":4687,"duration_ms":51525,"temperature":0.7,"pith_summary":"This paper tries to establish that the famous 2019 image of the M87* black hole shadow was not dictated by the data. Because the telescope data are extremely noisy and sparse, researchers had to make a choice at every processing stage—how to calibrate, which imaging algorithm to use, how to weight the regularization terms, what colors to display, and even whether to crop the frame. The published picture was actually the pixel-by-pixel average of three images made with different pipelines, shown in a custom orange color map chosen partly for its cultural associations. By re-running the collaboration's open-source imaging software with different options, the authors show that many plausible renderings exist, yet a ring-and-shadow structure persists across a broad set of them. Their central claim is that the most valuable scientific evidence is not the single advertised image but the demonstration of limited variability across those defensible choices.","feed_headline":"Black hole photo's real result: a ring that survives reprocessing","feed_subtitle":"The famous M87 picture was an average of three images; the evidence is the variation across them.","key_machinery":"The central mechanism is the variability survey: re-executing the EHT's open-source imaging workflow—calibration, CLEAN- and regularized-maximum-likelihood reconstruction, parameter grids, and color mapping—under alternative choices, and treating the spread of resulting images as the object of analysis. The key conceptual tool is the likelihood–posterior distinction: many images fit the observed visibilities equally well, so priors and regularizers decide among them; the paper then takes the ensemble of plausible outputs, rather than any single output, as the epistemic unit.","core_discovery":"The paper's central claim is that the EHT's M87* result is an ill-posed inverse problem: the visibility data underdetermine the image, so countless images fit the measurements equally well. At each stage—calibration pipeline, CLEAN versus regularized-maximum-likelihood reconstruction, regularizer weights chosen by a coarse parameter survey on synthetic 'ground truth' images, and final color and cropping decisions—researchers had to select among reasonable alternatives. The final publicized image was the simple average of three 'fiducial' images from three different pipelines, displayed with a color map chosen for perceptual and aesthetic reasons. By systematically varying algorithms, paramet","pith_inferences":["A quantitative extension: define the evidence as the distribution of ring diameters and brightness asymmetries across the tested parameter combinations; if that distribution is wider than the EHT's reported uncertainties, the 'limited variability' conclusion would need qualification.","A natural test: apply the same re-sampling to the larger public data releases from 2022 and to the later machine-learning-based reconstruction; if the sharper machine-learning image falls outside the original variability envelope, that would show the envelope depends on the reconstruction family chosen.","The same logic transfers to other inverse problems without ground truth, such as medical or planetary image reconstruction, where the defensible output may be a set of equally plausible reconstructions rather than a single 'best' image."],"forward_implications":["The single published image should not be treated as standalone evidence; the evidence is the range of plausible images and the structural stability across them.","Presenting an average of three pipeline images as 'the photo' conflates processing with observation; public communication should report the variability explicitly.","The EHT's selection of parameters using synthetic 'ground truth' means the images are only as trustworthy as the assumption that those simulations adequately represent the source regime.","For future targets such as Sagittarius A* or next-generation arrays, the collaboration could make the full set of plausible reconstructions the public result rather than a single image."],"fun_headline_variants":["Black hole photo is a composite; the real evidence is the variability","M87's image: many possible renderings, the ring is what survives","The black hole image was a choice among countless plausible reconstructions","Black hole photo: an average of three pipelines, not a single snapshot","What the black hole photo hides: the many decisions behind one image"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The demonstration of 'countless plausible ways' assumes that the public open-source software, archived data, and the parameter survey the authors ran faithfully reproduce the full space of choices the EHT actually faced; if the collaboration's internal versions allowed different options, the variability could be over- or under-stated.","fun_headline_variants_meta":{"raw":{"variants":["Black hole photo is a composite; the real evidence is the variability","M87's image: many possible renderings, the ring is what survives","The black hole image was a choice among countless plausible reconstructions","Black hole photo: an average of three pipelines, not a single snapshot","What the black hole photo hides: the many decisions behind one image"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000471,"raw_usage":{"total_tokens":2143,"prompt_tokens":673,"completion_tokens":1470,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":417,"completion_tokens_details":{"reasoning_tokens":1377}},"tokens_in":417,"tokens_out":1470,"duration_ms":10349,"temperature":1.0,"reasoning_tokens":1377,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T00:39:43.401114+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the public 2017 M87* data through the EHT's own imaging pipelines with a denser parameter sweep—including regularizer weights between the powers of ten the collaboration tested—and check whether any plausible combination removes the ring or changes its diameter by more than the reported uncertainty; if so, the limited-variability claim collapses. Alternatively, if the EHT were to release fully unprocessed correlator data and a unique image emerged from independent re-analysis, the underdetermination claim would be undercut.","supporting_citations":[],"review_version":1}