{"id":"cfe4ecf5-926b-4997-a89f-a69c27aaaf47","arxiv_id":"2412.11426","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"The LADUMA survey measures the local HI mass function and cosmic HI density with a new recovery-matrix method, yielding Omega_HI = 3.09e-4, consistent with previous surveys.","lead":"A deep MeerKAT survey of a single patch of sky provides a new census of neutral hydrogen in nearby galaxies, using a newly developed recovery-matrix method that corrects for how galaxies can be detected in the wrong mass bin. The measured cosmic hydrogen density agrees with earlier surveys, and the new method is the paper's main technical contribution.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The recovery-matrix completeness correction is only as trustworthy as the synthetic galaxy models, and the paper provides no end-to-end validation that these models match real HI galaxies in detectability.","rationale":"The central claim is a set of Schechter parameters and Omega_HI. All of these flow through the recovery matrix, so the fidelity of the synthetic sources to real HI galaxies is the load-bearing assumption. The paper's own limitation statement (Section 3.1.5) acknowledges the absence of an end-to-end mock test, and the LSS test covers only high masses. The off-diagonal mass migration is the essential correction in the new method, and it is the hardest to validate without an independent test. I agree with the reader's weakest_assumption. The concern is addressable, so the verdict remains CONDITIONAL rather than REJECT: a sensitivity analysis of the synthetic source model would either retire the concern or reveal a bias. The other issues raised by the reader (cosmic variance, low-mass-bin exclusion) are real but secondary; they affect the uncertainty budget and the alpha value through a transparent modeling choice, whereas the synthetic realism goes to the core of the method.","tokens_in":29280,"tokens_out":10581,"duration_ms":97576,"concrete_test":"Recompute the recovery matrix and the forward-model fit using an alternative synthetic source population: (a) draw sizes from the Wang et al. relation with its observed scatter and velocity widths from the empirical ALFALFA width-mass relation, or (b) replace Sersic profiles with empirical HI surface-density profiles from a resolved sample (e.g., LITTLE THINGS/Westerbork). If the recovered alpha, phi*, and log M* shift by more than the 1-sigma posterior uncertainties in Table 2, the headline parameters are not robust to the assumed source properties. If they shift by less, the synthetic-realism objection is substantially weakened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The new recovery-matrix method rests on a completeness correction derived from ~150,000 synthetic sources generated with GIPSY galmod using Sersic surface-density profiles, Persic/Courteau rotation curves, and the mean Wang et al. (2016) size-mass relation (Section 3.1.1). The paper concedes in Section 3.1.5 that no end-to-end validation of the method has been performed, and the large-scale-structure sensitivity check in Footnote 7 is restricted to log(MHI/Msun) >= 9, i.e., 75% of the observed sources, leaving the low-mass bins that most directly determine the faint-end slope alpha untested. The recovery matrix's off-diagonal elements, which correct for the systematic upward mass migration driven by threshold effects and pixel-based continuum subtraction, depend on the distribution of velocity widths and surface brightness profiles of the injected sources. If real HI galaxies are broader, fainter, or clumpier than these smooth models, the completeness correction will be biased, and with it alpha, phi*, and Omega_HI. Cross-checks do not fully retire this concern: the MML comparison uses the same injection/recovery selection function, and the agreement with previous surveys has large error bars that could mask a systematic offset. This is not a fatal flaw but the absence of a validation that directly probes the assumed source realism.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents the first HI mass function (HIMF) and cosmic HI density (Omega_HI) measurement from the LADUMA survey's low-redshift (0 <= z <= 0.088) spectral window, using a new 'recovery matrix' (RM) forward-modeling method. The method injects roughly 150,000 synthetic galaxies into the processed data cube, tracks their recovery across mass bins, and fits Schechter function parameters with a Poisson likelihood. The authors benchmark the RM method against a modified maximum likelihood (MML) method and compare with previous surveys (ALFALFA, HIPASS, AUDS, MIGHTEE). The best-fit parameters are phi* = 3.56e-3 Mpc^-3 dex^-1, alpha = -1.18, log(M*/M_sun) = 10.01, and Omega_HI = 3.09e-4, with the claim that the LADUMA volume is not an outlier in HI content.","tokens_in":29529,"tokens_out":6881,"duration_ms":54985,"significance":"The paper's methodological contribution is timely and potentially useful for future deep single-pointing HI surveys. The recovery matrix approach addresses a real problem, cross-bin mass migration, and the use of a Poisson likelihood for low-count bins is well justified in Appendix A. The authors are transparent about key limitations, including the absence of end-to-end validation. If the measurement holds, it provides a low-redshift anchor for LADUMA's higher-redshift program. However, the current sample is small (82 sources) and the statistical weight is dominated by a few bins, so the significance of the scientific claim is moderate rather than transformative.","major_comments":[{"comment":"The paper explicitly states that no end-to-end validation of the recovery-matrix method has been performed, yet the completeness correction rests entirely on the assumption that the synthetic sources (Sersic surface-density profiles, Persic/Courteau rotation curves, and the Wang et al. 2016 size-mass relation) have the same detectability and flux-recovery properties as real HI galaxies. The large-scale-structure sensitivity test in Footnote 7 covers only log(MHI/Msun) >= 9, leaving the low-mass bins that most directly determine the faint-end slope alpha untested. I recommend adding a sensitivity analysis that varies the input source models (e.g., velocity-width distribution, size-mass relation scatter, surface brightness profile) and re-derives the Schechter parameters, or a partial validation using real sources with known properties, before the quoted parameters can be regarded as robust against source-model uncertainties.","section":"Section 3.1.5"},{"comment":"The uncertainty requirement in Section 3.1.2 is stated only for the diagonal and below-diagonal recovery-matrix elements, with the claim that their fractional uncertainties are less than 10%. However, Figure 2 shows that the dominant mass-migration corrections are the above-diagonal elements (e.g., 24.7% of sources in the lowest mass bin are recovered one bin higher). Since the forward modeling specifically corrects for this upward migration, the paper must quantify the uncertainties of the above-diagonal elements as well, and confirm that they are propagated in the 512 Monte Carlo realizations described in Section 3.1.5; otherwise the quoted parameter uncertainties may be underestimated.","section":"Section 3.1.2"},{"comment":"Appendix B demonstrates that cosmic variance increases the uncertainty in the number of high-mass (log(MHI/Msun) >= 9.75) galaxies in the LADUMA volume by roughly a factor of two relative to Poisson expectations (standard deviation ~8.5 versus sqrt(16) = 4). The abstract, Table 2, and Section 4.2 quote uncertainties on Omega_HI and the Schechter parameters that do not include this contribution. Either propagate cosmic variance into the final uncertainties (or an explicit component of them), or clearly state in the abstract and conclusions that the quoted errors are conditional on the surveyed volume being representative of the cosmic mean. As written, the claim that the LADUMA volume is 'not an outlier' relies on error bars that may be too small.","section":"Appendix B and Section 4.2"},{"comment":"The decision to exclude the lowest-mass bin (two sources) from the analysis is based on an ad hoc 10% minimum accepted volume fraction, and the text states that including that bin would 'result in a significantly steeper value for alpha.' Since alpha is a headline result, the sensitivity of the fitted alpha and its uncertainty to this threshold should be quantified (e.g., by varying the threshold between about 5% and 20%), or a more principled criterion should be presented. Without such a test, the robustness of the reported faint-end slope to this modeling choice is not established.","section":"Section 5.1"}],"minor_comments":[{"comment":"The parameter uncertainties quoted in Section 6 differ from those in the abstract and Table 2 for the same quantities (e.g., phi* = 3.56^+1.79_-1.51 versus 3.56^+0.97_-1.92; alpha = -1.18 +/- 0.14 versus -1.18^+0.08_-0.19; log(M*/M_sun) = 10.01^+0.23_-0.17 versus 10.01^+0.31_-0.12). Please reconcile these values or explain the difference.","section":"Section 6 vs. Abstract/Table 2"},{"comment":"Section 4.2 quotes Omega_HI = 3.09^+0.58_-0.49e-4 for the recovery matrix method, while the abstract and Table 2 list Omega_HI = 3.09^+0.65_-0.47e-4 for the same quantity. This discrepancy should be fixed.","section":"Section 4.2"},{"comment":"The sentence stating that the Schechter parameters from the second iteration 'change by less than 1%' compared to the first iteration would be more informative if the actual changes in phi*, alpha, and log(M*) were reported, along with the convergence criterion used. As written, the claim is difficult to check.","section":"Section 3.1.5"},{"comment":"There is a typographical error in the phrase 'compeletness fractions' in the footnote to Section 5.4; it should read 'completeness fractions.'","section":"Section 5.4"}],"recommendation":"major_revision","confidential_remarks":"The paper is within the scope of the journal and the reader's conditional verdict is fair. The main risk is that the recovery-matrix completeness correction is not validated end-to-end, and the absence of a sensitivity analysis around the synthetic source models is the largest gap. The cosmic-variance issue is also important because the paper's central 'not an outlier' claim depends on the quoted error bars. These issues are fixable with additional analysis and explicit caveats, so major revision is appropriate rather than rejection. The inconsistency between the Section 6 and abstract/table uncertainties should be resolved during revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The useful thing to know first: this is a careful measurement of the local HI mass function from a single MeerKAT pointing, built around a genuinely new methodological piece, a 2D recovery matrix that tracks what fraction of sources injected in one mass bin get recovered in another, with the pixel-based continuum subtraction included in the injection loop. The Poisson forward model that handles zero-detection bins is a real improvement over the usual Gaussian-with-sqrt-N approach, and the appendix comparing likelihood choices is worth reading on its own. The headline numbers (phi* ~ 3.6e-3, alpha ~ -1.18, log M* ~ 10.0, Omega_HI ~ 3.1e-4) are consistent with ALFALFA and HIPASS within the large errors, and the RM and MML methods agree. The paper is honest about its soft spots. The quoted uncertainties do not include cosmic variance; Appendix B shows that variance roughly doubles the count uncertainty in the high-mass bins, and the authors say they cannot easily propagate that to Schechter parameters. That is a real limitation, and I would ask them to state it once more in the abstract or conclusions rather than leaving it in an appendix. Excluding the lowest-mass bin because its accessible volume is only about 2% of the survey volume is defensible, but the 10% volume-cut threshold in Section 5.3 does look somewhat ad hoc. The absence of end-to-end validation is the biggest soft spot: the recovery matrix is only as good as the synthetic source models (smooth Sersic profiles, a fixed size-mass relation), and the paper's argument that a single mock realization cannot validate the method due to Poisson noise is reasonable but leaves end-to-end accuracy unproven. The large-scale-structure check is also restricted to log M >= 9, leaving the faint-end bins untested. The MML comparison is not fully independent because it uses the same injection/recovery selection function. Minor: the quoted errors on the Schechter parameters differ between the abstract, the summary, and Figure 5. That is probably just a marginalized versus 3D-posterior choice, but it should be cleaned up. The central result holds up. The method is novel, the analysis is transparent, and the limitations are stated rather than hidden. This deserves a serious referee, and I would expect publication after moderate revision, ideally with the cosmic variance caveat front and center and at least an attempt at a small mock ensemble, however limited.","headline":"Genuinely new recovery-matrix method for the HI mass function, carefully executed and honestly caveated; deserves a serious referee and is likely publishable after moderate revision.","tokens_in":673,"tokens_out":1000,"would_cite":true,"duration_ms":39299,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The LADUMA survey's local HI mass function, measured with a new recovery-matrix method, matches previous low-redshift surveys, implying the volume is not an outlier in cosmic hydrogen content.","keywords":["neutral hydrogen mass function","cosmic HI density","MeerKAT","LADUMA","recovery matrix","forward modeling","21 cm line","Schechter function"],"falsifier":"Build a source-free version of the LADUMA cube, inject a galaxy population drawn from a known Schechter function (say $\\phi_\\ast = 4\\times10^{-3}$, $\\alpha = -1.25$, $\\log M_\\ast/M_\\odot = 10$), run the full pipeline, and repeat across enough realizations to beat Poisson scatter; if the median recovered parameters deviate from the input by more than the quoted uncertainties, the recovery-matrix correction is biased.","tokens_in":29087,"feed_emoji":"📡","tokens_out":5326,"duration_ms":46109,"temperature":0.7,"pith_summary":"The paper sets out to measure how many galaxies of each neutral-hydrogen mass live in the local universe, using the first data release of the deep MeerKAT LADUMA survey out to redshift 0.088. It claims that a new \"recovery matrix\" method, which tracks not just whether a synthetic galaxy is detected but which mass bin it lands in, corrects survey incompleteness more faithfully than the usual one-dimensional completeness corrections. With it, the survey finds a Schechter mass function with normalization $\\phi_\\ast = 3.56\\times10^{-3}\\,\\mathrm{Mpc^{-3}\\,dex^{-1}}$, faint-end slope $\\alpha = -1.18$, and knee mass $\\log(M_\\ast/M_\\odot)=10.01$, implying a cosmic hydrogen density $\\Omega_{\\mathrm{HI}}=3.09\\times10^{-4}$. The result matters because it independently confirms the local HI census from an untargeted, single-pointing deep survey, indicating that cosmic variance does not make this volume an outlier.","feed_headline":"Deep MeerKAT survey pins down cosmic hydrogen inventory","feed_subtitle":"The local HI census from 82 galaxies matches earlier surveys to within 1 sigma.","key_machinery":"The central object is the recovery matrix $R_{ij}$, the fraction of synthetic galaxies injected into mass bin $i$ that are recovered in mass bin $j$. Unlike a one-dimensional completeness vector, it captures the asymmetric tendency of sources to be recovered in a higher mass bin than the one they were injected into, a bias produced by source-finding thresholds and by the pixel-based continuum subtraction. The matrix is built from over 150,000 synthetic sources generated with a tilted-ring model, with Sersic surface-density profiles, Persic/Courteau rotation curves, the Wang et al. size-mass relation, and random inclinations and positions, and is then used in a forward model that predicts observed bin counts for trial Schechter parameters and compares them to data with a Poisson likelihood in an MCMC fit.","core_discovery":"The central claim is that the neutral atomic hydrogen mass function in the local volume probed by LADUMA is well described by a Schechter function with $\\phi_\\ast = 3.56\\times10^{-3}\\,\\mathrm{Mpc^{-3}\\,dex^{-1}}$, $\\alpha = -1.18$, $\\log(M_\\ast/M_\\odot)=10.01$, and correspondingly $\\Omega_{\\mathrm{HI}}=3.09\\times10^{-4}$, and that these values agree with earlier low-redshift surveys within the quoted uncertainties. The paper further claims that the recovery-matrix forward-modeling approach, benchmarked against a modified maximum likelihood method that uses the same injection/recovery selection function, yields mutually consistent parameters, and that the inclusion of a Poisson likelihood is required to avoid biased slopes and underestimated uncertainties in bins with few or zero detections. On this basis the paper concludes that LADUMA's cosmic volume is not an outlier in HI content.","pith_inferences":["The recovery-matrix idea transfers directly to other binned distribution measurements, such as optical luminosity functions or CO luminosity functions, wherever measurement scatter moves objects across bin boundaries asymmetrically.","Because the lowest-mass bin is excluded and the faint-end slope is constrained from a small accessible volume, the quoted $\\alpha$ should be read as applying to the mass range actually fit; a wider or deeper survey could reveal a steeper faint end.","A natural next test is to apply the recovery matrix to a mock universe with known large-scale structure, once sample sizes grow, to separate cosmic-variance noise from method bias.","If the upward mass-scattering tendency is generic, earlier surveys that used diagonal-only completeness corrections may have slightly overestimated the faint-end slope; reanalyzing their injection/recovery data with a matrix would settle this."],"forward_implications":["Future, deeper LADUMA data releases can use the same recovery-matrix machinery to push the HI mass function to higher redshift, where detection counts are lower and cross-bin corrections matter more.","Surveys with small numbers of detections can include zero-detection mass bins in the fit, tightening the constraint on the knee mass without ad hoc corrections.","The local cosmic HI density $\\Omega_{\\mathrm{HI}}=3.09\\times10^{-4}$ becomes a consistent anchor point that evolution studies can compare against at higher redshifts.","The agreement between the recovery-matrix and modified maximum likelihood methods, and with earlier surveys, supports the view that the single LADUMA pointing is not strongly biased by cosmic variance for the mass range probed.","Using a Poisson rather than Gaussian likelihood removes a systematic toward steeper low-mass slopes in low-count bins."],"supporting_citations":[{"why":"Supplies the synthetic-source injection strategy that the recovery-matrix tests build on.","marker":"(Gogate et al. 2020)"},{"why":"Provides the rotation-velocity profiles used to generate synthetic galaxy models.","marker":"(Persic et al. 1996)"},{"why":"Supplies the rotation-curve parameterization used for synthetic sources.","marker":"(Courteau 1997)"},{"why":"Provides the HI size-mass relation that fixes synthetic source sizes.","marker":"(Wang et al. 2016)"},{"why":"Defines the functional form fitted to the HI mass function.","marker":"(Schechter 1976)"},{"why":"Supplies ALFALFA comparison parameters and the initial recovery-matrix iteration seed.","marker":"(Jones et al. 2018)"},{"why":"Defines the modified maximum likelihood method and dftools used for benchmarking.","marker":"(Obreschkow et al. 2018)"},{"why":"Gives the HI mass conversion and the analytic $\\rho_{\\rm HI}$ formula used for $\\Omega_{\\rm HI}$.","marker":"(Meyer et al. 2017)"},{"why":"Supplies the SoFiA source finder that detects both real and injected sources, defining the recovery path.","marker":"(Westmeier et al. 2021)"}],"fun_headline_variants":["MeerKAT's LADUMA survey sharpens local cosmic HI census","New MeerKAT HI mass function agrees with older surveys","LADUMA survey pins down cosmic hydrogen density","MeerKAT's LADUMA delivers precise HI mass function","Cosmic hydrogen census tightened by MeerKAT's LADUMA"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The completeness correction is trustworthy only if the roughly 150,000 synthetic sources, built from Sersic profiles, Persic/Courteau rotation curves, and the Wang size-mass relation, are detected and flux-recovered the same way real HI galaxies are, including through the pixel-based continuum subtraction and SoFiA source finding.","fun_headline_variants_meta":{"raw":{"variants":["MeerKAT's LADUMA survey sharpens local cosmic HI census","New MeerKAT HI mass function agrees with older surveys","LADUMA survey pins down cosmic hydrogen density","MeerKAT's LADUMA delivers precise HI mass function","Cosmic hydrogen census tightened by MeerKAT's LADUMA"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000421,"raw_usage":{"total_tokens":2256,"prompt_tokens":1127,"completion_tokens":1129,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":743,"completion_tokens_details":{"reasoning_tokens":1052}},"tokens_in":743,"tokens_out":1129,"duration_ms":8524,"temperature":1.0,"reasoning_tokens":1052,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T14:56:56.140868+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Build a source-free version of the LADUMA cube, inject a galaxy population drawn from a known Schechter function (say $\\phi_\\ast = 4\\times10^{-3}$, $\\alpha = -1.25$, $\\log M_\\ast/M_\\odot = 10$), run the full pipeline, and repeat across enough realizations to beat Poisson scatter; if the median recovered parameters deviate from the input by more than the quoted uncertainties, the recovery-matrix correction is biased.","supporting_citations":[],"review_version":1}