{"id":"fadc8cd9-ab24-4172-85fa-b260551eaf06","arxiv_id":"1908.11493","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"For an LSST-like survey, four-bin tomographic weak lensing peak counts shrink the Omega_m-sigma_8 error contour by about a factor of five relative to 2D peak counts, and can simultaneously constrain photo-z bias to 10% and scatter to 5%.","lead":"This paper forecasts how much cosmological measurements improve when weak lensing signals from distant galaxies are split into redshift bins, and how errors in galaxy distance estimates affect the results. It finds four bins are near-optimal and that distance calibration bias matters more than scatter.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The numerical forecasts are self-consistent model calculations; the unvalidated off-fiducial derivatives, not the fiducial peak-count amplitude, carry the factor-5 gain and photo-z degradation factors.","rationale":"The paper is a careful forecast study: the model is described, ray-tracing validation is shown for the fiducial cosmology, the covariance is estimated from simulations with a Hartlap correction, and the Fisher/MCMC consistency is checked in Appendix A. Credit is due for the transparent acknowledgment that covariance scaling ignores super-survey modes and that the photo-z model is idealized. The concern I identify is not an internal inconsistency; it is a scope-of-validation gap. The central numbers -- factor 5 gain, sigma(z_bias)/z_bias ~10%, sigma(sigma_ph)/sigma_ph ~5%, degradation factors 2.2 and 1.8 -- are all derived from a model that is used both to create the mock data and to evaluate the likelihood. Validating the amplitude of the peak counts at one cosmology does not validate the derivatives with respect to cosmology and photo-z parameters, which control all contour areas. The factor-5 comparison between 2-D and 4-bin analyses could in principle be less sensitive to model error because both sides use the same model, but the absolute constraints and the photo-z degradation factors are directly proportional to the model's sensitivity. Because the paper frames these as quantitative design targets for LSST-like surveys, the proposed test would settle whether the numbers are robust or are artifacts of the theoretical peak-count model. This does not change the CONDITIONAL verdict; it strengthens the condition under which the numerical claims should be adopted.","tokens_in":18029,"tokens_out":5574,"duration_ms":55946,"concrete_test":"Run the ray-tracing pipeline of Sec. 3 at two additional cosmologies, e.g. (Omega_m, sigma_8) = (0.26, 0.82) and (0.30, 0.80), chosen near the 1-sigma contour, and compute the 2-D and 4-bin peak-count vectors for the same nu >= 4 bins. Compare the simulated count differences Delta N between cosmologies with the model predictions of Sec. 2. If the fractional difference in Delta N exceeds roughly 10-20%, the derivatives that set the contour sizes are not validated; replace the model-generated mock data in Fig. 4 with the simulated vectors and recompute the area ratio and photo-z degradation factors.","verdict_should_be":"UNCHANGED","load_bearing_attack":"In Sec. 4 the forecasts for Nbin = 0, 2, 4 and 8 use 'mock observational data generated directly from our model calculations' (Sec. 4, after Eq. 20), and the same Yuan et al. (2018) model supplies the likelihood prediction. The simulation checks in Figs. 1 and 2 validate the peak-count amplitude at the fiducial cosmology only. They do not validate the derivatives dN/dOmega_m, dN/dsigma_8, dN/dz_bias, and dN/dsigma_ph that determine the Fisher/MCMC contours, because no off-fiducial or photo-z-perturbed ray-tracing maps are compared with the model. The claimed factor-of-5 reduction in the 1-sigma contour area (Fig. 4) and the photo-z degradation factors (Figs. 6 and 8) are ratios of contours obtained from this same self-consistent model. If the model's response of high-peak counts to cosmology is inaccurate -- for instance, because the one-halo subtraction in Eq. (8) and the halo mass function are calibrated only at the fiducial point -- these headline numbers shift. The paper is transparent about the limitation, but the central quantitative claims are therefore conditional on the model's off-fiducial response, not empirically established.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper uses an analytic halo-based model for high weak-lensing peak counts to forecast tomographic peak-abundance constraints for an LSST-like survey. It validates the model's peak-count amplitude against ray-tracing simulations at the fiducial cosmology (Figs. 1-2), then runs MCMC forecasts with mock data vectors generated from the same model (Sec. 4) to compare 0/2/4/8-bin configurations, reporting a factor-of-5 reduction in the (Omega_m, sigma_8) 1-sigma contour area for 4 bins. It then uses a Fisher-matrix approach, calibrated by MCMC, to study simultaneous constraints on Omega_m, sigma_8 and the photo-z parameters z_bias and sigma_ph (Sec. 5), reporting ~10% constraints on z_bias, about 3-5% constraints on sigma_ph, degradation factors of ~2.2 and ~1.8 on the cosmological parameters, and spec-z calibration requirements. The paper is transparent that its forecasts use mock data generated from the same model that supplies the likelihood predictions, and it acknowledges the covariance-scaling approximation for large areas.","tokens_in":18301,"tokens_out":13242,"duration_ms":132082,"significance":"If the headline numbers are robust, the paper provides useful quantitative guidance for LSST/Euclid/CSST tomographic peak analyses and for photo-z calibration requirements. The analysis pipeline has genuine strengths: the model is checked against ray-tracing simulations at the fiducial cosmology, MCMC is used where the likelihood is non-Gaussian, the Hartlap correction is applied, and the Sherman-Morrison-based prior propagation is clearly laid out. The central limitation is that the reported factor-of-5 gain and the photo-z degradation factors are ratios of contours obtained from a self-consistent model calculation; the simulation checks do not exercise the off-fiducial parameter derivatives that drive these ratios. The results should therefore be read as model-conditional forecasts unless the derivatives are validated or the claims are explicitly qualified.","major_comments":[{"comment":"The headline quantitative results are produced by using the same Yuan et al. model to generate the mock data vectors and to evaluate the theoretical predictions in the chi-squared of Eq. (19), while the ray-tracing simulations in Figs. 1-2 validate the peak-count amplitude only at the fiducial cosmology. The factor-of-5 gain and the photo-z degradation factors are determined by the model's derivatives of the peak counts with respect to Omega_m, sigma_8, z_bias, and sigma_ph, and these derivatives are not checked against off-fiducial or photo-z-perturbed simulations. I recommend either adding such tests, even at a few off-fiducial points, or explicitly and prominently reframing the quantitative results in the abstract and conclusions as model-conditional forecasts rather than simulation-validated predictions.","section":"Sec. 4 and Sec. 5 (Eq. 19; Figs. 1-8)"},{"comment":"The abstract states that for surveys with area ~15000 deg^2 the 4-bin tomographic analysis reduces the error contours by a factor of 5, but the factor-of-5 result in Fig. 4 is computed in Sec. 4 for the 876 deg^2 effective simulation area using MCMC. The 15000 deg^2 Fisher analysis in Sec. 5 does not report the corresponding tomographic-versus-2D gain. If the ratio is assumed to be area-independent under the covariance scaling of Eq. (19), that assumption should be stated and justified; otherwise the abstract overstates the support for one of the paper's central claims.","section":"Abstract and Sec. 4"}],"minor_comments":[{"comment":"The abstract reports sigma(sigma_ph)/sigma_ph ~ 5%, while Sec. 6 reports sigma(sigma_ph) ~ 6e-4, which for the fiducial sigma_ph = 0.02 is about 3%. These numbers should be reconciled.","section":"Abstract and Sec. 6"},{"comment":"The symbol Nbin is used both for the number of redshift bins in Sec. 4 and for the number of peak-count data bins in the Hartlap correction factor. Using a separate symbol, such as N_data, for the data-vector dimension would avoid confusion.","section":"Eq. (20)"},{"comment":"Because the Fisher matrix is obtained by inverting the MCMC covariance in Eqs. (27)-(28), Fig. A1 is best described as a Gaussianity check of the posterior rather than an independent validation of the Fisher approximation against the MCMC calculation. Please rephrase the text to avoid overclaiming.","section":"Sec. 5.2 and Appendix A"},{"comment":"The covariance matrix used in Sec. 5 is scaled from simulations generated without photo-z errors and is assumed to be independent of z_bias and sigma_ph. A brief justification, or a quantitative estimate of the impact of this approximation on the photo-z constraints, would strengthen the analysis.","section":"Sec. 5"},{"comment":"The layout of Table 1, which combines two correlation matrices in one table using bold and non-bold entries, is difficult to parse. A separate sub-table or explicit row and column labels for the 2-bin and 4-bin cases would improve readability.","section":"Table 1"}],"recommendation":"major_revision","confidential_remarks":"This is a competent and clearly written forecast paper, but its central quantitative claims are model-conditional in a way that the abstract does not fully convey. The main issue is not internal inconsistency but the need to either validate the off-fiducial derivatives or qualify the claims. I would be comfortable with acceptance after a major revision that adds such validation or reframing, and after the abstract's area attribution and photo-z percentage are corrected."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the paper is a solid forecast study, and the stress-test note is basically right. The headline numbers are conditional on the model's off-fiducial response, not empirically established. But there's no load-bearing error; it's a well-done model-based forecast with real practical guidance.\n\nWhat's new: it quantifies how many tomographic bins actually help for high peaks (4 is nearly optimal, factor ~5 over 2D), and it shows that 4-bin peak tomography alone can constrain photo-z bias to ~10% and scatter to ~5% without priors, with corresponding degradation factors on Omega_m and sigma_8. It also converts degradation requirements into spec-z sample sizes (roughly 10^4 per bin for 1.5 degradation). That's a genuinely useful mapping for survey planning.\n\nWhat's done well: the pipeline is careful. Model-vs-simulation agreement is shown at the fiducial cosmology (Figs 1-2). MCMC is used for the small-area contours, Fisher is checked against MCMC in Appendix A, and the Hartlap correction is applied. The paper explicitly flags the covariance scaling limitation (no super-survey covariance) and the simple Gaussian photo-z model. No hidden circularity in the sense of hiding the model dependence – they state clearly that mock data are generated from the model.\n\nSoft spots: the main one is that the derivative response dN/dOmega_m, dN/dsigma_8, dN/dz_bias, dN/dsigma_ph is not validated against off-fiducial or photo-z-perturbed ray-tracing simulations. The factor-5 gain and degradation factors are contour ratios from the same self-consistent model. If the model's off-fiducial response is off, those numbers shift. The paper acknowledges this in spirit, but does not quantify the sensitivity. That's a moderate concern, not fatal. Also, the photo-z model has no outliers and no redshift-dependent error parameters beyond (1+z) scaling; the paper says so. The super-survey covariance issue is acknowledged but left unquantified.\n\nWho this is for: anyone doing WL peak tomography or photo-z calibration for LSST/Euclid/CSST. It's a useful design-guidance paper.\n\nRecommendation: yes, send to peer review. A good referee should push for an off-fiducial validation or at least a rough sensitivity test, but the paper is serious, honest, and worth engaging with.","headline":"Solid, transparent forecast study for tomographic WL peak counts; the factor-5 gain and photo-z degradation numbers are self-consistent model results, not empirically pinned, but the paper is honest and deserves peer review.","tokens_in":18863,"tokens_out":2366,"would_cite":true,"duration_ms":21630,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Tomographic lensing peak counts shrink cosmology error contours by a factor of five and can self-calibrate photometric redshift errors.","keywords":["weak lensing","cosmic shear","peak statistics","tomography","photometric redshifts","cosmological parameters","Fisher forecast"],"falsifier":"Generate ray-traced convergence maps at a non-fiducial cosmology, say $\\Omega_{\\rm m}=0.26,\\ \\sigma_8=0.84$ or $\\Omega_{\\rm m}=0.30,\\ \\sigma_8=0.80$, with source density 40 arcmin$^{-2}$, run the same four-bin high-peak analysis with and without photo-z errors, and compare the peak-count vectors to the model predictions; if discrepancies exceed the forecast's statistical errors scaled to 15,000 deg$^2$, the reported factor-5 gain and the 2.2/1.8 degradation numbers would need revision.","tokens_in":17827,"feed_emoji":"🔭","tokens_out":7415,"duration_ms":61545,"temperature":0.7,"pith_summary":"Using a halo-based theoretical model for high weak-lensing peaks, the paper asks how much cosmological information is gained by splitting source galaxies into redshift bins, and how severely photometric redshift (photo-z) errors degrade that gain. For a large future survey with about 40 galaxies per square arcminute and 15,000 square degrees, it finds that four-bin tomography shrinks the joint 1-$\\sigma$ uncertainty contour on $\\Omega_{\\rm m}$ and $\\sigma_8$ by roughly a factor of five relative to a single two-dimensional peak count. The paper further claims that the same four-bin peak data can constrain the photo-z bias parameter to about 10 percent and the photo-z scatter to about 5 percent without external priors, while the cosmological errors degrade by factors near 2.2 and 1.8. Because photo-z bias and scatter are major practical nuisances for future lensing surveys, this is a direct test of whether peak statistics can carry their own redshift calibration.","feed_headline":"Four redshift bins make lensing peaks five times stronger","feed_subtitle":"High peaks in weak-lensing maps can also measure their own photo-z errors, easing spectroscopic calibration.","key_machinery":"The machinery is a halo-based theoretical model for high weak-lensing peak abundances, extended by Yuan et al. (2018) from Fan et al. (2010). In halo regions the smoothed convergence is written as $K=K_{\\rm H}+K_{\\rm LSS}+N$, where $K_{\\rm H}$ is the contribution from massive halos, $K_{\\rm LSS}$ is a Gaussian random field from large-scale structure projection, and $N$ is Gaussian shape noise; the model then uses Gaussian random field theory, modulated by halo profiles and weighted by the halo mass function, to count peaks in halo and field regions separately. This model supplies the cosmology-dependent and photo-z-dependent peak counts used both to generate mock data vectors and to build the likelihood in the forecasts. The photo-z error enters through the conditional distribution $p(z_{\\rm ph}|z)$ with bias and scatter proportional to $(1+z)$, which shifts and broadens the true redshift distribution assigned to each tomographic bin.","core_discovery":"The central claim is that high convergence peaks (signal-to-noise ratio $\\nu\\geq 4$) in tomographic weak-lensing maps form a cosmological probe whose information content saturates at about four source-redshift bins, and whose sensitivity to photo-z errors is dominated by the bias rather than the scatter. Concretely, the paper forecasts that for source density $\\sim 40\\,{\\rm arcmin^{-2}}$, median redshift $\\sim 1$, and area $\\sim 15{,}000\\,{\\rm deg^2}$, four-bin tomography reduces the 1-$\\sigma$ area in the $\\Omega_{\\rm m}$--$\\sigma_8$ plane by a factor of 5 relative to two-dimensional analysis in the absence of photo-z errors. With photo-z errors modeled as a Gaussian conditional distribution with bias $z_{\\rm bias}(1+z)$ and scatter $\\sigma_{\\rm ph}(1+z)$ at fiducial values $z_{\\rm bias}=0.003$ and $\\sigma_{\\rm ph}=0.02$, the four-bin peak abundance alone constrains $z_{\\rm bias}$ to about $3\\times 10^{-4}$ and $\\sigma_{\\rm ph}$ to about $6\\times 10^{-4}$; $\\Omega_{\\rm m}$ degrades by a factor of 2.2 and $\\sigma_8$ by 1.8 relative to perfectly known photo-z parameters. The paper also establishes that the bias parameter is the bottleneck: a prior near $10^{-4}$ on $z_{\\rm bias}$ keeps degradation below 1.5, while priors on $\\sigma_{\\rm ph}$ have almost no effect in the four-bin case.","pith_inferences":["If the model's predicted response to cosmology holds away from the fiducial point, the factor-5 gain implies that high-peak tomography could substitute for part of the spectroscopic calibration effort, since four-bin peak data alone may meet photo-z requirements that otherwise need thousands of spectroscopic redshifts per bin.","The self-calibration result suggests a natural cross-check: comparing photo-z parameters inferred from peak counts with those from shear two-point analyses could expose unmodeled systematics such as intrinsic alignments or baryonic effects, which the paper notes can degenerate with photo-z errors.","The optimal bin number likely depends on survey depth and source density; for shallower surveys with lower galaxy density, shape noise may push the optimum below four bins, while deeper surveys could make eight bins worthwhile.","Because the forecasts use the same theoretical model for both mock data and likelihood, a stricter test would repeat the forecast with simulated peak maps at several non-fiducial cosmologies and photo-z error values; until then, the quantitative factors are conditional on the model's accuracy."],"forward_implications":["Four tomographic bins are near-optimal for high-peak analyses with a deep, dense source distribution; increasing to eight bins yields little additional constraining power.","In the ideal case with perfect photo-z information, four-bin peak tomography improves the $\\Omega_{\\rm m}$--$\\sigma_8$ 1-sigma contour area by a factor of about 5 over two-dimensional peak counts.","The same four-bin peak data can self-calibrate photo-z bias and scatter to roughly 10% and 5% respectively, at the cost of degrading cosmological constraints by factors of about 2.2 for $\\Omega_{\\rm m}$ and 1.8 for $\\sigma_8$.","Photo-z bias, not scatter, dominates the degradation; holding the bias prior near $10^{-4}$ limits degradation to about 1.5, which translates to roughly $10^4$ spectroscopic redshifts per calibration bin.","The model and methodology transfer directly to aperture-mass peaks constructed from shear, so the forecast framework extends beyond convergence peaks."],"supporting_citations":[{"why":"Supplies the halo-based theoretical model for high peak abundances, including large-scale structure projection and shape noise, that generates the forecast peak counts.","marker":"Yuan et al. (2018)"},{"why":"Introduces the original halo-based high-peak model that Yuan et al. extend, treating peaks in massive halo regions with Gaussian shape noise.","marker":"Fan et al. (2010)"},{"why":"Provides the source redshift distribution and survey specifications used for the mock maps and forecasts.","marker":"LSST Science Collaboration et al. (2009)"},{"why":"Supplies the N-body ray-tracing simulations used to validate the model and compute the peak-count covariance.","marker":"Liu et al. (2015b)"},{"why":"Gives the unbiased inverse-covariance correction applied in the MCMC forecasts for the bin-number comparison.","marker":"Hartlap et al. (2007)"},{"why":"Provides the Gaussian photo-z conditional distribution and the spec-z calibration relations used to model photo-z errors and derive spectroscopic requirements.","marker":"Ma et al. (2006)"},{"why":"Supplies the Sherman-Morrison-based formula that translates photo-z priors into posterior error degradations.","marker":"Amendola & Sellentin (2016)"}],"fun_headline_variants":["Lensing peak tomography: four bins beat 2D by fivefold","Photo-z bias dominates errors in lensing peak counts","Weak lensing peaks self-calibrate photo-z bias","Tomographic peaks: more bins than four yield little gain"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The forecast's numbers rest on the assumption that the halo-based peak-abundance model predicts how high-peak counts respond to cosmology and to photo-z errors correctly; the model is checked against ray-tracing simulations only at one fiducial cosmology, yet it is used both to make the mock data and to evaluate the likelihood.","fun_headline_variants_meta":{"raw":{"variants":["Lensing peak tomography: four bins beat 2D by fivefold","Photo-z bias dominates errors in lensing peak counts","Weak lensing peaks self-calibrate photo-z bias","Tomographic peaks: more bins than four yield little gain"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000719,"raw_usage":{"total_tokens":3391,"prompt_tokens":1270,"completion_tokens":2121,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":886,"completion_tokens_details":{"reasoning_tokens":2052}},"tokens_in":886,"tokens_out":2121,"duration_ms":14572,"temperature":1.0,"reasoning_tokens":2052,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T10:14:36.901157+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Generate ray-traced convergence maps at a non-fiducial cosmology, say $\\Omega_{\\rm m}=0.26,\\ \\sigma_8=0.84$ or $\\Omega_{\\rm m}=0.30,\\ \\sigma_8=0.80$, with source density 40 arcmin$^{-2}$, run the same four-bin high-peak analysis with and without photo-z errors, and compare the peak-count vectors to the model predictions; if discrepancies exceed the forecast's statistical errors scaled to 15,000 deg$^2$, the reported factor-5 gain and the 2.2/1.8 degradation numbers would need revision.","supporting_citations":[{"cited_title":"2018, ApJ, 857, 112","cited_arxiv_id":null,"evidence_quote":"Supplies the halo-based theoretical model for high peak abundances, including large-scale structure projection and shape noise, that generates the forecast peak counts."},{"cited_title":"2010, ApJ, 719, 1408","cited_arxiv_id":null,"evidence_quote":"Introduces the original halo-based high-peak model that Yuan et al. extend, treating peaks in massive halo regions with Gaussian shape noise."},{"cited_title":"2016, MNRAS, 457, 1490","cited_arxiv_id":null,"evidence_quote":"Supplies the Sherman-Morrison-based formula that translates photo-z priors into posterior error degradations."}],"review_version":1}