{"id":"0a825870-2f78-48d5-a1d1-80bea50f2951","arxiv_id":"1908.11508","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Survey-averaged Alcock-Paczynski parameters recover the BAO scale more accurately than single-redshift parameters in models with strong metric gradients, and a ~1% systematic remains that no constant AP scaling can remove.","lead":"This paper quantifies how well the standard Alcock-Paczynski correction recovers the baryon acoustic oscillation (BAO) distance scale when the assumed cosmology is far from the true one. It shows that survey-averaged correction parameters work better than parameters evaluated at a single redshift, and that a residual systematic near 1 percent remains even with the exact correction.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The unresolved ~1% systematic may be an artifact of shared fitting assumptions (standard-ruler and empirical shape), not a real AP-independent failure; the toy-model local-metric issue is secondary.","rationale":"The reader's verdict is CONDITIONAL, which is appropriate. I agree that the paper has a real soft spot, but I identify it as the unresolved origin of the 1% systematic rather than the local-metric approximation (2.4) for toy models. The headline claim in the abstract is the 1% systematic that cannot be attributed to any constant AP approximation. That claim is tested on large-scale smooth cosmologies (flat Lambda-CDM, Milne, timescape) where (2.4) is verified to high accuracy, so the toy-model approximation issue is not load-bearing for that claim. The load-bearing condition for the 1% claim is that the 'exact' AP scaling calculation (5.8) and the Lambda-CDM template cross-check are genuinely independent and free of shared artifacts. Both assume a redshift-independent BAO standard ruler, and both use fixed template shapes (Gaussian or Lambda-CDM power spectrum) that may be insufficiently flexible when the fiducial model is distant. The paper cites <0.5% BAO scale evolution in Lambda-CDM, which is not negligible compared to the 1% systematic for |alpha_bar-1|~0.1. A narrow-redshift-bin measurement of r(z) from the QPM mocks would directly test the standard-ruler assumption, and a flexible-r fit would show whether the bias is a modeling artifact. This is a concrete, feasible check that does not require new simulations. If the test shows r(z) is constant and the bias persists, the paper's central claim is strengthened; if r(z) varies substantially, the claim needs qualification. The paper's strengths include the mock-based validation, the analytic bounds in Section 3.2, and the cross-check with a different fitting procedure, and I do not see grounds for rejection. The verdict remains CONDITIONAL: the paper should be published with the source of the 1% residual addressed or explicitly tested against the standard-ruler assumption.","tokens_in":39164,"tokens_out":11104,"duration_ms":111218,"concrete_test":"Using the 1000 QPM CMASS mocks in the reference Lambda-CDM cosmology, measure the BAO scale r in narrow redshift bins (e.g., 6 bins of width Delta z ~ 0.045 spanning 0.43 < z < 0.7) with the same empirical Gaussian fitting function used in Section 5.2, calibrating each bin independently. If the recovered r(z) varies by more than ~0.5% (i.e., a significant fraction of the claimed 1% systematic), the redshift-independent standard-ruler assumption is violated at the relevant level. Then repeat the full-survey analysis of Section 5.3 with a modified fitting function where r is allowed to evolve linearly with redshift, r(z) = r0 + r1 (z - z_bar), and check whether the alpha bias shown in Figure 5 is reduced or disappears. If the bias persists with flexible r(z), the 1% systematic is not an artifact of standard-ruler evolution, and the paper's claim is supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim, from the abstract and Section 6, is that BAO scale recovery carries a ~1% systematic when the true model is far from the fiducial, and that this cannot be attributed to any constant AP approximation. The support is two-fold: (i) the 'exact' redshift-dependent AP calculation, Eq. (5.8), shows the same systematic as the constant AP fits (Figures 3b and 7b), and (ii) a Lambda-CDM template fitting cross-check shows a similar trend (Figure 5). Both of these routes share assumptions that are not independently tested: the BAO feature is a redshift-independent standard ruler r, the empirical Gaussian model (2.12) or the fixed Lambda-CDM template accurately describes the BAO shape under large distance distortions, and the approximate redshift integral (3.3) is accurate. The paper does not identify the source of the 1% residual, and it does not rule out that the residual arises from these shared modeling assumptions rather than from a genuine failure of the AP framework. In particular, Section 2.3 acknowledges that the standard-ruler approximation has shifts <~0.5% in Lambda-CDM at z>0.3, which is non-negligible relative to the claimed 1% for |alpha_bar-1|~0.1. If the effective BAO scale varies across the survey and couples to the fiducial model's distance mapping, a systematic of order 1% could emerge even with perfect AP scaling. The Lambda-CDM template cross-check shares the standard-ruler assumption, so it cannot distinguish this explanation from a real systematic. The reader's flagged concern about the local-metric approximation (2.4) for high-frequency toy models is valid, but it affects only the toy-model demonstration of modified AP superiority, not the headline 1% claim, because the large-scale smooth models satisfy (2.4) at the 10^-3 level. Thus the most load-bearing concern is the unresolved origin and potential artifact status of the 1% systematic.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper investigates the validity of the conventional constant Alcock-Paczyński (AP) scaling used in baryon acoustic oscillation (BAO) analyses. It derives analytic bounds on the difference between the standard approximation, in which the AP parameters α and ϵ are evaluated at the survey mean redshift, and a modified approximation in which they are replaced by survey-averaged values ᾱ and ϵ̄. The authors propose that published AP parameters should be reinterpreted as survey-averaged distance ratios, and they test this proposal on QPM mock catalogues for a range of ΛCDM, timescape, and toy-model cosmologies. They report systematic errors of about 1% in the recovery of the BAO scale when the fiducial and true cosmologies are distant, and show that these errors persist when the full redshift-dependent AP functions are used instead of constant parameters and that a similar trend appears with a ΛCDM template fitting procedure.","tokens_in":39359,"tokens_out":12716,"duration_ms":118772,"significance":"If correct, the paper has two significant consequences. First, the constant AP parameters reported in the literature are more accurately interpreted as survey-averaged ratios rather than values at a single effective redshift; this matters for models with large metric gradients. Second, BAO-derived distance measurements carry a model-dependent systematic at the percent level when extrapolated beyond the fiducial cosmology, a systematic that is not captured by the constant-AP approximation itself. The analytic bounds in Eqs. (3.9) and (3.14) are parameter-free and useful, the mock tests use 200–1000 QPM catalogues with estimated covariance, and the ΛCDM template cross-check provides a nontrivial robustness test. The main limitations are that the residual systematics are not traced to a specific source and that some claims rest on assumptions shared by the two fitting procedures.","major_comments":[{"comment":"The paper's central claim that the ~1% systematic 'cannot be attributed to any constant AP approximation' is supported by comparing fits with constant AP parameters to fits using the redshift-dependent functions (5.8). However, both fitting models assume a redshift-independent BAO scale r and both use the same approximate redshift integral (3.3). Section 2.3 acknowledges that the standard-ruler assumption has shifts below ~0.5% in ΛCDM at z>0.3, which is comparable to the claimed 1% effect at |ᾱ−1|~0.1. The ΛCDM template cross-check fixes the template shape and therefore shares the same standard-ruler assumption. Please quantify the impact of a redshift-dependent BAO scale, for example by fitting with a redshift-dependent r or by using mocks where the BAO scale variation is known, or explicitly scope the abstract and Section 6 claims to the standard-ruler ansatz.","section":"Sec. 5.3; Sec. 2.3"},{"comment":"The local-metric approximation (2.4) is verified in Section 2.1 only for FLRW and timescape models. It is then applied without verification to the toy models of Section 5.4, whose distance-redshift relations oscillate as A cos(fz+Φ) with f up to 50. For a radial pair at z≈0.5 with separation 100 Mpc/h, δz≈0.03, so f δz≈1.5 and the metric varies substantially across the pair. In this regime the 'exact' redshift-dependent AP predictions (5.8) used as the benchmark in Figure 7b are themselves approximate, and the statement that the modified constant AP scaling is efficient for all tested models rests on an unvalidated premise. Please verify (2.4) for the toy-model parameters, for instance by direct numerical geodesic integration, or restrict the toy-model conclusions to the regime where (2.4) is accurate.","section":"Sec. 5.4; Eq. (2.4)"},{"comment":"The statement that the level of systematic uncertainty is 'robust to the exact fitting method employed' is stronger than the evidence presented: only two fitting methods are compared, and both use a fixed template or a fixed standard ruler. The evidence supports robustness to the two fitting methods considered. Please rephrase the claim accordingly, or add a third, less parametric BAO extraction method.","section":"Abstract; Sec. 6"}],"minor_comments":[{"comment":"The paragraph following Figure 7a refers to 'figure 3b' when comparing the modified constant AP scaling to the exact AP scaling; this should be 'figure 7b'.","section":"Sec. 5.4"},{"comment":"The claimed verification that corrections to Eq. (2.4) are below 10^-3 for FLRW and timescape models is stated without quantitative details; please provide the numerical values or a reference to a figure or table.","section":"Sec. 2.1"},{"comment":"The text says that 1000 QPM mock catalogues are used for the reference model and 200 for the remaining trial cosmologies; please ensure the sample sizes are stated consistently in the figure captions and in the text describing the fits.","section":"Sec. 5.1"},{"comment":"There is a typo in the first paragraph of Section 1: 'use of of N-body mock catalogues' should read 'use of N-body mock catalogues'.","section":"Introduction"}],"recommendation":"major_revision","confidential_remarks":"The paper is generally well-executed and the central idea of survey-averaged AP parameters is a useful contribution. My main concern is that the headline ~1% systematic is not fully diagnosed and that the toy-model validation has a gap related to the local-metric approximation; both are addressable with additional tests or careful wording. The paper is within scope for JCAP, and I recommend major revision rather than rejection because the core derivations are sound and the claims are conditional on stated assumptions that can be made explicit."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This paper does something concrete: it proposes replacing the usual constant AP scaling parameters evaluated at the mean survey redshift with survey-averaged parameters (ᾱ, ε̄), derives analytic bounds on the difference between the two choices, and tests both on ΛCDM mocks and on toy models with large metric gradients. The mock work uses up to 1000 QPM catalogues with known true cosmology, and the authors cross-check their empirical fitting model against a ΛCDM template fitter. The analytic bounds in Eqs. (3.9) and (3.14) are real results, not placeholders, and the finding that the standard constant AP approximation breaks down for large metric gradients while the modified version stays accurate is well supported for the tested toy models.\n\nThe main substantive finding is the ~1% systematic in BAO scale recovery when the true model is far from the fiducial, which persists even when using the exact redshift-dependent AP scaling. The authors are careful to say this is not attributable to any constant AP approximation. I think that claim is as well supported as it can be given their setup, because the exact calculation shows the same shift. But the stress-test worry is legitimate: the exact calculation still shares the standard-ruler assumption and the empirical Gaussian shape (or the fixed template), so the residual could in principle come from those modeling choices rather than from a real failure of the AP framework. The authors do not identify the source of the 1%, and they acknowledge the standard-ruler approximation itself can shift the BAO scale by up to ~0.5% in ΛCDM. That is not a fatal flaw, but it means the headline quantitative claim is not yet as clean as the abstract suggests.\n\nA smaller soft spot: the local-metric approximation in Eq. (2.4) is verified for FLRW and timescape but not for the high-frequency toy models where the metric oscillates on scales comparable to pair separations. This matters for the toy-model demonstration of the modified AP scaling's superiority, but not for the main ~1% residual, which appears in smooth large-scale models.\n\nThe paper is honest about its limitations: the modified AP scaling's superiority is argued heuristically in Appendix B, the toy models are unphysical, and the analysis is restricted to spherically symmetric template metrics. None of these are hidden.\n\nOverall, this is a solid methodological study. The new AP scaling variant and the bounds are worth having, and the ~1% warning is important even if its ultimate origin remains unclear. A serious referee should engage with it. I would send it to review, and I would probably cite it in future BAO systematics work.","headline":"A careful, genuinely useful BAO systematics paper: the new survey-averaged AP scaling and the ~1% residual deserve referee time, though the residual's origin is not fully pinned down.","tokens_in":40157,"tokens_out":1645,"would_cite":true,"duration_ms":19294,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that the standard constant AP scaling of BAO measurements is accurate only when fiducial and true metric gradients are comparable, and that replacing the scaling parameters by their survey averages is more accurate; it…","keywords":["baryon acoustic oscillations","Alcock–Paczyński scaling","galaxy clustering","fiducial cosmology","BAO distance scale","timescape model","correlation function","cosmological model dependence"],"falsifier":"Recompute the 'exact' redshift-dependent AP predictions for the toy models $D(z)[1 + A\\cos(f z + \\Phi)]$ using full geodesic distances rather than the local-metric approximation. If the resulting BAO peak shifts differ from the paper's benchmark by more than about $0.1\\%$ at separations near $100\\,\\mathrm{Mpc}/h$ for the $f = 30$ and $f = 50$ cases, then the benchmark itself is approximate and the stated accuracy of the modified AP scaling for all test models is not established.","tokens_in":38800,"feed_emoji":"📏","tokens_out":11737,"duration_ms":95150,"temperature":0.7,"pith_summary":"This paper asks whether the standard Alcock–Paczyński scaling used to convert baryon acoustic oscillation measurements between cosmologies is trustworthy. It argues that the usual practice of evaluating the two AP scaling parameters at the survey's mean redshift is only valid when the fiducial and true metrics have similar gradients. It proposes replacing those values with redshift-averaged parameters, which it shows recovers the BAO scale accurately even for toy models with large metric oscillations. It also finds a residual systematic error of about 1% in the recovered BAO scale when the true model is far from the fiducial, independent of the fitting method, meaning BAO distance errors in the standard literature may be underestimated for models outside the concordance cosmology.","feed_headline":"Hidden ~1% BAO systematic appears when models stray from fiducial","feed_subtitle":"Standard AP scaling misses survey-averaged distance ratios; a reinterpretation fixes most of the error","key_machinery":"The engine of the argument is the small-separation geodesic distance approximation $D_T^2 \\approx g_{zz}(\\delta z)^2 + g_{\\theta\\theta}(\\delta\\Theta)^2$, which lets any spherically-symmetric template metric be parametrised relative to a fiducial one by the redshift-dependent AP functions $\\alpha(z)$ and $\\epsilon(z)$. From these, the paper constructs the exact redshift-dependent map of the 2-point correlation function and its wedge integrals, then expands the redshift average to first order; this expansion identifies the survey-averaged parameters $\\bar\\alpha$ and $\\bar\\epsilon$ as the correct constant parameters. The empirical Gaussian-plus-polynomial model of the BAO feature, fitted to the mean correlation function of mock catalogues, converts the analytic predictions into measured peak shifts and warping parameters.","core_discovery":"The paper's central claim is that the conventional constant Alcock–Paczyński approximation, which takes the two scaling functions at the survey's mean redshift, is only a special case of a more accurate constant choice: the survey-averaged values. In the spherically-symmetric template framework, a first-order expansion of the redshift-integrated correlation function identifies the survey-averaged parameters as the ones that enter the effective BAO template. On mock catalogues the two choices are nearly indistinguishable for smooth large-scale cosmologies, but the modified choice remains accurate for toy models with large metric oscillations where the conventional choice fails. The paper also finds a residual systematic error of about 1% in the recovered BAO scale when the true model differs from the fiducial by roughly ten percent in its distance measures, present even when the exact redshift-dependent AP functions are used, and similar for both the empirical Gaussian fit and the standard template fit; it therefore presents this as a systematic that must be added to the error budget when BAO distances are extrapolated beyond the fiducial cosmology.","pith_inferences":["The same Taylor-expansion argument suggests the survey-averaging correction applies to any redshift-dependent distance scale fitted through a correlation-function wedge, so the reinterpretation of AP parameters may carry over to other clustering statistics.","Because the ~1% residual appears even with exact redshift-dependent scaling, BAO-only constraints on models far from the fiducial cosmology are limited by a systematic of the same order as current statistical errors; combining probes reduces statistics, not this model-dependence.","A natural testable extension is to reverse the roles of true and fiducial models in the toy-model mock analysis; the paper expects its conclusions to hold under reversal but does not demonstrate it.","If the local-metric approximation fails for the high-frequency toy models, both the conventional and modified constant AP methods inherit the same benchmark error, so the modified method's reported advantage in those cases would need to be re-evaluated."],"forward_implications":["Published constant AP parameters from BAO analyses should be read as survey-averaged distance ratios, not as values at a single effective redshift.","For pairs of smooth large-scale cosmologies the conventional constant AP approximation is adequate, so existing concordance-model results do not need to be redone.","When a BAO measurement is used to constrain a model whose distance measures differ from the fiducial by about ten percent, a systematic error of roughly one percent in the acoustic scale should be added to the error budget.","The modified AP scaling can be applied to existing analyses by reinterpreting their fitted parameters, which matters for models with significant curvature gradients.","The standard approach breaks down when the true metric has large gradients relative to a smooth fiducial metric; in such cases evaluating AP parameters at the mean redshift misstates the effective distance ratios."],"supporting_citations":[{"why":"Original Alcock–Paczyński geometric test that the scaling methods extend.","marker":"[17]"},{"why":"Defines the constant AP parameters and the conventional constant-AP reparametrisation tested here.","marker":"[20]"},{"why":"Standard BAO application measuring angular diameter distance and Hubble parameter at z = 0.35 with constant AP scaling.","marker":"[21]"},{"why":"BOSS BAO analysis whose template and error budget the paper compares against when showing the ~1% systematics are robust to fitting procedure.","marker":"[22]"},{"why":"Supplies the mock catalogues and the conventional systematics baseline that the paper's larger-error claim is contrasted with.","marker":"[24]"},{"why":"Examines the impact of the fiducial cosmology assumption on BAO parameter inference, the comparison point for the paper's larger systematics.","marker":"[25]"},{"why":"Provides the spherically-symmetric template-metric framework, the empirical correlation-function model, and the redshift-dependent AP formulation on which the analysis is built.","marker":"[26]"}],"fun_headline_variants":["Survey-averaged AP scaling cuts BAO error in exotic models","BAO scale carries ~1% bias when fiducial is distant","New AP scaling handles extreme cosmologies; standard fails","Rethink AP scaling: average over survey, not mean redshift","AP scaling improved: survey-averaged parameters work"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is the small-separation approximation that treats the distance between galaxy pairs as if the metric components were constant across the pair; the paper verifies it for FLRW and timescape models but applies it without independent verification to the high-frequency toy models where the metric oscillates on scales comparable to the pair separation.","fun_headline_variants_meta":{"raw":{"variants":["Survey-averaged AP scaling cuts BAO error in exotic models","BAO scale carries ~1% bias when fiducial is distant","New AP scaling handles extreme cosmologies; standard fails","Rethink AP scaling: average over survey, not mean redshift","AP scaling improved: survey-averaged parameters work"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000483,"raw_usage":{"total_tokens":2453,"prompt_tokens":1083,"completion_tokens":1370,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":699,"completion_tokens_details":{"reasoning_tokens":1284}},"tokens_in":699,"tokens_out":1370,"duration_ms":11960,"temperature":1.0,"reasoning_tokens":1284,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T10:15:31.123790+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the 'exact' redshift-dependent AP predictions for the toy models $D(z)[1 + A\\cos(f z + \\Phi)]$ using full geodesic distances rather than the local-metric approximation. If the resulting BAO peak shifts differ from the paper's benchmark by more than about $0.1\\%$ at separations near $100\\,\\mathrm{Mpc}/h$ for the $f = 30$ and $f = 50$ cases, then the benchmark itself is approximate and the stated accuracy of the modified AP scaling for all test models is not established.","supporting_citations":[{"cited_title":"Alcock and B","cited_arxiv_id":null,"evidence_quote":"Original Alcock–Paczyński geometric test that the scaling methods extend."}],"review_version":1}