{"id":"9ddc8878-8823-4925-a149-d377f31ac2dd","arxiv_id":"2501.13458","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A new fitting technique uses energy-resolved X-ray variability to select the two varying spectral parameters in three competing models of Mrk 335, and it disfavors the relativistic reflection model.","lead":"This paper presents a technique to infer which two changing spectral parameters best explain the energy-dependent X-ray variability of the Seyfert galaxy Mrk 335. Applied to XMM-Newton data, it favors partial-covering and two-corona models over a relativistic reflection interpretation for this source.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Correlated Fvar errors undermine the chi-square ranking that supports the claim against the reflection model.","rationale":"The reader's weakest_assumption correctly identified that the model comparison relies on chi-square statistics computed from Fvar errors assumed independent across energy bands, and that the paper itself notes this could overestimate the errors. This is the single most load-bearing concern because the central claim against the reflection model depends on the ranking of parameter pairs, which is derived from those chi-square values. The paper's own admission that the fit statistics are 'somewhat low' and that this 'suggests an over estimation of errors' directly undermines the statistical significance of the pair rankings. Additionally, the two segment lengths (3 ks and 6 ks) are not independent samples, as they are derived from the same light curve, so the apparent change in best-fitting pair may reflect noise rather than a genuine need for more parameters. The concrete test proposed—using covariance-aware error estimates or rescaling errors to force reduced chi-square to 1—would directly assess whether the ranking and the 'more than two parameters' conclusion survive. This does not change the reader's CONDITIONAL verdict: the technique itself is a useful contribution, but the specific conclusion against the reflection model requires stronger statistical support. The paper deserves credit for a parameter-free derivation of the prediction formula, careful spectral fitting, and for explicitly acknowledging the limitation; however, that limitation is precisely what makes the central claim fragile.","tokens_in":22640,"tokens_out":7311,"duration_ms":67490,"concrete_test":"Recompute the chi-square statistics in Section 5.4 using a covariance matrix for Fvar estimates that accounts for correlated variability across energy bands, obtained from simulated light curves generated with the observed power spectrum and cross-spectrum (or the Ingram 2019 formalism). If the best-fitting pair for the 3 ks and 6 ks segments changes, or if the difference in chi-square between the top pairs falls below the level of statistical significance, the conclusion that more than two parameters are required for the reflection model is unsupported. A simpler check: rescale the error bars so that the best-fit reduced chi-square equals 1 for each model, then see if the same pairs remain ranked first.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that the blurred reflection model is disfavored rests on the inference that more than two parameters are required to explain the energy-dependent Fvar, which is inferred from the fact that the best-fitting parameter pair differs between the 3 ks and 6 ks segments (Table 6). This inference is statistically fragile. The chi-square values used to rank the pairs are computed under the assumption that Fvar errors in different energy bands are independent, an assumption the authors explicitly flag as potentially invalid ('an intrinsic correlation between the variabilities in different energy bands' could cause the errors to be overestimated). The observed reduced chi-square values are already below 1 for several fits (e.g., 16.93/21 and 17.89/21), indicating overestimated errors. With proper covariance-aware errors, the differences between the best and second-best pairs (e.g., 22.81 vs 23.91 for the 3 ks reflection case) are likely to become statistically insignificant, and the ranking could change. Since the 'more than two parameters' conclusion is based solely on this ranking, the central claim against the reflection model is not securely supported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper introduces a method to model the energy-dependent fractional rms variability amplitude Fvar(E) under the assumption that exactly two spectral parameters vary linearly and with a time-independent correlation CXY. The method is applied to a 133 ks XMM-Newton EPIC-pn observation of Mrk 335. Three spectral interpretations (partial covering, two-corona thermal Comptonization, and relativistic blurred reflection) are fit to the 0.3-10 keV time-averaged spectrum. For each model, all one- and two-parameter combinations are considered, and the pair minimizing Eq. (6) against the observed Fvar in 3 ks and 6 ks segments is selected. The authors find that a single parameter is insufficient for all models, that the partial covering and Comptonization models require variations in (f_c, Gamma_pl) and (Gamma_nthC, N_nthC) respectively, and that the reflection model selects different pairs for the two segment lengths. From this, combined with the extreme emissivity index and reflection fraction, they conclude that the blurred reflection model may not be a suitable description of Mrk 335.","tokens_in":22846,"tokens_out":7405,"duration_ms":67053,"significance":"The proposed method is simple and, in principle, useful: it converts spectral-model derivatives into a predicted rms spectrum and can rank variability scenarios as a diagnostic for models that are otherwise spectrally degenerate. The analytic expressions in Appendix A and the application to a real long XMM-Newton observation are positive features. If the statistical caveats can be resolved, the technique would be a valuable addition to the AGN variability toolbox. However, the paper's central astrophysical claim, that the reflection model is disfavored because more than two parameters are required, is not securely supported by the present analysis. The chi-square comparison is made under an independence assumption that the authors themselves flag as potentially wrong, and the differences between the best parameter pairs are often small. The significance of the method as a model discriminator is therefore currently limited by the statistical treatment rather than by the underlying idea.","major_comments":[{"comment":"The conclusion that the reflection model requires more than two parameters rests on the fact that the best-fitting pair differs between the 3 ks and 6 ks segments (Table 6). This difference is not shown to be statistically significant. For the 3 ks segment the best pair (Gamma_relx, log xi) has chi2/dof = 22.81/21, while the second-best pair (Gamma_relx, N_relx) has 23.91/21; a Delta chi2 of about 1.1 for 21 dof is negligible. More generally, several fits in Tables 4-6 have reduced chi2 below 1 (e.g., 16.93/21 and 17.89/21 for Model-II), indicating that the quoted errors are overestimated. The authors themselves note in the penultimate paragraph of Section 6 that intrinsic correlations between energy bands could cause exactly this overestimation. Because the chi2 ranking is the only quantitative evidence for the 'more than two parameters' claim, the claim is not securely established. A covariance-aware error estimate, or at least a Monte Carlo test of the pair ranking, is required before the reflection model can be declared disfavored.","section":"Section 6, Table 6"},{"comment":"The spectral fit quality is uneven across models, and the statement in Section 6 that 'the spectrum can be well represented by all the aforementioned models' is not consistent with the reported Model-I statistic of chi2/dof = 302.3/163 in Section 3.1. With 163 degrees of freedom this reduced chi2 corresponds to a very small p-value. Since the partial derivatives A_i and B_i in Eq. (5) are evaluated at the best-fit spectral parameters, a poor spectral fit makes the predicted Fvar shape less trustworthy for that model. The paper should either report the fit quality honestly, add a caveat to the comparison, or restrict the variability analysis to models that are statistically acceptable.","section":"Section 3.1, Section 6"},{"comment":"The use of the word 'predict' overstates what the method does. In Eq. (6) the parameters delta X, delta Y, and CXY are fitted to the observed F0 values, so the resulting FP is a best-fit model, not an independent prediction. This does not invalidate the approach as a model-comparison tool, but it changes the meaning of the chi2 values reported in Tables 4-6. The abstract's claim that the technique 'predicts the energy dependent fractional r.m.s' should be softened to 'models' or 'reproduces'.","section":"Eqs. (5)-(6), abstract"},{"comment":"The inference that 'more than two parameters are required to explain the data' is not directly tested. The analysis fits exactly two parameters for each model; no three-parameter variability fits are performed, and the two segment lengths are not treated jointly. A simpler interpretation of the different best pairs for 3 ks and 6 ks in Table 6 is that the variability pattern is timescale dependent, not that a two-parameter model is insufficient. To support the claim, the authors should fit a three-parameter model, or fit the same two-parameter model to both segments simultaneously, and compare the resulting statistics with the same error treatment.","section":"Section 5.4, Section 6"}],"minor_comments":[{"comment":"The partial derivatives are computed using a 10% parameter variation; for parameters with strongly nonlinear responses or parameters near their bounds (e.g., reflection fraction, covering fraction), 10% may violate the linear approximation of Eq. (4). Please verify linearity by also testing smaller variations, such as 1% and 5%.","section":"Section 5.1"},{"comment":"The text is inconsistent about whether sigma_i denotes the error on F0_i or on F0_i^2; Eq. (6) and the surrounding text should define this unambiguously.","section":"Section 5.1 and Eq. (6)"},{"comment":"Index1 is reported only as a lower limit (>8.30) while Section 6 refers to 'extreme values' of the emissivity index; please state explicitly whether the parameter is at or near the model boundary and discuss the possible effect on the variability derivatives.","section":"Table 3, Section 6"},{"comment":"The predicted curves are plotted continuously over 0.3-10 keV while the observed points are binned; please specify the energy binning used for the predicted curves or state explicitly that the curves are interpolated.","section":"Figure 4"},{"comment":"Please specify how the 3 ks and 6 ks segments are placed within the 133 ks light curve, since overlapping or differently aligned segments could introduce correlations between the two Fvar estimates.","section":"Section 4"}],"recommendation":"major_revision","confidential_remarks":"The technical core is sound, but the main astrophysical conclusion is presented more strongly than the statistics justify. I would encourage the editor to regard the paper primarily as a methods contribution with an illustrative application, and to require that the statistical caveats about correlated Fvar errors and the significance of pair rankings be addressed in revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The main value of this paper is the technique, not the Mrk 335 verdict. The idea of attributing the energy-dependent fractional rms to a pair of linearly varying spectral parameters, with a cross-correlation coefficient, is a clean extension of the standard Fvar framework. The numerical derivatives through xspec and the analytical appendix make the method reproducible in principle, and applying it to three competing spectral models is a sensible way to break degeneracy. Credit where due: the method is clearly explained, and the paper does not oversell it. The authors explicitly note that more than two parameters may be needed, that the specific spectral model matters, and that correlated errors between energy bands could inflate the quoted uncertainties.\n\nThe soft spots are real but not fatal. The chi-square comparison that drives the central claim is statistically fragile. The Fvar errors are assumed independent across bands; the paper itself acknowledges that inter-band correlation could overestimate them, and the low reduced chi-square values (e.g., 16.93/21, 17.89/21) suggest that already. With covariance-aware errors, the small differences between best and second-best pairs in the reflection model (22.81 vs. 23.91 for 3 ks) are likely to become insignificant. So the inference that \"more than two parameters are required\" is not securely supported by the ranking; it could be noise. The spectral fit for the partial covering model is also poor (302/163) and yet the paper says the spectrum is \"well represented\" by all models, which is an overstatement.\n\nThat said, the conclusion against the blurred reflection model does not rest solely on the chi-square ranking. The extreme values for emissivity index and reflection fraction are independently troubling, and the paper is honest about them. So the verdict is plausible, just not proven. A motivated reader could redo the comparison with covariance-aware errors or simulations, and that would settle the matter. No code or data products are provided, which would have strengthened the paper.\n\nWho should read this: anyone working on AGN variability and spectral model discrimination. The technique deserves to be in the literature, and the Mrk 335 result is a useful case study. I would send it to peer review with a request for robustness checks on the model comparison, but the core methodology is sound enough to engage with seriously.","headline":"A genuinely new Fvar-modelling technique, applied to Mrk 335; the reflection-model verdict is plausible but the chi-square ranking behind it is fragile.","tokens_in":23469,"tokens_out":1991,"would_cite":true,"duration_ms":20559,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that the energy-dependent X-ray variability of Mrk 335 can be used to identify which spectral parameters actually vary, and that the blurred reflection model fails this test.","keywords":["X-rays: galaxies","galaxies: Seyfert","methods: numerical","galaxies: individual: Mrk 335","fractional rms variability","soft X-ray excess","relativistic reflection","spectral degeneracy"],"falsifier":"A decisive check would be to recompute the Fvar errors including the covariance between energy bands, using cross-spectral or Monte Carlo light-curve simulations built from the fitted parameter variations, and then rerun the pair-selection for all three models on both segment lengths; if the blurred reflection model then yields the same best pair at 3 ks and 6 ks with a reduced chi-square near 1, the paper's conclusion that reflection is unsuitable would be refuted.","tokens_in":22335,"feed_emoji":"🔭","tokens_out":9923,"duration_ms":71974,"temperature":0.7,"pith_summary":"The paper presents a method for turning the energy dependence of X-ray variability into a test of which spectral parameters are varying, and applies it to a bright 133 ks XMM-Newton observation of the Seyfert galaxy Mrk 335. The observed fractional rms, Fvar, is not constant with energy, and each of three spectral models that fit the time-averaged spectrum is examined by predicting Fvar from a single or a pair of linearly correlated parameter variations. For the partial covering and warm/hot corona models the same parameter pair wins on both 3 ks and 6 ks timescales, whereas for the blurred reflection model the winning pair changes between timescales and no pair fits well. The authors conclude that at least two spectral parameters must vary, and that the blurred reflection model is not a suitable explanation for the Mrk 335 spectrum. The technique matters because it offers a route to break degeneracies among models that all describe the same time-averaged spectrum.","feed_headline":"Blurred reflection fails Mrk 335's variability test","feed_subtitle":"Two-parameter fitting finds stable variability drivers for absorbing and coronal models but none for reflection.","key_machinery":"The central machinery is a linear-response formula relating model count-rate derivatives to fractional rms. For each energy band the partial derivatives $\\partial R_i/\\partial X$ and $\\partial R_i/\\partial Y$ are computed numerically in the spectral fitting software used by the paper, by perturbing each parameter by 10% around the best-fit spectral model, then converted into $A_i$ and $B_i$. These enter the quadratic prediction for $F_{P,i}^2$ with a cross-correlation coefficient $C_{XY}$ that can represent correlated, uncorrelated, or anti-correlated variations ($C_{XY}=1,0,-1$) or a cosine of a phase lag. The predicted rms spectrum is fitted to the observed one by minimizing $\\chi^2$ with constrained optimisation, under the conditions $\\delta X^2,\\delta Y^2 \\ge 0$ and $-1 \\le C_{XY} \\le 1$. This machinery converts a time-averaged spectral fit into a temporal variability test without needing full light-curve simulations.","core_discovery":"The central claim is that the fractional rms spectrum can be predicted from a spectral model by assuming the variability is driven by two linearly correlated parameters, and that comparing this prediction with observed Fvar identifies the physical driver while discriminating between models. For an energy band $i$ the predicted squared fractional rms is written as $F_{P,i}^2 = A_i^2 \\delta X^2 + B_i^2 \\delta Y^2 + 2A_iB_iC_{XY}|\\delta X||\\delta Y|$, where $A_i$ and $B_i$ are dimensionless sensitivity coefficients obtained from numerical derivatives of the model count rate and $C_{XY}$ is the correlation between the parameter variations. Applied to Mrk 335, single-parameter variations fail for all three models, so at least two varying parameters are required. The best pairs are the covering fraction and power-law index for partial covering, and the warm corona photon index and normalisation for the two-corona interpretation, both stable across the 3 ks and 6 ks segments. For the relxillCp reflection model, however, the best pair is $\\Gamma_{\\mathrm{relx}}$ with $\\log\\xi$ at 3 ks but $\\Gamma_{\\mathrm{relx}}$ with normalisation at 6 ks, and the fit for the latter does not reproduce the observed dip near 0.7 keV. Together with the extreme fitted emissivity index (>8.3) and reflection fraction (~9.5), the paper argues that blurred reflection is not a suitable description for this source.","pith_inferences":["Inference: If the same analysis were run on multiple epochs of the same AGN, a model that is truly at work should yield the same parameter pair and similar correlation sign each time; the stability of the pair across segments could be used as a quantitative 'variability fingerprint' for each spectral interpretation.","Inference: The failure of the reflection model may not doom all reflection geometries, because only the coronal relxillCp version with tied emissivity indices was tested; a lamp-post or self-consistently computed emissivity model might behave differently.","Inference: The paper's caveat that correlated variability across energy bands could overestimate the Fvar errors means the absolute $\\chi^2$ values may not be reliable; however, the reflection model's instability across segments and its extreme fitted parameters are qualitative features that would likely persist under recalculated errors.","Inference: A testable extension would be to simulate light curves from the best-fitting pair, such as $f_c$ and $\\Gamma_{\\mathrm{pl}}$, with the fitted amplitudes and correlation, and then recover the input parameters with the same fitting procedure to validate that the method is unbiased."],"forward_implications":["At least two spectral parameters must vary: for all three models a single parameter always gives reduced $\\chi^2 > 45/21$, so one-parameter variability explanations are excluded for this observation.","For the partial covering model, variations in covering fraction $f_c$ and photon index $\\Gamma_{\\mathrm{pl}}$ with a small positive correlation ($C_{XY}\\sim 0.3$) reproduce the rms spectrum on both 3 ks and 6 ks timescales.","For the two-corona model, the warm corona's photon index and normalisation vary anti-correlated ($C_{XY}\\sim -0.5$), meaning the energy-dependent variability can be produced by the warm corona alone without requiring hot-corona variability.","For the blurred reflection model, no parameter pair is consistent across segment lengths; the 3 ks data prefer $\\Gamma_{\\mathrm{relx}}$ and $\\log\\xi$, while the 6 ks data prefer $\\Gamma_{\\mathrm{relx}}$ and normalisation, indicating that more than two parameters are varying.","The method is general and can be applied to other AGN datasets, including higher-energy NuSTAR observations, to differentiate between spectrally degenerate models."],"supporting_citations":[{"why":"Supplies the definition of Fvar, excess variance, and the error formula used to measure the observed rms spectra.","marker":"Vaughan et al. 2003"},{"why":"Provides the partial covering spectral prescription and the fitting context for the same Mrk 335 observation.","marker":"Grupe et al. 2008"},{"why":"Describes the relxill code used for the relativistic reflection spectral model in Model-III.","marker":"Dauser et al. 2014"},{"why":"Details the improved reflection model grid implemented in relxill that the reflection fit relies on.","marker":"García et al. 2014"},{"why":"Established that partial covering and reflection both describe the Mrk 335 spectrum, the spectral degeneracy that the variability method targets.","marker":"Gallo et al. 2015"},{"why":"Reported covering-fraction variability as the rms explanation on roughly a one-day timescale, which the paper contrasts with its shorter 3 and 6 ks result.","marker":"Yamasaki et al. 2016"},{"why":"Gives error formulae for correlated variability across energy bands, cited by the paper as a caveat on its low chi-square values.","marker":"Ingram 2019"},{"why":"Provides the constrained minimisation tool used to fit the parameter amplitudes and correlation within physical ranges.","marker":"Newville et al. 2016"}],"fun_headline_variants":["Reflection model flunks Mrk 335's rms spectrum","Mrk 335 variability rules out blurred reflection","Two-parameter rms test favors absorber and corona, not reflection","Fvar double-parameter test rejects blurred reflection for Mrk 335"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The comparison assumes that the measured Fvar errors in different energy bands are independent, although the paper notes that intrinsic correlations between variabilities in different bands would change the errors and could reorder which parameter pairs fit best.","fun_headline_variants_meta":{"raw":{"variants":["Reflection model flunks Mrk 335's rms spectrum","Mrk 335 variability rules out blurred reflection","Two-parameter rms test favors absorber and corona, not reflection","Fvar double-parameter test rejects blurred reflection for Mrk 335"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001327,"raw_usage":{"total_tokens":5809,"prompt_tokens":1134,"completion_tokens":4675,"prompt_tokens_details":{"cached_tokens":1024},"prompt_cache_hit_tokens":1024,"prompt_cache_miss_tokens":110,"completion_tokens_details":{"reasoning_tokens":4603}},"tokens_in":110,"tokens_out":4675,"duration_ms":1001081,"temperature":1.0,"reasoning_tokens":4603,"cache_read_input_tokens":1024,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T15:55:59.717900+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A decisive check would be to recompute the Fvar errors including the covariance between energy bands, using cross-spectral or Monte Carlo light-curve simulations built from the fitted parameter variations, and then rerun the pair-selection for all three models on both segment lengths; if the blurred reflection model then yields the same best pair at 3 ks and 6 ks with a reduced chi-square near 1, the paper's conclusion that reflection is unsuitable would be refuted.","supporting_citations":[{"cited_title":", author Edelson, R","cited_arxiv_id":null,"evidence_quote":"Supplies the definition of Fvar, excess variance, and the error formula used to measure the observed rms spectra."},{"cited_title":", author Mizumoto, M","cited_arxiv_id":null,"evidence_quote":"Reported covering-fraction variability as the rms explanation on roughly a one-day timescale, which the paper contrasts with its shorter 3 and 6 ks result."},{"cited_title":", author Stensitzki, T","cited_arxiv_id":null,"evidence_quote":"Provides the constrained minimisation tool used to fit the parameter amplitudes and correlation within physical ranges."}],"review_version":1}