{"id":"3b3db454-8566-4e76-8abe-014ee66d04a0","arxiv_id":"2606.02676","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Introduces standardised tail plots and normalised residual plots for extreme value regression models, claiming sample-size-independent uncertainty bounds for global and regional goodness-of-fit assessment.","lead":"The paper proposes two new visual diagnostics, the standardised tail plot and the normalised residual plot, for assessing goodness-of-fit in extreme value regression models used for extrapolating rare events. These tools aim to enable reliable model checks at both overall and local levels in the covariate space even when data amounts differ across regions.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Finite-sample performance of asymptotic uncertainty bounds for regional diagnostics unverified, especially with small local exceedance counts","rationale":"The reader's weakest assumption is precisely the finite-sample accuracy of the asymptotic approximation invoked for the uncertainty bounds. The concrete simulation test directly probes whether that assumption is strong enough for the headline claim about sample-size-independent diagnostics to be actionable.","tokens_in":1769,"tokens_out":305,"duration_ms":13998,"concrete_test":"Generate 500 Monte Carlo replicates from a known extreme-value regression model (e.g., GEV with linear location and scale on a 1-D covariate); for each replicate, fit the model, construct the standardised tail and normalised residual plots at total n = 200, 500, 2000 and at regional partitions with local exceedance counts 5–50; record empirical coverage of the proposed uncertainty bands and test whether coverage deviates systematically with n or local count.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim requires that the asymptotic distribution of normalised exceedance probabilities yields uncertainty bounds that are both useful and approximately independent of sample size (global or regional). Extreme-value regression models are typically fit with small numbers of tail observations; regional subsets of the covariate domain can have even fewer. The paper invokes this asymptotic result directly to justify the diagnostics, but without reported simulation checks or error bounds on the approximation, it is unclear whether the claimed independence and reliability hold at the sample sizes encountered in applications.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper proposes two novel visual diagnostics for extreme value regression models—the standardised tail plot and the normalised residual plot—derived from the asymptotic distribution of normalised exceedance probabilities. It claims that the resulting uncertainty bounds are approximately independent of sample size, enabling consistent global and regional goodness-of-fit assessment despite varying local sample sizes. The manuscript also discusses summary statistics for global and regional fit and illustrates the tools in two applications involving model comparison across thousands of candidate models.","tokens_in":1874,"tokens_out":367,"duration_ms":17323,"significance":"If the finite-sample behavior of the proposed bounds aligns with the asymptotic claims, the work would address an important gap in diagnostics for extreme-value regression models, especially for regional assessment on non-Euclidean or low-dimensional covariate domains and for scalable model comparison. The sample-size-independent uncertainty bounds, if reliable, represent a practical strength for extrapolation-focused applications.","major_comments":[{"comment":"The central claim (abstract) that uncertainty bounds derived from the asymptotic distribution of normalised exceedance probabilities are approximately independent of sample size and useful for regional diagnostics rests on the approximation holding in finite samples. No simulation studies, finite-sample error bounds, or checks with small local exceedance counts are referenced to support this, which is load-bearing for the regional-diagnostics contribution given that extreme-value regression fits typically involve limited tail observations.","section":"Abstract (asymptotic derivation and diagnostic construction)"}],"minor_comments":[{"comment":"The abstract states that the diagnostics are illustrated in 'two applications' but provides no detail on the data types, covariate domains, or model classes used; adding a brief description would improve clarity on scope.","section":null}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their constructive comments on the manuscript. We address the major comment below.","responses":[{"response":"The manuscript derives the sample-size independence of the uncertainty bounds from the limiting distribution of the normalised exceedance probabilities. We agree with the referee that this asymptotic justification does not automatically guarantee good finite-sample performance, particularly when local exceedance counts are small. The current version contains no simulation studies or finite-sample error bounds to quantify the approximation error. We will add a simulation study in the revised manuscript that examines the behaviour of the standardised tail plots and normalised residual plots for small to moderate local sample sizes, including scenarios representative of regional diagnostics.","revision_made":"yes","referee_comment":"[Abstract (asymptotic derivation and diagnostic construction)] The central claim (abstract) that uncertainty bounds derived from the asymptotic distribution of normalised exceedance probabilities are approximately independent of sample size and useful for regional diagnostics rests on the approximation holding in finite samples. No simulation studies, finite-sample error bounds, or checks with small local exceedance counts are referenced to support this, which is load-bearing for the regional-diagnostics contribution given that extreme-value regression fits typically involve limited tail observations."}],"tokens_in":1342,"tokens_out":263,"duration_ms":20339,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main contribution is two new visual diagnostics: the standardised tail plot and the normalised residual plot. They are built so that uncertainty bounds come from the asymptotic distribution of normalised exceedance probabilities and end up roughly independent of the number of observations used to build each plot. This is meant to let users compare model fit globally and in different regions of the covariate space even when those regions have very different tail sample sizes.\n\nThe approach addresses a clear practical need. Extreme value regression is used for extrapolation, often on awkward covariate domains, and existing diagnostics are mostly global or hard to scale. Being able to run consistent regional checks without the bounds shifting just because one region has fewer exceedances is useful, and the paper shows the plots in action for comparing thousands of candidate models and for getting design feedback.\n\nThe soft spot is the finite-sample behavior of the key approximation. Extreme value models are usually fit with small numbers of tail points, and regional subsets can have even fewer. The claim that the bounds are approximately sample-size independent depends on the asymptotics being accurate enough at those sizes. The abstract states the result, but there is no indication of simulation checks or quantification of approximation error. If the full paper contains those checks the concern shrinks; if not, it is the main thing that needs more evidence.\n\nThis is for people who fit and use extreme value regression in applications like risk analysis. A reader who needs better regional diagnostics will find something concrete to try. It deserves peer review because the gap is real and the proposal is specific enough to test.","headline":"The paper introduces two new plots for regional diagnostics in extreme value regression using asymptotic bounds on normalised exceedance probabilities, but the finite-sample reliability of those bounds is unverified.","tokens_in":2319,"tokens_out":391,"would_cite":false,"duration_ms":16483,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Standardised tail plots and normalised residual plots enable consistent goodness-of-fit assessment for extreme value regression models regardless of sample size variation.","keywords":["extreme value regression","goodness of fit","visual diagnostics","tail plot","residual plot","exceedance probabilities","asymptotic distribution","model comparison"],"falsifier":"Simulation experiments showing that the width or coverage of the uncertainty bounds changes substantially with sample size in moderate-sized datasets would falsify the usefulness of the sample-size independence.","tokens_in":2677,"feed_emoji":"","tokens_out":510,"duration_ms":20270,"temperature":0.7,"pith_summary":"The paper develops two visual diagnostics for extreme value regression models to address the shortage of reliable tools for assessing fit, especially for extrapolation and on complex covariate domains. The standardised tail plot and the normalised residual plot are constructed so that their uncertainty bounds come from the asymptotic distribution of normalised exceedance probabilities. This makes the bounds roughly independent of the number of observations, allowing direct comparison of model performance at global and regional scales even when sample sizes differ across covariate regions. The diagnostics are demonstrated in applications involving model comparison over thousands of candidates and yielding practical design guidance.","feed_headline":"Plots diagnose extreme value models with sample-size-independent bounds","feed_subtitle":"Standardised tail and normalised residual plots allow consistent global and regional fit checks even when observation counts vary across cov","key_machinery":"The asymptotic distribution of normalised exceedance probabilities, which supplies uncertainty bounds independent of sample size for the standardised tail plot and normalised residual plot.","core_discovery":"The central discovery is that uncertainty bounds derived from the asymptotic distribution of normalised exceedance probabilities are approximately independent of sample size. This property underpins the standardised tail plot and normalised residual plot, which permit visual comparison of goodness-of-fit both globally and within specific regions of the covariate domain for extreme value regression models.","pith_inferences":["These plots could be extended to other types of regression models that involve tail behavior.","Automated selection of extreme value models might incorporate these diagnostics as objective criteria.","Validation through simulation studies with controlled misspecification would test the finite-sample accuracy of the bounds."],"forward_implications":["Model comparison becomes feasible across thousands of candidate extreme value regression models.","Regional assessment reveals where model fit is inadequate even if global fit appears acceptable.","Diagnostics remain interpretable on low-dimensional or non-Euclidean covariate domains.","Summary statistics for global and regional goodness-of-fit can be derived from the plots."],"fun_headline_variants":["Extreme value regression diagnostics with size-independent bounds","Standardised tail plots enable consistent extreme value model checks","Normalised residuals support sample-size-independent fit diagnostics","Regional goodness-of-fit visualised for extreme value regressions"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The asymptotic distribution of normalised exceedance probabilities approximates the finite-sample behavior closely enough for the uncertainty bounds to be practically independent of sample size.","fun_headline_variants_meta":{"raw":{"variants":["Extreme value regression diagnostics with size-independent bounds","Standardised tail plots enable consistent extreme value model checks","Normalised residuals support sample-size-independent fit diagnostics","Regional goodness-of-fit visualised for extreme value regressions"]},"model":"grok-4.3","cost_usd":0.00534,"raw_usage":{"total_tokens":2586,"prompt_tokens":685,"num_sources_used":0,"completion_tokens":58,"cost_in_usd_ticks":53399500,"prompt_tokens_details":{"text_tokens":685,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1843,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":685,"tokens_out":58,"duration_ms":14266,"temperature":1.0,"reasoning_tokens":1843,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-28T13:14:15.907067+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Simulation experiments showing that the width or coverage of the uncertainty bounds changes substantially with sample size in moderate-sized datasets would falsify the usefulness of the sample-size independence.","supporting_citations":[],"review_version":1}