{"id":"53338c0b-9710-48bf-bd57-fa5ce96d7033","arxiv_id":"2411.13608","paper_version":2,"verdict":"REJECT","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"A new benchmarking framework uses extreme value theory and correlation shifts to compare solar plants under rare low-production conditions, but its scoring formula and weights are not yet reliable.","lead":"The paper introduces EVDBM, a scoring method that combines extreme value analysis with weighted benchmarking to compare how solar plants behave during rare low-production events. The method is applied to two Portuguese PV plants, but the current version contains a flawed probability formula and arbitrary weights.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Equation (6) defines the scaling 'probability' as (1/T_return) × return level, which is dimensioned and not a valid probability; because S_j multiplies the whole benchmarking score, the forward-looking adjustment and the B2-vs-B1 ranking rest on an undefined quantity.","rationale":"The paper's central claim is that EVA-based return values make benchmarking scores forward-looking and that the resulting EVDBM score shows the Joao plant is less resilient. For that claim to hold, the scaling factor S_j in Section 3.3 must be a well-defined quantitative adjustment. It is not: Eq. (6) multiplies a reciprocal return period by a return level with physical units, then calls the product a probability. Since S_j enters multiplicatively into every benchmarking score, the computed scores in Figure 8 and the B2-vs-B1 ordering are not interpretable as stated. The reader's weakest_assumption focuses on the preset weights in Table 8, which is a genuine robustness concern and is even acknowledged in the paper's Limitations section. But the invalid probability equation is more load-bearing because it breaks the core mechanism of the proposed method, regardless of which weights are chosen. A corrected approach might still produce useful rankings, but the current manuscript does not establish that its forward-looking scores are meaningful. The reader's overall REJECT verdict therefore remains appropriate, and no adjustment to that verdict is needed.","tokens_in":13545,"tokens_out":4263,"duration_ms":42992,"concrete_test":"Recompute the EVDBM scores for Section 4.3 from the fitted GPD using a valid formulation: either evaluate the actual tail probability P(X ≤ r_T) at each return level r_T estimated from the GPD, or, if the intended input is the return level itself, replace the scaling factor with S_j = E_c × r_T (dimensionally appropriate) and re-run the benchmarking. If the ordering of B_Zarco and B_Joao changes under either correction, or if the corrected S_j values differ nontrivially from those used in the paper, the reported resilience conclusion is an artifact of the ill-defined formula. A minimal unit check on Eq. (6) would also settle it: P_j must be dimensionless and between 0 and 1.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.3, Eq. (6), states P_j(X > x) = (1/T_return) × x(T)_CI, and Step 1 then treats 'Return Value', 'Lower CI', and 'Upper CI' as the probability inputs (a)-(c). A probability must be dimensionless and lie in [0,1]; the right-hand side is the reciprocal return period times a return level expressed in production units (kWh), which can be near zero, negative, or such that the product exceeds 1. This is not a minor notational slip: S_j = E_c × P_j in the scaling-factor equation is the only mechanism claimed to make scores 'forward-looking'. Because the scaling factor is undefined, the adjusted benchmarking scores in Figure 8 have no well-defined interpretation, and the statement that 'B2 (joao plant) is more sensitive to adverse conditions and is less resilient' is not supported by the calculation as written. The paper's own limitation section admits weight subjectivity, and the weights in Table 8 are explicitly 'dummy' and 'for testing purposes,' so the ranking is also not robust to a different reasonable weight choice; however, the invalid scaling probability is the more fundamental defect because it undermines the score before any weighting is applied.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes the Extreme Value Dynamic Benchmarking Method (EVDBM), a three-stage pipeline that (i) applies extreme value analysis (POT with a GPD fit) to a dependent variable, (ii) uses a novel 'DISC-Thresholding' algorithm to flag variables whose Pearson correlations change between extreme and normal periods, and (iii) computes a weighted benchmark score B_i for each use case, scaled by a factor S_j that combines historical extreme-event frequency with EVA return levels. The method is demonstrated on hourly production data from two Portuguese PV plants (Zarco and Joao), with the claims that the scores are 'forward-looking' and that the Joao plant (B2) is less resilient to adverse conditions. The paper also lists limitations concerning data quality, stationarity, and weight subjectivity.","tokens_in":13835,"tokens_out":6269,"duration_ms":64311,"significance":"If the central calculation were valid, the EVDBM idea would be a useful addition to the EVA benchmarking literature: it explicitly couples tail-risk estimation with multi-criteria performance comparison and is demonstrated on real operational data from two PV plants. The DISC-Thresholding procedure is simple and transparent, and the authors are honest about several limitations. However, as the manuscript stands, the load-bearing scaling factor in Eq. (6) is not a probability, the GPD fits are not documented, and the final ranking depends on unexamined 'dummy' weights. These issues prevent the results from supporting the paper's main claims, though they are, in principle, addressable in a major revision.","major_comments":[{"comment":"The quantity P_j(X > x) = (1/T_return) × x(T)_CI is not a probability: the right-hand side has units of production per time (e.g., kWh/year if return levels are in kWh), it can be negative or exceed 1, and the text's own example (1/T = 0.2) ignores the multiplicative return level. Because S_j = E_c × P_j in Eq. (7) is the only mechanism that makes the benchmark 'forward-looking,' the adjusted scores in Figure 8 and the B2 > B1 conclusion are undefined as written.","section":"Section 3.3, Eq. (6)"},{"comment":"The algorithm treats the return value and the two confidence-interval endpoints as three 'probabilities' P_j(X > x0), P(X > x_lower), and P(X > x_upper), but no transformation to a dimensionless quantity in [0,1] is specified. Consequently, B_i = S_j × Σ b(V_i) mixes a dimensionless weighted score with this ill-defined factor, and the three curves in Figure 8 have no clear probabilistic interpretation.","section":"Section 3.3, Step 1 items (a)-(c) and Eq. (9)"},{"comment":"The GPD fits are not documented with parameter estimates, threshold-selection details, or proper diagnostic checks; the reported R² = 0.997 and p-value = 0.000 are not defined, and p-values cannot be exactly zero. In addition, the confidence-interval columns are internally inconsistent: for Zarco's 1-year return value of 0.35, the reported 'Lower CI' is 1.66 and 'Upper CI' is 0.11, so the lower bound exceeds the upper bound. These values feed directly into Step 1(b)-(c), making the benchmark inputs unreliable.","section":"Section 4.1.2 / 4.2.2, Tables 3 and 6"},{"comment":"The variable weights in Table 8 are explicitly described as 'dummy' and chosen 'for testing purposes,' and the final ranking B2 > B1 is linear in these weights. No sensitivity analysis is provided, and the paper's own limitation section acknowledges that weighting subjectivity can bias comparisons. Therefore the statement in Section 4.3 that the Joao plant 'is more sensitive to adverse conditions and is less resilient' is not robust to reasonable alternative weight choices.","section":"Section 4.3, Table 8"},{"comment":"The DISC-Thresholding algorithm classifies correlation changes as 'significant' based only on 10th/90th percentiles of the empirical distribution of Δρ_ij, without any hypothesis test, confidence interval, or measure of sampling variability. The term 'significant' is therefore not statistically justified, and the identified HPC/HNC pairs in Figure 6 should be interpreted as data-driven threshold exceedances rather than significance findings.","section":"Section 3.2 / Section 4.1.3"}],"minor_comments":[{"comment":"The time range is inconsistent: the text in Section 4 states a 3-hour window from 13:00 to 16:00, while Table 1 lists 'Time-Range 13:00-14:00'.","section":"Section 4 and Table 1"},{"comment":"The equation numbering is confused: Step 2 labels b(V_i) = w_i × C_extreme(V_i) as Eq. (7), but Step 3 references Eqs. (6), (7), and (8) before introducing Eq. (9); Eq. (8) is never defined.","section":"Section 3.3"},{"comment":"The L-kurtosis and L-skewness formulas are duplicated verbatim in consecutive paragraphs, which disrupts the presentation.","section":"Section 2.1"},{"comment":"The text says that the return values are 'mostly negative,' but the values in Table 3 and Table 6 are all positive; this discrepancy should be clarified.","section":"Section 4.1.2 and Table 3"},{"comment":"The abbreviation 'HNN' is used in Figure 6 and in Section 4.1.3, but the abbreviation list defines 'HNC' as 'High Negative Correlation'; the notation should be unified.","section":"Appendix"}],"recommendation":"major_revision","confidential_remarks":"The manuscript has a promising conceptual framework and a real-data demonstration, but the current version is not publishable because the core scaling factor is dimensionally invalid and the final ranking is built on undocumented fits and arbitrary weights. I would encourage the authors to rework Section 3.3 with a proper probability or severity measure, document the EVA fits, and add a sensitivity analysis for the weights; after those changes, the paper could be reconsidered."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. The paper does something genuinely new: it couples EVT return levels with a correlation-shift screen and a weighted benchmark to compare PV plants under tail risk. The DISC algorithm is transparent—a percentile rule on correlation differences—and the data are real and cited. That part is fine.\n\nThe soft spot is load-bearing. Equation (6) defines the exceedance 'probability' as (1/T_return) × x(T)_CI. That is dimensionally inconsistent: the return level is in kWh, so the product has units of kWh/year and can be negative or exceed 1. This isn't a typo, because S_j = E_c × P_j multiplies the entire benchmarking score, so the forward-looking adjustment and the B2-vs-B1 ranking rest on an undefined quantity. The paper's own limitation section admits the weights are subjective, and the weights in Table 8 are explicitly 'dummy', so the ranking is also fragile to reasonable weight changes. On top of that, the GPD fits are asserted with R² and p-values but no parameter estimates or diagnostic plots—and the confidence intervals in Table 3 look inverted: the 'lower' CI is larger than the 'upper' CI.\n\nCredit where due: the empirical analysis is not sloppy in every respect. The authors identify a real problem—static benchmarking ignores tail behavior—and the overall pipeline is described clearly enough to reproduce and fix. The citation pattern is normal, with relevant EVT and energy references. But the central claim that the scores are 'forward-looking' is not supported by the math as written.\n\nWho is this for? Practitioners who want a ready-made EVT benchmark should wait for a corrected version. A referee could fix Eq. (6) with a proper exceedance probability (e.g., 1/T) or a hazard-based scaling, and then re-examine the weights. But as submitted, the paper has a fundamental defect in its core equation.\n\nMy recommendation: desk reject with an invitation to resubmit after fixing the scaling factor and adding sensitivity analysis. It does not deserve referee time in its current form.","headline":"The central scaling equation is dimensionally invalid, so the forward-looking benchmark and the B2>B1 ranking rest on an undefined quantity; the EVT-correlation pipeline is clearly described but needs a corrected probability and sensitivity analysis before it supports conclusions.","tokens_in":14352,"tokens_out":3105,"would_cite":false,"duration_ms":29505,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62G32","62H20"],"pacs":[],"model":"deepseek-v4-flash","headline":"Feeding extreme-value return levels into a weighted benchmark yields forward-looking resilience rankings of solar plants, and the paper applies this to two Portuguese PV sites.","keywords":["extreme value analysis","benchmarking","photovoltaic energy","return levels","peaks over threshold","correlation shifts","resilience","risk management"],"falsifier":"Recompute $B_i$ for both plants under several defensible weight vectors (e.g., equal weights, weights proportional to each variable's absolute correlation with production, or leave-one-out weights) and check whether Joao remains the higher-scoring plant; if the ordering flips, the paper's resilience conclusion collapses. A complementary out-of-sample check is to hold out the most recent years and test whether the plant with the higher EVDBM score actually exhibits more frequent or deeper low-production extremes in the holdout period.","tokens_in":13395,"feed_emoji":"☀️","tokens_out":7930,"duration_ms":75488,"temperature":0.7,"pith_summary":"The paper proposes a method, the Extreme Value Dynamic Benchmarking Method (EVDBM), that turns extreme value analysis into a forward-looking score for comparing how different systems will behave under rare adverse conditions. It claims that feeding extreme value theory return values, computed from a peaks-over-threshold fit, into a weighted sum of correlated weather variables produces a benchmark that reflects anticipated extremes rather than only past averages. Applied to two Portuguese photovoltaic plants, the method ranks one plant (Joao) as more sensitive and less resilient to low-production extremes, even though the plants look similar in ordinary statistics. The value of the claim is that asset owners and grid operators could use such scores to prioritize investment, storage, and risk management before the extreme events arrive.","feed_headline":"Extreme-value scores rank solar plants by future low-output risk","feed_subtitle":"Return-level forecasts, not history alone, decide which PV plant is less resilient.","key_machinery":"The load-bearing objects are (1) the Dynamic Identification of Significant Correlation (DISC)-Thresholding algorithm, which computes $\\Delta\\rho_{ij} = \\rho_{ij}^{\\text{extreme}} - \\rho_{ij}$ for every variable pair and flags changes above the 90th percentile or below the 10th percentile as High Positive or High Negative Correlation, and (2) the EVA-Driven Weighted Benchmarking score $B_i = S_j \\sum_i w_i C_{\\text{extreme}}(V_i)$, where $S_j = E_c \\times P(X>x)$ combines the normalized historical extreme-event frequency with the return-level exceedance probability, and $w_i$ are user-supplied weights. The first component tells the analyst which associations between weather variables and production actually change during extreme-low-production days; the second converts those extreme-condition statistics, scaled by projected return values and their confidence intervals, into a single comparable score per plant.","core_discovery":"The central claim is that benchmarking scores can be made forward-looking by multiplying a historical frequency factor $E_c$ by an exceedance probability $P(X>x)$ derived from the fitted extreme value distribution, and then weighting the extreme-condition statistics of related variables by pre-assigned importance weights $w_i$. In the two-plant photovoltaic study, this procedure assigns plant Joao (B2) a higher benchmarking score than plant Zarco (B1), which the paper interprets as evidence that Joao is more sensitive to adverse conditions and less resilient. The paper states this directly in Section 4.3: \"Since B2 has the highest EVDBM score it is clear that B2 (joao plant) is more sensitive to adverse conditions and is less resilient to fluctuations in related circumstances.\" If correct, the method converts historical extremes plus projected return values into a vulnerability ranking that static historical benchmarks cannot provide.","pith_inferences":["A direct testable extension is a weight-sensitivity analysis: recomputing $B_i$ under equal weights, data-driven weights, or weights derived from regression coefficients would show whether the Joao-less-resilient ranking is robust or an artifact of the 'dummy' proportions in Table 8.","The 90th/10th percentile thresholds in the DISC algorithm are preset; varying them would test whether the flagged high-correlation pairs, and therefore the variable weights used in benchmarking, are stable across threshold choices.","The same pipeline transfers to other extreme-low events, such as financial drawdowns or hospital admissions, where a return level plays the role of low PV production; the paper lists these as intended future applications but does not run them.","Because the paper acknowledges stationarity as only partially addressed, a natural next step is to make the exceedance probability time-dependent with climate or market covariates, turning the benchmark into an early-warning indicator rather than a static ranking."],"forward_implications":["Plant operators could use EVDBM scores to rank facilities by projected vulnerability before a rare low-output event occurs, not just after reviewing historical performance.","Because scores are computed from return values and their confidence intervals, the benchmark carries an explicit uncertainty range that widens with return period.","The DISC algorithm flags which pairwise correlations between weather variables and production shift most during extreme-low days, giving a diagnostic list of variables to monitor.","The same algorithmic pipeline can be reapplied to any dependent variable with extreme values and correlated covariates, making the score a generic benchmarking tool."],"supporting_citations":[{"why":"Supplies the extreme value theory foundations, including block maxima and generalized extreme value distributions, that justify the EVA step.","marker":"[18]"},{"why":"Supplies the asymptotic result that peaks over threshold follow the generalized Pareto distribution, which underlies the return-value calculations.","marker":"[19]"},{"why":"Supplies the Pearson correlation measure that the DISC-Thresholding algorithm uses to compare correlations between extreme and normal conditions.","marker":"[21]"},{"why":"Provides the prior circumstance-analysis process on which EVDBM builds and which the paper extends into benchmarking.","marker":"[28]"},{"why":"Provides the photovoltaic power production dataset used for both use cases, so the benchmark scores are computed on real data.","marker":"[31]"}],"fun_headline_variants":["New extreme-value method ranks solar plants by future low-output risk","EVDBM scores rank solar plants by extreme future output risk","Forward-looking benchmarking predicts vulnerable solar plants","Dynamic extreme-value model ranks solar plants by projected risk","Statistical method forecasts which PV plant is less resilient"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The final ranking rests on the manually preset variable weights in Table 8, which the paper calls \"dummy weights\" chosen \"for testing purposes\"; since the score is linear in those weights and no sensitivity analysis is given, a different reasonable weight set could reverse the conclusion that Joao is less resilient.","fun_headline_variants_meta":{"raw":{"variants":["New extreme-value method ranks solar plants by future low-output risk","EVDBM scores rank solar plants by extreme future output risk","Forward-looking benchmarking predicts vulnerable solar plants","Dynamic extreme-value model ranks solar plants by projected risk","Statistical method forecasts which PV plant is less resilient"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000645,"raw_usage":{"total_tokens":2960,"prompt_tokens":940,"completion_tokens":2020,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":556,"completion_tokens_details":{"reasoning_tokens":1944}},"tokens_in":556,"tokens_out":2020,"duration_ms":15561,"temperature":1.0,"reasoning_tokens":1944,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T17:07:31.938624+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute $B_i$ for both plants under several defensible weight vectors (e.g., equal weights, weights proportional to each variable's absolute correlation with production, or leave-one-out weights) and check whether Joao remains the higher-scoring plant; if the ordering flips, the paper's resilience conclusion collapses. A complementary out-of-sample check is to hold out the most recent years and test whether the plant with the higher EVDBM score actually exhibits more frequent or deeper low-production extremes in the holdout period.","supporting_citations":[{"cited_title":"Coles, J","cited_arxiv_id":null,"evidence_quote":"Supplies the extreme value theory foundations, including block maxima and generalized extreme value distributions, that justify the EVA step."},{"cited_title":"Xu, Proceedings of 2013 World Agricultural Outlook Conference, Springer, 2014","cited_arxiv_id":null,"evidence_quote":"Supplies the asymptotic result that peaks over threshold follow the generalized Pareto distribution, which underlies the return-value calculations."},{"cited_title":"Jebli, F.-Z","cited_arxiv_id":null,"evidence_quote":"Supplies the Pearson correlation measure that the DISC-Thresholding algorithm uses to compare correlations between extreme and normal conditions."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the prior circumstance-analysis process on which EVDBM builds and which the paper extends into benchmarking."},{"cited_title":"Sarmas, M","cited_arxiv_id":null,"evidence_quote":"Provides the photovoltaic power production dataset used for both use cases, so the benchmark scores are computed on real data."}],"review_version":1}