{"id":"ef80d885-cb7e-41dd-8229-67de15c955e2","arxiv_id":"2412.12236","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"The sustainability metric used in the Cool Copper Collider comparison is shown to be arithmetically wrong and mathematically ill behaved, making its rankings unreliable.","lead":"This comment shows that a sustainability metric proposed for ranking future Higgs factories contains an arithmetic error and is not a well-behaved measure of physics output. It also demonstrates that any multiplicative metric of this kind can be tuned to produce different rankings.","discovery_kind":"replication","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The strongest conclusion overreaches: exponent-induced re-ranking of W^x does not prove that every multiplicative metric is valueless, because a metric with a physically fixed normalization has no free exponent.","rationale":"The reader's weakest-assumption analysis already identifies the same load-bearing concern: the paper's conclusion that the multiplicative-metric strategy is 'valueless' depends on treating the metric exponent as unconstrained, which is not the same as showing that no principled metric can exist. I agree with that assessment. The central technical findings—the factor-100 hhh bug, the non-monotonic behavior of the original weighted average, and the inconsistency between kappa-0 and kappa-3 fits—are well supported and justify a conditional acceptance. My stress-test does not add a new objection; it sharpens the existing one by noting that a physically grounded definition of 'physics output' fixes both normalization and exponent, so the tunability shown in Fig. 4 is an argument against unnormalized metrics, not against all multiplicative metrics. Thus the reader's CONDITIONAL verdict stands unchanged.","tokens_in":5123,"tokens_out":6076,"duration_ms":66115,"concrete_test":"Build a fixed, physical output measure: for CLIC380, ILC250, ILC500, FCC-ee, and CEPC, compute the integrated luminosity (or run time) L_i required to reach a common target precision (e.g. Delta(kappa_hZZ) = 0.1% and Delta(kappa_hhh) = 5%) in a common kappa-0 or SMEFT fit including HL-LHC, using each collider's luminosity-scaling laws. Rank by E_i * L_i (electricity per unit of physics output) and compare with the Fig. 4 rankings. If this ranking is well-defined and stable under equivalent formulations (inverse variance versus inverse standard deviation), the arbitrary-exponent argument does not prove multiplicative strategies valueless.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The arithmetic-bug and non-monotonicity findings are solid, but the sweeping conclusion in the abstract and final paragraph overreaches. The demonstration that E_i * W_i^x changes ordering when x is treated as a free parameter does not establish that every multiplicative metric is valueless. If 'physics output' is given a physical definition—e.g. inverse variance from a global SMEFT fit, or the integrated luminosity needed to reach a fixed target precision—then the exponent and normalization are fixed by that definition; raising W to an arbitrary power is no longer a permitted reparametrization. Fig. 4's tunability therefore shows that the particular unnormalized W in Eq. (3), and any metric whose scale is unspecified, is fragile, not that the multiplicative strategy is in principle incapable of meaningful comparison. The paper itself gestures at the needed fix ('putting all colliders on an equal footing'), which undercuts the 'valueless' verdict.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript by Grojean and Janot is a comment on the sustainability metric proposed in Breidenbach et al., PRX Energy 2, 047001 (2023) (Ref. [1]). It makes three categories of claims: (i) the weighted average of Higgs-coupling precisions reported in Ref. [1] cannot be reproduced from the stated equations, and the disagreement is traced to an apparent factor-100 error in the Higgs self-coupling precision and an additional HZγ bounding condition; (ii) the estimator defined by Eqs. (1)-(2) is non-monotonic, as illustrated in Fig. 2, so improving a coupling precision can worsen the weighted average; and (iii) the general strategy of multiplying electricity consumption or carbon footprint by any such 'metric' is 'fragile at best, and valueless' because arbitrary powers of the metric change collider rankings, as shown in Fig. 4. The paper also lists inaccuracies in the input table of Ref. [1] and proposes the alternative quadratic metric W in Eq. (3).","tokens_in":5254,"tokens_out":6505,"duration_ms":60823,"significance":"If the arithmetic findings are correct, this is a useful and important correction to a published, high-profile comparison: the numerical conclusions of Ref. [1] appear to depend on a coding error and on an inconsistent treatment of the kappa-fit assumptions. The non-monotonicity demonstration in Fig. 2 is a mathematically solid and instructive result, independent of the private-communication bug, and it undermines the original interpretation of the weighted average as a 'precision per unit of physics output.' The proposed W metric with a common kappa-0 fit is a reasonable starting point for a fairer comparison. However, the manuscript's sweeping claim that the entire multiplicative-metric strategy is 'valueless' is not supported by the evidence. The exponent-tunability argument applies to unnormalized, dimensionless metrics of the form of Eq. (3), but not to a metric whose normalization is fixed by a physical definition, such as the integrated luminosity needed to reach a target precision.","major_comments":[{"comment":"The claim that the multiplicative-metric strategy is 'valueless when it comes to arguing in favour of or against such or such future Higgs factory' overreaches. Fig. 4 shows that rankings change as a function of the exponent x in W^x, but this is only because W in Eq. (3) has no fixed physical normalization. A metric defined by a physical requirement—for example, the integrated luminosity each collider needs to reach a specified coupling precision, or the inverse Fisher information from a global SMEFT fit—has a unique normalization, and raising it to an arbitrary power is not a permitted reparametrization. The paper's own final sentence, which recommends understanding 'with how much integrated luminosity they would all reach the same coupling precision,' defines precisely such a fixed metric, undercutting the 'valueless' verdict. The conclusion should be narrowed to the dimensionless, unnormalized 'improvement' metrics of the form of Eqs. (1)-(3).","section":"Abstract and final paragraph"},{"comment":"The reproduction of the original weighted averages rests on two pieces of information that are not fully disclosed in the manuscript: the authors' private communication with an author of Ref. [1] (including a screenshot of the input table) and an unspecified 'HZγ precision values be bounded from above by the HL-LHC value' instruction. To make the claimed factor-100 bug and the 'cannot reproduce' result independently checkable, the paper should include the exact input table (or a machine-readable listing), the values after the factor-100 correction, and the precise bounding rule used. As it stands, Table 1 is not reproducible from the information in the manuscript alone.","section":"Section 1 and Table 1"},{"comment":"The wording 'it is not a metric' conflates 'metric' in the mathematical sense of a distance function with 'metric' in the sense of a figure of merit, which is the sense used in Ref. [1]. The demonstrated non-monotonicity is a genuine and well-illustrated flaw, but it does not imply that the object ceases to be a metric in the mathematically strict sense; the correct statement is that the estimator is not a monotonic function of the coupling precisions. The rest of the paragraph and Fig. 2 are convincing and should be retained, but the terminology should be corrected.","section":"Item 2 and Fig. 2"}],"minor_comments":[{"comment":"The sentence 'We also demonstrates that' contains a grammar error and should read 'We also demonstrate that.'","section":"Abstract"},{"comment":"The phrase 'and and definitely ought to be included' contains a duplicated 'and' and should be corrected.","section":"Item 3"},{"comment":"The table heading contains a typo: 'T able 1' should be 'Table 1.'","section":"Table 1 heading"},{"comment":"The carbon intensity of 20 kg CO2e per MWh is a very low value; the ranking curves in Fig. 4 would rotate if a higher or regional carbon intensity were used. Since the conclusion is about exponent tunability, it would be helpful to state whether the qualitative message is robust to the carbon-intensity choice.","section":"Fig. 4 caption"},{"comment":"The argument that coupling-precision improvements relative to HL-LHC 'may not say much about the corresponding sensitivity to new physics' is presented as a general statement, but the supporting quantitative evidence is deferred to Ref. [6]. The manuscript should either state this explicitly as a qualitative expectation or present a concrete example, to avoid the impression of relying on an unpublished companion note.","section":"Item 4 and Refs. [5,6]"}],"recommendation":"major_revision","confidential_remarks":"The authors have a direct stake in the comparison through their companion paper Ref. [6], and the comment promotes that work as the preferred approach. This does not invalidate the technical findings, but the 'valueless' language in the abstract and conclusion appears stronger than the evidence warrants and should be moderated. In addition, the factor-100 bug finding is based on a private communication; if the authors wish to establish this fact in the published record, they should include the original input table in a supplementary file or a figure, so that the arithmetic claim is independently verifiable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this comment catches a real arithmetic bug and shows the original metric is non-monotonic in coupling precision, but the abstract's stronger claim that any multiplicative metric is 'valueless' doesn't follow from their own examples.\n\nThe factor-100 error in the hhh precision, traced with input from the original authors, is concrete and reproducible from the published table. Their Table 1 shows the ranking changes materially, and the non-monotonicity demonstration in Fig. 2 is a genuine mathematical observation about Eqs. 1-2. That alone invalidates the specific estimator used in the C3 sustainability paper. Credit is due for contacting the authors first and documenting the exchange.\n\nThe soft spot is the final verdict. Showing that W^x reorders colliders when x is a free parameter does not show that every multiplicative metric is meaningless. If a metric is defined from a physical target—inverse variance from a global SMEFT fit, or the luminosity needed to reach a fixed precision—the exponent and normalization are fixed by that definition. Fig. 4 demonstrates the fragility of an unnormalized weight, not the impossibility of a principled one. The paper even gestures at the fix ('putting all colliders on an equal footing'), which undercuts the 'valueless' conclusion.\n\nOne smaller practical issue: the paper does not ship a machine-readable table, so independent reproduction of Table 1 requires some digitizing effort. The equations are clear enough, though. The citations to the authors' own companion papers are not used as evidence for the central bug, so self-citation is not a concern here.\n\nWho is this for? Anyone weighing the C3 sustainability ranking, the broader collider-choice debate, and people who care about how metrics are constructed and abused. It deserves a serious referee: a comment that demonstrates a load-bearing arithmetic error in a published ranking should be reviewed, even if the general conclusion needs tempering. I would send it to review and ask the authors to scale back the sweeping claim—the specific critique is strong enough on its own.","headline":"Solid arithmetic-bug and non-monotonicity critique of a published sustainability metric; the broad 'valueless' conclusion overreaches.","tokens_in":5801,"tokens_out":1419,"would_cite":true,"duration_ms":14533,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A critical re-analysis shows the proposed sustainability metric for Higgs factories is not reproducible, is non-monotonic, and can be tuned to produce any collider ranking.","keywords":["Higgs factory","sustainability metric","carbon footprint","coupling precision","kappa fit","HL-LHC","collider comparison","reproducibility"],"falsifier":"Take the original input table, correct only the Higgs self-coupling values by the factor 100, and recompute Eqs. (1)-(2); if the result matches the published averages for the 550 GeV copper-collider entry and FCC-ee, the arithmetic-bug claim is falsified, and an independent non-arbitrary rule for choosing the exponent of any candidate metric would be needed to rebut the tunability objection.","tokens_in":4888,"feed_emoji":"⚛️","tokens_out":8069,"duration_ms":67281,"temperature":0.7,"pith_summary":"This comment tries to establish that the \"carbon footprint per unit of physics output\" metric proposed for future Higgs factories is mathematically and empirically unsound, and that no multiplicative precision weight can settle which collider to build. It does this by checking the arithmetic, finding corrected weighted averages that differ sharply from the published ones, and by pointing out several internal inconsistencies in the precision table. The deeper argument is structural: any metric of this kind has an adjustable scale that can reverse rankings at will. If the comment is correct, sustainability arguments for or against a specific Higgs factory should be based on an equal-physics-output comparison, using the integrated luminosity needed to reach the same coupling precision, rather than on a precision-weighted footprint. The reader should care because these numbers are being used in real decisions about very expensive future facilities.","feed_headline":"Correcting a factor-100 error flips Higgs factory green ranking","feed_subtitle":"Reproducing the original calculation gives a different winner, and the metric is not even monotonic in coupling precision.","key_machinery":"The central object is the weighted average $\\langle \\delta\\kappa/\\kappa\\rangle$ of Eq. (2), with per-coupling weight $w_i$ of Eq. (1) measuring the relative improvement of each Higgs-coupling precision at a future collider over the HL-LHC baseline. The paper shows that this object is not monotonic in the precisions: a better self-coupling measurement can worsen the average until the precision reaches roughly the percent level, and it cannot handle couplings invisible at HL-LHC. The proposed replacement, $W = 100 \\times [\\sum_i (\\delta\\kappa/\\kappa)_{\\mathrm{HL-LHC},i}^2 / (\\delta\\kappa/\\kappa)_{\\mathrm{HL-LHC+HF},i}^2]^{-1}$, is monotonic, but the comment then varies the exponent $x$ in $W^x$ to show that every ordering can be produced, which is the load-bearing demonstration against the multiplicative-metric strategy.","core_discovery":"The paper's central claim is that the metric introduced in the original sustainability study is neither a valid estimator nor a fair basis for comparing Higgs factory concepts. The comment reproduces the weighted-average calculation from the original table, finds values very different from those published, traces one large discrepancy to a factor-100 error in the Higgs self-coupling precisions, and shows that the estimator is not monotonic: improving a measured precision can make the weighted average worse. It further demonstrates that even a mathematically sound replacement metric, the quadratic combined improvement W, leaves the ranking arbitrary because raising W to any power preserves its formal properties while changing the ordering; hence the strategy of multiplying electricity or carbon figures by a single precision metric is, in the authors' words, fragile at best and valueless for choosing a Higgs factory.","pith_inferences":["Beyond the paper, the arbitrariness argument applies to any composite sustainability score built from a ratio of precisions, not just to this one; choosing weights, exponent, or baseline is a policy choice that should be made transparently.","Beyond the paper, the factor-100 error illustrates a broader reproducibility lesson: published footprint numbers should ship with their input tables and code, otherwise a single typo can flip a large infrastructure decision.","Beyond the paper, one could test the equal-luminosity normalization directly: if each collider is assigned the luminosity needed to reach a common precision, the electricity and carbon ranking may become stable and less sensitive to metric choices."],"forward_implications":["The published ranking that made the 550 GeV copper-collider option look best per unit of physics output does not survive the corrected arithmetic; the recalculated averages are 1.21 for both that option and FCC-ee rather than 0.45 and 0.59.","Any future comparison must use the same kappa-fit assumptions for every collider; mixing a kappa-0 fit for linear colliders with a kappa-3 fit for circular colliders gives linear colliders an artificial advantage.","A repaired quadratic metric W still leaves the ranking arbitrary, because any power W^x preserves monotonicity while changing the order; no scientific collider choice can be made until the exponent is fixed by an objective criterion.","The practical implication is to compare colliders at equal physics output, meaning the integrated luminosity needed to reach the same coupling precision, before contrasting electricity and carbon footprints."],"supporting_citations":[{"why":"Defines the metric and the coupling-precision table that the comment tries to reproduce, correct, and evaluate.","marker":"[1]"},{"why":"Defines the kappa-0 and kappa-3 fits and supplies the FCC-ee kappa-0 precision values used in the recalculations.","marker":"[2]"},{"why":"Is the original source of the precision table; its caption documents that different columns correspond to different kappa fits.","marker":"[3]"},{"why":"Provides globally consistent SMEFT fit results for all colliders, used to argue for an internally consistent comparison.","marker":"[5]"},{"why":"Outlines the equal-integrated-luminosity comparison the comment points to as the sound alternative.","marker":"[6]"},{"why":"Supplies the CLIC and ILC life-cycle assessment used for the carbon-footprint curves in the exponent-tunability figure.","marker":"[7]"},{"why":"Supplies the FCC construction carbon-footprint benchmark used for the same figure.","marker":"[8]"}],"fun_headline_variants":["Factor-100 error flips Higgs factory green ranking","Metric for collider sustainability called flawed and valueless","Non-monotonic precision metric derails Higgs factory comparisons","Comment shows sustainability metric misrepresents collider reality"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The critique stands on the assumption that the coupling-precision inputs used in the recalculation are the right ones, meaning that the original table, once the factor-100 error and the fit mismatch are corrected, fairly represents what each collider would measure; if those projections are themselves stale or not comparable, the corrected rankings would change, though the non-monotonicity and exponent-tuning objections would remain.","fun_headline_variants_meta":{"raw":{"variants":["Factor-100 error flips Higgs factory green ranking","Metric for collider sustainability called flawed and valueless","Non-monotonic precision metric derails Higgs factory comparisons","Comment shows sustainability metric misrepresents collider reality"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000238,"raw_usage":{"total_tokens":1440,"prompt_tokens":805,"completion_tokens":635,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":421,"completion_tokens_details":{"reasoning_tokens":570}},"tokens_in":421,"tokens_out":635,"duration_ms":6300,"temperature":1.0,"reasoning_tokens":570,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T14:21:13.444133+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the original input table, correct only the Higgs self-coupling values by the factor 100, and recompute Eqs. (1)-(2); if the result matches the published averages for the 550 GeV copper-collider entry and FCC-ee, the arithmetic-bug claim is falsified, and an independent non-arbitrary rule for choosing the exponent of any candidate metric would be needed to rebut the tunability objection.","supporting_citations":[],"review_version":1}