{"id":"b9d29452-6f11-44e4-a572-c0cb2381efcd","arxiv_id":"1908.04835","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"An exhaustive model selection over 511 Wilson-coefficient combinations for b to s l l decays finds that all surviving scenarios contain O9 (Delta C9 near -1.1 to -1.4), the only surviving one-operator scenario.","lead":"This paper exhaustively ranks 511 new-physics scenarios against B meson decay data using two statistical selection tools. Nearly every surviving scenario includes a left-handed vector muon coupling (O9), and it is the only one-operator scenario selected in the 'Moments' analysis.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"SM baseline is excluded from the 511-model set, so the claim that O9 is the only surviving one-operator scenario is untested against the 'no NP' hypothesis; the paper reports no SM chi2.","rationale":"Reader's weakest assumption matches mine: the SM is absent from the candidate pool. I think this is the single most load-bearing issue because it directly controls the interpretation of 'survive the test.' The numerical check is sharp: Eq. (12) with K=0 gives AICc_SM = chi2_SM, and model 2's AICc is known from Table II, so the comparison is a two-line computation once chi2_SM is available. The paper's own selected models have excellent p-values (about 57%), so it is entirely possible that SM chi2 is a few units larger and the NP conclusion survives; but that is an empirical fact the manuscript should report, not assume. The other concerns (missing repository, Model 2 wording contradiction, abstract lacking Moments/Likelihood qualification) are real but secondary: they affect reproducibility and precision of presentation, not the logical status of the NP-vs-SM comparison. I therefore do not move the reader's CONDITIONAL verdict; I would make the SM-baseline comparison a required condition for acceptance.","tokens_in":20625,"tokens_out":10819,"duration_ms":112955,"concrete_test":"Add the Standard Model (all NP Wilson coefficients fixed to zero, K=0) to the candidate set and rerun the exact same chi2 minimization and LOO-CV pipeline for the New Moments dataset. Compute chi2_SM and AICc_SM = chi2_SM via Eq. (12); compare with model 2's AICc ≈ 254.45 and with the selected two-operator models. If AICc_SM < 254.45, the abstract's claim fails because the data do not prefer any NP scenario. If AICc_SM > 254.45, report the value so the 'only surviving one-operator scenario' can be stated as a genuine preference over the SM. Also report the SM LOO-CV MSE to check whether cross-validation ranks the SM above all 511 models.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central conclusion ('only surviving one-operator scenario' with O9, and by implication that NP is needed) is established only within the 511 non-empty NP-operator combinations defined in Sec. IV A. The SM (all Wilson coefficients zero) is not a candidate. This omission is load-bearing because Eq. (12) makes the comparison quantitative: for the New Moments data, model 2 (O9, K=1, chi2_min=252.44, n=258) has AICc = 252.44 + 2 + 4/256 ≈ 254.45; the SM has K=0 and AICc = chi2_SM. If chi2_SM < 254.45, AICc prefers the SM over every NP model, so 'survives the test' would not mean 'preferred over no new physics.' The paper never reports chi2_SM, and the abstract's 'best explain the available data' and 'only surviving one-operator scenario' are unqualified. This is not a purely philosophical point: ref. [27] explicitly ranks NP models against the SM using an information criterion, and the present analysis itself states the hierarchy is relative to the best model in the chosen set. Adding the SM as a zero-parameter candidate is the natural null comparison for both AICc and leave-one-out cross-validation. Without it, the headline conclusion is conditional on NP being present, which is the very assumption the selection is meant to test.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper performs an exhaustive frequentist fit of 511 non-empty combinations of nine dimension-six Wilson coefficients to b→sℓℓ data, and ranks the resulting models with AICc and leave-one-out cross-validation for four dataset variants ('Old'/'New' LFUV data × 'Moments'/'Likelihood' angular data). The authors report that a small set of one-, two-, and three-operator scenarios survive, that all surviving scenarios contain the left-handed vector muon coupling O9, and that O9 is the only surviving one-operator scenario, with ΔC9 best-fit values near −1.1 to −1.4. They also study the impact of RK(*) data, check consistency with radiative B→Xsγ, B→K*γ, and Bs→φγ constraints, and provide best-fit Wilson coefficients, correlations, and ancillary model-index files.","tokens_in":20914,"tokens_out":3928,"duration_ms":41446,"significance":"If the central claim were properly benchmarked against the Standard Model, this would be a useful systematic study: the exhaustive coverage of 511 operator combinations, the simultaneous use of AICc and cross-validation, the separate treatment of Moments versus Likelihood angular data, and the radiative-decay consistency checks are genuine strengths. The ancillary 'models.json' file and the explicit AICc formula aid reproducibility. However, the headline conclusion is currently established only within the set of non-empty NP scenarios, not against the zero-parameter Standard Model, and one textual claim about 'all datasets' is contradicted by the paper's own table. The significance of the paper therefore depends on revisions that add the SM null comparison and qualify the dataset-specific claims.","major_comments":[{"comment":"The candidate set consists of the 511 non-empty combinations of the nine NP Wilson coefficients, explicitly excluding the Standard Model (all Wilson coefficients zero) as a candidate. This omission is load-bearing because AICc is a relative criterion: Eq. (12) with K=0 gives AICc(SM)=χ²_SM, whereas for the New Moments dataset the selected one-operator model 2 has χ²_min=252.44, n=258, K=1, so AICc=252.44+2+4/256≈254.45. If χ²_SM is below 254.45, the SM would be preferred over every NP model by AICc, and the statement that O9 is 'the only surviving one-operator scenario' would not imply that the data prefer new physics over no new physics. The paper never reports χ²_SM for any of the four datasets. I request that the SM be included as a K=0 candidate in both the AICc comparison and leave-one-out cross-validation, and that the abstract and summary be reworded so that 'survives' is understood as 'survives among the 511 NP scenarios' unless the SM is explicitly disfavored.","section":"Sec. IV A; Eq. (12); Abstract"},{"comment":"The text states 'Model 2 is clearly the better option for all the datasets' and that the single-operator scenario with ΔC9 is selected by the New data set. This is contradicted by Table III, where the selected models for the New Likelihood data are 132, 133, 130, 46, 47, 10, 257, 258, 131, and 265, and model 2 (the single-operator ΔC9 model) is absent. For the Likelihood data the selected models contain two or more operators, and the text itself later says that for Likelihood data 'an explanation of the observed data with a single operator is less plausible.' The abstract's unqualified statement that O9 is 'the only surviving one-operator scenario' is therefore valid, at best, for the New Moments dataset. Please qualify the claim by dataset and correct the 'all datasets' sentence.","section":"Sec. V A; Table III"},{"comment":"The post-processing step drops all scenarios that fail the Cramér–von Mises normality check on the pull distribution. The number of models dropped, the threshold used for the normality criterion, and the resulting model set are not reported. Since this filter is applied before the model selection and influences which models enter the AICc/cross-validation comparison, it can affect the central conclusion (e.g., the claim that only O9 survives as a one-operator scenario). Please report the number of dropped models for each dataset, the exact criterion, and, ideally, a robustness check showing that the selected models do not change under reasonable variations of the normality threshold.","section":"Sec. IV A(c)"}],"minor_comments":[{"comment":"There are two incomplete citations rendered as '[? ]' (one in the Introduction and one in Sec. IV B 2 about the relation between AIC and cross-validation). These need to be completed.","section":"Sec. I; Sec. IV B 2"},{"comment":"The caption of Fig. 4 says 'Same as fig. 4, but for the fit with Old Data' but should refer to Fig. 3. This makes the comparison of the two figures confusing.","section":"Fig. 4 caption"},{"comment":"The caption contains the typo 'avilable' instead of 'available'.","section":"Fig. 5 caption"},{"comment":"The description of the Differential Evolution optimizer is placed in a footnote that reads as an unfinished sentence or a fragment left from a footnote marker. Please integrate it into a complete sentence in the text or in the footnote itself.","section":"Sec. IV A, Footnote 4"},{"comment":"The sentence 'Evidently, AIC accounts for uncertainty in the data (-2Log(L)) and assumes that more parameters lead to a higher risk of overfitting (2k)' is imprecise: the 2K term is a penalty for parametric complexity, not an assumption about overfitting risk. Consider rephrasing to avoid conflating the penalty with a prior over models.","section":"Sec. V A"}],"recommendation":"major_revision","confidential_remarks":"The paper is a careful data-fitting study, and the central statistical machinery is sound. The main problem is that the headline conclusion is not benchmarked against the zero-parameter Standard Model, and the abstract overstates the dataset scope of the 'only surviving one-operator scenario' claim. Both issues are fixable within the manuscript's scope: add the SM as a candidate in AICc and LOOCV, report ΔAICc relative to the SM for each dataset, and qualify the 'all datasets' statement. I see no citation or novelty concerns; the self-citations are prior applications of the same AICc methodology and are not circular."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know before you read it. First, this is a genuinely useful model-selection exercise: 511 NP Wilson-coefficient combinations fitted to b→sℓℓ data, with leave-one-out cross-validation and AICc applied simultaneously, and the LHCb 'Method of Moments' and unbinned-likelihood angular datasets treated separately. Second, the central caveat is exactly where the stress-test lands: the Standard Model is not among the 511 candidates, so the claim that O9 is the only surviving one-operator scenario is only a statement about the best NP model, not about whether NP is preferred over no NP at all.\n\nWhat is new is the exhaustive scan and the joint selection criterion. Tables II and III give best-fit ΔC9 values around -1.1 to -1.4 across selected models, with correlations, and the ancillary models.json listing all combinations is a nice reproducibility gesture. The conclusion that all surviving models contain O9 is consistent with the global-fit literature, which is reassuring even though the exercise itself is new.\n\nThe soft spots are real but not fatal. The missing SM baseline is the main one. AICc needs a common candidate set; adding the SM as a zero-parameter model is the natural null. The paper never reports χ²_SM, and the abstract's 'best explain the available data' overstates the scope of the selection. This is a moderate limitation, not a fatal one, because the field already has strong independent hints of NP, but the claim should be explicitly conditional.\n\nThere is also an internal inconsistency: the text says 'Model 2 is clearly the better option for all the datasets,' but the tables show model 18 with a lower χ², a higher AICc weight, and a lower MSE for New Moments. If the intended point is that Model 2 is the best one-operator scenario, that is not what the sentence says.\n\nTwo smaller things. The normality-based dropping of scenarios is a post-hoc filter that could bias the candidate set; it is defensible but should be justified more carefully. And the analysis cites an 'OptEx' package 'under development' with no public URL, so the fits are not independently reproducible beyond the ancillary file.\n\nOverall: this is a solid, serious contribution with one load-bearing caveat (no SM baseline) and one confusing sentence about Model 2. It deserves peer review, and I would send it to a competent referee with a request to address the SM baseline and the abstract overstatement. Someone working on b→sℓℓ global fits will want to read it; I'd bring it to a reading group, and I'd cite it for the exhaustive scan even while noting the caveat.","headline":"A useful exhaustive NP model-selection scan whose O9-central result is solid within the chosen set but is never tested against the Standard Model baseline.","tokens_in":21539,"tokens_out":4266,"would_cite":true,"duration_ms":38810,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A model-selection scan over all 511 new-physics operator combinations for b→sℓℓ decays finds that every surviving model contains the left-handed vector-muon operator O9, which is also the only one-operator scenario that survives.","keywords":["b→sℓ+ℓ− decays","new physics","model selection","Akaike information criterion","cross-validation","effective operators","Wilson coefficients","lepton flavor universality"],"falsifier":"Run the same 511-model exercise with the Standard Model added as a 512th candidate (all Wilson coefficients fixed to zero); if AIC$_c$ or cross-validation selects it, the claim that $\\mathcal{O}_9$ is the only surviving one-operator scenario collapses. A second, independent check is to measure $R_{K^*}$ in the low $q^2$ bin $[0.045,1.1]\\,\\mathrm{GeV}^2$, or the angular observable $S_5$ in $B\\to K^*\\mu^+\\mu^-$, at higher precision and see whether the value returns to the Standard Model prediction, which would remove the tension that $\\Delta C_9\\approx -1.1$ fits are absorbing.","tokens_in":20376,"feed_emoji":"⚛️","tokens_out":15114,"duration_ms":130043,"temperature":0.7,"pith_summary":"This paper asks which combination of new-physics operators best explains the measured $b\\to s\\ell^+\\ell^-$ decay rates, angular distributions, and lepton-universality ratios. Instead of testing one or two Wilson coefficients at a time, the authors fit every non-empty combination of their operator set, 511 candidate models, and rank the survivors with the small-sample-corrected Akaike information criterion (AIC$_c$) and leave-one-out cross-validation. The central claim is that every model passing both filters contains the left-handed vector-muon operator $\\mathcal{O}_9$, and that $\\mathcal{O}_9$ by itself is the only single-operator scenario that survives; its best-fit coefficient is about $-1.1$ to $-1.4$. The paper also finds that the angular observables drive the selection, and that which multi-operator models survive depends on whether the angular data come from the principal-moment or maximum-likelihood analysis. The payoff is a short, testable shortlist of new-physics scenarios for the current $b\\to s\\ell\\ell$ anomalies.","feed_headline":"Only one single-operator model survives the full b→sℓℓ scan: O9","feed_subtitle":"If true, the b→sℓℓ anomaly shrinks to one operator, ΔC9 about -1.1, which future data can confirm or rule out.","key_machinery":"The load-bearing object is the complete candidate set: every non-empty combination of the new-physics Wilson coefficients, 511 scenarios, each fitted to the same data with a $\\chi^2$ statistic and then ranked by two competing criteria. The first is the corrected Akaike information criterion, $\\mathrm{AIC_c} = \\chi^2_{\\min} + 2K + 2K(K+1)/(n-K-1)$, with $K$ the number of new-physics parameters and $n$ the number of observables, converted to Akaike weights $w_i \\propto e^{-\\Delta_i/2}$. The second is leave-one-out cross-validation, summarized by its mean-squared prediction error. The selection rule combines them: keep models with $\\Delta\\mathrm{AIC_c}\\le 4$, plot them in the plane of cross-validation error against Akaike weight, and retain the low-error, high-weight cluster. This two-criterion filter is what turns 511 fits into a small surviving list, and it is why the single-operator $\\mathcal{O}_9$ outcome is a model-selection result rather than a single best-fit point.","core_discovery":"On the paper's own terms, the discovery is that exhaustive model selection over the 511 non-empty subsets of the operator set narrows the new-physics explanations of $b\\to s\\ell^+\\ell^-$ data to a small family with a common member. For the 'New Moments' dataset the survivors are one-, two-, and three-operator models, and the unique one-operator survivor is $\\mathcal{O}_9 = (\\bar{s}\\gamma_\\mu P_L b)(\\bar{\\mu}\\gamma^\\mu\\mu)$, with $\\Delta C_9$ between about $-1.1$ and $-1.4$; every multi-operator survivor also contains $\\Delta C_9$. With the 'Likelihood' angular dataset the same criteria select two- to five-operator models, again all containing $\\Delta C_9$, with a small $C'_7$ component in most. The paper reports that the angular observables play the dominant role in this selection, and that when only $R_{K^{(*)}}$ are considered, the axial-vector operator $\\mathcal{O}_{10}$ is the sole one-operator explanation, with a couple of two-operator combinations also capable of explaining the ratios while respecting $\\mathrm{Br}(B_s\\to\\mu^+\\mu^-)$.","pith_inferences":["If the Standard Model were admitted as a candidate model alongside the 511 new-physics scenarios, the information-theoretic ranking would test whether the data require new physics at all; the paper does not perform this comparison, so its 'only surviving one-operator scenario' is a claim about the best new-physics explanation, not about the existence of new physics.","The strong dataset dependence of the survivor list, single-operator $\\mathcal{O}_9$ with Moments data versus three-to-five-operator models with Likelihood data, suggests that the conclusion should be re-checked as the experimental analysis of angular observables evolves; future unbinned likelihood fits with more statistics could resolve which list is physical.","The same exhaustive-scan-plus-two-criteria procedure could be applied to the current $b\\to s\\ell\\ell$ data with a different operator basis, including tensor or four-quark operators, and the natural expectation is that the surviving set would remain anchored on $\\mathcal{O}_9$ if the data pull is genuine.","Because $\\Delta C_9\\approx -1.1$ is a coherent shift, any ultraviolet completion proposed to explain the anomalies, for instance a leptoquark or a heavy $Z'$, must produce the same left-handed vector coupling to muons to reproduce these model-selection results."],"forward_implications":["Every surviving model in the four data-set variants contains $\\Delta C_9$, so a positive experimental confirmation of a universal $\\Delta C_9$ shift near $-1.1$ would single out the $\\mathcal{O}_9$ operator as the carrier of the $b\\to s\\ell\\ell$ anomaly.","The $\\Delta C_9\\approx -1.1$ to $-1.4$ solution predicts specific distortions of the $B\\to K^*\\mu^+\\mu^-$ angular observables $A_{FB}$ and $S_5$ relative to the Standard Model, which the next round of high-statistics $B$-decay data can test directly.","Models containing $C'_7$ alongside $\\Delta C_9$ remain consistent with the measured radiative decays $B\\to X_s\\gamma$, $B\\to K^*\\gamma$, and $B_s\\to\\phi\\gamma$ within $1\\sigma$, so those decays cannot currently separate them, but more precise radiative measurements would.","Dropping $R_{K^{(*)}}$ from the fit shrinks the survivor list toward fewer operators, which implies the updated lepton-universality ratios are the input pushing the selection toward multi-operator scenarios.","The $\\Delta C_9=-\\Delta C_{10}$ alignment common in leptoquark scenarios fails the $\\Delta\\mathrm{AIC_c}\\le 4$ filter, so this particular class of new-physics models is not supported by the selection."],"supporting_citations":[{"why":"Defines the dimension-six operator basis and the angular-coefficient formalism on which every model scenario is built.","marker":"[23]"},{"why":"Supplies the measured $R_{K^*}$ values in the low and central $q^2$ bins that anchor the lepton-universality constraints.","marker":"[10]"},{"why":"Supplies the updated $R_K$ measurement that defines the 'New' dataset.","marker":"[11]"},{"why":"Supplies the 2019 Belle $R_{K^*}$ measurements added to the 'New' dataset.","marker":"[12]"},{"why":"Supplies the 2019 Belle $R_K$ measurement added to the 'New' dataset.","marker":"[13]"},{"why":"Supplies the binned $B^0\\to K^{*0}\\mu^+\\mu^-$ angular observables from the Method of Moments; the choice between this and the maximum-likelihood dataset changes the surviving model set.","marker":"[1]"},{"why":"Supplies the $B\\to K^*$ and $B_s\\to\\phi$ form factors used to compute the theoretical predictions.","marker":"[6]"},{"why":"Provides the small-sample-corrected AIC$_c$ statistic used in the information-theoretic ranking.","marker":"[79]"},{"why":"Establishes the asymptotic equivalence between AIC and cross-validation, the premise for pitting the two criteria.","marker":"[80]"}],"fun_headline_variants":["b→sℓℓ scan: ΔC9 is the sole one-operator survivor","Exhaustive b→sℓℓ selection leaves only vector muon coupling","ΔC9 alone survives the full b→sℓℓ model selection","Angular data crown O9 as the only one-operator NP","Model selection says b→sℓℓ needs one operator: ΔC9"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that new physics must be present: the 511 candidate models are all non-empty new-physics combinations, so the Standard Model (all Wilson coefficients zero) is never entered in the model-selection contest, and the procedure cannot detect that the data might prefer no new physics at all.","fun_headline_variants_meta":{"raw":{"variants":["b→sℓℓ scan: ΔC9 is the sole one-operator survivor","Exhaustive b→sℓℓ selection leaves only vector muon coupling","ΔC9 alone survives the full b→sℓℓ model selection","Angular data crown O9 as the only one-operator NP","Model selection says b→sℓℓ needs one operator: ΔC9"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00066,"raw_usage":{"total_tokens":3084,"prompt_tokens":1079,"completion_tokens":2005,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":695,"completion_tokens_details":{"reasoning_tokens":1905}},"tokens_in":695,"tokens_out":2005,"duration_ms":16182,"temperature":1.0,"reasoning_tokens":1905,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T13:32:42.668869+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same 511-model exercise with the Standard Model added as a 512th candidate (all Wilson coefficients fixed to zero); if AIC$_c$ or cross-validation selects it, the claim that $\\mathcal{O}_9$ is the only surviving one-operator scenario collapses. A second, independent check is to measure $R_{K^*}$ in the low $q^2$ bin $[0.045,1.1]\\,\\mathrm{GeV}^2$, or the angular observable $S_5$ in $B\\to K^*\\mu^+\\mu^-$, at higher precision and see whether the value returns to the Standard Model prediction, which would remove the tension that $\\Delta C_9\\approx -1.1$ fits are absorbing.","supporting_citations":[],"review_version":1}