{"id":"5b0f6d53-6de3-46cf-a580-d8ff5e2236a1","arxiv_id":"2607.23482","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Under LOBO on NASA cells, combined CC+CV charging indicators reach R²=0.874 and beat CC-only or CV-only, while 5-fold CV understates RMSE by about 119% on average.","lead":"Combining constant-current and constant-voltage charging features estimates battery health better than either alone under leave-one-battery-out tests on NASA cells. The work also shows ordinary cross-validation can roughly double reported error versus realistic cross-battery checks, with simple selection rules for maintenance systems.","discovery_kind":"extension","skeptic_critique":{"model":"moonshotai/kimi-k3","headline":"The headline complementarity ranking (Combined R²=0.874 vs t_CC-only 0.845) rests on pooled LOBO metrics over only 4 folds; no per-fold indicator-set breakdown is shown, so the central claim could be carried by one battery plus between-battery variance inflation.","rationale":"I agree with the reader's CONDITIONAL verdict and with the flagged generalization concern (4 LiCoO2 cells, one protocol, 24 °C), but I think the more immediately load-bearing issue is internal rather than external: the central complementarity claim is a pooled-metric comparison over 4 folds with no per-fold evidence that it is stable. This is related to the reader's \"four LOBO folds are few\" point but sharper — it questions not whether the result transfers to other chemistries, but whether the headline ranking is even established on this dataset. Importantly, the concern is cheap to settle with numbers the authors already have; it does not require new data. The paper otherwise shows good practice: sensitivity analysis on the detection threshold, nominal-capacity robustness check, fallback-rate reporting for the τ fit, per-battery error tables, and honest chronological-split limitations (negative R²). The 119% 5-fold-vs-LOBO gap claim is directionally unsurprising (random splits leak adjacent-cycle information from the same cell) and is a fair warning, though it is partly a straw-man baseline since grouped CV is the standard fix. Verdict stays CONDITIONAL: accept-shaped for the NASA-scoped comparison, but the strongest claim should be gated on the per-fold indicator-set breakdown, and Table 11's \"combined is best\" guidance should be softened until that check is shown.","tokens_in":15277,"tokens_out":2382,"duration_ms":99373,"concrete_test":"Recompute Table 7 per LOBO fold: for each held-out battery (B0005/6/7/18), train LightGBM on the other three and report RMSE and R² separately for CV-only, t_CC-only, and Combined. Check (a) whether Combined beats t_CC-only in at least 3 of 4 folds on within-fold R², and (b) the median per-fold ΔR². If the Combined advantage is concentrated in one fold or median ΔR² ≲ 0.01, the complementarity headline should be downgraded to \"pooled-effect, fold-inconsistent.\" As a 5-minute hygiene check, refit the z-score scaler inside each fold and confirm Table 7 numbers move <0.005 R².","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's strongest claim is the ordering in Table 7: Combined (CV + t_CC) R²=0.874 > t_CC-only 0.845 > CV-only 0.796, which grounds the \"complementary degradation information\" conclusion and the Table 11 guideline \"use combined when complete data available.\" Two weaknesses sit exactly here.\n\n(1) Table 7 reports only the pooled overall metrics. LOBO has just 4 folds, and the paper's own Table 5 shows enormous per-fold heterogeneity (per-battery R² from 0.609 to 0.896; RMSE 3.30 to 6.37). With N=4, a ΔR² of 0.029 between Combined and t_CC-only could easily be driven by a single fold (B0006 or B0018, the high-variance cells). No per-fold indicator-set results, no paired comparison, no uncertainty statement on the 0.874-vs-0.845 gap is given. The complementarity claim is therefore statistically under-supported even though the descriptive numbers are internally consistent.\n\n(2) The \"Overall R²\" is computed by pooling predictions across folds, where each fold's predictions come from a different model. Pooled R² credits the model with between-battery SOH variance, which is easier than within-battery tracking. The paper's own numbers hint at this: Table 5's mean per-fold R² is 0.769 but pooled is 0.796. The same inflation applies to the headline 0.874. For a maintenance-oriented claim about tracking degradation of an unseen battery, within-battery (per-fold) R² is the decision-relevant quantity.\n\nA secondary, smaller note: z-score standardization is described as applied \"before model training\" without stating the scaler is fit per LOBO fold; if fit on all 623 cycles, there is mild target-adjacent leakage (though for tree models with monotonic features the practical effect is small).","agreement_with_reader":"partial"},"referee_report":{"model":"moonshotai/kimi-k3","summary":"The manuscript compares health indicators extracted from the constant-current (CC) and constant-voltage (CV) phases of CC–CV charging for battery state-of-health (SOH) estimation. Using four NASA LiCoO2 18650 cells (B0005/6/7/18) and LightGBM (plus XGBoost, CatBoost, Random Forest baselines), the authors evaluate four CV-phase indicators (CV duration, CV/CC time ratio, current-decay time constant, CV charge throughput) and CC duration, individually and combined, under Leave-One-Battery-Out (LOBO) validation. Headline results: the combined set achieves pooled R² = 0.874 (RMSE 3.68%) versus 0.845 for t_CC alone and 0.796 for CV-only (Table 7), supporting a complementarity claim; and LOBO RMSE averages ~119% higher than 5-fold CV (Table 9), quantifying how random splits overestimate deployable accuracy. SHAP analysis ranks the CV/CC time ratio as dominant, and practical indicator-selection guidelines (Table 11) are offered.","tokens_in":15680,"tokens_out":3042,"duration_ms":90768,"significance":"If the results hold, the paper makes a useful, practice-oriented contribution: it quantifies, on public and widely used NASA data, the gap between conventional random-split cross-validation and cross-battery evaluation — a message the SOH literature needs, since many published sub-1% RMSE claims rest on random splits. The work is careful in several respects that deserve explicit credit: a voltage-threshold sensitivity analysis (Appendix C), a nominal-capacity sensitivity check (§3.1), a hyperparameter robustness check (§3.4), per-battery error statistics (Appendix B), a model-parity analysis showing the bottleneck is generalization rather than model capacity (Table 6), and honest limitations including a negative chronological-split result (§5.5). The 98.9% CV-detection success rate and the simple, differentiation-free indicator extraction are genuine practical strengths. The complementarity claim itself, however, currently rests on pooled metrics over only four folds and needs stronger statistical support before the Table 11 guideline built on it can be considered established.","major_comments":[{"comment":"The central complementarity claim — Combined R²=0.874 > t_CC-only 0.845 > CV-only 0.796 — rests on pooled LOBO metrics over only four folds, with no per-fold indicator-set breakdown. The paper's own Table 5 shows large per-fold heterogeneity (per-battery R² from 0.609 to 0.896). With N=4 folds, a ΔR² of 0.029 between Combined and t_CC-only could be carried by a single fold (B0006 or B0018). The authors should report per-fold metrics for each indicator set in Table 7, plus a paired comparison across folds (e.g., per-fold ΔRMSE with sign consistency), or temper the complementarity conclusion and the 'Complete CC-CV data available → Combined' recommendation in Table 11 accordingly.","section":"§4.4, Table 7"},{"comment":"The 'Overall R²' values appear to be computed by pooling predictions across folds, where each fold's predictions come from a differently-trained model. Pooled R² credits the model with between-battery SOH variance, which is an easier task than within-battery degradation tracking — the decision-relevant quantity for the maintenance use case the paper targets. The manuscript's own numbers reveal the discrepancy: Table 5 reports per-fold R² = 0.769 ± 0.110 while the text and Table 7 quote 0.796 for the same CV-only configuration. The authors should (i) state explicitly how 'Overall R²' is computed, (ii) report both pooled and mean per-fold (within-battery) R² for all indicator sets in Table 7, and (iii) verify that the Combined > t_CC-only ordering survives under the within-battery metric.","section":"§4.3.2/§4.4, Tables 5 and 7"},{"comment":"Numerical inconsistency: the threshold sensitivity analysis (§3.2, Table C1) reports the baseline 4.17 V configuration at LOBO RMSE = 4.235% and R² = 0.808, but Table 5 reports the same CV-only LOBO configuration at RMSE = 4.69% and R² = 0.796. These should be identical experiments. Please explain the discrepancy (different model configuration? different random seed? pooled vs. mean-per-fold computation?) and reconcile the two tables, since Appendix C is the basis for the robustness claim about threshold choice.","section":"§3.2 and Appendix C vs. §4.3.2, Table 5"},{"comment":"Feature standardization is described as 'z-score normalization before model training' without stating whether the scaler is fit on training folds only. If statistics are computed over all 623 samples before the LOBO split, test-battery information leaks into every fold. The effect is likely small for standardization, but given the paper's central message is evaluation rigor, the pipeline (fit scaler within each training fold, apply to test fold) should be stated explicitly and, if necessary, corrected.","section":"§3.4"}],"minor_comments":[{"comment":"Direct contradiction on SOH>100% cycles: the text states these are 'concentrated in B0006 (18 cycles, max 104.6%) and B0007 (20 cycles, max 101.7%)', but Table 1 lists B0006's SOH range as 69.8–99.6 and assigns the 104.6% maximum to B0018. Please correct whichever is wrong.","section":"§3.1 vs. Table 1"},{"comment":"Cycle-count inconsistency: the text states exponential fitting converged for '636 of 637 CV cycles', but §4.1 reports 623 valid cycles out of 630 charge cycles. The 637/630 discrepancy (and per-battery counts summing to 637) should be reconciled.","section":"§3.3, Indicator 3"},{"comment":"t_CV/t_CC and τ are both reported with r = -0.719 to three decimals. If this is a coincidence of rounding, fine, but given that τ is derived from the same current-decay profile, a note on the near-collinearity of the CV indicators (and its implications for the CV-only set) would strengthen the analysis.","section":"Table 2"},{"comment":"Two different SHAP summaries are given (mean |SHAP| = 7.839 for the ratio in Table 8; 6.77 ± 1.56 'under LOBO validation' in the text) without stating which model/validation produces the Table 8 values. Please clarify; presumably Table 8 is from a model trained on all data or under 5-fold CV.","section":"§4.5, Table 8"},{"comment":"The hyperparameter grid search is evaluated with 5-fold CV and the default configuration retained; this is reasonable, but note that selecting hyperparameters on 5-fold CV while reporting LOBO as primary is methodologically slightly mismatched. A sentence acknowledging this would suffice.","section":"§4.3.1/§3.4"},{"comment":"The limitations section is commendably honest. Consider adding that with only four LOBO folds, the ±1.26% RMSE spread and the 119% CV-to-LOBO gap are themselves noisy estimates; cross-dataset replication (already listed as future work) is the natural remedy.","section":"§5.5"},{"comment":"No code or extracted-indicator data availability statement is given. Given the public NASA dataset and the reproducibility-oriented framing, releasing the extraction and evaluation code would substantially increase the paper's usefulness.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The manuscript appears to carry publication metadata for Eksploatacja i Niezawodność 28(4), 2026 (DOI 10.17531/ein/220211), so this arXiv version may already be in press elsewhere; the editor may wish to confirm the submission's status and venue fit. Scientifically, the work is competent and honest but incremental: four cells, one chemistry, standard gradient boosting. Its main value is the evaluation-rigor message (LOBO vs. random splits). The load-bearing weakness is the statistical support for the headline Table 7 ordering — pooled metrics over four folds — which is fixable with per-fold reporting the authors almost certainly already have. I recommend major revision rather than rejection because the gap is one of analysis presentation, not of scope."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful bits here are the controlled LOBO comparison and the quantified evaluation gap, not a new indicator. On NASA B0005/6/7/18 they extract four known CV features plus t_CC, run LightGBM and three other learners, and show combined CC+CV beats either alone (pooled R² 0.874 vs 0.845 vs 0.796) while LOBO RMSE sits ~119% above 5-fold CV across models. That gap and the multi-model parity check are citable; the maintenance framing (4–5% RMSE is enough near an 80% threshold) is honest.\n\nWhat they do well: public data, 98.9% CV detection success, voltage-threshold sensitivity in Appendix C, SHAP ranking t_CV/t_CC first, and clear admission that early-cycle models fail late-cycle SOH. Labels come from discharge coulomb counting, so circularity is not an issue. Tables 5–9 and the appendices line up internally.\n\nSoft spots, in proportion. The headline complementarity claim rests on pooled LOBO metrics over only four folds with huge per-battery spread (R² 0.609–0.896). No per-fold indicator-set breakdown or uncertainty on the 0.029 R² gap, so one noisy cell could be carrying the ordering—stress-test is right that this is under-supported, not that the numbers are invented. Pooled R² also mixes between-battery variance; mean per-fold R² is lower. Everything is one chemistry, one temperature, one CC–CV protocol; Table 11 guidelines over-reach a bit relative to §5.5. No code. Hyperparameter and scaler details are minor.\n\nThis is for people who pick charging features for BMS/maintenance models and for anyone still reporting only random k-fold on multi-cell battery data. Not a methods breakthrough and not chemistry-agnostic law. I would send it to referees: the design is clean enough and the evaluation warning is worth having on the record, with a request for per-fold indicator tables and tighter scope language. Worth a cite when I need the LOBO gap number or a simple CC+CV baseline on NASA.","headline":"Solid NASA LOBO head-to-head of CC/CV indicators with a clear 5-fold vs LOBO warning; complementarity ranking is real but statistically thin on four folds.","tokens_in":16423,"tokens_out":546,"would_cite":true,"duration_ms":11015,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Combining CC and CV charging indicators beats either alone for battery health, and ordinary cross-validation overstates accuracy by about 119%.","keywords":["battery state-of-health","charging phase indicators","cross-battery validation","lithium-ion battery","machine learning","Leave-One-Battery-Out","CC-CV charging","LightGBM"],"falsifier":"Repeat the same LOBO indicator-set comparison on another public aging set (different chemistry or CC–CV cutoffs); if combined CC+CV no longer beats CC-only and CV-only, or the CV-vs-LOBO gap collapses, the central ranking and the overestimation claim fail.","tokens_in":16085,"feed_emoji":"🔋","tokens_out":1008,"duration_ms":19535,"temperature":0.7,"pith_summary":"This paper asks which simple features from a battery’s constant-current (CC) and constant-voltage (CV) charge phases actually predict State-of-Health when the model must work on a battery it has never seen. On the NASA aging cells, four CV-phase indicators plus CC duration are compared under Leave-One-Battery-Out validation. The combined CC+CV set wins (R² = 0.874), showing the two phases carry complementary aging signals—capacity fade in CC time, resistance and charge-acceptance changes in CV. The same models look far better under ordinary 5-fold cross-validation; LOBO RMSE is about 119% higher on average, so random splits overstate deployable accuracy. The authors turn that into practical selection rules: use both phases when full charge logs exist, fall back to CV-only when CC is incomplete or compute is tight, and treat ~4–5% RMSE as the realistic target for new cells.","feed_headline":"CC+CV charge features beat either alone for battery health","feed_subtitle":"Combined indicators hit R² 0.874 under battery-held-out tests; ordinary CV overstates accuracy by ~119%.","key_machinery":"Leave-One-Battery-Out (LOBO) comparison of indicator sets: CV duration, CV-to-CC time ratio, current-decay time constant τ, CV charge throughput, and CC duration, scored with LightGBM (and checked against other gradient-boosting models) plus SHAP importance, which ranks the CV-to-CC ratio highest.","core_discovery":"Under Leave-One-Battery-Out validation on four NASA LiCoO2 cells, the combined set of four CV-phase indicators plus CC phase duration achieves the best SOH estimates (R² = 0.874, RMSE 3.68%), beating CV-only (R² = 0.796) and CC-duration alone (R² = 0.845). That ranking shows CC and CV phases capture complementary degradation. Separately, LOBO RMSE averages about 119% higher than 5-fold cross-validation across models, so conventional splits substantially overestimate practical cross-battery accuracy.","pith_inferences":["If the 119% CV–LOBO gap is typical, many published sub-1% SOH numbers from mixed-cycle splits are not deployment-ready until re-checked with battery-held-out protocols.","The complementarity claim suggests multi-phase feature design may matter more than swapping among similar tree ensembles, which the paper’s model comparison already shows cluster tightly under LOBO.","A natural next test is whether the same CV-to-CC ratio stays top-ranked after temperature swings or after recalibrating the CV voltage threshold on CALCE/Oxford-style datasets."],"forward_implications":["When full CC–CV logs are available, prefer the five-indicator combined set for highest cross-battery SOH accuracy.","When only CV is logged or CC is adaptive/unstable, CV-only indicators remain usable without numerical differentiation.","Expect roughly 4–5% RMSE on unseen batteries of this type, not the ~2% suggested by random 5-fold CV.","SHAP ranking supports treating the dimensionless CV-to-CC time ratio as a primary, noise-robust health feature in BMS design.","Simple CV-duration thresholds can support field maintenance alarms without a full SOH model."],"fun_headline_variants":["CC+CV indicators top SOH estimates under LOBO battery hold-out","Combined charge-phase features beat CC or CV alone for battery SOH","LOBO shows CC+CV health indicators hit R² 0.874 on NASA cells","Cross-battery tests: CC and CV phases give complementary SOH signal","Conventional CV inflates SOH accuracy 119% versus true LOBO"],"cache_read_input_tokens":128,"weakest_assumption_plain":"That findings from four room-temperature NASA LiCoO2 18650 cells on one fixed CC–CV protocol are representative enough to guide indicator choice on other chemistries, temperatures, and BMS cutoff settings.","fun_headline_variants_meta":{"raw":{"variants":["CC+CV indicators top SOH estimates under LOBO battery hold-out","Combined charge-phase features beat CC or CV alone for battery SOH","LOBO shows CC+CV health indicators hit R² 0.874 on NASA cells","Cross-battery tests: CC and CV phases give complementary SOH signal","Conventional CV inflates SOH accuracy 119% versus true LOBO"]},"model":"grok-4.5","effort":"low","cost_usd":0.002762,"raw_usage":{"total_tokens":1026,"prompt_tokens":794,"num_sources_used":0,"completion_tokens":85,"cost_in_usd_ticks":27624000,"prompt_tokens_details":{"text_tokens":794,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":147,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":794,"tokens_out":85,"duration_ms":5019,"temperature":1.0,"reasoning_tokens":147,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-30T21:08:23.454226+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Repeat the same LOBO indicator-set comparison on another public aging set (different chemistry or CC–CV cutoffs); if combined CC+CV no longer beats CC-only and CV-only, or the CV-vs-LOBO gap collapses, the central ranking and the overestimation claim fail.","supporting_citations":[],"review_version":1}