{"id":"ca113b59-2798-4fd9-a3a7-e06acd04fc0d","arxiv_id":"2607.28115","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.5,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"A stage-aware linear–quadratic–exponential vibration target fits IMS degradation references better than anchored lines in point metrics and is learnable on held-out XJTU-SY bearings, with limited statistical and transfer strength.","lead":"Clock-linear remaining-life labels often disagree with how bearing vibration actually degrades. This paper builds a three-stage vibration-based target, shows it fits external IMS runs better than a fitted line in point estimates, and learns it on held-out XJTU-SY bearings with a compact dual-network model.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection beyond the reader's already-identified HI/stage circularity; the strongest claim is internally bounded and honestly caveated.","rationale":"The paper’s center of gravity matches the reader’s strongest_claim: a bounded, measurement-oriented demonstration that stage-dependent labels can describe a vibration-derived degradation reference better than a strong linear baseline on an external rig, plus learnability under a strict XJTU bearing hold-out, with explicit refusal of universal TTF or transfer claims. Evidence quality is mixed and reported as such (point estimates and BIC favor LQE; moving-block bootstrap does not; n=3). The circularity the reader flags—oriented PCA HI + chronological k-means repair into exactly three stages defining zb, then scoring LQE against that zb—is real but already the paper’s own framing of “descriptive target adequacy,” not an unacknowledged flaw that invalidates the stated comparison. No internal inconsistency, leakage into the IMS shape test, or over-claim beyond the abstract/conclusion was found that would justify moving off CONDITIONAL. A useful next check is functional-family robustness (splines/Gompertz on the same zb), which would tighten or qualify the stage-dependence interpretation without requiring new data. Verdict and confidence stay with the reader.","tokens_in":18314,"tokens_out":615,"duration_ms":14646,"concrete_test":"Re-fit the three official IMS runs replacing the LQE family with monotone cubic splines (or logistic/Gompertz) on the same zb reference and identical stage boundaries u1,u2; if mean RMSE/MAE gains vs. the anchored linear baseline fall below ~5% or ΔBIC ceases to favor the flexible form, the specific stage-aware advantage is family-dependent rather than evidence of stage dependence per se.","verdict_should_be":"UNCHANGED","load_bearing_attack":"No additional load-bearing concern lands that would overturn the strongest claim as stated. The claim is only that, on three documented IMS failed-bearing trajectories, a continuous LQE curve fits the same vibration-HI-derived reference better than the best anchored linear fit (RMSE/MAE gains, conservative ΔBIC), while bootstrap intervals cross zero and n=3 keeps the result descriptive. That comparison is well-specified in §2.6 and Table 6: both competitors approximate the identical zb=1−normalized oriented PCA HI, so the reported gains are real descriptive improvements of functional form against that shared surrogate, not a hidden claim about physical wear or clock-time TTF. The reader's weakest assumption (HI/k-means/three-stage construction defining the reference it then scores) is the genuine soft spot, but the manuscript already treats it as such—separating RQ1 target shape from RQ2 learnability, refusing universal nonlinear law and cross-domain claims, and reporting non-significant blocked bootstrap. Architecture/OWA and XJTU numbers are secondary and do not prop up the IMS shape claim.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The manuscript argues that clock-linear RUL labels can disagree with vibration-observed degradation and therefore separates target construction from prediction. A development-only pipeline builds an oriented PCA health indicator, repairs k-means stages into chronological early–middle–late segments, and fits a continuous linear–quadratic–exponential degradation-state target Rb(u) to the complementary normalized HI reference zb. Causal CNN–LSTM and Transformer branches learn this target from feature sequences, with validation-fitted two-expert OWA fusion. On a bearing-wise XJTU-SY hold-out (Bearings 1–4 development, Bearing 5 test per condition, excluded from all fitted steps), fused RMSE/MAE/R² are 0.0608/0.0392/0.9617, dominated by the Transformer. Independently, on three documented IMS failed-bearing runs, the stage-aware curve reduces RMSE by 3.6–18.2% (mean 10.2%) and MAE by 3.1–31.1% (mean 15.0%) versus the best anchored linear fit to the same zb, with conservative ΔBIC 128.8–368.1 favoring LQE, while moving-block bootstrap intervals cross zero. The authors explicitly limit claims to descriptive target adequacy and learnability, not a universal nonlinear TTF law or robust cross-domain prediction.","tokens_in":18612,"tokens_out":1484,"duration_ms":39157,"significance":"If the bounded claims hold, the paper’s main value is methodological rather than architectural: it treats the RUL label as a testable measurement choice, compares stage-aware and fitted-linear representations against a shared vibration-derived reference on a second rig, and keeps target validity, held-out learnability, and transfer as separate research questions with matching claim boundaries (Table 3). Strengths include leakage-audited development-only preprocessing, Bearing-5 exclusion from normalization/screening/training/early-stopping/OWA, a stricter linear baseline than 1−u (Eqs. 14–15), conservative five-parameter BIC, dependence-aware bootstrap with open non-significance, and transparent negative exploratory IMS leave-one-run-out results. That honesty and protocol design are useful for Measurement/prognostics practice even though n=3 IMS runs and non-significant blocked bootstrap keep RQ1 descriptive. The dual-scale OWA predictor is secondary; the contribution is a target-validity framework, not SOTA architecture dominance.","major_comments":[{"comment":"§2.4–2.6 and Table 6: RQ1’s reference zb is built from the same pipeline’s oriented PCA-HI, smoothing, three-stage k-means with chronological repair, and end-of-run min–max normalization (Eqs. 7–10). Relative gains of LQE vs Rlin(a*) on that shared zb are well-specified and not tautological, but they still test functional form against a surrogate the authors define. Stagewise results already show non-uniform superiority (e.g., IMS2-B1 middle-stage RMSE worse by 14.5%). A load-bearing sensitivity check is needed: alternative HIs (e.g., RMS/kurtosis-only, different q, unsmoothed, or monotone composite), two- vs four-stage repairs, and at least one other parametric family (logistic/Gompertz/monotone spline) on the same three official runs, reporting whether mean RMSE/MAE gains and sign of ΔBIC persist. Without this, “stage-dependent targets better describe the evaluated vibration-derived de","section":"§2.4–2.6, Table 6, Fig. 8"},{"comment":"§5 (Limitations) and Data Availability: the arXiv package omits final selected feature list, window/step, smoothing span, dropout/weight-decay, sequence length, seeds, and numerical OWA weights. For a measurement-oriented target-validity paper whose XJTU numbers and IMS fits are central evidence, these are not optional extras—they are part of the fitted model (§2.3). Either release a persistent-identifier artifact bundle with frozen preprocessing objects, stage boundaries, and exact configs, or demote quantitative XJTU/IMS point estimates to illustrative status until reproduction objects are available. This is fixable but currently load-bearing for verifiability.","section":"§5 Limitations; Data Availability"}],"minor_comments":[{"comment":"Table 9 vs text: OWA MAE equals Transformer MAE at four decimals (0.0392) while MSE/R² improve only slightly; state explicitly that fusion gain is negligible on MAE and avoid language that could be read as strong ensemble benefit.","section":"§3.3, Table 9–10"},{"comment":"§2.7: random 80/20 window-level validation within development bearings is correctly labeled optimistic; consider adding one blocked/group validation metric in the supplement so early-stopping behavior is auditable.","section":"§2.7"},{"comment":"Figs. 10–12: residual envelopes are descriptive only—add a one-line caption reminder that they are not calibrated prediction intervals (already noted in §5).","section":"Figs. 10–12"},{"comment":"Eq. (11): clarify whether u1, u2 are held fixed from clustering during least-squares for α,β,γ or jointly optimized; the conservative BIC counts five parameters, but the fitting paragraph should match that accounting exactly.","section":"§2.5–2.6, Eq. (11)"},{"comment":"Table 5 / exploratory extended run: keep the strong 56.4% RMSE figure strictly out of abstract/confirmatory aggregates (already done); ensure Fig. 9 caption remains the sole home for that result.","section":"Table 5, Fig. 9"},{"comment":"Minor prose: standardize “OWA” vs “OW A” spacing; fix “V alidation” line-break artifacts in §2.7–2.8.","section":"§2.7–2.8"}],"recommendation":"major_revision","confidential_remarks":"Fit for a measurement/prognostics journal is good if the authors deliver HI-sensitivity and a reproducibility bundle; the architecture/OWA material is thin relative to the target-validity framing and should not be sold as the novelty. No integrity red flags; caveats are unusually candid. I would not reject for “outside consensus”—the claim is carefully non-universal. Major revision is for verifiability and reference-construction sensitivity, not for overturning the stated descriptive comparison."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful part of this paper is not the CNN–LSTM/Transformer stack. It is the insistence that clock-linear RUL labels are a modeling choice, and the clean split of target shape (IMS) from learnability (XJTU Bearing-5 hold-out) from transfer (which they show fails).\n\nWhat is actually new is the protocol more than the pieces. Multistage degradation, PCA HIs, and dual temporal nets are all cited precedent. The stricter move is fitting both the linear–quadratic–exponential curve and a best anchored linear baseline to the same vibration-HI reference, using development-only screening, excluding every “5” bearing from fitted steps, conservative five-parameter BIC, and a moving-block bootstrap they report as crossing zero. That honesty is real. On three official IMS failures they get mean ~10% RMSE / ~15% MAE reductions and large ΔBIC; XJTU fused RMSE 0.0608 / R² 0.9617 is mostly the Transformer, with OWA almost cosmetic. They refuse universal TTF and cross-domain claims in the abstract and discussion.\n\nSoft spots match the reader’s, not worse. The reference zb is their own oriented PCA HI plus k-means/chronological repair and end-of-run min–max, so RQ1 is “better functional form for this surrogate,” not physical wear. n=3 keeps it descriptive; bootstrap non-significance is already on the page. Exact config/code is missing from the arXiv package. Architecture novelty is thin.\n\nFor people who build or audit bearing RUL benchmarks, this is worth a careful read. Math is elementary least squares and standard nets; citations look fair; data handling is above average for the subfield. I would send it to referees. I might cite the target-validity framing; I would not treat the IMS numbers as settled physics.","headline":"Honest target-vs-predictor separation with a real but descriptive IMS shape gain; useful hygiene, not a new RUL law.","tokens_in":19328,"tokens_out":458,"would_cite":true,"duration_ms":10726,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Clock-linear RUL labels disagree with vibration; a three-stage degradation-state target fits the measured reference better and can be learned on held-out bearings.","keywords":["bearing prognostics","remaining useful life","health indicator","target design","vibration measurement","Transformer","data fusion","stage-aware degradation"],"falsifier":"On additional complete failed-bearing runs from a different rig, check whether the same stage-aware curve still cuts RMSE and MAE versus the best anchored linear fit to that run’s vibration health reference, and whether blocked-bootstrap intervals for the squared-error gain stay above zero.","tokens_in":19131,"feed_emoji":"⚙️","tokens_out":845,"duration_ms":17874,"temperature":0.7,"pith_summary":"Bearing remaining-life work usually treats the training label as fixed, most often a straight line that falls with clock time. Vibration often stays nearly flat for a long stretch and then worsens quickly near failure, so that label can force a model to invent degradation the sensor has not yet shown. This paper splits the problem: first build a vibration-based health reference, mark early–middle–late stages, and fit a continuous linear–quadratic–exponential degradation-state target; then train compact sequence models to predict that target from past features only. On three official IMS failed-bearing runs the stage-aware curve beats the best anchored linear fit to the same reference by roughly 10% RMSE and 15% MAE on average. On a strict XJTU-SY hold-out the fused predictor reaches RMSE 0.0608 and R² 0.9617. The authors present this as a target-validity framework, not a universal physical law or a cure for cross-domain shift.","feed_headline":"Linear RUL labels miss how bearings actually vibrate toward failure","feed_subtitle":"A three-stage target fits vibration degradation better and stays learnable on held-out bearings","key_machinery":"The stage-aware degradation-state target: an oriented PCA health indicator is clustered into chronological early–middle–late stages, then fitted with a continuous linear–quadratic–exponential curve to the complementary normalized indicator, and compared against a best anchored linear fit to the same reference.","core_discovery":"For the evaluated vibration-derived degradation references, a continuous three-stage linear–quadratic–exponential target describes remaining-life state better than the strongest comparable global linear target, and that stage-aware target is learnable from causal feature sequences on bearings excluded from every fitted development step.","pith_inferences":["If target audits become standard, many published RUL gains may shrink once models are scored against condition-consistent labels rather than 1−u alone.","Online stage detection without full-run hindsight is the practical bottleneck between this retrospective supervision recipe and live maintenance use.","Comparing LQE against monotone splines or change-point families on the same HI reference would show whether three fixed pieces are special or just one workable nonlinear family."],"forward_implications":["RUL benchmarks should report target-shape checks against a vibration-derived reference, not only predictor error under a fixed clock-linear label.","Strong sequence models cannot fix a mismatched label; target design must be audited separately from architecture.","A better condition-consistent target still leaves cross-run and cross-fault shift unresolved, so transfer methods remain necessary.","Deployed systems should treat the output as a normalized degradation-state score that later needs machine-specific time calibration and uncertainty checks."],"fun_headline_variants":["Linear RUL labels miss staged vibration degradation in bearings","Stage-aware target beats linear fit on vibration-derived RUL states","Three-stage curve cuts RUL target error 10% vs best linear anchor","Held-out bearings learn stage-aware RUL from causal vibration features","Linear labels disagree with how bearing vibration actually degrades"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The whole case rests on treating a smoothed, oriented first principal component of screened vibration features, forced into three contiguous stages, as a faithful enough picture of real degradation that fitting curves to it tests whether labels are adequate.","fun_headline_variants_meta":{"raw":{"variants":["Linear RUL labels miss staged vibration degradation in bearings","Stage-aware target beats linear fit on vibration-derived RUL states","Three-stage curve cuts RUL target error 10% vs best linear anchor","Held-out bearings learn stage-aware RUL from causal vibration features","Linear labels disagree with how bearing vibration actually degrades"]},"model":"grok-4.5","effort":"low","cost_usd":0.00273,"raw_usage":{"total_tokens":1068,"prompt_tokens":865,"num_sources_used":0,"completion_tokens":73,"cost_in_usd_ticks":27304000,"prompt_tokens_details":{"text_tokens":865,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":130,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":865,"tokens_out":73,"duration_ms":3419,"temperature":1.0,"reasoning_tokens":130,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-31T17:24:02.973002+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"On additional complete failed-bearing runs from a different rig, check whether the same stage-aware curve still cuts RMSE and MAE versus the best anchored linear fit to that run’s vibration health reference, and whether blocked-bootstrap intervals for the squared-error gain stay above zero.","supporting_citations":[],"review_version":1}