REVIEW 3 major objections 5 minor 8 references
Marine Engine Fault Dataset: Open-Access Data under Controlled Reference and Fault Scenario Conditions
T0 review · 3 major / 5 minor · reviewed 2026-08-01 · deepseek-v4-flash
Pith's one-line read An open-access dataset from a real marine diesel engine testbed, with five physically implemented fault scenarios and multi-sensor time series, is presented as a benchmark for predictive-maintenance research.
desk verdict A genuinely useful open marine-engine fault dataset, honestly documented; the degradation-modelling claim is ahead of what the release demonstrably supports. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the dataset package itself: CSV time-series files from a water-brake testbed engine, organized into one reference-performance record and five scenario folders, together with a variable dictionary and dataset index. The experimental design carries the argument—a design-of-experiments reference campaign from 30% to 90% load establishes the healthy baseline, while each fault scenario begins with about an hour of stabilized fault-free operation before the intervention is introduced, so anomaly responses are observed against a controlled baseline. Two acquisition streams (low-frequency scans at about 1–2 s and high-frequency crank-angle-resolved cylinder pressure) are m
What would settle it
Install genuinely degraded parts (for example, a naturally clogged air filter or an eroded injector nozzle) on a similar engine and compare the resulting sensor signatures—filter pressure drop, compressor pressure ratio, cylinder peak pressure, exhaust temperature—against the corresponding scenario records; if the deviations diverge beyond sensor noise, the proxy-intervention assumption would be falsified.
Extended reading notes
Core claim
On the paper's own terms, the central claim is that controlled physical interventions on a real marine engine can generate fault data that are both realistic enough to be physically interpretable and structured enough to support benchmarking. The five anomaly classes were created by air injection into the cooling-water pump, tape restriction of the compressor air filter, a charge-air temperature setpoint change in the air cooler, injector nozzles with one or two welded-shut holes, and partial closure of an exhaust-stack valve to raise turbine back pressure. Multi-sensor records combine low-frequency scans of temperature, pressure, flow, and speed with crank-angle-resolved in-cylinder pressur
Load-bearing premise
The load-bearing premise is that the controlled interventions (air injection for pump cavitation, taped filter for clogging, a cooling-setpoint change for fouling, welded nozzle holes for injector clogging, and an exhaust-stack valve for turbine degradation) faithfully represent the real failure mechanisms they stand in for, so that patterns learned on this dataset transfer to actual engine faults.
Editorial extensions
If this is right
- Fault-detection and diagnosis methods can be benchmarked on real marine-engine data spanning cooling, air-intake, combustion, and turbocharging subsystems, not just on a single component.
- Because severity levels were staged (one versus two plugged injector holes) and loads varied, degradation-modelling studies can examine how anomaly signatures progress with severity and operating point.
- The reference-performance map over the 30–90% load range supports model-based and hybrid prognostics that need a healthy-baseline model of the engine.
- The causal-tree descriptions give researchers a physics-based expectation for each scenario, allowing data-driven findings to be checked against known response pathways.
- Users must treat the Anomaly State field at file level and verify variable correspondence with the reference file before comparing scenarios, which shapes how cross-scenario machine-learning pipelines should be built.
Reading between the lines
- Beyond the paper: the five interventions amount to a transferable recipe for fault-injection experiments on other engines, but each proxy's fidelity to genuine wear mechanisms is a separate question; the paper demonstrates internal consistency, not equivalence to naturally aged components.
- Beyond the paper: a natural next step is a transfer benchmark—train on these scenario files, then evaluate on data from a different engine or from the field—since the paper does not claim cross-engine generalization.
- Beyond the paper: the turbocharger energy-conversion ratio introduced in the validation section could serve as a standalone health indicator for operational monitoring, although the paper uses it only to interpret the turbine-degradation scenario.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This data descriptor introduces an open-access marine diesel engine fault dataset collected on a testbed at the National Maritime Research Institute. The dataset contains reference-performance measurements across 30–90% load and scenario-based fault records for five anomaly classes: cooling-water pump cavitation, compressor air-filter clogging, air-cooler fouling, injection-valve nozzle clogging, and turbine degradation. Faults were implemented by physical interventions, such as tape covering, welded nozzle holes, air injection, and exhaust backpressure changes. The released package includes CSV time-series files, a reference file, a variable dictionary, a dataset index, and a README. Technical validation plots show plausible trends in reference data and scenario-specific response patterns. The paper claims the dataset is a benchmark for anomaly detection, fault diagnosis, degradation modelling, and related condition-monitoring studies.
Significance. If the dataset delivers what is claimed, it is a valuable contribution to maritime PHM research, where publicly available system-level fault data are scarce. The strengths are the controlled testbed environment, multi-load operation, multi-sensor acquisition including crank-angle-resolved cylinder pressure, physical (rather than purely simulated) interventions, and the open release with structured metadata. The manuscript is honest about some limitations, such as the file-level nature of the Anomaly State field and the warning that files are not all column-compatible. However, the central claim of supporting 'degradation modelling' is not backed by the documented data structure, and the technical validation is qualitative rather than quantitative. These issues are fixable, but they require changes to the manuscript or the dataset release.
major comments (3)
- [§2.3, §2.4, §3, §5] The abstract and §1 present the dataset as a benchmark for 'degradation modelling' and RUL prediction, relying on staged fault development. However, §2.3 only states that failure development was 'discretized into fixed stages' without providing stage definitions, severity values, or transition times. §2.4 describes physical interventions but does not report an ordinal severity variable. §3 describes the file structure and the 'Anomaly State' field, and §5 explicitly warns that this field is 'not a globally uniform label' and should be interpreted at file level. No released severity/stage index or ordered progression trajectory is documented. As it stands, the data can support binary anomaly detection and fault diagnosis, but not degradation modelling or RUL benchmarking. Please either add explicit severity/stage metadata to the release and document it, or remove the degradation-modelling
- [§4 (Technical Validation)] The technical validation is entirely qualitative. For example, §4.2 concludes a 'clear inverse relationship' in Fig. 9 and a 'modest separation' in Fig. 11 based on visual inspection, without confidence intervals, effect sizes, or statistical separation tests. The claim that anomalies produce 'progressively distinguishable behaviour' is therefore not quantitatively supported. Since the dataset is intended as a benchmark, readers need objective measures of detectability and separation (e.g., distribution overlap metrics, trend slopes with uncertainties, or signal-to-noise ratios). Please add such quantitative validation or soften the validation-related claims.
- [§2.4 / §4.2] Several fault implementations are proxies for real degradation mechanisms: air injection for pump cavitation, tape for air-filter clogging, a charge-air temperature setpoint change for air-cooler fouling, welded nozzle holes for injector clogging, and an exhaust-stack valve for turbine degradation. The manuscript's abstract calls these 'controlled fault realization', which overstates the fidelity of the emulation. The causal trees in Fig. 7 are explicitly built as expected pathways, not validated against independent data, yet they are used as the interpretive framework. This creates a risk that anomaly labels will not transfer to real engine faults. Please add an explicit limitations paragraph about proxy validity and, where possible, compare observed signatures (e.g., pressure pulsations in Fig. 8) with known physical fingerprints from the literature.
minor comments (5)
- [§2.2] Typo: 'wheras' should be 'whereas'. Also, the text refers to 'Table 6' and 'Tables 4 and 5', but the manuscript only contains Tables 1–2 and Appendix Tables A1–A4; the in-text numbering appears inconsistent.
- [§4.2, Eq. (1)] Equation (1) is confusing: the text says 'Tc,in and Tc,out denote the turbine inlet and outlet temperatures' after already defining them as compressor variables. The turbine variables should be Tt,in and Tt,out. Also, the statement that the ratio is 'normalized with corresponding inlet temperatures at the reference condition' is not clearly reflected in the equation; please show the normalization explicitly.
- [§3 / §5] The usage note that 'Anomaly State' is not globally uniform is important, but it should appear earlier and be accompanied by a precise definition of the field's allowed values. The dataset index should also list the distinct values of this field per file so that users can avoid misinterpreting labels.
- [Fig. 4 and Fig. 12] The caption ordering in Fig. 4 appears as '(a) (c) (b)' while the panels are labeled (a)–(c); please align the caption with the panel order. In Fig. 12, the subcaptions should be placed consistently with the subfigures.
- [§7] The license is given as 'CC4.0'; please specify the exact license (e.g., CC-BY 4.0) and, if applicable, include the license text in the repository.
Circularity Check
No circular derivation: the paper is a data descriptor whose validation checks physical coherence rather than predicting from fitted inputs.
full rationale
The manuscript contains no fitted model, no equation whose output is defined in terms of its input, and no load-bearing self-citation chain. The central claim is that a new experimental dataset exists and is internally coherent. Section 4's technical validation compares measured trends (e.g., compressor outlet temperature vs. pressure ratio, turbine inlet temperature vs. compressor outlet pressure) against expected qualitative physics-based pathways; this is a sanity check of the released records, not a prediction derived from those records by construction. Equation (1) defines an auxiliary turbocharger energy conversion ratio from temperature measurements but is not used to derive the dataset's benchmark claim. The paper also explicitly states limitations (Section 5: 'The Anomaly State field should be interpreted at the file level, not as a globally uniform label'; pressure channels in volts; scenario-specific file structure; Section 1: dataset 'is not intended to replace field operational data'), which are scope caveats rather than circular steps. No self-citation is invoked to justify the dataset's validity, and no prior result by the same authors is used as the basis for the experimental methodology. Consequently, there is no circular step to flag.
Assumptions & free parameters
assumptions (4)
- domain assumption The propeller-law load profile and steady-state criteria (15-min stabilization, 1-hour baseline) produce representative, quasi-steady reference and fault data.
- domain assumption Emulated interventions are valid proxies for real degradation mechanisms.
- domain assumption Sensor accuracy, calibration, and synchronization are adequate for the claimed physical trends.
- ad hoc to paper The fault tree analysis constructed a posteriori reflects causal degradation pathways.
Cite this review
Pith. "Pith review of Marine Engine Fault Dataset: Open-Access Data under Controlled Reference and Fault Scenario Conditions." pith.science (2026). https://pith.science/paper/7UXH4AZV
@misc{pith2026260719444,
author = {Pith},
title = {Pith review of: Marine Engine Fault Dataset: Open-Access Data under Controlled Reference and Fault Scenario Conditions},
year = {2026},
howpublished = {\url{https://pith.science/paper/7UXH4AZV}},
note = {Machine review of arXiv:2607.19444}
}
read the original abstract
Open-access datasets for marine-engine predictive maintenance remain scarce, particularly those from controlled fault experiments with documented operating conditions, subsystem-level interventions and system-level measurements. This work presents the Marine Engine Fault Dataset, an openly available dataset from a turbocharged, intercooled three-cylinder marine diesel engine operated on a testbed under both reference and fault-scenario conditions. The experimental campaign combined a reference-performance program across the 30-90% load range with scenario-based tests in which abnormal conditions were introduced after stabilized fault-free operation, enabling controlled comparison between baseline and fault-affected behaviour. Five anomaly classes were implemented through physical interventions affecting major engine subsystems: cooling-water pump cavitation, compressor air-filter clogging, air-cooler fouling, injection-valve nozzle clogging and turbine degradation induced through increased exhaust-side restriction. The released data comprise multi-sensor time-series of operating, thermal, pressure, flow and combustion-related variables, with a separate reference-performance record and metadata for structured reuse. Technical validation shows that the reference measurements remain physically coherent across the operating range and that the imposed anomalies produce interpretable response patterns consistent with the affected subsystems, including progressively distinguishable behaviour where different severities were implemented. By combining controlled fault realization, multi-load operation and system-level measurements within a real marine-engine platform, the dataset provides a well-documented benchmark for anomaly detection, fault diagnosis, degradation modelling and related condition-monitoring studies in maritime machinery.
Reference graph
Works this paper leans on
-
[1]
Background & Summary Maritime transportation is the backbone of global trade, carrying more than 80% of world trade by volume. However, it is vulnerable to machinery-related failures, which were the leading cause of shipping incidents globally, accounting for more than half of incidents in recent annual records [1]. Although total vessel losses have decre...
2025
-
[3]
Data records The Marine Engine Fault Dataset is publicly available at 10.5281/zenodo.19857425 and contains the reference-performance and fault-scenario measurements acquired from the marine engine test platform described in Section 2. As shown in Fig. 5, the released package is organized into five scenario-specific folders; AC Clogging, AC Fouling, Inject...
-
[4]
Laurinavichyute, A., H. Yadav, and S. Vasishth, Share the code, not just the data: A case study of the reproducibility of articles published in the Journal of Memory and Language under the open data policy. Journal of Memory and Language, 2022. 125: p. 104332. 4. Su, H. and J. Lee, Machine learning approaches for diagnostics and prognostics of industrial ...
arXiv 2022
-
[5]
In addition, the high-speed data acquisition system provided crank-angle-resolved in-cylinder pressure measurements, from which key combustion and performance metrics were calculated and subsequently merged into the dataset at the corresponding low-frequency sampling instants. 2.3.Experimental design The experimental design comprised two main parts: refer...
-
[11]
Test-bed nomenclature, measured variables, and instrumentation Table A1
Appendix A. Test-bed nomenclature, measured variables, and instrumentation Table A1. Abbreviations used in the engine test-bed schematic. Symbol Description P Pressure measurement points T Temperature measurement points F Flow measurement points dP Pressure difference measurement points h Relative position measurement N Engine shaft speed measurement W Wa...
-
[14]
Reinforcement Learning-based Anomaly Detection for PHM applications
Khan, S., et al. Reinforcement Learning-based Anomaly Detection for PHM applications. in 2022 IEEE Aerospace Conference (AERO). 2022. IEEE. 15. Cao, R., et al., Model-constrained deep learning for online fault diagnosis in Li-ion batteries over stochastic conditions. Nature Communications, 2025. 16(1): p. 1651. 16. Zio, E., Data-driven prognostics and hea...
arXiv 2022
-
[27]
DeCastro, and J.S
Frederick, D.K., J.A. DeCastro, and J.S. Litt, User's guide for the commercial modular aero-propulsion system simulation (C-MAPSS). 2007. 28. Goebel, K.F., Management of uncertainty in sensor validation, sensor fusion, and diagnosis of mechanical systems using soft computing techniques. 1996: University of California, Berkeley. 29. Society, P., 2010 PHM S...
2007
-
[40]
Computers & industrial engineering, 2019
Carvalho, T.P., et al., A systematic literature review of machine learning methods applied to predictive maintenance. Computers & industrial engineering, 2019. 137: p. 106024. 41. Khan, S. and T. Yairi, A review on the application of deep learning in system health management. Mechanical systems and signal processing, 2018. 107: p. 241-265. 42. Dataset, I....
2019
Reviewed August 1, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.