REVIEW 4 major objections 7 minor 84 references
Maturity Framework for Enhancing Machine Learning Quality
T0 review · 4 major / 7 minor · reviewed 2026-08-08 · deepseek-v4-flash
Pith's one-line read This paper claims that machine-learning system quality can be turned into a measurable, governable quantity through a 0–100 quality score and a five-level maturity framework, with empirical evidence from a company-wide rollout at…
desk verdict A useful, open-sourced ML quality framework with an honest rollout narrative, but the abstract oversells the empirical validation: Figures 1-3 track compliance with the framework's own Q score, not independent quality gains. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the quality score $Q$ together with the maturity ladder defined by the requirement table. $Q$ compresses the whole assessment into a single number: for each of $N$ sub-characteristics the assessor records a gap $g_s \in \{0,1,2\}$, and $Q = 100\left(1 - \frac{\sum_s g_s}{2N}\right)$; this is the quantity shown improving over time in the paper's figures. The seven characteristics — utility, economy, robustness, modifiability, productionizability, comprehensibility, and responsibility — are the grid on which requirements are placed, and Table 1 spells out minimal and full requirements per sub-characteristic with the maturity levels they unlock. Business criticality (proof of concept, production non-critical, production critical) sets the target level, so the same requirement table produces different obligations for a toy experiment and a revenue-critical system.
What would settle it
Run an external validation: take a set of production ML systems outside Booking.com, have two independent teams score each with the open-source package, and compare the resulting $Q$ scores and maturity levels against independently measured outcomes such as outage frequency, retraining lag, and business metrics over six months. The central claim would collapse if high-scoring systems fail as often as low-scoring ones, or if the two teams disagree on most gap judgments.
Extended reading notes
Core claim
The central claim is that ML quality can be assessed rigorously enough to drive governance: each sub-characteristic has a minimal and a full requirement, the gap between current state and requirement is coded as no gap (0), small gap (1), or large gap (2), and the quality score $Q = 100 \left(1 - \frac{\sum_s g_s}{g_{\text{large}}\,N}\right)$ summarizes the system in one number between 0 and 100. The framework then maps required sub-characteristics to five maturity levels, so a system's maturity is not a global label but a profile of which requirements are met, and business criticality — proof-of-concept, production non-critical, or production critical — decides which maturity level is actually required. The authors assert empirical validation from the Booking.com rollout: quality improved across nearly all sub-characteristics, with the largest gains in accuracy, testability, readability, understandability, and maintainability; monitoring was identified as the biggest gap and addressed by a central observability platform; legacy systems with low scores were shown by A/B benchmarking to be non-inferior to simple baselines and were retired; and efficiency work cut one critical system's training time from 7 hours to under 1 hour and another system's hyperparameter tuning from 20 hours to under 1 hour.
Load-bearing premise
The load-bearing assumption is that the 0–100 quality score, built from yes/no judgments about requirements, really measures how good an ML system is, and that the thresholds tuned at Booking.com apply elsewhere.
Editorial extensions
If this is right
- Because the assessment is packaged as open-source code, an organization can generate the same standardized reports and folder-structured, versioned quality history without building the framework from scratch.
- The split between minimal and full requirements gives teams an incremental path: they can close small gaps first and move up one maturity level at a time rather than facing an all-or-nothing standard.
- The business-criticality tiering means the framework does not impose the same burden on every system; low-stakes experiments can stay lightweight while critical systems must meet the full ladder.
- The reported shift in the quality-score distribution implies that visibility alone — monthly reports and a governance dashboard — can change team priorities, even before new tooling is built.
- If the automated registry-based assessment works as described, quality evaluation can run monthly across hundreds of systems with a human needed only for code-readability and modularity judgments.
Reading between the lines
- The claimed validity rests on the assumption that the score $Q$ measures something real; a natural test the paper does not run is an inter-rater reliability study, where independent evaluators score the same system and their $Q$ values are compared.
- The thresholds in Table 1 are presented as principled but were calibrated to Booking.com's internal distributions; organizations with different scale, domain risk, or regulatory pressure will likely need to recalibrate them, and the paper gives no procedure for doing so.
- The before/after evidence is observational, not experimental: teams chose which gaps to close, so the improvement trend may partly reflect selection rather than the framework's causal effect; a controlled rollout to matched teams would test this.
- For LLM-based systems the paper describes adjustments but does not yet define GenAI-specific sub-characteristics; a concrete extension would add requirements such as hallucination rate, prompt-injection resistance, and evaluation-data freshness, and check whether $Q$ still tracks business outcomes.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper proposes an ML quality assessment and maturity framework developed at Booking.com. Seven quality characteristics (Utility, Economy, Robustness, Modifiability, Productionizability, Comprehensibility, Responsibility) are decomposed into sub-characteristics, each scored against a checklist of minimal and full requirements in Table 1. Gap values of 0/1/2 are aggregated into a quality score Q in Eq. (1), and a maturity framework with five levels and three business-criticality tiers maps expected quality standards per system. The authors describe an open-source Python package with automated and semi-automated gap inference from an ML registry, a two-year company-wide rollout, organizational challenges and lessons learned, and domain-specific adaptations for LLMs, causal ML, and CV/NLP systems. Empirical evidence is presented in Figures 1-3 as before/after quality scores, maturity levels, and compliance percentages, together with business-value anecdotes in Section 6.3. The central claim, stated in the abstract and Section 8, is that the framework is open-sourced, validated with empirical evidence, and produced significant quality improvements and business value at scale.
Significance. If the validation claims were fully supported, this would be a genuinely reusable governance artifact for industrial ML: the granular checklist-based quality model is grounded in ISO/IEC 9126 and ML-specific literature; the system-level maturity definitions with criticality-based expectations are more actionable than organization-level maturity models; and the open-source release with versioned, reproducible report generation is a concrete engineering contribution. The paper also provides candid process lessons and quantified efficiency examples (e.g., training time reduced from 7 hours to under 1 hour in Section 6.3.3), which give the framework practical credibility. The main weakness is that the reported validation is largely self-referential: the outcome measure Q is built from the same requirement gaps that the framework recommends closing, so Figures 1-3 primarily demonstrate compliance with the framework's own rules. The significance of the contribution is therefore conditional on either anchoring Q to independent quality or business outcomes, or substantially tempering the 'validated with empirical evidence' claim.
major comments (4)
- [§3.2 Eq. (1); §7 Figures 1-3; abstract and §8] The validation evidence is self-referential. Q in Eq. (1) is defined as the normalized sum of gap values g_s that are assigned exactly by checking the requirements in Table 1, and the framework's recommendations are designed to close those same gaps. Figures 1 and 3 therefore show that systems improved on the framework's own requirements after the framework was applied; this is expected by construction and does not by itself establish improvement in ML quality or business value. Section 8's claim of 'significant quality improvements across all the aspects, resulting in business value' and the abstract's 'validated with empirical evidence' consequently overstate what the data can support. I recommend correlating Q (or its changes) against independent outcomes, such as model performance on fresh test data, incident counts, latency, or business metrics from A/B tests, or comparing with systems not subject to the framework. The conclusion's own concession that 'further empirical validation' is needed supports this reading.
- [§6.1.4; §7 Figures 1-2] The reported aggregate improvement may be partly attributable to population attrition. Section 6.1.4 states that the rollout led to the cleanup of 'dozens of ML systems,' and Figure 1 is restricted to a 'selected subset' of systems that voluntarily participated in the rollout. If low-scoring systems were decommissioned or excluded between the before and after measurements, the distribution shift in Figure 2 could reflect changes in which systems constitute the population rather than genuine quality gains. Please report a fixed-cohort analysis in which the same systems are measured at both time points, state the number of systems contributing to each distribution in Figure 2, and specify the inclusion criteria for the subset shown in Figure 1.
- [§7; §6.1.5] No statistical support is provided for the claimed improvements. Figures 1-3 contain no confidence intervals, significance tests, or effect sizes; Section 7's phrase 'significant improvement' appears to be used in the informal sense, and Section 8 then elevates it to 'significant quality improvements across all the aspects.' Separately, Section 6.1.5 discloses that several requirements and thresholds were calibrated from quantiles of Booking.com's internal distributions, which makes Q and the compliance percentages in Figure 3 organization-relative rather than absolute quality measures. The claim in Section 4.1 that the maturity levels work 'out of the box for most of the large scale organizations with no safety-critical applications' would be more credible with a sensitivity analysis showing how results change under plausible threshold variations, or with explicit recalibration guidance.
- [§5.1] The reliability of the measurement instrument is unquantified. Section 5.1 states that readability and modularity still require a human in the loop, and in the semi-automated mode practitioners fill in a survey whose answers determine gap values. These subjective judgments enter Q in Eq. (1) with the same weight as automated metrics such as test coverage, yet no inter-rater reliability, evaluation protocol, or rubric calibration is reported. Reporting agreement statistics (e.g., Cohen's kappa) on a sample of systems, or at least documenting the rater guidance provided to teams, would substantially strengthen the construct validity of Q.
minor comments (7)
- [§3.2, Table 1] The symbols used in Table 1 to denote minimal and full requirements do not render in the submitted manuscript (Section 3.2 refers to them but they appear as blank spaces), which makes the requirements table difficult to interpret; please provide a clearly visible legend.
- [§7, Figures 1-3] Sample sizes are not reported: Figure 1 does not state how many systems were in the voluntary rollout or why those particular systems were selected; Figure 2 does not give the number of systems per month; Figure 3 does not report how many systems are in the before and after groups.
- [§7, Figure 2] The violin plot's x-axis spans May through November 2023, while Section 6 describes a two-year rollout; please clarify the relationship between the rollout timeline and the measurement window.
- [Appendix A, Figures 4-5] The phrases 'The ML system is fullon' and 'the model is fullon' appear to be jargon or typos; please clarify the intended wording.
- [§2] Typo: 'They key differences' should be 'The key differences.'
- [§3.2, Eq. (1)] Equation (1) assigns equal weight to every sub-characteristic regardless of business criticality; since the maturity framework elsewhere differentiates by criticality, a brief discussion of how criticality weighting would interact with the per-level requirements would help.
- [§3.1, [77]] The initial attribute shortlist is anchored in part to a Wikipedia list ([77]); a more citable source such as ISO/IEC 25010 for the starting set of quality attributes would strengthen the related-work grounding.
Circularity Check
The claimed empirical validation is largely self-referential: Figures 1–3 measure compliance with the framework's own Q score and thresholds, and the quality model is imported from the authors' prior work, though Section 6.3 provides some independent business grounding.
-
self definitional
[Section 3.2, Eq. (1); Section 7, Figure 1 and Figure 3]
"By assigning the following numerical value to the three gap types: no gap: 0 / small gap: 1 / large gap: 2. we can introduce the quality score Q as: Q = 100(1 − ΣN s gs glarge N) ... The quality score ranging between 0 and 100, can be used to measure ML system quality, to track improvements and to compare different systems. ... Figure 1 shows the quality improvement of several ML systems over a number of cycles of assessments and implementation of the recommended best practices."
Q is a linear transform of the same binary gap judgments g_s that Table 1 defines and that the framework's recommendations are designed to remove. Consequently, an increase in Q is, by construction, an increase in compliance with the framework's own rubric, not an independent measure of ML quality or business value. The abstract's claim that the method is 'validated with empirical evidence' therefore rests on the self-defined score, and Figures 1 and 3 demonstrate rule-following rather than externally anchored quality improvement.
-
fitted input called prediction
[Section 6.1.5 and Section 7, Figure 3]
"We substantiated our requirements and thresholds by taking them from the quantiles of the empirical distribution of the ML systems in the company (e.g. for training cost) or from well-known industry standards (e.g. benchmarking against simple baselines) [34]. ... For each quality sub-characteristic Figure 3 shows the percentage of systems complying with the requirements (without technical gaps) before and after the rollout of the framework."
Several thresholds that determine whether a sub-characteristic has a gap were fitted to the quantiles of the very same systems whose improvement is then reported. The before/after compliance gain in Figure 3 is therefore partly a mechanical consequence of setting the bar relative to the initial population distribution, so the reported 'quality improvement' is not independent of the way the requirements were calibrated.
1 more flagged steps
-
self citation load bearing
[Section 3 opening; Section 5, ML Quality Python package]
"To define quality, we use a refined version of the quality model tailored for ML systems, introduced in [20]. ... The generation and prioritization of the recommendations is based on our framework described in [20]."
The quality model that supplies the sub-characteristics and requirements, and the recommendation priorities that drive the measured improvements, are imported from reference [20], whose authorship overlaps with the present paper. No independent derivation or external validation of that prior framework is given here; the central assessment method therefore leans on a load-bearing self-citation rather than on evidence established in this paper.
full rationale
The paper's central empirical validation is an internal loop: Q (Eq. 1) is defined as a sum over the same binary gap judgments that Table 1 prescribes and that the framework's recommendations are designed to remove; Figures 1–3 then report increases in Q and in the fraction of systems with 'no technical gaps' after teams applied those recommendations. This makes the headline 'validated with empirical evidence' and 'significant quality improvements across all aspects' substantially self-referential: the measured quantity is compliance with the framework's own rubric. The calibration of several thresholds to quantiles of the company's own distribution (Section 6.1.5) means the before/after comparison in Figure 3 is also relative to norms derived from the same population, not to an external standard. Additionally, the quality model and the recommendation priorities are imported from the authors' prior work [20] via load-bearing self-citation, without independent derivation in this paper. There is some non-circular grounding: Section 6.3 gives concrete business incidents (e.g., non-inferiority A/B results showing stale models, efficiency improvements reducing training time from 7 to under 1 hour) that do not reduce to Q. Because of that independent content, the paper is not fully circular; but the principal aggregate evidence is self-referential, so the score is 6.
Assumptions & free parameters
free parameters (5)
- Gap values 0, 1, 2 in Eq. 1 =
0, 1, 2 (glarge = 2)
- Test coverage thresholds =
20% minimal, 80% full
- Pipeline failure resilience thresholds =
Up to 30% failed pipelines per quarter minimal, at most 10% full
- Business criticality cutoffs =
66th percentile of request volume, more than 4 dependent teams, more than 1% yearly revenue, strategic importance
- Expected maturity levels per criticality =
Proof of concept = 1, production non-critical = 3, production critical = 5
assumptions (4)
- domain assumption The seven quality characteristics form a complete decomposition of ML system quality.
- domain assumption Quality gaps can be meaningfully quantized to no/small/large and mapped to the linear score Q in Eq. 1.
- ad hoc to paper The maturity level requirements in Table 1 are appropriate for large organizations without safety-critical applications.
- domain assumption The empirical distribution of internal Booking.com systems is a valid basis for setting quality thresholds.
Cite this review
Pith. "Pith review of Maturity Framework for Enhancing Machine Learning Quality." pith.science (2026). https://pith.science/paper/V7OBQDXS
@misc{pith2026250215758,
author = {Pith},
title = {Pith review of: Maturity Framework for Enhancing Machine Learning Quality},
year = {2026},
howpublished = {\url{https://pith.science/paper/V7OBQDXS}},
note = {Machine review of arXiv:2502.15758}
}
read the original abstract
With the rapid integration of Machine Learning (ML) in business applications and processes, it is crucial to ensure the quality, reliability and reproducibility of such systems. We suggest a methodical approach towards ML system quality assessment and introduce a structured Maturity framework for governance of ML. We emphasize the importance of quality in ML and the need for rigorous assessment, driven by issues in ML governance and gaps in existing frameworks. Our primary contribution is a comprehensive open-sourced quality assessment method, validated with empirical evidence, accompanied by a systematic maturity framework tailored to ML systems. Drawing from applied experience at Booking.com, we discuss challenges and lessons learned during large-scale adoption within organizations. The study presents empirical findings, highlighting quality improvement trends and showcasing business outcomes. The maturity framework for ML systems, aims to become a valuable resource to reshape industry standards and enable a structural approach to improve ML maturity in any organization.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[20]
Georgios Christos Chouliaras, Kornel Kiełczewski, Amit Beka, David Konopnicki, and Lucas Bernardi. 2023. Best Practices for Machine Learning Systems: An Industrial Framework for Analysis and Optimization. arXiv preprint arXiv:2306.13662 (2023)
work page Pith review arXiv 2023
-
[1]
Machine Learning operations maturity model
2023. Machine Learning operations maturity model. https:// learn.microsoft.com/en-us/azure/architecture/ai-ml/guide/mlops- maturity-model. Accessed: 2023-11-24
2023
-
[2]
MLOps: Continuous delivery and automation pipelines in machine learning
2023. MLOps: Continuous delivery and automation pipelines in machine learning. https://cloud.google.com/architecture/mlops-continuous- delivery-and-automation-pipelines-in-machine-learning . Accessed: 2023-11-24
2023
-
[3]
Rama Akkiraju, Vibha Sinha, Anbang Xu, Jalal Mahmud, Pritam Gundecha, Zhe Liu, Xiaotong Liu, and John Schumacher. 2018. Characterizing machine learning process: A maturity framework. arXiv:1811.04871 [cs.LG]
work page Pith review arXiv 2018
-
[4]
Omar Y. Al-Jarrah, Paul D. Yoo, Sami Muhaidat, George K. Karagiannidis, and Kamal Taha. 2015. Efficient Machine Learning for Big Data: A Review. Big Data Research 2, 3 (2015), 87–93. https://doi.org/10.1016/j.bdr. 2015.04.001 Big Data, Analytics, and High-Performance Computing
-
[5]
Rafa Al-Qutaish and Alain Abran. 2011. A Maturity Model of Software Product Quality. Journal of Research and Practice in Information Technology 43 (11 2011), 307–327
2011
-
[6]
Md Abdullah Al Alamin and Gias Uddin. 2021. Quality Assurance Chal- lenges for Machine Learning Software Applications During Software De- velopment Life Cycle Phases. arXiv:2105.01195 [cs.SE]
work page Pith review arXiv 2021
-
[7]
Javier Albert and Dmitri Goldenberg. 2022. E-Commerce Promotions Per- sonalization via Online Multiple-Choice Knapsack with Uplift Modeling. In Proceedings of the 31st ACM International Conference on Information & Knowledge Management. 2863–2872
2022
Show all 84 references
-
[8]
Saleema Amershi, Andrew Begel, Christian Bird, Robert DeLine, Harald Gall, Ece Kamar, Nachiappan Nagappan, Besmira Nushi, and Thomas Zimmermann. 2019. Software Engineering for Machine Learning: A Case Study. In 2019 IEEE/ACM 41st International Conference on Software Engi- neer...
2019
-
[9]
Ioannis Arapakis, Souneil Park, and Martin Pielot. 2021. Impact of Re- sponse Latency on User Behaviour in Mobile Web Search. In Proceedings of the 2021 Conference on Human Information Interaction and Retrieval (CHIIR ’21). ACM. https://doi.org/10.1145/3406522.3446038
2021
-
[10]
Vaishak Belle and Ioannis Papantonis. 2021. Principles and Practice of Explainable Machine Learning. Frontiers in Big Data 4 (2021). https: //doi.org/10.3389/fdata.2021.688969
2021
-
[11]
Lucas Bernardi, Themistoklis Mavridis, and Pablo Estevez. 2019. 150 Suc- cessful Machine Learning Models: 6 Lessons Learned at Booking.com. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining. 1743–1751
2019
-
[12]
Christian Bird, Nachi Nagappan, Brendan Murphy, Harald Gall, and Premkumar Devanbu. 2011. Don’t Touch My Code! Examining the Effects of Ownership on Software Quality. In Proceedings of the the eighth joint meeting of the European Software Engineering Conference and the ACM SIG...
2011
-
[13]
B. W. Boehm, J. R. Brown, and M. Lipow. 1976. Quantitative Evaluation of Software Quality. In Proceedings of the 2nd International Conference on Software Engineering (San Francisco, California, USA) (ICSE ’76). IEEE Computer Society Press, Washington, DC, USA, 592–605
1976
-
[14]
Eric Breck, Shanqing Cai, Eric Nielsen, Michael Salib, and D. Sculley. 2017. The ML Test Score: A Rubric for ML Production Readiness and Technical Debt Reduction. In Proceedings of IEEE Big Data
2017
-
[15]
Eric Breck, Marty Zinkevich, Neoklis Polyzotis, Steven Whang, and Sudip Roy. 2019. Data Validation for Machine Learning. In Proceedings of SysML. https://mlsys.org/Conferences/2019/doc/2019/167.pdf
2019
-
[16]
Cavano and James A
Joseph P . Cavano and James A. McCall. 1978. A Framework for the Measurement of Software Quality. SIGSOFT Softw. Eng. Notes 3, 5 (jan 1978), 133–139. https://doi.org/10.1145/953579.811113
1978
-
[17]
Jiyoo Chang and Christine Custis. 2022. Understanding Implemen- tation Challenges in Machine Learning Documentation. In Proceedings of the 2nd ACM Conference on Equity and Access in Algorithms, Mecha- nisms, and Optimization (, Arlington, VA, USA,) (EAAMO ’22). Associ- ation f...
2022
-
[18]
Celia Chen, Reem Alfayez, Kamonphop Srisopha, Barry Boehm, and Lin Shi. 2017. Why Is It Important to Measure Maintainability and What Are the Best Ways to Do It?. In 2017 IEEE/ACM 39th International Conference on Software Engineering Companion (ICSE-C). 377–378. https://doi.or...
2017
-
[19]
Shir Chorev, Philip Tannor, Dan Ben Israel, Noam Bressler, Itay Gabbay, Nir Hutnik, Jonatan Liberman, Matan Perlmutter, Yurii Romanyshyn, and Lior Rokach. 2022. Deepchecks: A library for testing and validating machine learning models and data. The Journal of Machine Learning R...
2022
-
[21]
Ralph B D’Agostino Sr, Joseph M Massaro, and Lisa M Sullivan. 2003. Non- inferiority trials: design concepts and issues–the encounters of academic consultants in statistics. Statistics in medicine 22, 2 (2003), 169–186
2003
-
[22]
Patricia Gomes Rêgo de Almeida, Carlos Denner dos Santos, and Josiva- nia Silva Farias. 2021. Artificial intelligence regulation: a framework for governance. Ethics and Information Technology 23, 3 (2021), 505–525
2021
-
[23]
Jack B. Dennis. 1975. Modularity. Springer Berlin Heidelberg, Berlin, Heidelberg, 128–182. https://doi.org/10.1007/3-540-07168-7_77
1975 doi
-
[24]
Ren Bin Lee Dixon. 2023. A principled governance for emerging AI regimes: lessons from China, the European Union, and the United States. AI and Ethics 3, 3 (2023), 793–810
2023
-
[25]
Geoff Dromey
R. Geoff Dromey. 1995. A Model for Software Product Quality. IEEE Trans. Softw. Eng.21, 2 (feb 1995), 146–162. https://doi.org/10.1109/32. 345830
1995 doi
-
[26]
The OWASP Foundation. 2024. OWASP Top 10 for Large Lan- guage Model Applications. https://owasp.org/www-project-top-10- for-large-language-model-applications/ . [Online; accessed 17-May- 2024]
2024
-
[27]
João Gama, Indrundefined Žliobaitundefined, Albert Bifet, Mykola Pech- enizkiy, and Abdelhamid Bouchachia. 2014. A Survey on Concept Drift Adaptation. ACM Comput. Surv. 46, 4, Article 44 (mar 2014), 37 pages. https://doi.org/10.1145/2523813
2014 doi
-
[28]
Görkem Giray. 2021. A software engineering perspective on engineering machine learning systems: State of the art and challenges.Journal of Systems and Software 180 (2021), 111031. https://doi.org/10.1016/j.jss.2021. 111031
2021 doi
-
[29]
Dmitri Goldenberg, Kostia Kofman, Javier Albert, Sarai Mizrachi, Adam Horowitz, and Irene Teinemaa. 2021. Personalization in Practice: Methods and Applications. In Proceedings of the 14th International Conference on Web Search and Data Mining
2021
-
[30]
Dmitri Goldenberg and Pavel Levin. 2021. Booking.com Multi-Destination Trips Dataset. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR ’21) . https: //doi.org/10.1145/3404835.3463240
2021
-
[31]
Dmitri Goldenberg, Chana Ross, Shir Meir Lador, Lin Lee Cheong, Pan- pan Xu, Elena Sokolova, Amit Mandelbaum, Irina Vasilinetc, Ankit Jain, Amit Weil Modlinger, and Saloni Potdar. 2023. The Second Workshop on Applied Machine Learning Management. In Proceedings of the 29th ACM ...
2023
-
[32]
Robert B. Grady. 1992. Practical Software Metrics for Project Management and Process Improvement. Prentice-Hall, Inc., USA
1992
-
[33]
Humphrey
W.S. Humphrey. 1988. Characterizing the software process: a maturity framework. IEEE Software 5, 2 (1988), 73–79. https://doi.org/10.1109/ 52.2014
1988
-
[34]
O’Reilly Media, Inc
Chip Huyen. 2022. Designing machine learning systems. " O’Reilly Media, Inc."
2022
-
[35]
ISO/IEC 9126. 2001. ISO/IEC 9126. Software engineering – Product quality. ISO/IEC
2001
-
[36]
Shibo Jie, Zhi-Hong Deng, and Ziheng Li. 2022. Alleviating Representa- tional Shift for Continual Fine-tuning. arXiv:2204.10535 [cs.CV]
2022 arXiv
-
[37]
Meenu Mary John, Helena Holmström Olsson, and Jan Bosch. 2021. To- wards MLOps: A Framework and Maturity Model. In 2021 47th Euromicro Conference on Software Engineering and Advanced Applications (SEAA). 1–8. https://doi.org/10.1109/SEAA53835.2021.00050
2021
-
[38]
Raphael Lopez Kaufman, Jegar Pitchforth, and Lukas Vermeer. 2017. De- mocratizing online controlled experiments at Booking. com. arXiv preprint arXiv:1710.08217 (2017)
2017 arXiv
-
[39]
Thomas Kluyver, Benjamin Ragan-Kelley, Fernando Pérez, Brian Granger, Matthias Bussonnier, Jonathan Frederic, Kyle Kelley, Jessica Hamrick, Ja- son Grout, Sylvain Corlay, Paul Ivanov, Damián Avila, Safia Abdalla, and Carol Willing. 2016. Jupyter Notebooks – a publishing format...
2016
-
[40]
J.C. Knight. 2002. Safety critical systems: challenges and directions. In Proceedings of the 24th International Conference on Software Engineering. ICSE
2002
-
[41]
Ron Kohavi, Alex Deng, Roger Longbotham, and Ya Xu. 2014. Seven Rules of Thumb for Web Site Experimenters. In Proceedings of the 20th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (New York, New York, USA) (KDD ’14). ACM, New York, NY, USA, 1857–18...
2014
-
[42]
Ron Kohavi, Alex Deng, and Lukas Vermeer. 2022. A/b testing intuition busters: Common misunderstandings in online controlled experiments. In Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 3168–3177
2022
-
[43]
Anastasiia Kornilova and Lucas Bernardi. 2021. Mining the stars: learn- ing quality ratings with user-facing explanations for vacation rentals. In Proceedings of the 14th ACM International Conference on Web Search and Data Mining. 976–983
2021
-
[44]
Dominik Kreuzberger, Niklas Kühl, and Sebastian Hirschl. 2023. Machine Learning Operations (MLOps): Overview, Definition, and Architecture. IEEE Access 11 (2023), 31866–31879. https://doi.org/10.1109/ACCESS. 2023.3262138
2023
-
[45]
Lucy Ellen Lwakatare, Aiswarya Raj, Ivica Crnkovic, Jan Bosch, and Helena Holmström Olsson. 2020. Large-scale machine learning sys- tems in real-world industrial settings: A review of challenges and so- lutions. Information and Software Technology 127 (2020), 106368. https: //...
2020
-
[46]
Michael R. Lyu. 2007. Software Reliability Engineering: A Roadmap. In Future of Software Engineering (FOSE ’07). 153–170. https://doi.org/10. 1109/FOSE.2007.24
2007
-
[47]
Mateo-Casali, Francisco Fraile, Andrés Boza, and Artem Nazarenko
Miguel A. Mateo-Casali, Francisco Fraile, Andrés Boza, and Artem Nazarenko. 2023. Maturity Model for Analysis of Machine Learning Oper- ations in Industry. https://doi.org/10.1007/978-3-031-27915-7_57
2023 doi
-
[48]
Ninareh Mehrabi, Fred Morstatter, Nripsuta Saxena, Kristina Lerman, and Aram Galstyan. 2021. A Survey on Bias and Fairness in Machine Learning. ACM Comput. Surv. 54, 6, Article 115 (jul 2021), 35 pages. https: //doi.org/10.1145/3457607
2021 doi
-
[49]
Microsoft. 2023. How to Evaluate LLMs: A Complete Metric Framework. https://www.microsoft.com/en-us/research/group/experimentation- platform-exp/articles/how-to-evaluate-llms-a-complete-metric- framework/. [Online; accessed 17-May-2024]
2023
-
[50]
Jose Miguel, David Mauricio, and Glen Rodriguez. 2014. A Review of Software Quality Models for the Evaluation of Software Products. Inter- national journal of Software Engineering & Applications 5 (11 2014), 31–54. https://doi.org/10.5121/ijsea.2014.5603
2014
-
[51]
Sanjay Misra, Adewole Adewumi, Nicholas Omoregbe, and Broderick Crawford. 2016. A systematic literature review of open source software quality assessment models. SpringerPlus 2016 (11 2016), 1936. https: //doi.org/10.1186/s40064-016-3612-4
2016 doi
-
[52]
Siba N. Mohanty. 1979. Models and Measurements for Quality Assessment of Software. ACM Comput. Surv. 11, 3 (sep 1979), 251–275. https://doi. org/10.1145/356778.356783
1979
-
[53]
Sasi Kumar Murakonda and Reza Shokri. 2020. ML Privacy Meter: Aiding Regulatory Compliance by Quantifying the Privacy Risks of Machine Learning. arXiv:2007.09339 [cs.CR]
2020 arXiv
-
[54]
Lawrence
Andrei Paleyes, Raoul-Gabriel Urma, and Neil D. Lawrence. 2022. Chal- lenges in Deploying Machine Learning: A Survey of Case Studies. ACM Comput. Surv. 55, 6, Article 114 (dec 2022), 29 pages. https://doi.org/ 10.1145/3533378
2022 doi
-
[55]
Jiantao Pan. 1999. Software testing. Dependable Embedded Systems 5, 2006 (1999), 1
1999
-
[56]
Mark Paulk, William Curtis, Mary Beth Chrissis, and Charles We- ber. 1993. Capability Maturity Model for Software (Version 1.1) . Tech- nical Report CMU/SEI-93-TR-024. https://insights.sei.cmu.edu/ library/capability-maturity-model-for-software-version-11/ Ac- cessed: 2023-Nov-30
1993
-
[57]
Shachaf Poran, Gil Amsalem, Amit Beka, and Dmitri Goldenberg. 2022. With One Voice: Composing a Travel Voice Assistant from Repurposed Models. In Companion Proceedings of the Web Conference 2022. 383–387
2022
-
[58]
Adarsh Prasad. 2022. Towards Robust and Resilient Machine Learning. (4 2022). https://doi.org/10.1184/R1/19552420.v1
2022 doi
-
[59]
Chana Ross, Tomer Ovadia, Jake Mooney, Amit Meitin, Eytan Kabilou, Mush Kabalo, and Dmitri Goldenberg. 2022. Democratizing Travel Per- sonalization via Central Recommendation Platform. (2022)
2022
-
[60]
Raghad Baker Sadiq, Nurhizam Safie, Abdul Hadi Abd Rahman, and Shidrokh Goudarzi. 2021. Artificial intelligence maturity model: a system- atic literature review. PeerJ Computer Science 7 (2021), e661
2021
-
[61]
Christopher Samiullah. 2020. Monitoring Machine Learning Models in Production. https://christophergs.com/machine%20learning/2020/03/ 14/how-to-monitor-machine-learning-models/ Accessed on December 22, 2023
2020
-
[62]
Santhanam
P . Santhanam. 2020. Quality Management of Machine Learning Systems. arXiv:2006.09529 [cs.SE]
2020 arXiv
-
[63]
Simone Scalabrino, Mario Linares-Vásquez, Rocco Oliveto, and Denys Poshyvanyk. 2018. A comprehensive model for code readability. Journal of Software: Evolution and Process 30 (06 2018). https://doi.org/10.1002/ smr.1958
2018
-
[64]
David Sculley, Gary Holt, Daniel Golovin, Eugene Davydov, Todd Phillips, Dietmar Ebner, Vinay Chaudhary, Michael Young, Jean-Francois Crespo, and Dan Dennison. 2015. Hidden technical debt in machine learning systems. Advances in neural information processing systems 28 (2015)
2015
-
[65]
Alex Serban, Koen van der Blom, Holger Hoos, and Joost Visser. 2020. Adoption and Effects of Software Engineering Best Practices in Machine Learning. In Proceedings of the 14th ACM / IEEE International Symposium on Empirical Software Engineering and Measurement (ESEM) (Bari, I...
2020
-
[66]
Alex Serban, Koen van der Blom, Holger Hoos, and Joost Visser. 2022. Software engineering for machine learning. https://se-ml.github.io/
2022
-
[67]
Ghorbani
Ali Shiravi, Hadi Shiravi, Mahbod Tavallaee, and Ali A. Ghorbani. 2012. Toward developing a systematic approach to generate benchmark datasets for intrusion detection. Computers & Security 31, 3 (2012), 357–374. https: //doi.org/10.1016/j.cose.2011.12.012
2012 doi
-
[68]
Julien Siebert, Lisa Joeckel, Jens Heidrich, Koji Nakamichi, Kyoko Ohashi, Isao Namba, Rieko Yamamoto, and Mikio Aoyama. 2020. Towards Guide- lines for Assessing Qualities of Machine Learning Systems. Springer Interna- tional Publishing, 17–31. https://doi.org/10.1007/978-3-03...
2020 doi
-
[69]
Stefan Studer, Thanh Binh Bui, Christian Drescher, Alexander Hanuschkin, Ludwig Winkler, Steven Peters, and Klaus-Robert Müller. 2021. Towards CRISP-ML(Q): A Machine Learning Process Model with Quality Assur- ance Methodology. Machine Learning and Knowledge Extraction 3, 2 (20...
2021 doi
-
[70]
Symeonidis, E
G. Symeonidis, E. Nerantzis, A. Kazakis, and G. A. Papakostas. 2022. MLOps – Definitions, Tools and Challenges. arXiv:2201.00162 [cs.LG]
2022 arXiv
-
[71]
Michael Veale and Frederik Zuiderveen Borgesius. 2021. Demystifying the Draft EU Artificial Intelligence Act—Analysing the good, the bad, and the unclear elements of the proposed approach. Computer Law Review International 22, 4 (2021), 97–112
2021
-
[72]
Fengjun Wang, Moran Beladev, Ofri Kleinfeld, Elina Frayerman, Tal Shachar, Eran Fainman, Karen Lastmann Assaraf, Sarai Mizrachi, and Benjamin Wang. 2023. Text2Topic: Multi-Label Text Classification System for Efficient Topic Detection in User Generated Content with Zero-Shot C...
2023
-
[73]
Fengjun Wang, Sarai Mizrachi, Moran Beladev, Guy Nadav, Gil Am- salem, Karen Lastmann Assaraf, and Hadas Harush Boker. 2023. MuMIC– Multimodal Embedding for Multi-Label Image Classification with Tem- pered Sigmoid. In Proceedings of the AAAI Conference on Artificial Intelligen...
2023
-
[74]
Shuhei Watanabe. 2023. Tree-Structured Parzen Estimator: Understanding Its Algorithm Components and Their Roles for Better Empirical Perfor- mance. arXiv:2304.11127 [cs.LG]
2023 arXiv
-
[75]
Roy Wendler. 2012. The maturity of maturity model research: A systematic mapping study. Information and Software Technology 54, 12 (2012), 1317–
2012
-
[76]
Wieder, J.M
P . Wieder, J.M. Butler, W. Theilmann, and R. Yahyapour. 2011. Service Level Agreements for Cloud Computing. Springer New York. https://books. google.nl/books?id=z306GUfFL5gC
2011
-
[77]
Wikipedia. 2022. List of system quality attributes — Wikipedia, The Free Encyclopedia. https://en.wikipedia.org/wiki/List_of_system_ quality_attributes. [Online; accessed 30-April-2022]
2022
-
[78]
Wikipedia. 2024. Critical System — Wikipedia, The Free Encyclopedia. https://en.wikipedia.org/wiki/Critical_system. [Online; accessed 11- January-2024]
2024
-
[79]
Meng Yan, Xin Xia, Xiaohong Zhang, Ling Xu, Dan Yang, and Shanping Li
-
[80]
Fahri Anıl Yerlikaya and ¸ Serif Bahtiyar. 2022. Data poisoning attacks against machine learning algorithms. Expert Systems with Applications 208 (2022), 118101. https://doi.org/10.1016/j.eswa.2022.118101
2022
-
[81]
Xing, Hao Zhang, Joseph E
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P . Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. arXiv:2306.05685 [cs.CL] Maturity ...
2023 arXiv
-
[1339]
https://doi.org/10.1016/j.infsof.2012.07.007 Special Section on Software Reliability and Security
2012 doi
-
[2002]
Castelli, Chouliaras and Goldenberg
547–550. Castelli, Chouliaras and Goldenberg
-
[2019]
https://api.semanticscholar
Software quality assessment model: a systematic mapping study.Sci- ence China Information Sciences 62 (2019). https://api.semanticscholar. org/CorpusID:53378706
2019
Reviewed August 8, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.