Pith. sign in

REVIEW 4 major objections 7 minor 84 references

Maturity Framework for Enhancing Machine Learning Quality

T0 review · 4 major / 7 minor · reviewed 2026-08-08 · deepseek-v4-flash

Pith's one-line read This paper claims that machine-learning system quality can be turned into a measurable, governable quantity through a 0–100 quality score and a five-level maturity framework, with empirical evidence from a company-wide rollout at…

desk verdict A useful, open-sourced ML quality framework with an honest rollout narrative, but the abstract oversells the empirical validation: Figures 1-3 track compliance with the framework's own Q score, not independent quality gains. read the letter →

arxiv 2502.15758 v1 pith:V7OBQDXS submitted 2025-02-12 cs.LG cs.CYcs.SE

classification cs.LGcs.CYcs.SE
keywords machinelearningqualitymaturityframeworkassessmentMLOpsgovernancescorereproducibilitybusinesscriticalityMLregistry
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper claims that the quality of a machine-learning system can be turned from a vague ideal into a measurable, governable quantity. It defines seven quality characteristics — utility, economy, robustness, modifiability, productionizability, comprehensibility, and responsibility — split into sub-characteristics scored by whether minimal or full requirements are met, and aggregates the scores into a single quality score between 0 and 100. On top of that score it builds a five-level maturity ladder, with the level a system is expected to reach set by its business criticality. The authors report applying this framework at Booking.com across hundreds of production systems with an open-source Python package, and they present before/after evidence that quality scores rose, the low-quality tail shrank, and the rollout produced concrete business improvements such as defunct systems being retired and retraining pipelines being made much faster. A sympathetic reader would care because the paper offers a concrete, reusable method — not just a checklist — and claims empirical validation at scale, which is exactly what most ML governance frameworks lack.

What carries the argument

The load-bearing object is the quality score $Q$ together with the maturity ladder defined by the requirement table. $Q$ compresses the whole assessment into a single number: for each of $N$ sub-characteristics the assessor records a gap $g_s \in \{0,1,2\}$, and $Q = 100\left(1 - \frac{\sum_s g_s}{2N}\right)$; this is the quantity shown improving over time in the paper's figures. The seven characteristics — utility, economy, robustness, modifiability, productionizability, comprehensibility, and responsibility — are the grid on which requirements are placed, and Table 1 spells out minimal and full requirements per sub-characteristic with the maturity levels they unlock. Business criticality (proof of concept, production non-critical, production critical) sets the target level, so the same requirement table produces different obligations for a toy experiment and a revenue-critical system.

What would settle it

Run an external validation: take a set of production ML systems outside Booking.com, have two independent teams score each with the open-source package, and compare the resulting $Q$ scores and maturity levels against independently measured outcomes such as outage frequency, retraining lag, and business metrics over six months. The central claim would collapse if high-scoring systems fail as often as low-scoring ones, or if the two teams disagree on most gap judgments.

Watch

Extended reading notes

Core claim

The central claim is that ML quality can be assessed rigorously enough to drive governance: each sub-characteristic has a minimal and a full requirement, the gap between current state and requirement is coded as no gap (0), small gap (1), or large gap (2), and the quality score $Q = 100 \left(1 - \frac{\sum_s g_s}{g_{\text{large}}\,N}\right)$ summarizes the system in one number between 0 and 100. The framework then maps required sub-characteristics to five maturity levels, so a system's maturity is not a global label but a profile of which requirements are met, and business criticality — proof-of-concept, production non-critical, or production critical — decides which maturity level is actually required. The authors assert empirical validation from the Booking.com rollout: quality improved across nearly all sub-characteristics, with the largest gains in accuracy, testability, readability, understandability, and maintainability; monitoring was identified as the biggest gap and addressed by a central observability platform; legacy systems with low scores were shown by A/B benchmarking to be non-inferior to simple baselines and were retired; and efficiency work cut one critical system's training time from 7 hours to under 1 hour and another system's hyperparameter tuning from 20 hours to under 1 hour.

Load-bearing premise

The load-bearing assumption is that the 0–100 quality score, built from yes/no judgments about requirements, really measures how good an ML system is, and that the thresholds tuned at Booking.com apply elsewhere.

Editorial extensions

If this is right

  • Because the assessment is packaged as open-source code, an organization can generate the same standardized reports and folder-structured, versioned quality history without building the framework from scratch.
  • The split between minimal and full requirements gives teams an incremental path: they can close small gaps first and move up one maturity level at a time rather than facing an all-or-nothing standard.
  • The business-criticality tiering means the framework does not impose the same burden on every system; low-stakes experiments can stay lightweight while critical systems must meet the full ladder.
  • The reported shift in the quality-score distribution implies that visibility alone — monthly reports and a governance dashboard — can change team priorities, even before new tooling is built.
  • If the automated registry-based assessment works as described, quality evaluation can run monthly across hundreds of systems with a human needed only for code-readability and modularity judgments.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The claimed validity rests on the assumption that the score $Q$ measures something real; a natural test the paper does not run is an inter-rater reliability study, where independent evaluators score the same system and their $Q$ values are compared.
  • The thresholds in Table 1 are presented as principled but were calibrated to Booking.com's internal distributions; organizations with different scale, domain risk, or regulatory pressure will likely need to recalibrate them, and the paper gives no procedure for doing so.
  • The before/after evidence is observational, not experimental: teams chose which gaps to close, so the improvement trend may partly reflect selection rather than the framework's causal effect; a controlled rollout to matched teams would test this.
  • For LLM-based systems the paper describes adjustments but does not yet define GenAI-specific sub-characteristics; a concrete extension would add requirements such as hallucination rate, prompt-injection resistance, and evaluation-data freshness, and check whether $Q$ still tracks business outcomes.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 7 minor

Summary. This paper proposes an ML quality assessment and maturity framework developed at Booking.com. Seven quality characteristics (Utility, Economy, Robustness, Modifiability, Productionizability, Comprehensibility, Responsibility) are decomposed into sub-characteristics, each scored against a checklist of minimal and full requirements in Table 1. Gap values of 0/1/2 are aggregated into a quality score Q in Eq. (1), and a maturity framework with five levels and three business-criticality tiers maps expected quality standards per system. The authors describe an open-source Python package with automated and semi-automated gap inference from an ML registry, a two-year company-wide rollout, organizational challenges and lessons learned, and domain-specific adaptations for LLMs, causal ML, and CV/NLP systems. Empirical evidence is presented in Figures 1-3 as before/after quality scores, maturity levels, and compliance percentages, together with business-value anecdotes in Section 6.3. The central claim, stated in the abstract and Section 8, is that the framework is open-sourced, validated with empirical evidence, and produced significant quality improvements and business value at scale.

Significance. If the validation claims were fully supported, this would be a genuinely reusable governance artifact for industrial ML: the granular checklist-based quality model is grounded in ISO/IEC 9126 and ML-specific literature; the system-level maturity definitions with criticality-based expectations are more actionable than organization-level maturity models; and the open-source release with versioned, reproducible report generation is a concrete engineering contribution. The paper also provides candid process lessons and quantified efficiency examples (e.g., training time reduced from 7 hours to under 1 hour in Section 6.3.3), which give the framework practical credibility. The main weakness is that the reported validation is largely self-referential: the outcome measure Q is built from the same requirement gaps that the framework recommends closing, so Figures 1-3 primarily demonstrate compliance with the framework's own rules. The significance of the contribution is therefore conditional on either anchoring Q to independent quality or business outcomes, or substantially tempering the 'validated with empirical evidence' claim.

major comments (4)
  1. [§3.2 Eq. (1); §7 Figures 1-3; abstract and §8] The validation evidence is self-referential. Q in Eq. (1) is defined as the normalized sum of gap values g_s that are assigned exactly by checking the requirements in Table 1, and the framework's recommendations are designed to close those same gaps. Figures 1 and 3 therefore show that systems improved on the framework's own requirements after the framework was applied; this is expected by construction and does not by itself establish improvement in ML quality or business value. Section 8's claim of 'significant quality improvements across all the aspects, resulting in business value' and the abstract's 'validated with empirical evidence' consequently overstate what the data can support. I recommend correlating Q (or its changes) against independent outcomes, such as model performance on fresh test data, incident counts, latency, or business metrics from A/B tests, or comparing with systems not subject to the framework. The conclusion's own concession that 'further empirical validation' is needed supports this reading.
  2. [§6.1.4; §7 Figures 1-2] The reported aggregate improvement may be partly attributable to population attrition. Section 6.1.4 states that the rollout led to the cleanup of 'dozens of ML systems,' and Figure 1 is restricted to a 'selected subset' of systems that voluntarily participated in the rollout. If low-scoring systems were decommissioned or excluded between the before and after measurements, the distribution shift in Figure 2 could reflect changes in which systems constitute the population rather than genuine quality gains. Please report a fixed-cohort analysis in which the same systems are measured at both time points, state the number of systems contributing to each distribution in Figure 2, and specify the inclusion criteria for the subset shown in Figure 1.
  3. [§7; §6.1.5] No statistical support is provided for the claimed improvements. Figures 1-3 contain no confidence intervals, significance tests, or effect sizes; Section 7's phrase 'significant improvement' appears to be used in the informal sense, and Section 8 then elevates it to 'significant quality improvements across all the aspects.' Separately, Section 6.1.5 discloses that several requirements and thresholds were calibrated from quantiles of Booking.com's internal distributions, which makes Q and the compliance percentages in Figure 3 organization-relative rather than absolute quality measures. The claim in Section 4.1 that the maturity levels work 'out of the box for most of the large scale organizations with no safety-critical applications' would be more credible with a sensitivity analysis showing how results change under plausible threshold variations, or with explicit recalibration guidance.
  4. [§5.1] The reliability of the measurement instrument is unquantified. Section 5.1 states that readability and modularity still require a human in the loop, and in the semi-automated mode practitioners fill in a survey whose answers determine gap values. These subjective judgments enter Q in Eq. (1) with the same weight as automated metrics such as test coverage, yet no inter-rater reliability, evaluation protocol, or rubric calibration is reported. Reporting agreement statistics (e.g., Cohen's kappa) on a sample of systems, or at least documenting the rater guidance provided to teams, would substantially strengthen the construct validity of Q.
minor comments (7)
  1. [§3.2, Table 1] The symbols used in Table 1 to denote minimal and full requirements do not render in the submitted manuscript (Section 3.2 refers to them but they appear as blank spaces), which makes the requirements table difficult to interpret; please provide a clearly visible legend.
  2. [§7, Figures 1-3] Sample sizes are not reported: Figure 1 does not state how many systems were in the voluntary rollout or why those particular systems were selected; Figure 2 does not give the number of systems per month; Figure 3 does not report how many systems are in the before and after groups.
  3. [§7, Figure 2] The violin plot's x-axis spans May through November 2023, while Section 6 describes a two-year rollout; please clarify the relationship between the rollout timeline and the measurement window.
  4. [Appendix A, Figures 4-5] The phrases 'The ML system is fullon' and 'the model is fullon' appear to be jargon or typos; please clarify the intended wording.
  5. [§2] Typo: 'They key differences' should be 'The key differences.'
  6. [§3.2, Eq. (1)] Equation (1) assigns equal weight to every sub-characteristic regardless of business criticality; since the maturity framework elsewhere differentiates by criticality, a brief discussion of how criticality weighting would interact with the per-level requirements would help.
  7. [§3.1, [77]] The initial attribute shortlist is anchored in part to a Wikipedia list ([77]); a more citable source such as ISO/IEC 25010 for the starting set of quality attributes would strengthen the related-work grounding.

Circularity Check

3 steps flagged · score 6.0 of 10

The claimed empirical validation is largely self-referential: Figures 1–3 measure compliance with the framework's own Q score and thresholds, and the quality model is imported from the authors' prior work, though Section 6.3 provides some independent business grounding.

  1. self definitional [Section 3.2, Eq. (1); Section 7, Figure 1 and Figure 3]
    "By assigning the following numerical value to the three gap types: no gap: 0 / small gap: 1 / large gap: 2. we can introduce the quality score Q as: Q = 100(1 − ΣN s gs glarge N) ... The quality score ranging between 0 and 100, can be used to measure ML system quality, to track improvements and to compare different systems. ... Figure 1 shows the quality improvement of several ML systems over a number of cycles of assessments and implementation of the recommended best practices."

    Q is a linear transform of the same binary gap judgments g_s that Table 1 defines and that the framework's recommendations are designed to remove. Consequently, an increase in Q is, by construction, an increase in compliance with the framework's own rubric, not an independent measure of ML quality or business value. The abstract's claim that the method is 'validated with empirical evidence' therefore rests on the self-defined score, and Figures 1 and 3 demonstrate rule-following rather than externally anchored quality improvement.

  2. fitted input called prediction [Section 6.1.5 and Section 7, Figure 3]
    "We substantiated our requirements and thresholds by taking them from the quantiles of the empirical distribution of the ML systems in the company (e.g. for training cost) or from well-known industry standards (e.g. benchmarking against simple baselines) [34]. ... For each quality sub-characteristic Figure 3 shows the percentage of systems complying with the requirements (without technical gaps) before and after the rollout of the framework."

    Several thresholds that determine whether a sub-characteristic has a gap were fitted to the quantiles of the very same systems whose improvement is then reported. The before/after compliance gain in Figure 3 is therefore partly a mechanical consequence of setting the bar relative to the initial population distribution, so the reported 'quality improvement' is not independent of the way the requirements were calibrated.

1 more flagged steps
  1. self citation load bearing [Section 3 opening; Section 5, ML Quality Python package]
    "To define quality, we use a refined version of the quality model tailored for ML systems, introduced in [20]. ... The generation and prioritization of the recommendations is based on our framework described in [20]."

    The quality model that supplies the sub-characteristics and requirements, and the recommendation priorities that drive the measured improvements, are imported from reference [20], whose authorship overlaps with the present paper. No independent derivation or external validation of that prior framework is given here; the central assessment method therefore leans on a load-bearing self-citation rather than on evidence established in this paper.

full rationale

The paper's central empirical validation is an internal loop: Q (Eq. 1) is defined as a sum over the same binary gap judgments that Table 1 prescribes and that the framework's recommendations are designed to remove; Figures 1–3 then report increases in Q and in the fraction of systems with 'no technical gaps' after teams applied those recommendations. This makes the headline 'validated with empirical evidence' and 'significant quality improvements across all aspects' substantially self-referential: the measured quantity is compliance with the framework's own rubric. The calibration of several thresholds to quantiles of the company's own distribution (Section 6.1.5) means the before/after comparison in Figure 3 is also relative to norms derived from the same population, not to an external standard. Additionally, the quality model and the recommendation priorities are imported from the authors' prior work [20] via load-bearing self-citation, without independent derivation in this paper. There is some non-circular grounding: Section 6.3 gives concrete business incidents (e.g., non-inferiority A/B results showing stale models, efficiency improvements reducing training time from 7 to under 1 hour) that do not reduce to Q. Because of that independent content, the paper is not fully circular; but the principal aggregate evidence is self-referential, so the score is 6.

Assumptions & free parameters 5 free parameters · 4 assumptions · 0 invented entities

The framework rests on the validity and completeness of the seven quality characteristics, the quantization of gaps, and the specific maturity thresholds. None of these are derived; they are justified by literature, practitioner feedback, and internal Booking.com distributions. No new physical entities are introduced.

free parameters (5)
  • Gap values 0, 1, 2 in Eq. 1 = 0, 1, 2 (glarge = 2)
    The linear mapping of no/small/large gaps to numbers is chosen to normalize Q between 0 and 100; the weights are not derived from data or theory.
  • Test coverage thresholds = 20% minimal, 80% full
    Table 1 sets unit test coverage requirements; the paper says thresholds were substantiated from internal quantiles and industry standards, but no derivation or external benchmark is given.
  • Pipeline failure resilience thresholds = Up to 30% failed pipelines per quarter minimal, at most 10% full
    These thresholds for the Resilience sub-characteristic in Table 1 are hand-selected and not derived from a reliability model.
  • Business criticality cutoffs = 66th percentile of request volume, more than 4 dependent teams, more than 1% yearly revenue, strategic importance
    Section 4.2 defines criticality by these cutoffs; the 66th percentile and other thresholds are policy choices without stated statistical justification.
  • Expected maturity levels per criticality = Proof of concept = 1, production non-critical = 3, production critical = 5
    Section 4.2 assigns required maturity levels; levels 2 and 4 are left as intermediate progress markers, so the mapping is a governance choice, not an empirical calibration.
assumptions (4)
  • domain assumption The seven quality characteristics form a complete decomposition of ML system quality.
    Section 3.1 adopts the model from [20] and ISO 9126; completeness was checked qualitatively through practitioner interviews, not proven.
  • domain assumption Quality gaps can be meaningfully quantized to no/small/large and mapped to the linear score Q in Eq. 1.
    Equal weights across sub-characteristics and information loss from quantization are assumed without sensitivity analysis.
  • ad hoc to paper The maturity level requirements in Table 1 are appropriate for large organizations without safety-critical applications.
    Section 4.1 says the exact choice depends on organizational needs; the presented mapping is a specific policy not derived from first principles.
  • domain assumption The empirical distribution of internal Booking.com systems is a valid basis for setting quality thresholds.
    Section 6.1.5 says thresholds were taken from quantiles of the internal distribution, which may encode the status quo rather than a quality target.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Maturity Framework for Enhancing Machine Learning Quality." pith.science (2026). https://pith.science/paper/V7OBQDXS

@misc{pith2026250215758,
  author       = {Pith},
  title        = {Pith review of: Maturity Framework for Enhancing Machine Learning Quality},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/V7OBQDXS}},
  note         = {Machine review of arXiv:2502.15758}
}
read the original abstract

With the rapid integration of Machine Learning (ML) in business applications and processes, it is crucial to ensure the quality, reliability and reproducibility of such systems. We suggest a methodical approach towards ML system quality assessment and introduce a structured Maturity framework for governance of ML. We emphasize the importance of quality in ML and the need for rigorous assessment, driven by issues in ML governance and gaps in existing frameworks. Our primary contribution is a comprehensive open-sourced quality assessment method, validated with empirical evidence, accompanied by a systematic maturity framework tailored to ML systems. Drawing from applied experience at Booking.com, we discuss challenges and lessons learned during large-scale adoption within organizations. The study presents empirical findings, highlighting quality improvement trends and showcasing business outcomes. The maturity framework for ML systems, aims to become a valuable resource to reshape industry standards and enable a structural approach to improve ML maturity in any organization.

Figures

Figures reproduced from arXiv: 2502.15758 by the authors.

Figure 1
Figure 1. Quality (Y-axis) and maturity (markers) score of a [PITH_FULL_IMAGE:figures/full_fig_p007_1.png] view at source ↗
Figure 3
Figure 3. Percentage of systems complying with the require [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figure 2
Figure 2. Violin plot of the ML quality score for different [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Example of a report for a fully mature system. No technical gaps are present, all the fulfilled quality attributes are [PITH_FULL_IMAGE:figures/full_fig_p011_4.png]
Figure 5
Figure 5. Figure 5: Example of a report for a system of maturity level 1. The gaps to be fulfilled to pass to the next maturity level are [PITH_FULL_IMAGE:figures/full_fig_p012_5.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

84 extracted references · 58 canonical work pages

  1. [20]

    Georgios Christos Chouliaras, Kornel Kiełczewski, Amit Beka, David Konopnicki, and Lucas Bernardi. 2023. Best Practices for Machine Learning Systems: An Industrial Framework for Analysis and Optimization. arXiv preprint arXiv:2306.13662 (2023)

  2. [1]

    Machine Learning operations maturity model

    2023. Machine Learning operations maturity model. https:// learn.microsoft.com/en-us/azure/architecture/ai-ml/guide/mlops- maturity-model. Accessed: 2023-11-24

  3. [2]

    MLOps: Continuous delivery and automation pipelines in machine learning

    2023. MLOps: Continuous delivery and automation pipelines in machine learning. https://cloud.google.com/architecture/mlops-continuous- delivery-and-automation-pipelines-in-machine-learning . Accessed: 2023-11-24

  4. [3]

    Rama Akkiraju, Vibha Sinha, Anbang Xu, Jalal Mahmud, Pritam Gundecha, Zhe Liu, Xiaotong Liu, and John Schumacher. 2018. Characterizing machine learning process: A maturity framework. arXiv:1811.04871 [cs.LG]

  5. [4]

    Al-Jarrah, Paul D

    Omar Y. Al-Jarrah, Paul D. Yoo, Sami Muhaidat, George K. Karagiannidis, and Kamal Taha. 2015. Efficient Machine Learning for Big Data: A Review. Big Data Research 2, 3 (2015), 87–93. https://doi.org/10.1016/j.bdr. 2015.04.001 Big Data, Analytics, and High-Performance Computing

  6. [5]

    Rafa Al-Qutaish and Alain Abran. 2011. A Maturity Model of Software Product Quality. Journal of Research and Practice in Information Technology 43 (11 2011), 307–327

  7. [6]

    Md Abdullah Al Alamin and Gias Uddin. 2021. Quality Assurance Chal- lenges for Machine Learning Software Applications During Software De- velopment Life Cycle Phases. arXiv:2105.01195 [cs.SE]

  8. [7]

    Javier Albert and Dmitri Goldenberg. 2022. E-Commerce Promotions Per- sonalization via Online Multiple-Choice Knapsack with Uplift Modeling. In Proceedings of the 31st ACM International Conference on Information & Knowledge Management. 2863–2872

Show all 84 references
  1. [8]

    Saleema Amershi, Andrew Begel, Christian Bird, Robert DeLine, Harald Gall, Ece Kamar, Nachiappan Nagappan, Besmira Nushi, and Thomas Zimmermann. 2019. Software Engineering for Machine Learning: A Case Study. In 2019 IEEE/ACM 41st International Conference on Software Engi- neer...

  2. [9]

    Ioannis Arapakis, Souneil Park, and Martin Pielot. 2021. Impact of Re- sponse Latency on User Behaviour in Mobile Web Search. In Proceedings of the 2021 Conference on Human Information Interaction and Retrieval (CHIIR ’21). ACM. https://doi.org/10.1145/3406522.3446038

  3. [10]

    Vaishak Belle and Ioannis Papantonis. 2021. Principles and Practice of Explainable Machine Learning. Frontiers in Big Data 4 (2021). https: //doi.org/10.3389/fdata.2021.688969

  4. [11]

    Lucas Bernardi, Themistoklis Mavridis, and Pablo Estevez. 2019. 150 Suc- cessful Machine Learning Models: 6 Lessons Learned at Booking.com. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining. 1743–1751

  5. [12]

    Christian Bird, Nachi Nagappan, Brendan Murphy, Harald Gall, and Premkumar Devanbu. 2011. Don’t Touch My Code! Examining the Effects of Ownership on Software Quality. In Proceedings of the the eighth joint meeting of the European Software Engineering Conference and the ACM SIG...

  6. [13]

    B. W. Boehm, J. R. Brown, and M. Lipow. 1976. Quantitative Evaluation of Software Quality. In Proceedings of the 2nd International Conference on Software Engineering (San Francisco, California, USA) (ICSE ’76). IEEE Computer Society Press, Washington, DC, USA, 592–605

  7. [14]

    Eric Breck, Shanqing Cai, Eric Nielsen, Michael Salib, and D. Sculley. 2017. The ML Test Score: A Rubric for ML Production Readiness and Technical Debt Reduction. In Proceedings of IEEE Big Data

  8. [15]

    Eric Breck, Marty Zinkevich, Neoklis Polyzotis, Steven Whang, and Sudip Roy. 2019. Data Validation for Machine Learning. In Proceedings of SysML. https://mlsys.org/Conferences/2019/doc/2019/167.pdf

  9. [16]

    Cavano and James A

    Joseph P . Cavano and James A. McCall. 1978. A Framework for the Measurement of Software Quality. SIGSOFT Softw. Eng. Notes 3, 5 (jan 1978), 133–139. https://doi.org/10.1145/953579.811113

  10. [17]

    Jiyoo Chang and Christine Custis. 2022. Understanding Implemen- tation Challenges in Machine Learning Documentation. In Proceedings of the 2nd ACM Conference on Equity and Access in Algorithms, Mecha- nisms, and Optimization (, Arlington, VA, USA,) (EAAMO ’22). Associ- ation f...

  11. [18]

    Celia Chen, Reem Alfayez, Kamonphop Srisopha, Barry Boehm, and Lin Shi. 2017. Why Is It Important to Measure Maintainability and What Are the Best Ways to Do It?. In 2017 IEEE/ACM 39th International Conference on Software Engineering Companion (ICSE-C). 377–378. https://doi.or...

  12. [19]

    Shir Chorev, Philip Tannor, Dan Ben Israel, Noam Bressler, Itay Gabbay, Nir Hutnik, Jonatan Liberman, Matan Perlmutter, Yurii Romanyshyn, and Lior Rokach. 2022. Deepchecks: A library for testing and validating machine learning models and data. The Journal of Machine Learning R...

  13. [21]

    Ralph B D’Agostino Sr, Joseph M Massaro, and Lisa M Sullivan. 2003. Non- inferiority trials: design concepts and issues–the encounters of academic consultants in statistics. Statistics in medicine 22, 2 (2003), 169–186

  14. [22]

    Patricia Gomes Rêgo de Almeida, Carlos Denner dos Santos, and Josiva- nia Silva Farias. 2021. Artificial intelligence regulation: a framework for governance. Ethics and Information Technology 23, 3 (2021), 505–525

  15. [23]

    Jack B. Dennis. 1975. Modularity. Springer Berlin Heidelberg, Berlin, Heidelberg, 128–182. https://doi.org/10.1007/3-540-07168-7_77

  16. [24]

    Ren Bin Lee Dixon. 2023. A principled governance for emerging AI regimes: lessons from China, the European Union, and the United States. AI and Ethics 3, 3 (2023), 793–810

  17. [25]

    Geoff Dromey

    R. Geoff Dromey. 1995. A Model for Software Product Quality. IEEE Trans. Softw. Eng.21, 2 (feb 1995), 146–162. https://doi.org/10.1109/32. 345830

  18. [26]

    The OWASP Foundation. 2024. OWASP Top 10 for Large Lan- guage Model Applications. https://owasp.org/www-project-top-10- for-large-language-model-applications/ . [Online; accessed 17-May- 2024]

  19. [27]

    João Gama, Indrundefined Žliobaitundefined, Albert Bifet, Mykola Pech- enizkiy, and Abdelhamid Bouchachia. 2014. A Survey on Concept Drift Adaptation. ACM Comput. Surv. 46, 4, Article 44 (mar 2014), 37 pages. https://doi.org/10.1145/2523813

  20. [28]

    Görkem Giray. 2021. A software engineering perspective on engineering machine learning systems: State of the art and challenges.Journal of Systems and Software 180 (2021), 111031. https://doi.org/10.1016/j.jss.2021. 111031

  21. [29]

    Dmitri Goldenberg, Kostia Kofman, Javier Albert, Sarai Mizrachi, Adam Horowitz, and Irene Teinemaa. 2021. Personalization in Practice: Methods and Applications. In Proceedings of the 14th International Conference on Web Search and Data Mining

  22. [30]

    Dmitri Goldenberg and Pavel Levin. 2021. Booking.com Multi-Destination Trips Dataset. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR ’21) . https: //doi.org/10.1145/3404835.3463240

  23. [31]

    Dmitri Goldenberg, Chana Ross, Shir Meir Lador, Lin Lee Cheong, Pan- pan Xu, Elena Sokolova, Amit Mandelbaum, Irina Vasilinetc, Ankit Jain, Amit Weil Modlinger, and Saloni Potdar. 2023. The Second Workshop on Applied Machine Learning Management. In Proceedings of the 29th ACM ...

  24. [32]

    Robert B. Grady. 1992. Practical Software Metrics for Project Management and Process Improvement. Prentice-Hall, Inc., USA

  25. [33]

    Humphrey

    W.S. Humphrey. 1988. Characterizing the software process: a maturity framework. IEEE Software 5, 2 (1988), 73–79. https://doi.org/10.1109/ 52.2014

  26. [34]

    O’Reilly Media, Inc

    Chip Huyen. 2022. Designing machine learning systems. " O’Reilly Media, Inc."

  27. [35]

    ISO/IEC 9126. 2001. ISO/IEC 9126. Software engineering – Product quality. ISO/IEC

  28. [36]

    Shibo Jie, Zhi-Hong Deng, and Ziheng Li. 2022. Alleviating Representa- tional Shift for Continual Fine-tuning. arXiv:2204.10535 [cs.CV]

  29. [37]

    Meenu Mary John, Helena Holmström Olsson, and Jan Bosch. 2021. To- wards MLOps: A Framework and Maturity Model. In 2021 47th Euromicro Conference on Software Engineering and Advanced Applications (SEAA). 1–8. https://doi.org/10.1109/SEAA53835.2021.00050

  30. [38]

    Raphael Lopez Kaufman, Jegar Pitchforth, and Lukas Vermeer. 2017. De- mocratizing online controlled experiments at Booking. com. arXiv preprint arXiv:1710.08217 (2017)

  31. [39]

    Thomas Kluyver, Benjamin Ragan-Kelley, Fernando Pérez, Brian Granger, Matthias Bussonnier, Jonathan Frederic, Kyle Kelley, Jessica Hamrick, Ja- son Grout, Sylvain Corlay, Paul Ivanov, Damián Avila, Safia Abdalla, and Carol Willing. 2016. Jupyter Notebooks – a publishing format...

  32. [40]

    J.C. Knight. 2002. Safety critical systems: challenges and directions. In Proceedings of the 24th International Conference on Software Engineering. ICSE

  33. [41]

    Ron Kohavi, Alex Deng, Roger Longbotham, and Ya Xu. 2014. Seven Rules of Thumb for Web Site Experimenters. In Proceedings of the 20th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (New York, New York, USA) (KDD ’14). ACM, New York, NY, USA, 1857–18...

  34. [42]

    Ron Kohavi, Alex Deng, and Lukas Vermeer. 2022. A/b testing intuition busters: Common misunderstandings in online controlled experiments. In Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 3168–3177

  35. [43]

    Anastasiia Kornilova and Lucas Bernardi. 2021. Mining the stars: learn- ing quality ratings with user-facing explanations for vacation rentals. In Proceedings of the 14th ACM International Conference on Web Search and Data Mining. 976–983

  36. [44]

    Dominik Kreuzberger, Niklas Kühl, and Sebastian Hirschl. 2023. Machine Learning Operations (MLOps): Overview, Definition, and Architecture. IEEE Access 11 (2023), 31866–31879. https://doi.org/10.1109/ACCESS. 2023.3262138

  37. [45]

    Lucy Ellen Lwakatare, Aiswarya Raj, Ivica Crnkovic, Jan Bosch, and Helena Holmström Olsson. 2020. Large-scale machine learning sys- tems in real-world industrial settings: A review of challenges and so- lutions. Information and Software Technology 127 (2020), 106368. https: //...

  38. [46]

    Michael R. Lyu. 2007. Software Reliability Engineering: A Roadmap. In Future of Software Engineering (FOSE ’07). 153–170. https://doi.org/10. 1109/FOSE.2007.24

  39. [47]

    Mateo-Casali, Francisco Fraile, Andrés Boza, and Artem Nazarenko

    Miguel A. Mateo-Casali, Francisco Fraile, Andrés Boza, and Artem Nazarenko. 2023. Maturity Model for Analysis of Machine Learning Oper- ations in Industry. https://doi.org/10.1007/978-3-031-27915-7_57

  40. [48]

    Ninareh Mehrabi, Fred Morstatter, Nripsuta Saxena, Kristina Lerman, and Aram Galstyan. 2021. A Survey on Bias and Fairness in Machine Learning. ACM Comput. Surv. 54, 6, Article 115 (jul 2021), 35 pages. https: //doi.org/10.1145/3457607

  41. [49]

    Microsoft. 2023. How to Evaluate LLMs: A Complete Metric Framework. https://www.microsoft.com/en-us/research/group/experimentation- platform-exp/articles/how-to-evaluate-llms-a-complete-metric- framework/. [Online; accessed 17-May-2024]

  42. [50]

    Jose Miguel, David Mauricio, and Glen Rodriguez. 2014. A Review of Software Quality Models for the Evaluation of Software Products. Inter- national journal of Software Engineering & Applications 5 (11 2014), 31–54. https://doi.org/10.5121/ijsea.2014.5603

  43. [51]

    Sanjay Misra, Adewole Adewumi, Nicholas Omoregbe, and Broderick Crawford. 2016. A systematic literature review of open source software quality assessment models. SpringerPlus 2016 (11 2016), 1936. https: //doi.org/10.1186/s40064-016-3612-4

  44. [52]

    Siba N. Mohanty. 1979. Models and Measurements for Quality Assessment of Software. ACM Comput. Surv. 11, 3 (sep 1979), 251–275. https://doi. org/10.1145/356778.356783

  45. [53]

    Sasi Kumar Murakonda and Reza Shokri. 2020. ML Privacy Meter: Aiding Regulatory Compliance by Quantifying the Privacy Risks of Machine Learning. arXiv:2007.09339 [cs.CR]

  46. [54]

    Lawrence

    Andrei Paleyes, Raoul-Gabriel Urma, and Neil D. Lawrence. 2022. Chal- lenges in Deploying Machine Learning: A Survey of Case Studies. ACM Comput. Surv. 55, 6, Article 114 (dec 2022), 29 pages. https://doi.org/ 10.1145/3533378

  47. [55]

    Jiantao Pan. 1999. Software testing. Dependable Embedded Systems 5, 2006 (1999), 1

  48. [56]

    Mark Paulk, William Curtis, Mary Beth Chrissis, and Charles We- ber. 1993. Capability Maturity Model for Software (Version 1.1) . Tech- nical Report CMU/SEI-93-TR-024. https://insights.sei.cmu.edu/ library/capability-maturity-model-for-software-version-11/ Ac- cessed: 2023-Nov-30

  49. [57]

    Shachaf Poran, Gil Amsalem, Amit Beka, and Dmitri Goldenberg. 2022. With One Voice: Composing a Travel Voice Assistant from Repurposed Models. In Companion Proceedings of the Web Conference 2022. 383–387

  50. [58]

    Adarsh Prasad. 2022. Towards Robust and Resilient Machine Learning. (4 2022). https://doi.org/10.1184/R1/19552420.v1

  51. [59]

    Chana Ross, Tomer Ovadia, Jake Mooney, Amit Meitin, Eytan Kabilou, Mush Kabalo, and Dmitri Goldenberg. 2022. Democratizing Travel Per- sonalization via Central Recommendation Platform. (2022)

  52. [60]

    Raghad Baker Sadiq, Nurhizam Safie, Abdul Hadi Abd Rahman, and Shidrokh Goudarzi. 2021. Artificial intelligence maturity model: a system- atic literature review. PeerJ Computer Science 7 (2021), e661

  53. [61]

    Christopher Samiullah. 2020. Monitoring Machine Learning Models in Production. https://christophergs.com/machine%20learning/2020/03/ 14/how-to-monitor-machine-learning-models/ Accessed on December 22, 2023

  54. [62]

    Santhanam

    P . Santhanam. 2020. Quality Management of Machine Learning Systems. arXiv:2006.09529 [cs.SE]

  55. [63]

    Simone Scalabrino, Mario Linares-Vásquez, Rocco Oliveto, and Denys Poshyvanyk. 2018. A comprehensive model for code readability. Journal of Software: Evolution and Process 30 (06 2018). https://doi.org/10.1002/ smr.1958

  56. [64]

    David Sculley, Gary Holt, Daniel Golovin, Eugene Davydov, Todd Phillips, Dietmar Ebner, Vinay Chaudhary, Michael Young, Jean-Francois Crespo, and Dan Dennison. 2015. Hidden technical debt in machine learning systems. Advances in neural information processing systems 28 (2015)

  57. [65]

    Alex Serban, Koen van der Blom, Holger Hoos, and Joost Visser. 2020. Adoption and Effects of Software Engineering Best Practices in Machine Learning. In Proceedings of the 14th ACM / IEEE International Symposium on Empirical Software Engineering and Measurement (ESEM) (Bari, I...

  58. [66]

    Alex Serban, Koen van der Blom, Holger Hoos, and Joost Visser. 2022. Software engineering for machine learning. https://se-ml.github.io/

  59. [67]

    Ghorbani

    Ali Shiravi, Hadi Shiravi, Mahbod Tavallaee, and Ali A. Ghorbani. 2012. Toward developing a systematic approach to generate benchmark datasets for intrusion detection. Computers & Security 31, 3 (2012), 357–374. https: //doi.org/10.1016/j.cose.2011.12.012

  60. [68]

    Julien Siebert, Lisa Joeckel, Jens Heidrich, Koji Nakamichi, Kyoko Ohashi, Isao Namba, Rieko Yamamoto, and Mikio Aoyama. 2020. Towards Guide- lines for Assessing Qualities of Machine Learning Systems. Springer Interna- tional Publishing, 17–31. https://doi.org/10.1007/978-3-03...

  61. [69]

    Stefan Studer, Thanh Binh Bui, Christian Drescher, Alexander Hanuschkin, Ludwig Winkler, Steven Peters, and Klaus-Robert Müller. 2021. Towards CRISP-ML(Q): A Machine Learning Process Model with Quality Assur- ance Methodology. Machine Learning and Knowledge Extraction 3, 2 (20...

  62. [70]

    Symeonidis, E

    G. Symeonidis, E. Nerantzis, A. Kazakis, and G. A. Papakostas. 2022. MLOps – Definitions, Tools and Challenges. arXiv:2201.00162 [cs.LG]

  63. [71]

    Michael Veale and Frederik Zuiderveen Borgesius. 2021. Demystifying the Draft EU Artificial Intelligence Act—Analysing the good, the bad, and the unclear elements of the proposed approach. Computer Law Review International 22, 4 (2021), 97–112

  64. [72]

    Fengjun Wang, Moran Beladev, Ofri Kleinfeld, Elina Frayerman, Tal Shachar, Eran Fainman, Karen Lastmann Assaraf, Sarai Mizrachi, and Benjamin Wang. 2023. Text2Topic: Multi-Label Text Classification System for Efficient Topic Detection in User Generated Content with Zero-Shot C...

  65. [73]

    Fengjun Wang, Sarai Mizrachi, Moran Beladev, Guy Nadav, Gil Am- salem, Karen Lastmann Assaraf, and Hadas Harush Boker. 2023. MuMIC– Multimodal Embedding for Multi-Label Image Classification with Tem- pered Sigmoid. In Proceedings of the AAAI Conference on Artificial Intelligen...

  66. [74]

    Shuhei Watanabe. 2023. Tree-Structured Parzen Estimator: Understanding Its Algorithm Components and Their Roles for Better Empirical Perfor- mance. arXiv:2304.11127 [cs.LG]

  67. [75]

    Roy Wendler. 2012. The maturity of maturity model research: A systematic mapping study. Information and Software Technology 54, 12 (2012), 1317–

  68. [76]

    Wieder, J.M

    P . Wieder, J.M. Butler, W. Theilmann, and R. Yahyapour. 2011. Service Level Agreements for Cloud Computing. Springer New York. https://books. google.nl/books?id=z306GUfFL5gC

  69. [77]

    Wikipedia. 2022. List of system quality attributes — Wikipedia, The Free Encyclopedia. https://en.wikipedia.org/wiki/List_of_system_ quality_attributes. [Online; accessed 30-April-2022]

  70. [78]

    Wikipedia. 2024. Critical System — Wikipedia, The Free Encyclopedia. https://en.wikipedia.org/wiki/Critical_system. [Online; accessed 11- January-2024]

  71. [79]

    Meng Yan, Xin Xia, Xiaohong Zhang, Ling Xu, Dan Yang, and Shanping Li

  72. [80]

    Fahri Anıl Yerlikaya and ¸ Serif Bahtiyar. 2022. Data poisoning attacks against machine learning algorithms. Expert Systems with Applications 208 (2022), 118101. https://doi.org/10.1016/j.eswa.2022.118101

  73. [81]

    Xing, Hao Zhang, Joseph E

    Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P . Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. arXiv:2306.05685 [cs.CL] Maturity ...

  74. [1339]

    https://doi.org/10.1016/j.infsof.2012.07.007 Special Section on Software Reliability and Security

  75. [2002]

    Castelli, Chouliaras and Goldenberg

    547–550. Castelli, Chouliaras and Goldenberg

  76. [2019]

    https://api.semanticscholar

    Software quality assessment model: a systematic mapping study.Sci- ence China Information Sciences 62 (2019). https://api.semanticscholar. org/CorpusID:53378706

Pith tools

Reviewed August 8, 2026 · model on record in the stance chip above.