Pith. sign in

REVIEW 2 cited by

Test & Evaluation Best Practices for Machine Learning-Enabled Systems

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.06800 v1 pith:Q4VGRPE2 submitted 2023-10-10 cs.SE cs.LG

classification cs.SEcs.LG
keywords ml-enabledsoftwaresystemslifecyclesystemcomponentacrosspractices
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Machine learning (ML) - based software systems are rapidly gaining adoption across various domains, making it increasingly essential to ensure they perform as intended. This report presents best practices for the Test and Evaluation (T&E) of ML-enabled software systems across its lifecycle. We categorize the lifecycle of ML-enabled software systems into three stages: component, integration and deployment, and post-deployment. At the component level, the primary objective is to test and evaluate the ML model as a standalone component. Next, in the integration and deployment stage, the goal is to evaluate an integrated ML-enabled system consisting of both ML and non-ML components. Finally, once the ML-enabled software system is deployed and operationalized, the T&E objective is to ensure the system performs as intended. Maintenance activities for ML-enabled software systems span the lifecycle and involve maintaining various assets of ML-enabled software systems. Given its unique characteristics, the T&E of ML-enabled software systems is challenging. While significant research has been reported on T&E at the component level, limited work is reported on T&E in the remaining two stages. Furthermore, in many cases, there is a lack of systematic T&E strategies throughout the ML-enabled system's lifecycle. This leads practitioners to resort to ad-hoc T&E practices, which can undermine user confidence in the reliability of ML-enabled software systems. New systematic testing approaches, adequacy measurements, and metrics are required to address the T&E challenges across all stages of the ML-enabled system lifecycle.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. From Tea Leaves to System Maps: A Survey and Framework on Context-aware Machine Learning Monitoring

    cs.SE 2025-06 conditional novelty 6.0 of 10

    A systematic review of 94 studies proposes C-SAR, a three-dimensional framework (System, Aspect, Representation) describing how contextual information is used in ML monitoring, with 20 recurring patterns.

  2. Approach Towards Semi-Automated Certification for Low Criticality ML-Enabled Airborne Applications

    cs.SE 2025-01 reject novelty 5.0 of 10

    A semi-automated certification framework for DO-178C Level D ML systems is demonstrated on a YOLOv8 vehicle detector, producing a Moderate Assurance score of 74.7.

Pith tools