{"id":"a5e481ed-d9be-4889-a338-082ce61e5659","arxiv_id":"2507.19489","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"MAIA is a modular open-source platform for collaborative medical AI development that packages Kubernetes, MONAI, and clinical imaging tools into one workspace.","lead":"This paper presents MAIA, an open-source Kubernetes-based platform that bundles data management, model training, annotation, and clinical deployment tools for medical AI projects. It is a systems paper that reports deployments at KTH and Karolinska University Hospital, with two medical imaging case studies.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Clinical-deployment claim rests on a single unmeasured pipeline; Section 6's own limitation statement admits this.","rationale":"The reader's weakest assumption correctly identifies the clinical-deployment validation gap as the load-bearing uncertainty. My stress-test finds no independent fatal flaw: the open-source repository, the described HPC integrations, and the two case studies are genuine supporting evidence for the platform's development-lifecycle claims. The specific weak point is the MONAI Deploy/PACS loop in §5.3, which is described at architectural level and asserted as performed, but without measurements or failure analysis. Section 6's own limitation statement confirms this is not yet systematically validated. A concrete integration test with real-world DICOM variability would settle whether the deployment claim holds. Since the reader already issued a conditional verdict and this concern does not invalidate the paper, \"UNCHANGED\" is the appropriate outcome.","tokens_in":12768,"tokens_out":2767,"duration_ms":29346,"concrete_test":"Run an end-to-end test from a DICOM-emitting PACS simulator (or staging PACS) sending a diverse set of real-world series (varied modalities, orientations, slice thickness, plus at least one intentionally non-conformant DICOM study) into MAIA's Orthanc, with MONAI Deploy auto-triggering on each incoming study. Record success rate, per-study latency, and whether the generated DICOM SEG is accepted and viewable in the downstream PACS. Repeat for at least 100 studies. If success rate is not near 100% with documented failure modes, the clinical-deployment claim should be scaled back to \"demonstrated in a controlled test\" rather than \"supports clinical workflow deployment.\"","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that MAIA supports the full AI lifecycle, including \"deployment in clinical workflows\" (Abstract; §5.3), depends on the RADIANCE → Orthanc → MONAI Deploy → PACS path functioning under real clinical constraints. Section 5.3 reports this \"was performed for the brain metastases project,\" but gives no data on runtime, failure rates, data governance, or how the DICOM SEG output is accepted by the clinical PACS. Section 6 explicitly states \"current clinical testing is limited, and MAIA must further demonstrate its scalability and adaptability across diverse settings and workflows.\" Because this is a platform paper, an architecture diagram alone cannot establish the deployment claim. Without evidence that the MONAI Deploy pipeline reliably ingests heterogeneous DICOM from a live PACS and returns usable results, the statement \"supports real-world use cases... in clinical environments\" is stronger than the presented evidence supports.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces MAIA, an open-source Kubernetes-based platform for collaborative medical AI development and deployment. The architecture is built around namespaces that host a suite of tools for data management (MinIO, Orthanc, OHIF), annotation (Label Studio, MONAI Label), model development (JupyterHub, MLFlow, Kubeflow), and clinical integration (MONAI Deploy, DICOM SEG back to PACS). The paper describes workflows for AI development, clinical data import via the RADIANCE pipeline at Karolinska, and active learning, and reports deployments at KTH and Karolinska University Hospital. Two case studies are presented (vertebral metastasis segmentation in CT and brain metastasis segmentation in MRI), followed by a brief report of a clinical deployment pipeline. The paper concludes with an acknowledgement of limited clinical testing and an open call for collaboration.","tokens_in":12890,"tokens_out":4658,"duration_ms":48671,"significance":"If the platform operates as described, MAIA could be a valuable open-source infrastructure for multidisciplinary medical AI research and for bridging the gap between research and clinical workflows. The manuscript's strengths include the open availability of code and documentation, a coherent modular architecture built on standard Kubernetes tooling, integration with national HPC resources, and evidence of real deployments in both academic and clinical settings. The case studies, while anecdotal, indicate that the platform has been used in genuine research and clinical environments. However, the significance is limited by the lack of quantitative evaluation of the platform's reliability and scalability, and by the thin evidence supporting the clinical deployment claim, which is load-bearing for the paper's abstract and title.","major_comments":[{"comment":"The abstract's claim that MAIA \"supports real-world use cases... in clinical environments\" is not supported by systematic evidence. The only clinical deployment evidence is the brain-metastasis pipeline description in §5.3, which reports no measurements of runtime, success rate, data volume, or verification that the DICOM SEG output was accepted by the clinical PACS. Section 6 explicitly concedes that \"current clinical testing is limited.\" As a platform paper, this gap is load-bearing: the deployment claim is one of the central contributions. The authors should either add operational metrics (for example, processing times, failure rates, transfer volumes, PACS interoperability tests) or temper the claims to describe a preliminary integration rather than routine clinical support.","section":"§5.3, §6, Abstract"},{"comment":"The reported sensitivity increase from 55% to 82% after expanding the dataset from 30 to 150 cases is the only quantitative evidence that MAIA's active-learning workflow improves model performance, but no confidence intervals, dataset splits, cross-validation, or statistical tests are provided. Without these details, the \"substantial improvement\" claim cannot be evaluated, and the improvement may be attributable to changes in dataset composition or annotation protocol rather than the MAIA workflow itself. Please provide a full experimental setup, including evaluation on a held-out test set and measures of variability.","section":"§5.1"},{"comment":"The brain metastasis case study reports no segmentation performance numbers (for example, Dice coefficient, Hausdorff distance, or sensitivity/specificity) on any held-out set; the text only states that the model was developed and that the deployment pipeline was executed. Given that this case is used to demonstrate MAIA's support for the full development and deployment lifecycle, the absence of any quantitative outcome makes it impossible to assess whether the workflow produces usable models. I recommend including at least a summary of validation results or explicitly stating that quantitative validation is outside the scope and directing the reader to a future report.","section":"§5.2"},{"comment":"The related-work comparison makes several strong claims about Kaapana (for example, \"lacks robust project isolation,\" \"restricted deployment flexibility,\" \"absence of native integration with CI/CD\") without providing evidence or citations. While these limitations may be accurate, the paper should either support them with references or soften the language to indicate that these are the authors' assessments; otherwise, the comparison appears anecdotal and undermines the reproducibility of the positioning argument.","section":"§1.1"}],"minor_comments":[{"comment":"In the sentence \"MAIA's architecture incudes two internal layers,\" the word \"incudes\" should be \"includes.\"","section":"§2.1.5"},{"comment":"The bullet for Prometheus is formatted as \"Prometheus 38 A monitoring system...\" and should read \"Prometheus: A monitoring system...\" or end the fragment with a period for consistency with other list items.","section":"§2.1.6"},{"comment":"Step 5 states that \"Figure 7 illustrates the core components of the described workflow,\" but Figure 7 shows the generic AI development workflow; the active-learning workflow described in this case is more accurately represented by Figures 8 or 9. Please correct the cross-reference.","section":"§5.1, Step 5"},{"comment":"The text refers to \"BraTS-METs 2025,\" while the cited reference [12] describes the BraTS-METS 2023 challenge. Please align the terminology and ensure the correct year is used in both the text and the reference.","section":"§5.2"},{"comment":"The description of XNAT as \"a tool designed to integrate AI-based applications into clinical deployment scenarios\" is imprecise; XNAT is primarily an imaging data management and sharing platform. Please rephrase to reflect its actual role in the workflow.","section":"§2.1.3"}],"recommendation":"major_revision","confidential_remarks":"This manuscript is a systems/platform description rather than a clinical validation study. The central architectural claims are plausible and the open-source availability is a strength, but the evidence for the clinical deployment claim is thin, and the two case studies need more rigorous quantitative reporting or a clearly stated scope limitation. I would encourage the editor to ask the authors to add operational metrics for the deployment pipeline and to either strengthen or temper the abstract's claims accordingly. The paper may also benefit from a clearer statement of what is not yet validated, beyond the brief acknowledgement in Section 6."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read the paper if you care about medical AI infrastructure. MAIA is not vaporware: the GitHub repo, docs, and online instance exist, and the architecture description is concrete. The genuinely new bits are the federation-of-clusters design, the MAIA-HPC module wrapping NAISS supercomputers, and the GPU booking system. Those go beyond just gluing together MLflow, Orthanc, OHIF, Kubeflow, and MONAI Deploy.\n\nWhat the paper does well is show how components fit together, and the two case studies are real clinical pilots: vertebral metastasis segmentation with active learning, and brain metastasis segmentation on BraTS-METS data. The 55% to 82% sensitivity jump is interesting, but it is not evidence of platform value—it likely reflects more training data. No confidence intervals, no dataset details, no cross-validation. That is a soft spot but not disqualifying for a systems paper.\n\nThe bigger issue is the clinical-deployment claim. Section 5.3 says the RADIANCE -> Orthanc -> MONAI Deploy -> PACS pipeline was performed for the brain mets project, but gives no runtime, failure rates, or how the DICOM SEG is accepted by the clinical PACS. Section 6 explicitly concedes clinical testing is limited. The abstract says MAIA 'supports real-world use cases... in clinical environments,' which is stronger than the evidence. That needs to be toned down or substantiated.\n\nAlso check reference [15]. It is cited for Kaapana, but the title is 'Joint Imaging Platform for Federated Clinical Data Analytics.' That might be a different system. The authors should correct the citation or clarify.\n\nThe central architecture story holds up. The platform is plausible, and the federation and HPC integration appear genuinely useful. But this is a systems report, not a clinical validation. For a reader adopting or building similar infrastructure, it is worth reading. For a clinical reader, it will not convince.\n\nRecommendation: send to peer review. The engineering is real, the open-source artifact is reproducible, and the problems it addresses are important. The review should press for either concrete deployment metrics or a revised abstract that does not overclaim. Conditional accept, not rejection.","headline":"A real open-source medical AI platform that deserves refereeing, but the paper overstates clinical deployment evidence it does not provide.","tokens_in":13428,"tokens_out":3831,"would_cite":false,"duration_ms":39384,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"MAIA is an open-source, Kubernetes-based platform that aims to cover the full medical-AI lifecycle—from cohort selection and annotation through HPC training and active learning to deployment back into the clinical PACS.","keywords":["medical AI platform","Kubernetes","open source","medical imaging","clinical workflow integration","active learning","CI/CD","HPC integration"],"falsifier":"Deploy MAIA at a second hospital with a different PACS vendor and a different data-governance regime; the central claim fails if exporting a pseudonymized imaging cohort or returning DICOM SEG results to the clinical system requires re-engineering rather than configuring the existing RADIANCE/MONAI Deploy pathway.","tokens_in":12579,"feed_emoji":"🩻","tokens_out":6410,"duration_ms":64446,"temperature":0.7,"pith_summary":"The paper introduces MAIA, an open-source platform built on Kubernetes that is designed to let clinicians, researchers, and AI developers collaborate on medical-imaging AI projects inside one isolated, modular environment. The central claim is that MAIA can support the full AI lifecycle in a way that accelerates translation into clinical practice, by integrating data management, annotation, model training, experiment tracking, high-performance computing, and deployment into hospital workflows. The authors describe the architecture, the workflows it enables, and two real hospital use cases in which segmentation models for vertebral and brain metastases were developed, refined by radiologist feedback, and deployed. A sympathetic reader would take the contribution to be a reproducible infrastructure blueprint plus evidence that it functions in academic and clinical settings, not a claim that any single model outperforms clinical baselines.","feed_headline":"MAIA: one open-source platform for medical AI from lab to clinic","feed_subtitle":"Kubernetes-based and modular, it covers annotation, training, HPC jobs, and DICOM deployment in one isolated workflow.","key_machinery":"The carrying mechanism is the MAIA Namespace, a per-project Kubernetes namespace that bundles a complete, standard toolchain: the MAIA Workspace (JupyterHub with JupyterLab, remote desktop, and SSH), MLflow for experiment tracking, MinIO for file storage, Orthanc with the OHIF viewer for DICOM handling, Kubeflow for pipelines, XNAT and MONAI Deploy for clinical integration, and Label Studio for annotation. Running above this are MAIA Core (ArgoCD, monitoring, networking, GPU operator) and MAIA Admin (dashboard, Harbor, Keycloak, Rancher), plus the MAIA-HPC and GPU-booking submodules. The design pattern is to compose mature open-source components behind a lightweight control plane so that each project is isolated yet shares the same workflow surface; the workflows then reduce to choosing which modules to connect.","core_discovery":"The discovery the paper asserts is that a single open-source platform can serve as the connective tissue for medical-AI work: project-scoped Kubernetes namespaces bundled with a JupyterHub workspace, MLflow, MinIO, Orthanc/OHIF, Kubeflow, XNAT, MONAI Label, MONAI Deploy, and Label Studio, all synchronized by ArgoCD and federated across clusters. MAIA adds connectors—a pseudonymized clinical export workflow, an HPC bridge, and a GPU booking system—so that data can flow from the hospital PACS into an isolated project, models can be trained locally or on supercomputers, radiologists can correct predictions in their usual viewers, and final results return to the PACS as DICOM SEG. The two case studies (a vertebral-metastasis CT segmentation model whose sensitivity rose from 55% to 82% as the annotated dataset grew from 30 to 150 cases, and a brain-metastasis MRI segmentation model trained on BraTS data and run through a MONAI Deploy pipeline) are presented as evidence that the platform works in real clinical environments.","pith_inferences":["If the platform generalizes, other hospitals could adopt the same component stack, but the hard part is likely the clinic-specific data-governance workflow (the paper's RADIANCE-style export); that piece may need local reimplementation per site.","The active-learning loop is described with future automatic retraining; if that automation matures, deployed models would improve with every radiologist review, turning static deployment into continuous learning.","A direct test of the 'accelerated translation' claim would be to measure time from data request to deployed model at a second site; the paper does not yet provide such comparative evidence.","The sustainability of an open-source clinical platform depends on continued institutional commitment; the paper itself flags community support and broader clinical testing as open challenges."],"forward_implications":["If MAIA works as claimed, a hospital can run the whole AI lifecycle inside its own firewall, with pseudonymized data exported from clinical PACS into project-isolated namespaces.","Radiologists can iteratively refine models via active learning directly in familiar viewers, reducing annotation effort as models improve.","Researchers can offload heavy training to HPC systems without leaving the platform, and GPUs are time-shared via booking to raise utilization.","New projects can be spun up with standard tools and CI/CD updates, shortening student and researcher onboarding and enabling reproducible experiments.","Deployed models can return results as DICOM SEG to the clinical PACS, making AI output visible to downstream clinical tools."],"supporting_citations":[{"why":"Supplies the MONAI Label and MONAI Deploy frameworks that carry the active-learning and clinical-deployment workflows in MAIA.","marker":"[4]"},{"why":"Prior open-source platform whose limitations (isolation, Kubernetes flexibility, CI/CD, HPC, MONAI integration) define the gap MAIA targets.","marker":"[15]"},{"why":"Related medical-imaging AI platform with MONAI integration and a use case; its closed, single-use-case nature motivates MAIA's open, modular design.","marker":"[6]"},{"why":"Provides the publicly available benchmark dataset family used for the brain-metastasis segmentation model training.","marker":"[11]"},{"why":"Defines the brain-metastasis segmentation task and dataset used in the second case study.","marker":"[12]"},{"why":"Supplies clinical motivation and systematic-review context for malignant bone-lesion segmentation, framing the first case study.","marker":"[14]"}],"fun_headline_variants":["MAIA: open-source glue for clinical AI, from PACS to prediction","MAIA: one platform to run medical AI from research to radiology","One open platform for medical AI: from annotation to DICOM","MAIA: Kubernetes-based platform linking clinicians and AI teams","MAIA: one open-source platform for clinical AI collaboration"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim rests on the assumption that the two hospital case studies and the described PACS/HPC integrations are representative enough to show MAIA generalizes to other clinical workflows and institutions.","fun_headline_variants_meta":{"raw":{"variants":["MAIA: open-source glue for clinical AI, from PACS to prediction","MAIA: one platform to run medical AI from research to radiology","One open platform for medical AI: from annotation to DICOM","MAIA: Kubernetes-based platform linking clinicians and AI teams","MAIA: one open-source platform for clinical AI collaboration"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000692,"raw_usage":{"total_tokens":3129,"prompt_tokens":938,"completion_tokens":2191,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":554,"completion_tokens_details":{"reasoning_tokens":2111}},"tokens_in":554,"tokens_out":2191,"duration_ms":16159,"temperature":1.0,"reasoning_tokens":2111,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T13:17:06.996032+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Deploy MAIA at a second hospital with a different PACS vendor and a different data-governance regime; the central claim fails if exporting a pseudonymized imaging cohort or returning DICOM SEG results to the clinical system requires re-engineering rather than configuring the existing RADIANCE/MONAI Deploy pathway.","supporting_citations":[{"cited_title":"Joint Imaging Platform for Federated Clinical Data Analytics","cited_arxiv_id":null,"evidence_quote":"Prior open-source platform whose limitations (isolation, Kubernetes flexibility, CI/CD, HPC, MONAI integration) define the gap MAIA targets."},{"cited_title":"Integrating Artificial Intelligence Tools in the Clinical Research Setting: The Ovarian Cancer Use Case","cited_arxiv_id":null,"evidence_quote":"Related medical-imaging AI platform with MONAI integration and a use case; its closed, single-use-case nature motivates MAIA's open, modular design."},{"cited_title":"Moawad et al","cited_arxiv_id":null,"evidence_quote":"Defines the brain-metastasis segmentation task and dataset used in the second case study."},{"cited_title":"Deep learning image segmentation approaches for malignant bone lesions: a systematic review and meta-analysis","cited_arxiv_id":null,"evidence_quote":"Supplies clinical motivation and systematic-review context for malignant bone-lesion segmentation, framing the first case study."}],"review_version":1}