{"id":"26dbb6a7-cf6b-413d-a40a-2ac67c452023","arxiv_id":"2505.03599","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A systematic PRISMA-structured survey that taxonomizes medical image-to-mesh reconstruction into template, statistical shape, generative, and implicit models, with a meta-analysis of reported metrics.","lead":"This paper surveys deep learning methods that turn medical scans into 3D mesh models of organs, organizing the field into four method families. It is useful as a reference map for researchers building simulation-ready anatomical models for in silico medicine.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Meta-analysis pooling of heterogeneous cardiac/brain MR studies does not support the headline method ranking; a matched-dataset reanalysis is needed.","rationale":"The reader's weakest assumption, that published results are directly comparable, is exactly the load-bearing concern. I agree with the reader. The meta-analysis in §10 is the only quantitative support for the headline ranking, and it violates basic meta-analytic requirements: effect definitions are not homogeneous, metric implementations differ, and study-level variability is uncontrolled. The paper's descriptive taxonomy, loss-function and metric summaries, and dataset compilation have independent reference value, so the appropriate verdict remains CONDITIONAL: accept the survey as a reference map, but require that the comparative ranking be removed or re-derived from matched evaluations. Since the reader's verdict already captures this, no change to the verdict is needed.","tokens_in":53491,"tokens_out":3525,"duration_ms":38056,"concrete_test":"Re-run the meta-analysis on a matched subset: select only cardiac MR studies that evaluated on the same dataset (e.g., UK Biobank) with the same anatomical structure (e.g., LV myocardium) and the same evaluation code; recompute Dice and Hausdorff distance for each method on a common test set, then fit a mixed-effects model with study as a random effect and method category as moderator. If the category coefficient loses statistical significance or the ranking reorders, the §10 ranking is an artifact of pooling heterogeneous studies.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline ranking (implicit > generative > statistical shape model > template model) is stated as a key contribution in §1.3-§1.4, §10, and the Conclusion, but it is derived from the meta-analysis in §10, which pools published numbers across studies that differ in dataset, ground-truth generation, metric implementation, and evaluated substructures. Tables 8 and 9 mix rows from UK Biobank, MMWHS, ACDC, ADNI, HCP, and private data, and Figures 20-21 report per-category medians without any heterogeneity statistic, confidence interval, or significance test. For example, the cardiac Dice rows in Table 8 include whole-heart, left-atrium, LV, RV, and myocardium values from different papers, and the Hausdorff distances are computed with different mesh resolutions and correspondence schemes. Consequently, the observed ordering could reflect study difficulty or evaluation code rather than method capability. This is load-bearing because the survey explicitly frames the ranking as evidence for the superiority of implicit and generative models and uses it to recommend method choices in §11 and §13. The taxonomy and method descriptions remain useful, but the comparative claim is unsupported without controlling for these confounds.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This survey reviews deep learning-based methods for direct medical image-to-mesh reconstruction. It proposes a four-category taxonomy (template models, statistical shape models, generative models, implicit models) with twelve subcategories, organizes the literature by anatomy and modality, catalogs loss functions and evaluation metrics, curates a list of public datasets, and reports a meta-analysis of published cardiac and brain MRI results from which it derives a relative ranking: implicit model > generative model > statistical shape model > template model. The descriptive sections closely follow the cited papers and provide a broad map of the field. The comparative ranking is presented as a key contribution but rests on a meta-analysis that pools heterogeneous published numbers without controlling for dataset, substructure, ground-truth generation, or evaluation protocol.","tokens_in":53832,"tokens_out":6034,"duration_ms":56984,"significance":"If taken as a structured literature map, the paper is a useful contribution: the taxonomy is mostly clear, the loss/metric tables are broad, and the dataset summary is convenient for newcomers to the field. The paper is also honest in parts, acknowledging that no absolute superiority exists between methods in the caption of Figure 20. However, the paper's headline claim—the four-way method ranking—is not supported by the evidence as presented. Because that ranking appears in the abstract, Section 10, the discussion, and the conclusion, the central comparative claim needs substantial reworking. The survey can be valuable after either a matched-dataset reanalysis or an explicit downgrade of the ranking to a qualitative observation.","major_comments":[{"comment":"The meta-analysis does not control for the factors that make published results non-comparable. The cardiac Dice rows in Table 8 come from different datasets and different anatomical structures: MeshDeformNet and HeartFFDNet are evaluated on whole-heart CT/MR labels, Attar et al. and MCSI-Net on UK Biobank biventricular meshes, and Xu et al. on private cardiac MR contours. The table aggregates 'Myo', 'LA', 'LV', 'RA', and 'RV' cells as if they were interchangeable. The brain rows in Table 9 mix ADNI, HCP, OASIS, dHCP, and private scans, with ground truths produced by different pipelines (e.g., FreeSurfer meshes vs. manually constructed surfaces). Hausdorff and Chamfer values also depend on mesh resolution, vertex sampling, and correspondence schemes, none of which are controlled. Figures 20-21 report per-category medians without any heterogeneity statistic, confidence interval, or significance test, and the caption of Figure 20 explicitly disclaims 'no absolute superiority or inferiority between the methods.' In the absence of a matched-dataset reanalysis or at least a per-dataset stratification, the ranking stated in Section 10 and the Conclusion is not supported.","section":"Section 10, Tables 8-9, Figs 20-21"},{"comment":"The text and the figure caption contradict each other. The caption of Figure 20 says 'there is no absolute superiority or inferiority between the methods,' while Section 10 concludes with a strict ranking: implicit > generative > statistical shape model > template model. Both statements cannot stand in their current form. The paper should either present the ranking as a qualitative tendency supported by the per-method distributions, or provide a statistical model that justifies the ordering with appropriate uncertainty quantification.","section":"Section 10 and Fig. 20"},{"comment":"The taxonomy labels in Table 8 are internally inconsistent with the body text. Section 3.1 and Table 1 classify MR-Net as a conditioned deformation method, but Table 8 lists MR-Net as 'T- Registration' under both Hausdorff and Mean Distance. Additionally, Table 8 includes classical non-deep baselines (CPD, GMMREG, FFD, dDemons) inside the template-registration category, even though Section 1.3 limits the survey's scope to deep learning-based end-to-end image-to-mesh reconstruction. Both issues bias the per-category aggregation and should be corrected before the comparison is used to support any ranking.","section":"Table 8 and Section 3.1"},{"comment":"The paper claims PRISMA adherence and a 'study-based statistical approach,' but the meta-analysis reporting is incomplete. There is no description of the search strategy, inclusion/exclusion criteria, screening decisions, data extraction form, risk-of-bias assessment, or heterogeneity analysis. The cited Julian et al. [2019] is a clinical meta-analysis and does not serve as a methodological guideline for medical-image meta-analysis. If the comparison remains, it should be labeled a narrative overview rather than a meta-analysis; if the authors wish to call it a meta-analysis, the PRISMA-compliant reporting items need to be added.","section":"Sections 1.5 and 10"}],"minor_comments":[{"comment":"The sentence ending 'for advancing diagnostic and therapeutic techniques.for advanc-' is duplicated and truncated; please fix the wording.","section":"Section 1.1"},{"comment":"In the Hausdorff block, the row 'MeshDeformNet Kong and Shadden [2021]' should cite Kong et al. [2021], while the row 'HeartFFDNet Kong and Shadden [2023]' should cite Kong and Shadden [2021]; the citation years for MeshDeformNet and HeartFFDNet appear swapped.","section":"Table 8"},{"comment":"The text says 'Common types of implicit models summarized in Table 15' but the corresponding table is numbered Table 4; please correct the cross-reference.","section":"Section 6"},{"comment":"The 'Download Link' column contains no actual links; either provide URLs for each dataset or remove the column.","section":"Table 7"},{"comment":"The sentence 'Song et al. [2020] used denoising diffusion implicit models to iteratively refine reconstructed slices' appears to attribute the DDIM usage to Song et al., but in context this is part of the DMCVR pipeline (He et al. [2023]); please clarify the attribution.","section":"Section 5.4"},{"comment":"The label 'Voxe Grid' should read 'Voxel Grid'.","section":"Figure 3"},{"comment":"The in-text citation 'Maxime et al. [2022]' is inconsistent with the reference entry 'Di Folco Maxime, Moceri Pamela, Clarysse Patrick, and Duchateau Nicolas'; please harmonize the citation style.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The core risk is the comparative claim. The descriptive parts of the survey are serviceable and could support publication after revision, but the headline ranking should not appear in the abstract or conclusion unless it is backed by a matched-dataset comparison or a properly conducted meta-analysis. I would ask the authors to either perform a fair comparison on a common dataset (for example, UK Biobank for cardiac and ADNI/HCP for cortical surfaces) or to relabel Section 10 as a qualitative comparison and remove the ranking from the abstract and conclusion."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This survey will be a useful shelf item for anyone entering medical image-to-mesh reconstruction. The four-way taxonomy (template, statistical shape, generative, implicit) with twelve subcategories is a legitimate organizing contribution, and the compiled tables of losses, metrics, and public datasets are genuinely handy. The descriptive sections closely follow the cited papers and look accurate. The PRISMA framing and the explicit registration number are a plus; the paper is a real systematic review, not just a narrative stroll through the literature.\n\nThe soft spot is Section 10, and it is load-bearing. The paper promotes a headline ranking — implicit > generative > statistical shape model > template model — in the introduction, the conclusion, and the future-directions section. That ranking comes from a meta-analysis that pools published numbers across different datasets (UK Biobank, ACDC, MMWHS, ADNI, HCP, plus private data), different substructures (whole heart, LV, RV, LA, myocardium, pial vs white matter), and different metric implementations. Tables 8 and 9 mix rows in exactly that way, and Figures 20–21 show per-category medians with no heterogeneity statistics, confidence intervals, or significance tests. The paper claims in Section 10 that only studies with the same objective and modality were included to ensure comparability, but the tables themselves do not satisfy that criterion: a Dice for the left atrium from one paper and a Dice for the whole heart from another are not the same measurement. The caption of Figure 20 even concedes that “there is no absolute superiority or inferiority between the methods,” which sits awkwardly with the strong ordering asserted elsewhere. The ranking is used to recommend method choices in Sections 11 and 13, so this is not a minor caveat.\n\nWho should read this: grad students and researchers who need a structured map of methods, losses, metrics, and datasets. I would cite it for the taxonomy and the tables, not for the comparative ranking. It deserves a serious referee rather than a desk reject, because the field needs this kind of organized reference, but the meta-analysis should be redone on matched datasets with consistent evaluation protocols, or at minimum be rewritten as a descriptive summary with the ranking removed or heavily qualified.","headline":"A useful reference map for medical image-to-mesh reconstruction, but the headline method ranking rests on a meta-analysis that mixes incomparable studies and should be heavily revised before the paper is used as evidence.","tokens_in":54239,"tokens_out":2394,"would_cite":true,"duration_ms":25159,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This systematic review argues that deep learning-based medical image-to-mesh reconstruction splits into four method families, and that published cardiac and brain MRI results rank implicit models first, followed by generative models…","keywords":["image-to-mesh reconstruction","deep learning","mesh generation","template deformation","statistical shape model","generative model","implicit neural representation","medical imaging survey"],"falsifier":"Run template, statistical, generative, and implicit methods on one public cardiac MRI dataset and one cortical MRI dataset with identical training and test splits, ground-truth meshes, and metric code; if template or statistical models match or beat implicit/generative models on Dice and surface distance, the survey's ranking fails.","tokens_in":53309,"feed_emoji":"🧠","tokens_out":8286,"duration_ms":76338,"temperature":0.7,"pith_summary":"This survey organizes the fast-moving field of deep learning-based medical image-to-mesh reconstruction, where CT, MR, or ultrasound images are turned directly into the 3D meshes needed for simulation and in silico trials. The authors argue that all end-to-end approaches fall into four families—template deformation, statistical shape models, generative models, and implicit models—further divided into twelve subcategories by pipeline and feature representation. They also map the loss functions, evaluation metrics, and public datasets used across the field, and pool published cardiac MRI and brain MRI results into a meta-analysis. The paper's comparative claim is a relative ranking: implicit models first, then generative models, then statistical shape models, then template models, based on reported Dice, Hausdorff, Chamfer, and surface-distance numbers. The claim matters because it gives practitioners a structured way to choose a reconstruction approach and tells the field where the published evidence currently points.","feed_headline":"Implicit models top a four-way ranking for medical mesh reconstruction","feed_subtitle":"Review of brain and cardiac MRI studies ranks four deep-learning mesh approaches, with implicit surfaces on top.","key_machinery":"The load-bearing object is the four-part taxonomy, with each family defined by how the surface is represented and produced. Template models use a fixed initial mesh refined by learned vertex displacements; statistical shape models compress shape variation into a linear (PCA) or non-linear latent space; generative models build shapes from learned distributions, including point-cloud completion; implicit models represent the surface as the zero level set of a learned function, typically a signed distance function, occupancy probability, neural ODE flow, or volumetric density field, and convert it to a mesh by isosurface extraction. The taxonomy does the organizing work of the survey, turning dozens of papers into comparable families. The comparative ranking is carried by a meta-analysis that pools published results by anatomy and imaging modality—cardiac MRI and cortical MRI—across standard metrics such as Dice similarity, Hausdorff distance, Chamfer distance, mean distance, and average symmetric surface distance, with the stated assumption that studies sharing anatomy and modality are comparable by clinical standards.","core_discovery":"On the paper's own terms, the discovery is that a scattered set of reconstruction techniques can be organized into four coherent paradigms with distinct strengths. Template models deform a hand-built mesh under image guidance; statistical shape models project images into low-dimensional shape spaces; generative models synthesize point clouds or meshes from image latents through VAEs, GANs, completion networks, or diffusion; implicit models learn continuous fields such as signed distance, occupancy, neural ODE flows, or radiance fields, then extract the mesh as an isosurface. The meta-analysis covers cardiac MRI with template, statistical, and generative families and cortical MRI with template, generative, and implicit families, using common metrics. The reported numbers place implicit models at the low-error end for cortex, generative models ahead on cardiac Dice and Hausdorff distance, and template models trailing on these tasks; joining the two comparisons yields the overall ranking implicit > generative > statistical shape > template. The authors explicitly call this ranking relative rather than absolute, since particular methods can win on specific anatomies, data qualities, or tasks.","pith_inferences":["Because the meta-analysis pools numbers from studies with different datasets, ground-truth definitions, and preprocessing, the ranking should be read as a statement about current reporting rather than a controlled comparison; a single benchmark with identical splits could reorder the middle of the ranking.","The same logic that favors implicit models on cortex suggests they are untested candidates for cardiac mesh reconstruction, where the meta-analysis had no implicit entries; applying SDF or neural-ODE methods to a public cardiac cohort would directly extend the ranking.","The taxonomy implies that template and statistical models may remain competitive in low-data regimes because their priors encode anatomy explicitly, a consequence the paper's discussion supports but its meta-analysis does not test.","A useful next experiment would be to use the paper's loss-and-metric classification as a reporting checklist and run one method from each family on the same anatomy, since the face validity of the ranking depends more on such a study than on further pooling."],"forward_implications":["On this evidence, a practitioner starting a cardiac or cortical MRI reconstruction task would look first at implicit or generative models for raw geometric accuracy, and at template or statistical models when fixed topology, stability, or small training sets are priorities.","The twelve-subcategory taxonomy gives the field a shared vocabulary, so new methods can be positioned by pipeline and output representation rather than by name alone.","The loss and metric classification implies that reported accuracy depends on metric choice and regularization as much as on architecture, so future comparisons should state which representation each number refers to.","The survey's own future-directions section predicts that new continuous representations such as Gaussian splatting and multi-modal fusion will push mesh fidelity and efficiency further.","The meta-analysis also exposes a gap: no diffusion-based end-to-end image-to-mesh pipeline exists yet, so this family is currently used for image enhancement or data augmentation before meshing."],"supporting_citations":[{"why":"Supplies whole-heart template-deformation Dice and Hausdorff numbers used in the cardiac meta-analysis.","marker":"Kong et al. [2021]"},{"why":"Provides HeartFFDNet template-deformation results that anchor the cardiac comparison.","marker":"Kong and Shadden [2021]"},{"why":"Provides linear statistical shape model cardiac reconstruction results used in the ranking.","marker":"Attar et al. [2019a]"},{"why":"Provides MCSI-Net linear statistical shape model results across cardiac substructures in the meta-analysis.","marker":"Xia et al. [2022]"},{"why":"Provides MV-HybridVNet variational autoencoder generative results on cardiac MRI metrics.","marker":"Gaggion et al. [2024]"},{"why":"Provides interpolation-based generative cardiac results in the meta-analysis.","marker":"Xu et al. [2019]"},{"why":"Provides DeepCSR implicit SDF and occupancy results on cortical MRI metrics.","marker":"Cruz et al. [2021]"},{"why":"Provides CortexODE neural ODE results on cortical surface distances in the meta-analysis.","marker":"Ma et al. [2023]"},{"why":"Provides Vox2Cortex template-deformation results on cortical MRI metrics.","marker":"Bongratz et al. [2022]"},{"why":"Provides TopoFit template-deformation cortical results used in the comparison.","marker":"Hoopes et al. [2022]"}],"fun_headline_variants":["Implicit edges out generative in four-way mesh ranking","Implicit wins cortex, generative wins cardiac in mesh survey","Implicit models take top overall in mesh reconstruction survey","Medical mesh survey ranks implicit first among four AI approaches"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The meta-analysis treats published numbers from different studies as comparable whenever the anatomy and imaging modality match, so the ranking could reflect dataset splits, ground-truth construction, or metric definitions rather than true method differences.","fun_headline_variants_meta":{"raw":{"variants":["Implicit edges out generative in four-way mesh ranking","Implicit wins cortex, generative wins cardiac in mesh survey","Implicit models take top overall in mesh reconstruction survey","Medical mesh survey ranks implicit first among four AI approaches"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001232,"raw_usage":{"total_tokens":5063,"prompt_tokens":947,"completion_tokens":4116,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":563,"completion_tokens_details":{"reasoning_tokens":4052}},"tokens_in":563,"tokens_out":4116,"duration_ms":30574,"temperature":1.0,"reasoning_tokens":4052,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T23:46:27.288146+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run template, statistical, generative, and implicit methods on one public cardiac MRI dataset and one cortical MRI dataset with identical training and test splits, ground-truth meshes, and metric code; if template or statistical models match or beat implicit/generative models on Dice and surface distance, the survey's ranking fails.","supporting_citations":[],"review_version":1}