Pith. sign in

REVIEW 3 major objections 5 minor 53 references

Treating fine-grained ear CT segmentation as a short latent anatomical evolution, guided by hierarchical add/remove actions, cuts boundary error by more than 43%.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-31 05:48 UTC pith:5L6TFNLF

load-bearing objection Solid niche methods paper with a real nested-label ear CT dataset; the ~43% HD95 headline is real on their split but over-attributed to multi-step world-model reasoning. the 3 major comments →

arxiv 2607.28487 v1 pith:5L6TFNLF submitted 2026-07-30 cs.CV

AuricularWorld: Hierarchical Action-Guided World Modeling for Fine-Grained Auricular Structure Segmentation from CT Scans

classification cs.CV
keywords auricular segmentationworld modelmulti-label segmentationCT imagelatent dynamicshierarchical actionsRSSMcartilage
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Fine ear anatomy in CT is hard to segment: the ear is tiny in the volume, cartilage edges are thin and irregular, contrast with soft tissue is weak, and clinical labels nest skin-covered subunits with their cartilage-only counterparts. This paper argues that a single feed-forward map from image to mask is the wrong formulation for that problem. Instead it embeds a small recurrent state-space model in the middle of a standard encoder–decoder, fuses multi-scale features into an anatomical observation, and runs a three-step latent rollout in which hierarchical add and remove actions progressively correct the intermediate representation before high-resolution decoding. A balanced, foreground-masked action loss teaches those transitions under sparse foreground, missing groups, and add/remove imbalance. On a dedicated multi-label auricular CT benchmark the method raises overlap modestly and roughly halves the worst-boundary error relative to a strong convolutional baseline, especially on small, complex, and cartilaginous structures.

Core claim

Fine-grained auricular CT segmentation is better cast as an action-conditioned latent anatomical evolution than as one-shot image-to-mask prediction: a three-step RSSM rollout inside the feature space, driven by hierarchical add/remove actions and returned to a high-resolution decoder, consistently improves overlap and cuts HD95 by about 43% on small, irregular, nested ear structures.

What carries the argument

AuricularWorld: multi-scale observation fusion into a recurrent state-space model that performs a three-step action-conditioned latent rollout, supervised by a foreground-masked balanced hierarchical action objective (atomic/canonical/global add and remove maps).

Load-bearing premise

That a short, prior-mean latent rollout at one intermediate resolution, steered only by learned hierarchical add/remove maps, is doing genuine structural correction rather than mainly adding capacity and heavily reweighted auxiliary losses on a single-center cohort.

What would settle it

An ablation that keeps the same extra capacity, auxiliary heads, and loss weights but replaces the action-conditioned RSSM transitions with non-action or scrambled action updates should erase most of the HD95 drop on thin cartilage and complex-boundary structures; if the gain remains, the central mechanism claim fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Nested skin–cartilage ear labels can be trained as exclusive atomic classes and recovered after inference without sacrificing hierarchical structure.
  • Boundary error on thin, low-contrast cartilage subunits becomes the main measurable win from short latent refinement, not just volumetric Dice.
  • Feed-forward medical segmenters can host a small latent world-model block without discarding skip connections or full-resolution decoding.
  • A publicly released fine-grained auricular CT set with complementary skin and cartilage annotations becomes a dedicated benchmark for this task.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The same hierarchical add/remove latent rollout may transfer to other nested small-structure problems (e.g., vessel walls, airway subunits, costal cartilage) where feed-forward nets leak across weak interfaces.
  • If the action maps stay spatially localized on initial errors across datasets, they could serve as an interpretable audit trail of what the network is ‘correcting’ at each step.
  • Multi-center or non-Philips CT would be the natural next stress test: the large HD95 gain is most informative if it survives scanner and protocol shift.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper introduces AuricularWorld, an nnU-Net-based framework that inserts a three-step action-conditioned RSSM into an intermediate decoder latent space for fine-grained multi-label auricular CT segmentation. Multi-scale encoder and partial-decoder features form an observation that initializes latent dynamics; hierarchical add/remove actions over 73 anatomical groups update a ConvGRU state, and the refined latent is decoded with preserved skip connections. A foreground-masked balanced action loss addresses sparsity, missing groups, and add/remove imbalance. The authors also release a dedicated single-center dataset (193 patients, 35 atomic labels recomposed into nested canonical structures). On a 32-case test set, AuricularWorld reports Dice 78.42±0.09% and HD95 1.379±0.014 mm versus nnU-Net 77.75% / 2.426 mm (~43% relative HD95 reduction), with best mean Dice on 31/35 structures.

Significance. Fine-grained nested auricular CT segmentation is clinically relevant for reconstructive planning and is underexplored relative to major-organ benchmarks. The complementary skin-covered and cartilage annotations, atomic-to-canonical evaluation protocol, and public-release commitment are genuine contributions. Casting segmentation as short latent anatomical evolution with explicit hierarchical actions is a coherent alternative to pure feed-forward prediction, and the experimental package (patient-level splits, three seeds, structure-wise and category-wise tables, incremental ablations, action rollout visualizations, hyperparameter sweeps) is more careful than typical medical segmentation papers. If the HD95 gains can be shown to arise from multi-step latent correction rather than extra capacity and intermediate supervision on a few unstable structures, the work would be a solid methods-plus-benchmark contribution for challenging small-structure segmentation.

major comments (3)
  1. [§5.6.1–5.6.2, Tables 2–5] The central attribution of the ~43% HD95 reduction to three-step action-conditioned world-model reasoning is under-supported by the ablations and structure-wise tables. Table 3 shows MS-RSSM alone moves HD95 from 2.426 to 1.406 mm (essentially the full gain); balanced weighting then worsens HD95 to 1.537 before FG recovers 1.379. Tables 4–5 show macro HD95 is dominated by a handful of baseline-unstable structures (External Auditory Canal 5.999±2.075→1.181; Intertragic Notch 5.804±4.301→1.388; Ear Lobule 4.618±2.805→1.364; Auricular Cartilage 3.305±2.167→0.791), while overall Dice rises only +0.67 pp. Please report median HD95, structure-wise win rates excluding empty/catastrophic cases, and a breakdown of how much of the macro gain comes from the top-k worst baseline structures, so the headline claim is not driven by seed-unstable surface outliers.
  2. [§4.3, §5.4, Table 3] There is no capacity-matched non-recurrent control that retains multi-scale observation fusion, three auxiliary segmentation heads, and the hierarchical action losses, but replaces the ConvGRU prior/posterior rollout (Eqs. 5–7) with feed-forward residual blocks of comparable parameter/FLOP count. Without this control, an equally plausible explanation of Table 3 is that extra intermediate supervision and multi-scale fusion stabilize rare large surface errors, rather than that iterative latent dynamics perform anatomical reasoning. This control is load-bearing for the paper’s methodological claim in §1 and §4.3.
  3. [§3.2, §4.4.2, Abstract, §6] Evaluation is confined to n=32 test cases from a single center and two Philips scanners (§3.2, Table 1). Given the large seed variance of baseline HD95 (Table 2: nnU-Net 2.426±1.014) and the many free design choices in the action objective (hierarchy priors 0.6/0.3/0.1, λ_add=2, ρ_bg=0.02, absence/patch coefficients; §4.4.2), external validity of the 43% HD95 claim is limited. At minimum, add a leave-one-scanner or bootstrap confidence analysis on HD95, and discuss generalization risk explicitly in the conclusion. If multi-center data are unavailable, temper the abstract/conclusion language that presents latent world-model reasoning as broadly validated for challenging medical segmentation.
minor comments (5)
  1. [Table 1, §3.2, §5.1] Dataset counts are inconsistent across sections: Abstract/§3 report 193 patients and 198 auricles; Table 1 lists 185 patients / 194 auricles; §5.1 states 202 volumes with 121/41/32 split. Please reconcile and state exclusion criteria once.
  2. [Figure 2, Table 2] Figure 2 omits SwinUNETR and UNETR despite their inclusion in Table 2; either show them or note why qualitative panels are restricted.
  3. [Eq. (14), §4.4.2, Figure 5] Eq. (14) and §5.3 give loss weights 1.0/0.2/0.1/0.2, while Figure 5 sweeps λ_act and λ_reg; briefly state whether λ_reg in Fig. 5 is the same as ρ_bg=0.02 in §4.4.2 to avoid notation drift.
  4. [References] Several related-work citations have future-dated years (e.g., 2025–2026 venues). Verify bibliographic metadata before camera-ready.
  5. [§4.3, §4.5] Clarify inference path once in the main text: posterior mean only at t=0, prior means for t=1..T (Eq. 7), no GT or action targets at test time—this is stated but easy to miss given the training-time action construction.

Circularity Check

0 steps flagged

No significant circularity: held-out supervised segmentation with standard teacher-style action targets; inference uses no GT.

full rationale

AuricularWorld is a conventional supervised encoder–decoder segmentation paper. The central empirical claim (Dice 78.42%, HD95 1.379 mm on a patient-held-out test set; ~43% HD95 reduction vs nnU-Net) is measured after training, not obtained by redefining a fitted quantity as a prediction. Hierarchical add/remove action targets (Eq. 9) are built from detached current low-resolution probabilities versus ground truth only during training—standard corrective/teacher-style supervision—while inference explicitly rolls out three prior-mean steps without GT or action targets (Sec. 4.3, 4.5). Loss weights, group priors (0.6/0.3/0.1), λ_add/λ_remove, and foreground spatial weights are declared hyperparameters, not reported as derived predictions. Self-citations (e.g., nnMamba and other Gong et al. works) appear in related-work and baseline comparisons and do not supply a uniqueness theorem or load-bearing premise that forces the method or the headline metric. The derivation chain does not reduce by construction to its inputs; any debate about causal attribution of the HD95 drop (capacity vs. world-model reasoning) is a correctness/ablation concern, not circularity.

Axiom & Free-Parameter Ledger

8 free parameters · 6 axioms · 4 invented entities

The central empirical claim rests on standard deep-learning and world-model machinery plus many hand-chosen training constants and one private single-center dataset. No new physical law is postulated; the invented pieces are methodological constructs (hierarchical anatomical actions, multi-scale observation for RSSM-in-decoder). Load-bearing domain assumptions include that atomic decomposition preserves clinical hierarchy and that short latent rollouts without GT at test time transfer.

free parameters (8)
  • rollout_steps_T = 3
    Number of action-conditioned latent transitions fixed to 3 across all experiments; not derived.
  • loss_weights_final_aux_KL_action = 1.0, 0.2, 0.1, 0.2
    Complete objective uses fixed coefficients 1.0 / 0.2 / 0.1 / 0.2 chosen by authors.
  • hierarchy_prior_masses_atomic_canonical_global = 0.6 / 0.3 / 0.1
    Action channel priors allocated 0.6/0.3/0.1 by hand across hierarchy levels.
  • lambda_add_lambda_remove = 2 and 1
    Add corrections weighted twice remove corrections in the balanced action objective.
  • lambda_reg_background_spatial_weight = 0.02
    Non-foreground voxels down-weighted to 0.02; selected via sensitivity sweep.
  • lambda_act_action_loss_weight = 0.2
    Action supervision coefficient selected at 0.2 after sweep over {0,0.05,0.2,0.8}.
  • group_absence_and_bg_patch_coefficients = 0.1
    Missing anatomical groups and near-empty patches receive 0.1 presence/patch coefficients rather than zero or one.
  • RSSM_channel_dims_and_latent_resolution = 128 / 128 / 32 at 32^3
    Observation 128-d at 32³; deterministic 128 / stochastic 32 channels—architectural knobs fixed without theoretical necessity.
axioms (6)
  • domain assumption Encoder–decoder skip networks (nnU-Net configuration) are a valid base mapping from CT patches to multi-class masks.
    Entire system is built on nnU-Net preprocessing, encoder, and high-res decoder (Section 4.1, 5.2).
  • domain assumption A deterministic ConvGRU RSSM with Gaussian prior/posterior in feature space can represent useful anatomical hypothesis transitions under discrete-time actions (PlaNet-style).
    Invoked in Section 2.3 and 4.3; dynamics are assumed learnable from short supervised rollouts, not proved.
  • domain assumption Nested clinical auricular annotations can be losslessly converted to 35 mutually exclusive atomic labels and deterministically recomposed to canonical structures for fair multi-class training/eval.
    Section 3.2 and 5.1; evaluation metrics are defined only after this reconversion.
  • ad hoc to paper Soft add/remove maps A_add=Y(1-P), A_remove=(1-Y)P are the correct supervisory semantics for latent anatomical corrections.
    Defined in Eq. (9); central to claiming 'hierarchical anatomical actions' rather than generic aux losses.
  • domain assumption Patient-level split of a single-center retrospective cohort is sufficient to support generalization claims for the method.
    All tables use one internal test set of 32 cases; no external cohort (Table 1, Section 5).
  • standard math Standard reparameterized KL-regularized latent Gaussians and Dice+CE segmentation losses are appropriate training principles.
    Eqs. (6), (13), (14); conventional VAE/segmentation practice.
invented entities (4)
  • Hierarchical anatomical actions (add/remove over 73 groups) no independent evidence
    purpose: Provide structured semantic supervision for each latent transition instead of implicit layer-wise feature evolution.
    G=35 atomic + 35 canonical + 3 global groups with soft add/remove targets; core named mechanism of AuricularWorld.
  • Multi-scale anatomical observation o for RSSM init no independent evidence
    purpose: Fuse encoder stages E0–E4, bottleneck, and partial decoder D2 into the latent world-model input.
    Eq. (3); paper-specific interface between U-Net features and RSSM.
  • Foreground-masked balanced hierarchical action objective no independent evidence
    purpose: Reweight sparse foreground, missing groups, and add/remove imbalance so latent transitions learn meaningful corrections.
    Eqs. (11)–(12); engineered loss unique to this work’s training recipe.
  • Fine-grained nested auricular CT dataset (35 atomic labels) no independent evidence
    purpose: Benchmark complementary skin-covered subunits and cartilage structures absent from public CT datasets.
    Section 3; clinically annotated private cohort to be released later.

pith-pipeline@v1.2.0-daily-grok45 · 24619 in / 4528 out tokens · 77136 ms · 2026-07-31T05:48:13.384747+00:00 · methodology

0 comments
read the original abstract

Fine-grained segmentation of auricular structures in CT is challenging because the ear occupies a small image region, cartilage boundaries are highly irregular, and interfaces between cartilage and surrounding soft tissues are often ambiguous. Clinical annotations may also include both composite structures containing cartilage and adjacent skin and their corresponding cartilage-only regions, producing nested and overlapping labels. We propose a world-model-based segmentation framework that enables iterative anatomical reasoning beyond conventional feed-forward prediction. Built on an encoder-decoder architecture, the framework introduces a deterministic recurrent state-space model into the intermediate latent space. Multi-scale encoder features and partially decoded representations are fused to form a structural observation that initializes the latent dynamics. During inference, the model performs a three-step latent rollout without ground-truth guidance. Hierarchical anatomical actions update the recurrent state and progressively refine the latent representation. The resulting latent trajectory is projected back into the decoder and combined with high-resolution features to produce the final segmentation. To learn reliable latent transitions, we introduce a balanced hierarchical action objective that addresses foreground sparsity, missing anatomical groups, and imbalance between add and remove operations. Extensive experiments show that the proposed framework consistently improves segmentation accuracy and reduces HD95 by more than 43% for small, irregular, and overlapping auricular structures in CT. These results demonstrate the effectiveness of latent world-model reasoning for challenging medical image segmentation.

Figures

Figures reproduced from arXiv: 2607.28487 by Haifan Gong, Haiyue Jiang, Jingwen Yang, Keying Zhang, Lin Lin, Luoyao Kang, Runmeng Cui, Senmao Wang, Yunjia Bao.

Figure 1
Figure 1. Figure 1: Overview of the proposed AuricularWorld framework. (A) Overall architecture. The input 3D CT volume is processed by the nnU-Net encoder and the first part of the decoder to obtain an intermediate decoded feature. Multi-scale encoder representations and the decoded feature are fused into an observation feature, which is refined by the RSSM latent world model. The resulting latent feature is passed to the re… view at source ↗
Figure 2
Figure 2. Figure 2: Qualitative comparison on representative test cases. Columns show the input CT, ground truth, and [PITH_FULL_IMAGE:figures/full_fig_p023_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Category-wise comparison of AuricularWorld and external segmentation methods on test set. The [PITH_FULL_IMAGE:figures/full_fig_p024_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Visualization of the three-step RSSM action-guided refinement process. Each row shows one [PITH_FULL_IMAGE:figures/full_fig_p027_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Sensitivity of AuricularWorld to the region weighting coe [PITH_FULL_IMAGE:figures/full_fig_p028_5.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

53 extracted references · 7 linked inside Pith

  1. [1]

    J. P. Rodríguez-Arias, A. Gutiérrez Venturini, M. M. Pampín Martínez, E. Gómez García, J. M. Muñoz Caro, M. San Basilio, M. Martín Pérez, J. L. Ce- brián Carretero, Microtia ear reconstruction with patient-specific 3d models—a segmentation protocol, Journal of clinical medicine 11 (13) (2022) 3591

  2. [2]

    E. N. Mohamed, A. Elshahat, H. E.-D. Hany, F. R. Shafik, R. Lashin, Segmenta- tion of the 3d printed mirror image auricular model to ease sculpture of the costal cartilages in total auricular aesthetic reconstruction, Asian Journal of Surgery 46 (12) (2023) 5429–5437

  3. [3]

    W. Zhu, Y . Huang, L. Zeng, X. Chen, Y . Liu, Z. Qian, N. Du, W. Fan, X. Xie, Anatomynet: Deep learning for fast and fully automated whole-volume segmen- tation of head and neck anatomy, Medical Physics 46 (2) (2019) 576–589

  4. [4]

    S. Wang, H. Gong, R. Cui, B. Wan, Z. Hu, H. Yang, J. Zhou, H. Jiang, L. Lin, Costal cartilage segmentation with topology guided deformable mamba: Method and benchmark, Expert Systems with Applications 300 (2026) 130085

  5. [5]

    L. Kang, H. Gong, X. Wan, H. Li, Visual-attribute prompt learning for progres- 29 sive mild cognitive impairment prediction, in: Medical Image Computing and Computer Assisted Intervention – MICCAI 2023, Springer, 2023, pp. 547–557

  6. [6]

    Isensee, P

    F. Isensee, P. F. Jaeger, S. A. Kohl, J. Petersen, K. H. Maier-Hein, nnu-net: a self-configuring method for deep learning-based biomedical image segmentation, Nature methods 18 (2) (2021) 203–211

  7. [7]

    H. Gong, G. Chen, R. Wang, X. Xie, M. Mao, Y . Yu, F. Chen, G. Li, Multi- task learning for thyroid nodule segmentation with thyroid region prior, in: 2021 IEEE 18th International Symposium on Biomedical Imaging (ISBI), 2021, pp. 257–261

  8. [8]

    H. Gong, J. Chen, G. Chen, H. Li, G. Li, F. Chen, Thyroid region prior guided attention for ultrasound segmentation of thyroid nodules, Computers in Biology and Medicine 155 (2023) 106389

  9. [9]

    H. Gong, W. Huang, H. Zhang, Y . Wang, X. Wan, H. Shen, G. Li, H. Li, Intensity confusion matters: An intensity-distance guided loss for bronchus segmentation, in: 2024 IEEE International Conference on Multimedia and Expo (ICME), 2024, pp. 1–6

  10. [10]

    Z. Xu, H. Gong, X. Wan, H. Li, Asc: Appearance and structure consistency for unsupervised domain adaptation in fetal brain mri segmentation, in: Medical Im- age Computing and Computer Assisted Intervention – MICCAI 2023, Springer, 2023, pp. 325–335

  11. [11]

    Huang, H

    W. Huang, H. Gong, H. Zhang, Y . Wang, X. Wan, G. Li, H. Li, H. Shen, Bc- net: Bronchus classification via structure guided representation learning, IEEE Transactions on Medical Imaging 44 (1) (2025) 489–498

  12. [12]

    H. Gong, B. Wan, L. Kang, X. Wan, L. Zhang, H. Li, Boundary as the bridge: Toward heterogeneous partially-labeled medical image segmentation and land- mark detection, IEEE Transactions on Medical Imaging 44 (7) (2025) 2747–2756. doi:10.1109/TMI.2025.3548919. 30

  13. [13]

    Ronneberger, P

    O. Ronneberger, P. Fischer, T. Brox, U-net: Convolutional networks for biomedi- cal image segmentation, in: International Conference on Medical image comput- ing and computer-assisted intervention, Springer, 2015, pp. 234–241

  14. [14]

    J. Chen, Y . Lu, Q. Yu, X. Luo, E. Adeli, Y . Wang, L. Lu, A. L. Yuille, Y . Zhou, Transunet: Transformers make strong encoders for medical image segmentation, arXiv preprint arXiv:2102.04306 (2021)

  15. [15]

    H.-Y . Zhou, J. Guo, Y . Zhang, X. Han, L. Yu, L. Wang, Y . Yu, nnformer: V ol- umetric medical image segmentation via a 3d transformer, IEEE transactions on image processing 32 (2023) 4036–4045

  16. [16]

    H. Gong, L. Kang, Y . Wang, Y . Wang, X. Wan, X. Wu, H. Li, nnmamba: 3d biomedical image segmentation, classification and landmark detection with state space model, in: 2025 IEEE 22nd International Symposium on Biomedical Imag- ing (ISBI), 2025, pp. 1–5

  17. [17]

    D. Ha, J. Schmidhuber, World models, arXiv preprint arXiv:1803.10122 2 (3) (2018) 440

  18. [18]

    Y . Song, Y . Lu, L. Chen, Y . Luo, Hierarchical multi-scale enhanced transformer for medical image segmentation, IEEE Journal of Biomedical and Health Infor- matics 29 (12) (2025) 8917–8927. doi:10.1109/JBHI.2024.3515477

  19. [19]

    H. Wang, Y . Qi, W. Liu, K. Guo, W. Lv, Z. Liang, DPGNet: A boundary-aware medical image segmentation framework via uncertainty perception, IEEE Journal of Biomedical and Health Informatics (2025). doi:10.1109/JBHI.2025.3601025

  20. [20]

    Zhang, Y

    Y . Zhang, Y . Wang, Y . Wan, Q. Zhao, L. Zhao, B. Li, L. Zhang, Z. Chen, A carv- ing hierarchical information integration network for medical image segmentation, Pattern Recognition 171 (2026) 112291. doi:10.1016/j.patcog.2025.112291

  21. [21]

    C. Yu, Y . Li, J. Li, Z. Zhao, T. Zhang, Rethinking feature interactions for med- ical image segmentation: A unified hierarchical aggregation framework with boundary guidance, IEEE Journal of Biomedical and Health Informatics (2026). doi:10.1109/JBHI.2026.3677486. 31

  22. [22]

    Z. Zhu, H. Wang, G. Qi, Y . Li, N. Mazur, Y . Liu, H. Li, B. Cong, L. Bai, A survey on lightweight technology of neural networks for medical image segmentation, Pattern Recognition 179 (2026) 113870. doi:10.1016/j.patcog.2026.113870

  23. [23]

    Vaswani, N

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, I. Polosukhin, Attention is all you need, in: Advances in Neural In- formation Processing Systems, V ol. 30, 2017, pp. 5998–6008

  24. [24]

    A. Gu, T. Dao, Mamba: Linear-time sequence modeling with selective state spaces, in: First Conference on Language Modeling, 2024

  25. [25]

    T. Dao, A. Gu, Transformers are SSMs: Generalized models and efficient algo- rithms through structured state space duality, in: Proceedings of the 41st Inter- national Conference on Machine Learning, V ol. 235, PMLR, 2024, pp. 10041– 10071

  26. [26]

    C. Fan, H. Yu, Y . Huang, L. Wang, Z. Yang, X. Jia, Slicemamba with neural archi- tecture search for medical image segmentation, IEEE Journal of Biomedical and Health Informatics 29 (10) (2025) 7446–7458. doi:10.1109/JBHI.2025.3564381

  27. [27]

    C. Yu, H. Zhang, C. Pu, S. Lv, J. Yu, X. Wu, D. Ruan, H. Xuan, Y . Yan, Graph- Mamba: Graph-driven spatial order-aware mamba for medical image segmenta- tion, Pattern Recognition 171 (2026) 112231. doi:10.1016/j.patcog.2025.112231

  28. [28]

    C. Wang, Y . Xie, Q. Chen, Y . Zhou, Q. Wu, A comprehensive analysis of mamba for 3d volumetric medical image segmentation, Pattern Recognition 173 (2026) 112701. doi:10.1016/j.patcog.2025.112701

  29. [29]

    Cheng, Z

    Y . Cheng, Z. Liu, S. Tamura, Deep hierarchy-aware segmentation: A novel frame- work for mris brain tumor segmentation, IEEE Transactions on Medical Imaging 45 (5) (2026) 2023–2038. doi:10.1109/TMI.2025.3645821

  30. [30]

    Y . Yang, J. Zhuang, G. Sun, R. Wang, J. Su, Boundary-guided contrastive learning for semi-supervised medical image segmentation, IEEE Transactions on Medical Imaging 44 (7) (2025) 2973–2988. doi:10.1109/TMI.2025.3556482. 32

  31. [31]

    Huang, S

    Y . Huang, S. Li, Z. Guo, Q. Mei, Z. Han, X. Wang, H. Wang, Boundary feature alignment for semi-supervised medical image segmentation, Pattern Recognition 170 (2026) 111946. doi:10.1016/j.patcog.2025.111946

  32. [32]

    Zhou, Boundary-aware and cross-modal fusion network for enhanced multi- modal brain tumor segmentation, Pattern Recognition 165 (2025) 111637

    T. Zhou, Boundary-aware and cross-modal fusion network for enhanced multi- modal brain tumor segmentation, Pattern Recognition 165 (2025) 111637. doi:10.1016/j.patcog.2025.111637

  33. [33]

    H. Gong, H. Liu, Y . Wang, X. Liu, X. Wan, Q. Shi, H. Li, Fetal cerebellum landmark detection based on 3d mri: Method and benchmark, IEEE Journal of Biomedical and Health Informatics 29 (8) (2025) 5712–5721

  34. [34]

    Mussi, M

    E. Mussi, M. Servi, F. Facchini, R. Furferi, L. Governi, Y . V olpe, A novel ear elements segmentation algorithm on depth map images, Computers in Biology and Medicine 129 (2021) 104157

  35. [35]

    Servi, E

    M. Servi, E. Mussi, R. Magherini, M. Carfagni, R. Furferi, Y . V olpe, U-net for auricular elements segmentation: a proof-of-concept study, in: 2021 43rd Annual International Conference of the IEEE Engineering in Medicine & Biology Society (EMBC), IEEE, 2021, pp. 2712–2716

  36. [36]

    Q. Wang, Y . Wang, X. Zhou, Q. Zhang, Three-dimensional auricular subunit mod- els for cartilage framework fabrication: our preliminary experience, Journal of Craniofacial Surgery 33 (4) (2022) 1111–1115

  37. [37]

    M. T. Ross, M. Antico, K. L. McMahon, J. Ren, S. K. Powell, A. K. Pandey, M. C. Allenby, D. Fontanarosa, M. A. Woodruff, Ultrasound imaging offers promising alternative to create 3-d models for personalised auricular implants, Ultrasound in Medicine & Biology 48 (3) (2022) 450–459

  38. [38]

    Hafner, T

    D. Hafner, T. Lillicrap, I. Fischer, R. Villegas, D. Ha, H. Lee, J. Davidson, Learn- ing latent dynamics for planning from pixels, in: International conference on machine learning, PMLR, 2019, pp. 2555–2565

  39. [39]

    Hafner, T

    D. Hafner, T. Lillicrap, J. Ba, M. Norouzi, Dream to control: Learning behaviors by latent imagination, in: International Conference on Learning Representations. 33

  40. [40]

    Z. Chen, Z. Cong, Z. Jin, W. Fan, D. Zhou, Q. Ai, H. Gong, C. Liao, X. Liu, C. Wang, Medical world models in healthcare: Foundations, applications, and challenges for trustworthy clinical translation, arXiv preprint arXiv:2607.25242 (2026). arXiv:2607.25242

  41. [41]

    L. Kang, Y . Zhang, J. Shan, H. Gong, Q. Ding, S. S. Cheng, Dreamreg: Belief-driven world model for 2d-3d ultrasound registration, arXiv preprint arXiv:2606.18825 (2026)

  42. [42]

    Y . Yue, Y . Wang, C. Tao, P. Liu, S. Song, G. Huang, Chexworld: Exploring image world modeling for radiograph representation learning, in: Proceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 20778–20788

  43. [43]

    Yang, Z.-Y

    Y . Yang, Z.-Y . Wang, Q. Liu, S. Sun, K. Wang, R. Chellappa, Z. Zhou, A. Yuille, L. Zhu, Y .-D. Zhang, et al., Medical world model: Generative simulation of tumor evolution for treatment planning, arXiv preprint arXiv:2506.02327 (2025)

  44. [44]

    Moritz, E

    M. Moritz, E. Topol, P. Rajpurkar, Coordinated ai agents for advancing health- care, Nature Biomedical Engineering 9 (2025) 432–438

  45. [45]

    Q. Kong, Y . Zhao, H. Gong, L. Kang, J. Fu, L. Li, B. Wan, P. Wang, X. Li, Y . Wang, J. Zhang, Y . Yu, X. Yang, X. Zuo, H. Wang, Y . Li, Ai agent-based dis- covery of d-enantiomeric antimicrobial peptides against multidrug-resistant bac- terial infection, Biomaterials 329 (2026) 123927

  46. [46]

    Schmidgall, R

    S. Schmidgall, R. Ziaei, C. Harris, J. W. Kim, E. P. Reis, J. Jopling, M. Moor, Agentclinic: A multimodal benchmark for tool-using clinical ai agents, npj Digi- tal Medicine 9 (2026) 499

  47. [47]

    X. Luo, W. Liao, J. Xiao, J. Chen, T. Song, X. Zhang, K. Li, D. N. Metaxas, G. Wang, S. Zhang, Word: A large scale dataset, benchmark and clinical ap- plicable study for abdominal organ segmentation from ct image, Medical Image Analysis 82 (2022) 102642

  48. [48]

    Y . Ji, H. Bai, C. Ge, J. Yang, Y . Zhu, R. Zhang, Z. Li, L. Zhang, W. Ma, X. Wan, P. Luo, Amos: A large-scale abdominal multi-organ benchmark for versatile med- 34 ical image segmentation, in: Advances in Neural Information Processing Sys- tems, V ol. 35, 2022, pp. 36722–36732

  49. [49]

    Wasserthal, H.-C

    J. Wasserthal, H.-C. Breit, M. T. Meyer, M. Pradella, D. Hinck, A. W. Sauter, T. Heye, D. T. Boll, J. Cyriac, S. Yang, M. Bach, M. Segeroth, Totalsegmenta- tor: Robust segmentation of 104 anatomic structures in ct images, Radiology: Artificial Intelligence 5 (5) (2023) e230024

  50. [50]

    Podobnik, P

    G. Podobnik, P. Strojan, P. Peterlin, B. Ibragimov, T. Vrtovec, Han-seg: The head and neck organ-at-risk ct and mr segmentation dataset, Medical Physics 50 (3) (2023) 1917–1927

  51. [51]

    Hatamizadeh, Y

    A. Hatamizadeh, Y . Tang, V . Nath, D. Yang, A. Myronenko, B. Landman, H. R. Roth, D. Xu, Unetr: Transformers for 3d medical image segmentation, in: Pro- ceedings of the IEEE/CVF winter conference on applications of computer vision, 2022, pp. 574–584

  52. [52]

    Hatamizadeh, V

    A. Hatamizadeh, V . Nath, Y . Tang, D. Yang, H. Roth, D. Xu, Swin unetr: Swin transformers for semantic segmentation of brain tumors in mri images, arXiv preprint arXiv:2201.01266 (2022)

  53. [53]

    M. J. Cardoso, W. Li, R. Brown, N. Ma, E. Kerfoot, Y . Wang, B. Murrey, A. My- ronenko, C. Zhao, D. Yang, et al., Monai: An open-source framework for deep learning in healthcare, arXiv preprint arXiv:2211.02701 (2022). 35