Pith. sign in

REVIEW 1 cited by

Are Vision xLSTM Embedded UNet More Reliable in Medical 3D Image Segmentation?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.16993 v3 pith:2Z73T67K submitted 2024-06-24 eess.IV cs.CV

classification eess.IVcs.CV
keywords segmentationvision-xlstmcnnsmedicalu-vixlstmvisionblockscapture
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The development of efficient segmentation strategies for medical images has evolved from its initial dependence on Convolutional Neural Networks (CNNs) to the current investigation of hybrid models that combine CNNs with Vision Transformers (ViTs). There is an increasing focus on creating architectures that are both high-performing and computationally efficient, capable of being deployed on remote systems with limited resources. Although transformers can capture global dependencies in the input space, they face challenges from the corresponding high computational and storage expenses involved. This research investigates the integration of CNNs with Vision Extended Long Short-Term Memory (Vision-xLSTM)s by introducing the novel U-VixLSTM. The Vision-xLSTM blocks capture the temporal and global relationships within the patches extracted from the CNN feature maps. The convolutional feature reconstruction path upsamples the output volume from the Vision-xLSTM blocks to produce the segmentation output. Our primary objective is to propose that Vision-xLSTM forms an appropriate backbone for medical image segmentation, offering excellent performance with reduced computational costs. The U-VixLSTM exhibits superior performance compared to the state-of-the-art networks in the publicly available Synapse, ISIC and ACDC datasets. Code provided: https://github.com/duttapallabi2907/U-VixLSTM

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DKT2: Revisiting Applicable and Comprehensive Knowledge Tracing in Large-Scale Data

    cs.LG 2025-01 conditional novelty 5.0 of 10

    DKT2, an xLSTM-based model with Rasch embeddings and an IRT-style decomposition, generally beats 18 knowledge tracing baselines on three large datasets, though not on every metric or task.

Pith tools