REVIEW 3 major objections 6 minor 6 cited by
CRISP-SAM2: SAM2 with Cross-Modal Interaction and Semantic Prompting for Multi-Organ Segmentation
T0 review · 3 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read CRISP-SAM2, a SAM2 variant that replaces clicks and boxes with organ descriptions, reports higher average DSC and NSD than all compared baselines on seven public multi-organ CT datasets.
desk verdict Solid engineering with strong ablations, but the text-only SOTA claim is not yet established because the number of text sentences was chosen using segmentation performance, and no validation split is documented. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is a pipeline of four additions to SAM2. A Cross-Modal Semantic Interaction module runs two levels of multi-head cross-attention between CLIP text features and image features to produce CS-Features, which are injected into the frozen Hiera image encoder by a lightweight CS-Injector. A Semantic Prompt Projector then generates sparse and dense prompt embeddings from the CS-Features and four-scale image features, replacing the original point/box prompt encoder. The mask decoder receives extra learnable CS-Output Tokens, and a Local Refiner predicts a refined mask from the updated tokens and CS-Features. Finally, a similarity-sorting self-updating memory bank replaces SAM2's first-in-first-out memory so that only high-quality CT slices guide later slices in the volume.
What would settle it
Re-run the seven-dataset benchmark with the sentence count fixed ahead of time, or tuned only on a held-out validation split, and compare DSC and NSD against Table 1; if the gap over CAT and ZePT shrinks to noise, the headline advantage comes from test-set text selection rather than from the model architecture.
Extended reading notes
Core claim
The paper's central claim is that cross-modal semantic features, obtained by progressively fusing text and image embeddings, can take over the role of geometric prompts inside SAM2 and improve detail-sensitive segmentation in 3D. CRISP-SAM2 reports higher average DSC and NSD than the eleven compared baselines across the seven M3D-Seg sub-datasets, with the largest gaps on MSD-Spleen, Pancreas-CT, LUNA16, AbdomenCT-1k and WORD, and near-parity plus NSD leads on FLARE22 and AMOS22. The authors conclude that text-only semantic prompting is a viable substitute for points and boxes in multi-organ medical segmentation, and that the injected text semantics plus local refinement and a similarity-sorted memory bank fix the boundary and small-structure errors common in SAM2-style models.
Load-bearing premise
The reported gains rest on the assumption that the number of text sentences used per organ (two to six) was chosen without consulting test-set results; Appendix B says the sentence count was increased when segmentation performance was worse, and the paper does not describe a separate validation set for that choice.
Editorial extensions
If this is right
- If the reported numbers hold, multi-organ segmentation can be run from organ descriptions alone, removing the need for expert point or box prompts at inference.
- The NSD gains over DSC gains indicate the largest benefit is boundary and small-structure accuracy, not just gross overlap.
- Because text length can be increased for harder cases, the same model can be steered at inference by supplying more descriptive sentences.
- Treating CT volumes as video-like sequences with a similarity-sorted memory bank makes SAM2's temporal machinery applicable to 3D medical scans.
- Text-guided models are reported to beat purely visual SAM and SAM2 variants by wide margins, which strengthens the case for multimodal prompting in medical segmentation.
Reading between the lines
- If the claim holds, organ descriptions could be generated automatically from structured radiology reports, turning CRISP-SAM2 into a report-driven segmenter that needs no human interaction at all.
- The sentence-count selection rule in Appendix B means an independent re-benchmark with a frozen text length is needed to know how much of the gain is architectural versus test-time text tuning.
- The same cross-modal and memory-sorting machinery could plausibly transfer to other foundation models or to MRI volumes, but the paper only demonstrates CT data.
- The zero-shot results suggest text prompts generalize to unseen organs more weakly than geometric prompts; a natural follow-up is training with organ descriptions that vary in style to close that gap.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes CRISP-SAM2, an extension of SAM2 for text-guided 3D multi-organ segmentation. The method introduces a progressive two-level cross-attention module that fuses visual and textual features, a Semantic Prompt Projector that generates sparse and dense prompt embeddings from cross-modal semantics instead of geometric prompts, a Local Refiner for boundary refinement, and a similarity-sorting self-updating memory bank tailored to 3D medical volumes. The model is evaluated on seven sub-datasets from M3D-Seg against ten baselines, reporting higher average DSC and NSD on most datasets, with ablations supporting each component. The code is released.
Significance. The central claim is that purely textual organ descriptions can replace geometric prompts in 3D multi-organ segmentation without sacrificing accuracy, which would be an practically useful result. The paper is strong on experimental breadth: seven datasets, multiple baselines, component ablations, and auxiliary analyses of text length, quality, type, and zero-shot generalization. The authors are also transparent about several limitations, including modest zero-shot performance and computational cost. The main scientific value, if the protocol concerns are resolved, is in the architecture study showing how cross-modal semantic prompting can be injected into a video-segmentation foundation model for volumetric medical imaging. However, the headline claim depends on how the text prompt lengths were chosen and on the comparability of the modified label sets across baselines, so the current evidence is conditional.
major comments (3)
- [Appendix B, last paragraph] The manuscript states: "The number of sentences is based on the performance of the segmentation prediction; the worse the performance, the more sentences we add to assist the segmentation." This is the only disclosure about the selection of text lengths, and the paper describes an 8:2 train/test split with no separate validation set. If sentence counts were increased after evaluating on the test split, the reported DSC/NSD values are optimistically biased and CRISP-SAM2 is not a fixed text-prompted model. The magnitude matters: Table 9 shows that moving from Short to Long texts changes DSC by 1.15 on AbdomenCT-1k and 1.48 on AMOS22, whereas the claimed margins over the closest text-guided baselines on the five lead datasets are roughly 0.93--1.27% DSC. Please specify exactly how sentence counts were determined, use a validation split or a pre-registered text budget, and report results under fixed text budgets (e.g., 2, 4, and 6 sentences) to establish that the claimed superiority is not an artifact of test-set feedback.
- [Section 4.2 and Table 1] All performance comparisons are point estimates without error bars, confidence intervals, or significance tests. On FLARE22 and AMOS22 the method is only 0.12 and 0.15 DSC behind the best baseline, and FLARE22 has just 10 test volumes. Given the small test sets and the small margins, the claim that CRISP-SAM2 "outperforms" all compared models is not statistically supported. Please add bootstrapped confidence intervals or paired significance tests (e.g., Wilcoxon signed-rank or paired bootstrap) for the DSC/NSD differences, especially on the datasets where the lead is under 1%.
- [Section 4.1 and Appendix B] The evaluation uses modified label sets relative to the original benchmarks: Pancreas-CT excludes the tumor label, LUNA16 excludes trachea and nodules, WORD excludes adrenal/colon/intestine/rectum, AMOS22 excludes prostate, and AbdomenCT-1k splits kidney into left/right. The paper does not state whether the published baseline numbers are reproduced under exactly the same label protocol or whether all baselines were retrained on the same modified masks. Without this clarification, the "state-of-the-art on seven public datasets" claim is ambiguous, because the task difficulty is changed by the exclusions and mergers. Please report which models were trained under the modified protocol, which numbers are quoted from external publications, and ideally provide results on the full original label sets to allow direct comparison with prior work.
minor comments (6)
- [Section 2.2 and Section 3.1] Reference numbering is inconsistent: SAM2 is cited as [75] in most places, but Section 2.2 and Section 3.1 refer to "SAM2 [62]", where [62] is MedSAM. Please correct the citations.
- [Abstract and Introduction] The abstract and introduction say "based on SAM2", but the Introduction text near Figure 1 states "based on SAM as shown in Figure 1". This should be "SAM2" for accuracy.
- [Appendix E.3] There is a typo: "decreases both DCS and NSD" should read "DSC". Also, the numbers "1.32% on DSC and 1.13% on NSD" and "4.32% and 4.99%" are not clearly tied to a specific dataset; specify whether these are averages across the two ablation datasets or per-dataset values.
- [Table 4] The checkmark/blank pattern in the Prompt columns is ambiguous; the first data row appears to correspond to no prompt, but this is not labeled. Please make prompt combinations explicit in the table header (e.g., "none", "text only", "text+point", "text+bbox", "text+point+bbox").
- [Section 3.3, Eq. (4)] Equation (4) uses LN (layer normalization) for both inputs, while the surrounding text and Figure 2 use BN (batch normalization). Please clarify which normalization is applied where, and keep notation consistent.
- [Figure 3] The method names are not shown consistently across all panels; for example, row (a) and row (b) labels appear only once rather than above each subplot. Adding a clear method label to every row or column would improve readability.
Circularity Check
No circular derivation; Appendix B text-length selection is a test-set protocol concern, not a circular reduction.
full rationale
CRISP-SAM2 is an empirical systems paper with no formal derivation chain to audit. Its performance claims rest on measured DSC/NSD over seven public datasets from M3D-Seg against external baselines (SAM2, MedSAM, CT-SAM3D, CAT, SegVol, ZePT, etc.), with text descriptions generated by external GPT-4o and provided by the M3D-Seg dataset. The architecture is adapted from SAM2 using standard cross-attention, injector, projector, and refiner components; none of the reported equations (Eqs. 1-9) define the target metric in terms of the model's own output, and the self-citations in the paper (e.g., [106], [107]) are prior CLIP-SAM and joint-attention applications that are not load-bearing for the central claim. The only flagged item is Appendix B: 'The number of sentences is based on the performance of the segmentation prediction; the worse the performance, the more sentences we add to assist the segmentation.' Because no separate validation split is described for this choice, the reported test-set margins could be optimistically biased if sentence counts were selected after seeing test performance. However, this is a data-selection/protocol weakness rather than a circular reduction: the reported DSC/NSD values are empirical measurements, not quantities that equal the sentence-count choices by construction, and the core model is not fitted to or derived from its own benchmark output. Accordingly, no circular step is established, and the circularity score is 0.
Assumptions & free parameters
free parameters (3)
- Memory IoU threshold (thresh) =
0.6
- Initial loss weights alpha/beta/gamma =
0.6/0.2/0.2
- Text sentence count =
2-6 sentences selected based on performance
assumptions (3)
- domain assumption SAM2 pretrained weights provide a suitable initialization for medical segmentation
- domain assumption GPT-4o-generated textual descriptions contain spatial and shape cues consistent across training and test images
- ad hoc to paper Slices with higher mean cosine similarity to other slices are more informative for memory-guided segmentation
Cite this review
Pith. "Pith review of CRISP-SAM2: SAM2 with Cross-Modal Interaction and Semantic Prompting for Multi-Organ Segmentation." pith.science (2026). https://pith.science/paper/RF27WLFR
@misc{pith2026250623121,
author = {Pith},
title = {Pith review of: CRISP-SAM2: SAM2 with Cross-Modal Interaction and Semantic Prompting for Multi-Organ Segmentation},
year = {2026},
howpublished = {\url{https://pith.science/paper/RF27WLFR}},
note = {Machine review of arXiv:2506.23121}
}
read the original abstract
Multi-organ medical segmentation is a crucial component of medical image processing, essential for doctors to make accurate diagnoses and develop effective treatment plans. Despite significant progress in this field, current multi-organ segmentation models often suffer from inaccurate details, dependence on geometric prompts and loss of spatial information. Addressing these challenges, we introduce a novel model named CRISP-SAM2 with CRoss-modal Interaction and Semantic Prompting based on SAM2. This model represents a promising approach to multi-organ medical segmentation guided by textual descriptions of organs. Our method begins by converting visual and textual inputs into cross-modal contextualized semantics using a progressive cross-attention interaction mechanism. These semantics are then injected into the image encoder to enhance the detailed understanding of visual information. To eliminate reliance on geometric prompts, we use a semantic prompting strategy, replacing the original prompt encoder to sharpen the perception of challenging targets. In addition, a similarity-sorting self-updating strategy for memory and a mask-refining process is applied to further adapt to medical imaging and enhance localized details. Comparative experiments conducted on seven public datasets indicate that CRISP-SAM2 outperforms existing models. Extensive analysis also demonstrates the effectiveness of our method, thereby confirming its superior performance, especially in addressing the limitations mentioned earlier. Our code is available at: https://github.com/YU-deep/CRISP_SAM2.git.
Figures
Figures from the paper (6 more)
Forward citations
Cited by 6 Pith papers
-
GeoMag: A Vision-Language Model for Pixel-level Fine-Grained Remote Sensing Image Parsing
A remote sensing vision-language model that uses task-aware resolution adjustment and attention-based cropping to perform pixel-level segmentation alongside image- and region-level tasks.
-
Dual form Complementary Masking for Domain-Adaptive Image Segmentation
The paper proposes complementary masking consistency for UDA segmentation and reports empirical gains, but its theoretical proof contains a direct internal contradiction.
-
Visual Instance-aware Prompt Tuning
ViaPT generates instance-aware prompts per image, fuses them with dataset-level prompts, and applies PCA compression to outperform VPT-Deep and other PEFT baselines on FGVC, HTA, and VTAB-1k.
-
MARL-MambaContour: Unleashing Multi-Agent Deep Reinforcement Learning for Active Contour Optimization in Medical Image Segmentation
A multi-agent Soft Actor-Critic framework with adaptive entropy and a Mamba policy network iteratively moves contour points to segment organs, reporting higher Dice and boundary scores on five datasets.
-
Small Lesions-aware Bidirectional Multimodal Multiscale Fusion Network for Lung Disease Classification
MMCAF-Net, a CNN-KAN multimodal fusion network with 3D multi-scale attention and cross-attention fusion, reports improved lung cancer subtype classification on Lung-PET-CT-Dx, but the gains are not statistically suppo...
-
FADE: Adversarial Concept Erasure in Flow Models
FADE combines adversarial training with trajectory preservation to erase concepts from diffusion models, reporting state-of-the-art erasure on Stable Diffusion benchmarks, but the evidence is incomplete and the theore...
Reference graph
Works this paper leans on
-
[1]
Fan Bai, Yuxin Du, Tiejun Huang, Max Q-H Meng, and Bo Zhao. 2024. M3d: Advancing 3d medical image analysis with multi-modal large language models. arXiv preprint arXiv:2404.00578(2024)
arXiv 2024
-
[2]
Jinhe Bi, Yujun Wang, Haokun Chen, Xun Xiao, Artur Hecker, Volker Tresp, and Yunpu Ma. 2024. Visual Instruction Tuning with 500x Fewer Parameters through Modality Linear Representation-Steering.arXiv preprint arXiv:2412.12359(2024)
arXiv 2024
-
[3]
Jinhe Bi, Yifan Wang, Danqi Yan, Xun Xiao, Artur Hecker, Volker Tresp, and Yunpu Ma. 2025. Prism: Self-pruning intrinsic selection method for training-free multimodal data selection.arXiv preprint arXiv:2502.12119(2025)
arXiv 2025
-
[4]
Guohui Cai, Ruicheng Zhang, Hongyang He, Zeyu Zhang, Daji Ergu, Yuanzhouhan Cao, Jinman Zhao, Binbin Hu, Zhinbin Liao, Yang Zhao, et al
-
[5]
Keyan Chen, Chenyang Liu, Hao Chen, Haotian Zhang, Wenyuan Li, Zhengxia Zou, and Zhenwei Shi. 2024. RSPrompter: Learning to prompt for remote sensing instance segmentation based on visual foundation model.IEEE Transactions on Geoscience and Remote Sensing(2024), 1–17
2024
-
[6]
Tianrun Chen, Ankang Lu, Lanyun Zhu, Chaotao Ding, Chunan Yu, Deyi Ji, Zejian Li, Lingyun Sun, Papa Mao, and Ying Zang. 2024. Sam2-adapter: Evaluat- ing & adapting segment anything 2 in downstream tasks: Camouflage, shadow, medical image segmentation, and more.arXiv preprint arXiv:2408.04579(2024)
arXiv 2024
-
[7]
Tianrun Chen, Lanyun Zhu, Chaotao Deng, Runlong Cao, Yan Wang, Shangzhan Zhang, Zejian Li, Lingyun Sun, Ying Zang, and Papa Mao. 2023. Sam-adapter: Adapting segment anything in underperformed scenes. InProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) Workshops. 3367– 3375
2023
-
[8]
Yinda Chen, Wei Huang, Xiaoyu Liu, Shiyu Deng, Qi Chen, and Zhiwei Xiong
Show all 131 references
-
[9]
Yinda Chen, Che Liu, Xiaoyu Liu, Rossella Arcucci, and Zhiwei Xiong. 2024. Bimcv-r: A landmark dataset for 3d ct text-image retrieval. InInternational Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI). Springer, 124–134
2024
-
[10]
InICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
Learning multiscale consistency for self-supervised electron microscopy instance segmentation. InICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 1566–1570
2024
-
[11]
Zhimin Chen, Longlong Jing, Yingwei Li, and Bing Li. 2023. Bridging the domain gap: Self-supervised 3d scene understanding with foundation models.Advances in Neural Information Processing Systems (NeurIPS)36 (2023), 79226–79239
2023
-
[12]
Yi-Chia Chen, Wei-Hua Li, Cheng Sun, Yu-Chiang Frank Wang, and Chu-Song Chen. 2025. SAM4MLLM: Enhance Multi-Modal Large Language Model for Referring Expression Segmentation. InEuropean Conference on Computer Vision (ECCV). Springer, 323–340
2025
-
[13]
Zhimin Chen and Bing Li. 2025. DyConfidMatch: Dynamic thresholding and re-sampling for 3D semi-supervised learning.Pattern Recognition159 (2025), 111154
2025
-
[14]
Zhimin Chen, Longlong Jing, Liang Yang, Yingwei Li, and Bing Li. 2023. Class-level confidence based 3d semi-supervised learning. InProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (W ACV). 633–642
2023
-
[15]
Fenghua Cheng, Xue Li, Haoyang Wu, Jiangcheng Sang, and Wenqi Zhao. 2024. Recent Advances on Multi-modal Dialogue Systems: A Survey. InInternational Conference on Advanced Data Mining and Applications. Springer, 33–47
2024
-
[16]
Zhen Chen, Qing Xu, Xinyu Liu, and Yixuan Yuan. 2024. UN-SAM: Univer- sal Prompt-Free Segmentation for Generalized Nuclei Images.arXiv preprint arXiv:2402.16663(2024)
2024 arXiv
-
[17]
Kenneth Clark, Bruce Vendt, Kirk Smith, John Freymann, Justin Kirby, Paul Koppel, Stephen Moore, Stanley Phillips, David Maffitt, Michael Pringle, et al
-
[18]
Yangming Cheng, Liulei Li, Yuanyou Xu, Xiaodi Li, Zongxin Yang, Wenguan Wang, and Yi Yang. 2023. Segment and track anything.arXiv preprint arXiv:2305.06558(2023)
2023 arXiv
-
[19]
Daiheng Gao, Shilin Lu, Shaw Walters, Wenbo Zhou, Jiaming Chu, Jie Zhang, Bang Zhang, Mengxi Jia, Jian Zhao, Zhaoxin Fan, et al. 2024. EraseAnything: Enabling Concept Erasure in Rectified Flow Transformers.arXiv preprint arXiv:2412.20413(2024)
2024 arXiv
-
[20]
Eli Gibson, Francesco Giganti, Yipeng Hu, Ester Bonmati, Steve Bandula, Kur- inchi Gurusamy, Brian Davidson, Stephen P Pereira, Matthew J Clarkson, and Dean C Barratt. 2018. Automatic multi-organ segmentation on abdominal CT with dense V-networks.IEEE Transactions on Medical I...
2018
-
[21]
Yuxin Du, Fan Bai, Tiejun Huang, and Bo Zhao. 2025. Segvol: Universal and inter- active volumetric medical image segmentation.Advances in Neural Information Processing Systems (NeurIPS)37 (2025), 110746–110783
2025
-
[22]
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Ab- hishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schel- ten, Alex Vaughan, et al . 2024. The llama 3 herd of models.arXiv preprint arXiv:2407.21783(2024)
2024 arXiv
-
[23]
2025.\A𝐿𝐿𝑀 4𝐴𝐷𝐷 : Unlocking the Capabilities of Audio Large Language Models for Audio Deepfake Detection
Hao Gu, Jiangyan Yi, Chenglong Wang, Jianhua Tao, Zheng Lian, Jiayi He, Yong Ren, Yujie Chen, and Zhengqi Wen. 2025.\A𝐿𝐿𝑀 4𝐴𝐷𝐷 : Unlocking the Capabilities of Audio Large Language Models for Audio Deepfake Detection. arXiv preprint arXiv:2505.11079(2025)
2025 arXiv
-
[24]
Shreyank N Gowda and David A Clifton. 2025. CC-SAM: SAM with Cross- feature Attention and Context for Ultrasound Image Segmentation. InEuropean Conference on Computer Vision (ECCV). Springer, 108–124
2025
-
[25]
Heng Guo, Jianfeng Zhang, Jiaxing Huang, Tony CW Mok, Dazhou Guo, Ke Yan, Le Lu, Dakai Jin, and Minfeng Xu. 2025. Towards a comprehensive, efficient and promptable anatomic structure segmentation model using 3d whole-body ct scans. InProceedings of the AAAI Conference on Artif...
2025
-
[26]
Jianyuan Guo, Kai Han, Han Wu, Yehui Tang, Xinghao Chen, Yunhe Wang, and Chang Xu. 2022. Cmt: Convolutional neural networks meet vision transform- ers. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 12175–12185
2022
-
[27]
Hao Guo, Zihan Ma, Zhi Zeng, Minnan Luo, Weixin Zeng, Jiuyang Tang, and Xiang Zhao. 2025. Each Fake News is Fake in its Own Way: An Attribution Multi- Granularity Benchmark for Multimodal Fake News Detection. InProceedings of the AAAI Conference on Artificial Intelligence, Vol...
2025
-
[28]
Jihong Hu, Yinhao Li, Hao Sun, Yu Song, Chujie Zhang, Lanfen Lin, and Yen- Wei Chen. 2024. LGA: A Language Guide Adapter for Advancing the SAM Model’s Capabilities in Medical Image Segmentation. InInternational Conference on Medical Image Computing and Computer-Assisted Interv...
2024
-
[29]
Jian Hu, Jiayi Lin, Shaogang Gong, and Weitong Cai. 2024. Relax Image-Specific Prompt Requirement in SAM: A Single Generic Prompt for Segmenting Camou- flaged Objects. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 38. 12511–12518
2024
-
[30]
Weizhao He, Yang Zhang, Wei Zhuo, Linlin Shen, Jiaqi Yang, Songhe Deng, and Liang Sun. 2024. APSeg: Auto-Prompt Network for Cross-Domain Few-Shot Semantic Segmentation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 23762–23772
2024
-
[31]
Yuzhi Huang, Chenxin Li, Zixu Lin, Hengyu Liu, Haote Xu, Yifan Liu, Yue Huang, Xinghao Ding, Xiaotong Tu, and Yixuan Yuan. 2024. P2sam: Probabilistically prompted sams are efficient segmentator for ambiguous medical images. In Proceedings of the 32nd ACM International Conferen...
2024
-
[32]
Zhongzhen Huang, Yankai Jiang, Rongzhao Zhang, Shaoting Zhang, and Xiaofan Zhang. 2025. Cat: Coordinating anatomical-textual prompts for multi-organ and tumor segmentation.Advances in Neural Information Processing Systems (NeurIPS)37 (2025), 3588–3610
2025
-
[33]
Xiaorui Huang, Gen Luo, Chaoyang Zhu, Bo Tong, Yiyi Zhou, Xiaoshuai Sun, and Rongrong Ji. 2024. Deep Instruction Tuning for Segment Anything Model. InProceedings of the 32nd ACM International Conference on Multimedia (ACM MM). 905–914
2024
-
[34]
Sergey Ioffe and Christian Szegedy. 2015. Batch normalization: Accelerating deep network training by reducing internal covariate shift. InInternational Conference on Machine Learning (ICML). pmlr, 448–456
2015
-
[35]
Yuanfeng Ji, Haotian Bai, Chongjian Ge, Jie Yang, Ye Zhu, Ruimao Zhang, Zhen Li, Lingyan Zhanng, Wanling Ma, Xiang Wan, et al. 2022. Amos: A large-scale abdominal multi-organ benchmark for versatile medical image segmentation. Advances in Neural Information Processing Systems ...
2022
-
[36]
Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, et al. 2024. Gpt-4o system card.arXiv preprint arXiv:2410.21276(2024)
2024 arXiv
-
[37]
Lei Ke, Mingqiao Ye, Martin Danelljan, Yu-Wing Tai, Chi-Keung Tang, Fisher Yu, et al. 2024. Segment anything in high quality.Advances in Neural Information Processing Systems (NeurIPS)36 (2024), 29914–29934. Conference acronym ’MM, October 27–31, 2025, Dublin, Ireland Xinlei Yu et al
2024
-
[38]
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al
-
[39]
Yankai Jiang, Zhongzhen Huang, Rongzhao Zhang, Xiaofan Zhang, and Shaoting Zhang. 2024. Zept: Zero-shot pan-tumor segmentation via query-disentangling and self-prompting. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 11386–11397
2024
-
[40]
Bin Li, Bin Sun, Shutao Li, Encheng Chen, Hongru Liu, Yixuan Weng, Yongping Bai, and Meiling Hu. 2024. Distinct but correct: generating diversified and entity-revised medical response.Science China Information Sciences67, 3 (2024), 132106
2024
-
[41]
Bin Li, Yixuan Weng, Fei Xia, and Hanjun Deng. 2024. Towards better Chinese- centric neural machine translation for low-resource languages.Computer Speech & Language84 (2024), 101566
2024
-
[42]
Feng Li, Hao Zhang, Peize Sun, Xueyan Zou, Shilong Liu, Chunyuan Li, Jianwei Yang, Lei Zhang, and Jianfeng Gao. 2025. Segment and Recognize Anything at Any Granularity. InEuropean Conference on Computer Vision (ECCV). Springer, 467–484
2025
-
[43]
Youngwan Lee, Jonghee Kim, Jeffrey Willette, and Sung Ju Hwang. 2022. Mpvit: Multi-path vision transformer for dense prediction. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 7287–7296
2022
-
[44]
Quanjiang Li, Tingjin Luo, Mingdie Jiang, Zhangqi Jiang, Chenping Hou, and Feijiang Li. 2025. Semi-Supervised Multi-View Multi-Label Learning with View- Specific Transformer and Enhanced Pseudo-Label. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 39. 18...
2025
-
[45]
Quanjiang Li, Tingjin Luo, Mingdie Jiang, Jiahui Liao, and Zhangqi Jiang. 2024. Deep Incomplete Multi-View Network Semi-Supervised Multi-Label Learning with Unbiased Loss. InProceedings of the 32nd ACM International Conference on Multimedia. 9048–9056
2024
-
[46]
Quanjiang Li, Tingjin Luo, and Jiahui Liao. 2025. Theory-Inspired Deep Multi- View Multi-Label Learning with Incomplete Views and Noisy Labels. InPro- ceedings of the Computer Vision and Pattern Recognition Conference (CVPR). 20706–20715
2025
-
[47]
Leyang Li, Shilin Lu, Yan Ren, and Adams Wai-Kin Kong. 2025. Set you straight: Auto-steering denoising trajectories to sidestep unwanted concepts.arXiv preprint arXiv:2504.12782(2025)
2025 arXiv
-
[48]
Shutao Li, Bin Li, Bin Sun, and Yixuan Weng. 2024. Towards visual-prompt temporal answer grounding in instructional video.IEEE Transactions on Pattern Analysis and Machine Intelligence(2024)
2024
-
[49]
Shijie Li, Yunbin Tu, Qingyuan Xiang, and Zheng Li. 2024. MAGIC: Rethinking Dynamic Convolution Design for Medical Image Segmentation. InProceedings of the 32nd ACM International Conference on Multimedia (ACM MM). 9106–9115
2024
-
[50]
Wenxue Li, Xinyu Xiong, Peng Xia, Lie Ju, and Zongyuan Ge. 2024. TP-DRSeg: improving diabetic retinopathy lesion segmentation with explicit text-prompts assisted SAM. InInternational Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI). Springer, 743–753
2024
-
[51]
Shengze Li, Jianjian Cao, Peng Ye, Yuhan Ding, Chongjun Tu, and Tao Chen. 2025. ClipSAM: CLIP and SAM collaboration for zero-shot anomaly segmentation. Neurocomputing618 (2025), 129122
2025
-
[52]
Zixu Li, Zhiwei Chen, Haokun Wen, Zhiheng Fu, Yupeng Hu, and Weili Guan
-
[53]
Zixu Li, Zhiheng Fu, Yupeng Hu, Zhiwei Chen, Haokun Wen, and Liqiang Nie. 2025. Finecir: Explicit parsing of fine-grained modification semantics for composed image retrieval.arXiv preprint arXiv:2503.21309(2025)
2025 arXiv
-
[54]
Zihan Li, Yunxiang Li, Qingde Li, Puyang Wang, Dazhou Guo, Le Lu, Dakai Jin, You Zhang, and Qingqi Hong. 2023. Lvit: language meets vision transformer in medical image segmentation.IEEE Transactions on Medical Imaging (TMI)43, 1 (2023), 96–107
2023
-
[55]
Yonglin Li, Jing Zhang, Xiao Teng, Long Lan, and Xinwang Liu. 2023. Ref- sam: Efficiently adapting segmenting anything model for referring video object segmentation.arXiv preprint arXiv:2307.00997(2023)
2023 arXiv
-
[56]
Xvyuan Liu, Xiangfei Qiu, Xingjian Wu, Zhengyu Li, Chenjuan Guo, Jilin Hu, and Bin Yang. 2025. Rethinking Irregular Time Series Forecasting: A Simple yet Effective Baseline.arXiv preprint arXiv:2505.11250(2025)
2025
-
[57]
Ilya Loshchilov and Frank Hutter. 2017. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101(2017)
2017 arXiv
-
[58]
Shilin Lu, Yanzhu Liu, and Adams Wai-Kin Kong. 2023. Tf-icon: Diffusion-based training-free cross-domain image composition. InProceedings of the IEEE/CVF International Conference on Computer Vision (CVPR). 2294–2305
2023
-
[59]
Shilin Lu, Zilan Wang, Leyang Li, Yanzhu Liu, and Adams Wai-Kin Kong. 2024. Mace: Mass concept erasure in diffusion models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 6430–6440
2024
-
[60]
Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Qing Jiang, Chunyuan Li, Jianwei Yang, Hang Su, et al. 2025. Grounding dino: Marry- ing dino with grounded pre-training for open-set object detection. InEuropean Conference on Computer Vision (ECCV). Springer, 38–55
2025
-
[61]
Xiangde Luo, Wenjun Liao, Jianghong Xiao, Jieneng Chen, Tao Song, Xiaofan Zhang, Kang Li, Dimitris N Metaxas, Guotai Wang, and Shaoting Zhang. 2022. WORD: A large scale dataset, benchmark and clinical applicable study for abdominal organ segmentation from CT image.Medical Imag...
2022
-
[62]
Jun Ma, Yuting He, Feifei Li, Lin Han, Chenyu You, and Bo Wang. 2024. Segment anything in medical images.Nature Communications15, 1 (2024), 654
2024
-
[63]
Jun Ma, Yao Zhang, Song Gu, Cheng Ge, Shihao Ma, Adamo Young, Cheng Zhu, Kangkang Meng, Xin Yang, Ziyan Huang, et al. 2023. Unleashing the strengths of unlabeled data in pan-cancer abdominal organ quantification: the flare22 challenge.arXiv preprint arXiv:2308.05862(2023)
2023 arXiv
-
[64]
Jun Ma, Yao Zhang, Song Gu, Cheng Zhu, Cheng Ge, Yichi Zhang, Xingle An, Congcong Wang, Qiyuan Wang, Xin Liu, et al . 2021. Abdomenct-1k: Is abdominal organ segmentation a solved problem?IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI)44, 10 (2021), 6695–6714
2021
-
[65]
Shilin Lu, Zihan Zhou, Jiayou Lu, Yuanzhi Zhu, and Adams Wai-Kin Kong
-
[66]
Robust watermarking using generative priors against image editing: From benchmarking to advances.arXiv preprint arXiv:2410.18775(2024)
2024 arXiv
-
[67]
Jay N Paranjape, Nithin Gopalakrishnan Nair, Shameema Sikder, S Swaroop Vedula, and Vishal M Patel. 2024. Adaptivesam: Towards efficient tuning of sam for surgical scene segmentation. InAnnual Conference on Medical Image Understanding and Analysis. Springer, 187–201
2024
-
[68]
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gre- gory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, L...
2019
-
[69]
Haotian Qian, Yinda Chen, Shengtao Lou, Fahad Shahbaz Khan, Xiaogang Jin, and Deng-Ping Fan. 2024. Maskfactory: Towards high-quality synthetic data generation for dichotomous image segmentation.Advances in Neural Information Processing Systems (NeurIPS)37 (2024), 66455–66478
2024
-
[70]
Xiangfei Qiu, Hanyin Cheng, Xingjian Wu, Jilin Hu, and Chenjuan Guo. 2025. A Comprehensive Survey of Deep Learning for Multivariate Time Series Forecast- ing: A Channel Strategy Perspective.arXiv preprint arXiv:2502.10721(2025)
2025
-
[71]
Stanislav Nikolov, Sam Blackwell, Alexei Zverovitch, Ruheena Mendes, Michelle Livne, Jeffrey De Fauw, Yojan Patel, Clemens Meyer, Harry Askham, Bernardino Romera-Paredes, et al . 2018. Deep learning to achieve clinically applicable segmentation of head and neck anatomy for rad...
2018 arXiv
-
[72]
Ting Pan, Lulu Tang, Xinlong Wang, and Shiguang Shan. 2025. Tokenize anything via prompting. InEuropean Conference on Computer Vision (ECCV). Springer, 330–348
2025
-
[73]
Xiangfei Qiu, Xingjian Wu, Yan Lin, Chenjuan Guo, Jilin Hu, and Bin Yang
-
[74]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sand- hini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al
-
[75]
Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Ronghang Hu, Chaitanya Ryali, Tengyu Ma, Haitham Khedr, Roman Rädle, Chloe Rolland, Laura Gustafson, Eric Mintun, Junting Pan, Kalyan Vasudev Alwala, Nicolas Carion, Chao-Yuan Wu, Ross Girshick, Piotr Dollar, and Christoph Feichtenh...
2025
-
[76]
Tianhe Ren, Shilong Liu, Ailing Zeng, Jing Lin, Kunchang Li, He Cao, Jiayu Chen, Xinyu Huang, Yukang Chen, Feng Yan, et al. 2024. Grounded sam: Assembling open-world models for diverse visual tasks.arXiv preprint arXiv:2401.14159 (2024)
2024 arXiv
-
[77]
Jensen, Zhenli Sheng, and Bin Yang
Xiangfei Qiu, Jilin Hu, Lekui Zhou, Xingjian Wu, Junyang Du, Buang Zhang, Chenjuan Guo, Aoying Zhou, Christian S. Jensen, Zhenli Sheng, and Bin Yang
-
[78]
InProceedings Of The Vldb Endowment (PVLDB)
TFB: Towards Comprehensive and Fair Benchmarking of Time Series Forecasting Methods. InProceedings Of The Vldb Endowment (PVLDB). 2363– 2377
-
[79]
Xiangfei Qiu, Zhe Li, Wanghui Qiu, Shiyan Hu, Lekui Zhou, Xingjian Wu, Zhengyu Li, Chenjuan Guo, Aoying Zhou, Zhenli Sheng, et al . 2025. Tab: Unified benchmarking of time series anomaly detection methods.arXiv preprint arXiv:2506.18046(2025)
2025 arXiv
-
[80]
Amber L Simpson, Michela Antonelli, Spyridon Bakas, Michel Bilello, Keyvan Farahani, Bram Van Ginneken, Annette Kopp-Schneider, Bennett A Landman, Geert Litjens, Bjoern Menze, et al . 2019. A large annotated medical image dataset for the development and evaluation of segmentat...
2019 arXiv
-
[81]
InSIGKDD
DUET: Dual Clustering Enhanced Multivariate Time Series Forecasting. InSIGKDD. 1185–1196
-
[82]
Chenyue Song, Chen Hui, Wei Zhang, Haiqi Zhu, Shaohui Liu, Hong Huang, and Feng Jiang. 2025. BPCLIP: A Bottom-up Image Quality Assessment from Distortion to Semantics Based on CLIP.arXiv preprint arXiv:2506.17969(2025)
2025 arXiv
-
[83]
Chenyue Song, Wei Zhang, Feng Jiang, Deqin Zheng, Ruichen Gao, and Haiqi Zhu. 2024. ADTAH: Neuron 3D Reconstruction Via Adaptive Distance Trans- formation and Adaptive Hessian Matrix. In2024 IEEE International Conference on Bioinformatics and Biomedicine (BIBM). IEEE, 3711–3716
2024
-
[84]
Alexandros Stergiou and Ronald Poppe. 2022. Adapool: Exponential adaptive pooling for information-retaining downsampling.IEEE Transactions on Image Processing (TIP)32 (2022), 251–266
2022
-
[85]
Mingxing Tan, Ruoming Pang, and Quoc V Le. 2020. Efficientdet: Scalable and efficient object detection. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 10781–10790
2020
-
[86]
Chaitanya Ryali, Yuan-Ting Hu, Daniel Bolya, Chen Wei, Haoqi Fan, Po-Yao Huang, Vaibhav Aggarwal, Arkabandhu Chowdhury, Omid Poursaeed, Judy Hoffman, et al. 2023. Hiera: A hierarchical vision transformer without the bells- and-whistles. InInternational Conference on Machine Le...
2023
-
[87]
Arnaud Arindra Adiyoso Setio, Alberto Traverso, Thomas De Bel, Moira SN Berens, Cas Van Den Bogaard, Piergiorgio Cerello, Hao Chen, Qi Dou, Maria Evelina Fantacci, Bram Geurts, et al. 2017. Validation, comparison, and combination of algorithms for automatic detection of pulmon...
2017
-
[88]
Chao Shang, Zichen Song, Heqian Qiu, Lanxiao Wang, Fanman Meng, and Hongliang Li. 2024. Prompt-Driven Referring Image Segmentation with Instance Contrasting. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 4124–4134
2024
-
[89]
Yan Wang, Yuyin Zhou, Wei Shen, Seyoun Park, Elliot K Fishman, and Alan L Yuille. 2019. Abdominal multi-organ segmentation with organ-attention net- works and statistical fusion.Medical Image Analysis (MedIA)55 (2019), 88–102
2019
-
[90]
Chenyue Song, Chen Hui, Qing Lin, Wei Zhang, Siqiao Li, Shengping Zhang, Haiqi Zhu, Zhixuan Li, Shaohui Liu, Feng Jiang, et al. 2025. LVPNet: A Latent- variable-based Prediction-driven End-to-end Framework for Lossless Compres- sion of Medical Images.arXiv preprint arXiv:2506....
2025 arXiv
-
[91]
Wangyu Wu, Tianhong Dai, Zhenhong Chen, Xiaowei Huang, Fei Ma, and Jimin Xiao. 2025. Generative Prompt Controlled Diffusion for weakly supervised semantic segmentation.Neurocomputing638 (2025), 130103
2025
-
[92]
Wangyu Wu, Xianglin Qiu, Siqi Song, Zhenhong Chen, Xiaowei Huang, Fei Ma, and Jimin Xiao. 2025. Prompt categories cluster for weakly supervised semantic segmentation. InProceedings of the Computer Vision and Pattern Recognition Conference (CVPR) Workshops. 3198–3207
2025
-
[93]
Wangyu Wu, Siqi Song, Xianglin Qiu, Xiaowei Huang, Fei Ma, and Jimin Xiao
-
[94]
Xiangyu Wu, Qing-Yuan Jiang, Yang Yang, Yi-Feng Wu, Qing-Guo Chen, and Jianfeng Lu. 2024. Tai++: Text as image for multi-label image classification by co-learning transferable prompt.arXiv preprint arXiv:2405.06926(2024)
2024 arXiv
-
[95]
Haoyu Wang, Sizheng Guo, Jin Ye, Zhongying Deng, Junlong Cheng, Tianbin Li, Jianpin Chen, Yanzhou Su, Ziyan Huang, Yiqing Shen, et al. 2025. Sam-med3d: towards general-purpose segmentation models for volumetric medical images. InEuropean Conference on Computer Vision (ECCV). S...
2025
-
[96]
Haoxiang Wang, Pavan Kumar Anasosalu Vasu, Fartash Faghri, Raviteja Vemula- palli, Mehrdad Farajtabar, Sachin Mehta, Mohammad Rastegari, Oncel Tuzel, and Hadi Pouransari. 2024. SAM-CLIP: Merging Vision Foundation Models Towards Semantic and Spatial Understanding. InProceedings...
2024
-
[97]
Yihao Wang, Meng Yang, and Rui Cao. 2024. Fine-grained Semantic Alignment with Transferred Person-SAM for Text-based Person Retrieval. InProceedings of the 32nd ACM International Conference on Multimedia (ACM MM). 5432–5441
2024
-
[98]
Xi Xiao, Zhengji Li, Wentao Wang, Jiacheng Xie, Houjie Lin, Swalpa Kumar Roy, Tianyang Wang, and Min Xu. 2025. TD-RD: A Top-Down Benchmark with Real- Time Framework for Road Damage Detection.arXiv preprint arXiv:2501.14302 (2025)
2025 arXiv
-
[99]
Xiaobao Wei, Jiajun Cao, Yizhu Jin, Ming Lu, Guangyu Wang, and Shanghang Zhang. 2024. I-medsam: Implicit medical image segmentation with segment anything. InEuropean Conference on Computer Vision. Springer, 90–107
2024
-
[100]
Zhaozhi Xie, Bochen Guan, Weihao Jiang, Muyang Yi, Yue Ding, Hongtao Lu, and Lei Zhang. 2024. PA-SAM: Prompt Adapter SAM for High-Quality Image Segmentation. In2024 IEEE International Conference on Multimedia and Expo (ICME). 1–6
2024
-
[101]
Yi Xin, Junlong Du, Qiang Wang, Zhiwen Lin, and Ke Yan. 2024. VMT-Adapter: Parameter-Efficient Transfer Learning for Multi-Task Dense Scene Understand- ing. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 38. 16085–16093
2024
-
[102]
Yi Xin, Junlong Du, Qiang Wang, Ke Yan, and Shouhong Ding. 2024. MmAP: Multi-modal Alignment Prompt for Cross-domain Multi-task Learning. InPro- ceedings of the AAAI Conference on Artificial Intelligence, Vol. 38. 16076–16084
2024
-
[103]
InCompanion Proceedings of the ACM on Web Conference
Image fusion for cross-domain sequential recommendation. InCompanion Proceedings of the ACM on Web Conference. 2196–2202
-
[104]
Zhen Yao and Mooi Choo Chuah. 2024. Event-guided low-light video semantic segmentation.arXiv preprint arXiv:2411.00639(2024)
2024
-
[105]
Xingjian Wu, Xiangfei Qiu, Hongfan Gao, Jilin Hu, Bin Yang, and Chenjuan Guo. 2025. K 2VAE: A Koopman-Kalman Enhanced Variational AutoEncoder for Probabilistic Time Series Forecasting.arXiv preprint arXiv:2505.23017(2025)
2025 arXiv
-
[106]
Xingjian Wu, Xiangfei Qiu, Zhengyu Li, Yihang Wang, Jilin Hu, Chenjuan Guo, Hui Xiong, and Bin Yang. 2025. CATCH: Channel-Aware multivariate Time Series Anomaly Detection via Frequency Patching. InInternational Conference on Learning Representations (ICLR)
2025
-
[107]
Xiangyu Wu, Feng Yu, Qing-Guo Chen, Yang Yang, and Jianfeng Lu. 2025. Multi- Label Test-Time Adaptation with Bound Entropy Minimization.arXiv preprint arXiv:2502.03777(2025)
2025 arXiv
-
[108]
Haobo Yuan, Xiangtai Li, Chong Zhou, Yining Li, Kai Chen, and Chen Change Loy. 2025. Open-vocabulary SAM: Segment and recognize twenty-thousand classes interactively. InEuropean Conference on Computer Vision (ECCV). Springer, 419–437
2025
-
[109]
Xi Xiao, Yunbei Zhang, Thanh-Huy Nguyen, Ba-Thinh Lam, Janet Wang, Jihun Hamm, Tianyang Wang, Xingjian Li, Xiao Wang, Hao Xu, et al. 2025. Describe Anything in Medical Images.arXiv preprint arXiv:2505.05804(2025)
2025 arXiv
-
[110]
Pengyu Zeng, Wen Gao, Jizhizi Li, Jun Yin, Jiling Chen, and Shuai Lu. 2025. Automated residential layout generation and editing using natural language and images.Automation in Construction174 (2025), 106133
2025
-
[111]
Pengyu Zeng, Wen Gao, Jun Yin, Pengjian Xu, and Shuai Lu. 2024. Residential floor plans: Multi-conditional automatic generation using diffusion models. Automation in Construction162 (2024), 105374
2024
-
[112]
Zhi Zeng, Minnan Luo, Xiangzheng Kong, Huan Liu, Hao Guo, Hao Yang, Zihan Ma, and Xiang Zhao. 2024. Mitigating World Biases: A Multimodal Multi-View Debiasing Framework for Fake News Video Detection. InProceedings of the 32nd ACM International Conference on Multimedia. 6492–6500
2024
-
[113]
Yi Xin, Siqi Luo, Haodi Zhou, Junlong Du, Xiaohong Liu, Yue Fan, Qing Li, and Yuntao Du. 2024. Parameter-efficient fine-tuning for pre-trained vision models: A survey.arXiv preprint arXiv:2402.02242(2024)
2024
-
[114]
Ruicheng Zhang, Haowei Guo, Zeyu Zhang, Puxin Yan, and Shen Zhao. 2025. GAMED-Snake: Gradient-aware Adaptive Momentum Evolution Deep Snake Model for Multi-organ Segmentation.arXiv preprint arXiv:2501.12844(2025)
2025 arXiv
-
[115]
Zhen Yao, Xiaowen Ying, and Mooi Choo Chuah. 2025. Rethinking RGB-Event Semantic Segmentation with a Novel Bidirectional Motion-enhanced Event Representation.arXiv preprint arXiv:2505.01548(2025)
2025
-
[116]
Xinlei Yu, Ahmed Elazab, Ruiquan Ge, Hui Jin, Xinchen Jiang, Gangyong Jia, Qing Wu, Qinglei Shi, and Changmiao Wang. 2024. ICH-SCNet: Intracerebral Hemorrhage Segmentation and Prognosis Classification Network Using CLIP- guided SAM mechanism. In2024 IEEE International Conferen...
2024
-
[117]
Xinlei Yu, Ahmed Elazab, Ruiquan Ge, Jichao Zhu, Lingyan Zhang, Gangyong Jia, Qing Wu, Xiang Wan, Lihua Li, and Changmiao Wang. 2025. ICH-PRNet: a cross-modal intracerebral haemorrhage prognostic prediction method using joint-attention interaction mechanism.Neural Networks184 ...
2025
-
[118]
Yuyin Zhou, Zhe Li, Song Bai, Chong Wang, Xinlei Chen, Mei Han, Elliot Fishman, and Alan L Yuille. 2019. Prior-aware neural network for partially- supervised multi-organ segmentation. InProceedings of the IEEE/CVF Interna- tional Conference on Computer Vision (ICCV). 10672–10681
2019
-
[119]
Wenxi Yue, Jing Zhang, Kun Hu, Yong Xia, Jiebo Luo, and Zhiyong Wang. 2024. Surgicalsam: Efficient class promptable surgical instrument segmentation. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 38. 6890–6898
2024
-
[120]
Xueyan Zou, Jianwei Yang, Hao Zhang, Feng Li, Linjie Li, Jianfeng Wang, Lijuan Wang, Jianfeng Gao, and Yong Jae Lee. 2024. Segment everything everywhere all at once.Advances in Neural Information Processing Systems (NeurIPS)36 (2024), 19769–19782. Conference acronym ’MM, Octob...
2024
-
[123]
Jielu Zhang, Zhongliang Zhou, Gengchen Mai, Mengxuan Hu, Zihan Guan, Sheng Li, and Lan Mu. 2023. Text2seg: Remote sensing image semantic segmen- tation via text-guided visual foundation models.arXiv preprint arXiv:2304.10597 (2023)
2023
-
[125]
Ruicheng Zhang, Yu Sun, Zeyu Zhang, Jinai Li, Xiaofan Liu, Au Hoi Fan, Haowei Guo, and Puxin Yan. 2025. MARL-MambaContour: Unleashing Multi-Agent Deep Reinforcement Learning for Active Contour Optimization in Medical Image Segmentation.arXiv preprint arXiv:2506.18679(2025)
2025 arXiv
-
[126]
Yuxuan Zhang, Tianheng Cheng, Rui Hu, Lei Liu, Heng Liu, Longjin Ran, Xiaoxin Chen, Wenyu Liu, and Xinggang Wang. 2024. Evf-sam: Early vision-language fu- sion for text-prompted segment anything model.arXiv preprint arXiv:2406.20076 (2024)
2024 arXiv
-
[127]
Haochen Zhao, Hui Meng, Deqian Yang, Xiaozheng Xie, Xiaoze Wu, Qingfeng Li, and Jianwei Niu. 2024. GuidedNet: Semi-Supervised Multi-Organ Segmenta- tion via Labeled Data Guide Unlabeled Data. InProceedings of the 32nd ACM International Conference on Multimedia (ACM MM). 886–895
2024
-
[129]
Jiayuan Zhu, Abdullah Hamdi, Yunli Qi, Yueming Jin, and Junde Wu. 2024. Medical sam 2: Segment medical images as video via segment anything model 2.arXiv preprint arXiv:2408.00874(2024)
2024 arXiv
-
[131]
In this stage, the initial learning rate is2𝑒−4 to get superior global optimal solutions
Stage-II: we further add other components of our model into optimization for 280 epochs, except the four pretrained transformer blocks in the image encoder, and the pretrained text and image encoders in the Cross-Modal Semantic Interaction module. In this stage, the initial le...
2025
-
[2013]
The Cancer Imaging Archive (TCIA): maintaining and operating a public information repository.Journal of Digital Imaging26 (2013), 1045–1057
2013
-
[2021]
InInternational Conference on Machine Learning (ICML)
Learning transferable visual models from natural language supervision. InInternational Conference on Machine Learning (ICML). PMLR, 8748–8763
-
[2023]
InProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)
Segment anything. InProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 4015–4026
-
[2024]
Msdet: Receptive field enhanced multiscale detection for tiny pulmonary nodule.arXiv preprint arXiv:2409.14028(2024)
2024 arXiv
-
[2025]
InProceedings of the AAAI Conference on Artificial Intelligence, Vol
Encoder: Entity mining and modification relation binding for composed image retrieval. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 39. 5101–5109
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.