{"id":"4ce4313d-451f-49cc-a16f-2ef87b146fd4","arxiv_id":"2505.06118","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A systematic review of 23 deep-learning lymph node segmentation studies finds that CNN-based methods and ultrasound images currently yield the highest average Dice scores.","lead":"This paper is a systematic review of 23 studies on using deep learning to automatically outline lymph nodes in CT, PET/CT, MRI, and ultrasound images. It compares the accuracy of different model families and imaging types, and identifies data scarcity and annotation cost as the main barriers to progress.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Technique ranking in Figure 17 is confounded by modality mix; CNN-group studies are mostly ultrasound while encoder-decoder-group studies are mostly CT/PET, so the 'CNN best' claim may not survive within-modality comparison.","rationale":"After reading the paper and the reader's verdict, the most load-bearing part of the central claim is the quantitative comparison of techniques. The paper's strongest_claim explicitly states that CNN methods achieve the highest mean Dice and ultrasound the highest modality mean. Both conclusions are based on simple means over Table II. The reader correctly identifies the pooled Dices as heterogeneous. In stress-testing, I found an even sharper threat: the CNN group is disproportionately composed of ultrasound studies, and the encoder-decoder group is dominated by CT/PET. Since ultrasound is the best-performing modality, the technique ranking may simply mirror the modality mix. A within-modality stratification shows the CNN advantage shrinks to within-noise. The paper's own limitations section concedes dataset variation biases comparisons, but it does not perform any stratified analysis. Therefore the quantitative ranking should be presented with a caveat or removed; the qualitative overview of architectures and challenges remains useful. This is consistent with the reader's CONDITIONAL verdict.","tokens_in":28230,"tokens_out":7201,"duration_ms":62187,"concrete_test":"Using Table II, build a two-way table of per-study Dice by technique and modality and compute within-modality means (e.g., CNN-US vs Encoder-Decoder-US; CNN-CT vs Encoder-Decoder-CT). If the CNN advantage reverses or falls within one standard error in every modality, then Figure 17's 'CNN best' conclusion is an artifact of modality mix and should be withdrawn.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that CNN-based methods (mean Dice 0.836) and ultrasound (mean Dice 0.857) are currently the most accurate rests on unadjusted pooling of per-study Dice. The more specific, load-bearing flaw is that the technique ranking in Figure 17 is confounded by imaging modality. Of the four CNN-group studies, two are ultrasound (CFS-FCN 0.858; Zhang et al. 0.895) and one MRI; the encoder-decoder group contains six CT and two PET/CT studies, modalities whose overall means are 0.750 and 0.716. Because ultrasound Dice are systematically higher, the CNN group's composition can explain its top rank. Within ultrasound, CNN mean Dice is 0.877 (n=2) vs encoder-decoder mean 0.850 (n=6); within CT, CNN has one study (0.821) vs encoder-decoder mean 0.797 (n=6). These differences are small and not statistically defensible. The paper acknowledges dataset variation in Section IV.F.3, yet presents Figures 17-18 as quantitative evidence. The modality ranking is similarly vulnerable to dataset difficulty and LN fraction in the image.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a systematic review of deep learning methods for lymph node (LN) segmentation in medical imaging. Following a PRISMA-style workflow, the authors searched five databases, identified 198 unique records, and included 23 studies. The review categorizes methods into CNN, encoder-decoder, transformer, object-detection-assisted, and loss-function groups, and compares reported performance (mainly Dice) across techniques and across imaging modalities (CT, PET/CT, MRI, ultrasound). The authors report that CNN-based methods and ultrasound achieve the highest mean Dice scores (0.836 and 0.857, respectively) and claim this is the first comprehensive systematic review on this topic. The paper also discusses challenges, including data scarcity, annotation cost, and clinical integration, and proposes future research directions.","tokens_in":28476,"tokens_out":6439,"duration_ms":57991,"significance":"A well-conducted systematic review of LN segmentation would be valuable given the clinical importance of automated nodal staging. The strengths of the manuscript are its broad literature coverage, the inter-reviewer agreement assessment (kappa = 0.793), the detailed summary of included studies (Table II), and the structured description of architectures and loss functions. The qualitative synthesis, such as the prevalence of U-Net variants and the diversity of private datasets, is genuinely useful. However, the quantitative claims that CNN methods and ultrasound imaging are currently the most accurate are not supported by the evidence as presented, because the pooled means are unadjusted for dataset difficulty, modality, anatomical site, and evaluation protocol, a concern the authors themselves acknowledge in Section IV.F.3. The value of the review therefore lies primarily in its cataloging and qualitative synthesis rather than in its comparative performance rankings.","major_comments":[{"comment":"The claim that CNN-based methods achieve the highest mean Dice (0.836) is confounded by imaging modality. Within the CNN group, two of four studies are ultrasound (0.858 and 0.895), while the encoder-decoder group contains six CT and two PET/CT studies whose modalities have lower overall means (0.750 and 0.716 in Figure 18). A stratified view shows within-ultrasound CNN mean Dice of 0.877 (n=2) versus encoder-decoder mean of 0.850 (n=6), and within-CT CNN Dice of 0.821 (n=1) versus encoder-decoder mean of 0.797 (n=6); these differences are small and not statistically testable with the reported data. The conclusion in Section IV.A that \"CNN-based methods are currently the most accurate\" and the corresponding abstract statement should therefore be either removed or replaced by a modality-stratified analysis.","section":"Section IV.A / Figure 17"},{"comment":"The PRISMA flow diagram is internally inconsistent. The figure shows 198 unique records, then \"Full-text articles assessed for eligibility (n=23),\" and then \"Articles excluded, with reasons (n=175)\" beneath the full-text assessment box. If only 23 full-text articles were assessed, they cannot have generated 175 full-text exclusions. The accompanying text states \"Among the 198 full-text publications, 175 were removed,\" which implies that 198 full-text publications were screened, not 23. This discrepancy makes the screening totals unreproducible and needs to be corrected consistently in both the text and the figure.","section":"Figure 1 and Section II.B"},{"comment":"The modality ranking that places ultrasound first (mean Dice 0.857) is based on the same kind of unadjusted pooling across studies as the technique ranking. The studies differ in the proportion of the image occupied by LNs, anatomical sites, imaging protocols, annotation quality, and dataset sizes; the discussion attributes ultrasound's advantage to image characteristics, but the same pattern could arise from dataset selection bias. At a minimum, the authors should show per-study Dice values with dataset size and site annotations, and should refrain from presenting Figure 18 as evidence of a modality ranking unless a within-modality comparison on comparable datasets is possible.","section":"Section IV.B / Figure 18"},{"comment":"The manuscript cites PRISMA 2020 [19] but does not report a risk-of-bias or quality assessment of the included studies, which PRISMA 2020 explicitly requires. This omission is especially relevant because the review pools quantitative results from heterogeneous studies. The authors should either add a risk-of-bias assessment or clearly state in Section IV.F that it was not performed, and temper the \"systematic review\" claims accordingly.","section":"Sections II and IV.F"}],"minor_comments":[{"comment":"The paragraph describing MA-Net contains the typo \"It it highly generalizable\"; it should read \"It is highly generalizable.\"","section":"Section III.B"},{"comment":"The auto-LNDS row contains \"Testin: 1,192,\" which should be \"Test_in: 1,192\" to match the internal/external dataset notation used elsewhere in the table.","section":"Table II"},{"comment":"The legend below Figure 16 lists categories such as \"CNN CT Encoder-Decoder PET/CT\" without a clear mapping of colors or markers to individual studies; please add a proper legend key or use distinct markers with a caption explaining the mapping.","section":"Figure 16"},{"comment":"The abstract states \"this is the first study\" without the qualifier \"to the best of our knowledge\" that appears in the Introduction; the abstract should use the same cautious wording.","section":"Abstract and Section I"},{"comment":"The sentence claiming that the best-performing methods achieve Dice scores \"more than twice as high\" as the worst-performing methods is imprecise because the group means 0.836 and 0.612 differ by a factor of 1.37; if the comparison is between individual studies (0.935 versus 0.409), that should be stated explicitly.","section":"Section IV.A"},{"comment":"In the unit row of Table I, the notation \"/\" is ambiguous; it should be written as \"dimensionless\" rather than \"/\" to avoid confusion with division.","section":"Table I"}],"recommendation":"major_revision","confidential_remarks":"The review includes two studies (Zhang et al. [25] and [26]) in which one co-author (M. Ying) is also an author of this review; the \"no conflict of interest\" statement may warrant expansion to disclose this relationship. In addition, the manuscript's primary contribution is descriptive, so editors should consider whether the venue is appropriate for a systematic review without novel experimental results. If the authors can fix the PRISMA flow inconsistency and temper or stratify the performance claims, the review could be a useful reference for the community."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things about this one. First, it is genuinely the first systematic review I know of that is specifically about deep-learning lymph node segmentation, and the authors do a decent job of taxonomy: CNN, encoder-decoder, transformer, object-detection-assisted, loss-function work, with a table that clinicians and incoming graduate students will find useful. Second, the quantitative claims in Figures 17-18 should not be trusted as comparisons, and the paper's own text half-admits this.\n\nWhat it does well: the PRISMA-style search is documented across five databases, inclusion/exclusion criteria are explicit, Cohen's kappa is reported, and the limitations section is candid, especially about private datasets and metric heterogeneity. The discussion of data scarcity, annotation cost, and clinical integration is sensible and not padded. For a reader entering this niche, the paper is a serviceable map of 23 studies and their reported Dice scores.\n\nThe soft spots are real. The PRISMA flow diagram is internally confusing: it says 23 full-text articles were assessed for eligibility but then lists 175 exclusions \"with reasons\" that sum to 198 with the included studies. The numbers can be reverse-engineered, but a PRISMA flow should not require reverse-engineering. More importantly, the technique ranking in Figure 17 is confounded by modality mix. The CNN group is mostly ultrasound studies; the encoder-decoder group is heavily CT and PET/CT, and ultrasound Dice scores run systematically higher. So the claim that \"CNN-based methods are currently the most accurate\" is largely an artifact of which modalities those studies happened to use. Within-modality, the differences are small and not statistically defensible. The modality ranking in Figure 18 has the same weakness: ultrasound images have a higher typical lymph-node fraction, so difficulty is not matched across modalities. Object-detection-assisted methods are judged on two studies, one of which scored 0.409; calling that group's mean \"significantly lower\" is over-reading.\n\nThe citation pattern is fine, including the self-citations to Zhang et al., which are relevant prior work and not hidden. The aggregated Dice values are externally anchored in published studies, so this is not a circularity problem. The central failing is analytic discipline, not integrity.\n\nWho gets value from this paper? A researcher wanting an organized entry point into LN segmentation methods and a list of open problems. Who should not get value from it yet? Anyone citing Figures 17-18 as evidence about which architecture or modality is best.\n\nMy recommendation: this deserves a serious referee, but with a clear request for revision: fix the PRISMA counting, present the Dice aggregations as descriptive summaries rather than comparative findings, and either add a modality-stratified analysis or drop the ranking language entirely. I would engage with it after those changes.","headline":"A useful structural map of the LN-segmentation literature, but its headline quantitative rankings (CNN best, ultrasound best) rest on confounded pooling and should be reframed before the review is published.","tokens_in":28990,"tokens_out":1457,"would_cite":false,"duration_ms":17293,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Deep-learning lymph node segmentation is most accurate with CNN-based methods and ultrasound imaging, according to a 23-study systematic review.","keywords":["lymph node segmentation","deep learning","systematic review","Dice similarity coefficient","convolutional neural networks","encoder-decoder networks","medical imaging modalities","ultrasound imaging"],"falsifier":"Re-run the representative architectures on shared public lymph-node datasets with fixed training/test splits and identical preprocessing and evaluation software; if the ordering among CNN, encoder-decoder, transformer, and object-detection-assisted methods changes—or if CT matches or beats ultrasound—the paper's pooled mean-Dice rankings are an artifact of comparing incompatible studies.","tokens_in":28076,"feed_emoji":"🩻","tokens_out":7642,"duration_ms":71056,"temperature":0.7,"pith_summary":"Automatic lymph node segmentation is the step that would make computer-aided cancer detection and staging practical, and this paper asks which deep-learning recipe works best for it. After screening 411 records down to 23 studies, the review argues that convolutional-network methods currently average the highest Dice score (0.836) and that ultrasound images support the best modality-level results (0.857). It also reports that encoder-decoder networks, especially U-Net variants, are the most widely used and that transformer and object-detection-assisted approaches lag in adoption and accuracy. The value of establishing this is that clinicians and researchers get a map of what to build on, and where the open problems—scarce labeled data, shape variability, modality differences—actually sit.","feed_headline":"CNNs and ultrasound top lymph node segmentation","feed_subtitle":"A systematic review of 23 studies finds CNN methods (mean Dice 0.836) and ultrasound imaging (0.857) lead the field.","key_machinery":"The load-bearing quantitative object is the Dice similarity coefficient (DSC), defined as $2|P \\cap G|/(|P|+|G|)$ for predicted pixels $P$ and ground-truth pixels $G$. Because DSC is the only metric reported by all 23 studies, the review uses it as the common yardstick to rank technique families and imaging modalities, and it is the basis of the mean Dice comparisons. The other machinery is the selection funnel: five databases are searched with a fixed query, 411 records are screened by two reviewers, and 23 studies survive inclusion criteria requiring a fully automated deep-learning method, a CT/PET/MRI/US modality, and quantitative segmentation results. The funnel determines which evidence enters the pooled comparisons.","core_discovery":"The paper claims to be the first systematic review dedicated to deep-learning lymph node segmentation. Its central finding is comparative: aggregating the Dice similarity coefficients reported by 23 included studies, CNN-based methods achieve the highest mean score among technique families (0.836), ahead of encoder-decoder methods (0.802), loss-function-focused studies (0.760), transformers (0.736), and object-detection-assisted methods (0.612). By imaging modality, ultrasound leads with a mean of 0.857, followed by MRI (0.792), CT (0.750), and PET/CT (0.716). Within each modality the best reported result is an encoder-decoder variant except for MRI, where a Mask R-CNN-based detection-assisted method wins, and the review attributes the difference to dataset size, augmentation, and skip connections. The authors themselves caution that pooling scores from heterogeneous private datasets may bias the comparisons, and they identify data scarcity, annotation cost, and limited use of newer architectures as the principal barriers, proposing multimodal fusion, transfer learning, and large pre-trained models as the way forward.","pith_inferences":["Editorial inference: The modality ranking may reflect how tightly an image is framed around the node rather than intrinsic image quality; ultrasound images are typically zoomed onto the node, whereas CT and PET frames contain large surrounding anatomy. A fair cross-modality test should crop CT/PET regions to comparable fields of view.","Editorial inference: The pooled technique ranking is sensitive to the unbalanced representation of each category; only one transformer study and two detection-assisted studies feed the means, so a few new results could overturn the ordering.","Editorial inference: A natural extension is a meta-regression that weights each study by dataset size and patient count, and tests heterogeneity in Dice across sites and annotation protocols; this would show whether the headline comparisons survive statistical pooling.","Editorial inference: If large pre-trained segmentation models are fine-tuned on lymph node data, they may reduce the data-scarcity barrier identified here, but their performance on small, irregular nodes in CT/PET would need direct comparison against the reviewed CNN and encoder-decoder baselines."],"forward_implications":["CNN-based segmentation (mean Dice 0.836) is, on current evidence, the most accurate deep-learning family for lymph node segmentation, slightly ahead of encoder-decoder networks (0.802).","Ultrasound imaging appears to offer the most favorable setting for lymph node segmentation (mean Dice 0.857), ahead of MRI (0.792), CT (0.750), and PET/CT (0.716).","U-Net and its variants dominate practice, and the best results in CT, PET/CT, and ultrasound all use encoder-decoder designs with skip connections, large datasets, and data augmentation.","Transformer and object-detection-assisted methods are currently underrepresented and underperform CNNs, so their potential in lymph node segmentation remains largely untested.","Future progress is expected to come from multimodal fusion (for example grayscale plus Doppler ultrasound), transfer learning, and fine-tuning large pre-trained models, not from a single architecture."],"supporting_citations":[{"why":"Supplies the systematic-review reporting standard that structures the search, screening, and inclusion flow.","marker":"[19]"},{"why":"Provides the CNN-based segmentation result on a public CT dataset that anchors the CNN category's mean.","marker":"[23]"},{"why":"Reports the highest CT Dice (0.935) with a modified UNet++ on the largest datasets, setting the review's top score.","marker":"[35]"},{"why":"Holds the best ultrasound result (Dice 0.924) with MUNet, supporting the modality ranking.","marker":"[44]"},{"why":"Gives the best PET/CT result (Dice 0.770) via DiSegNet, defining that modality's ceiling.","marker":"[37]"},{"why":"Supplies the best MRI result (Dice 0.815) with a Mask R-CNN detection-assisted pipeline.","marker":"[60]"},{"why":"Represents the only transformer-based method, DE-Net, used to estimate transformer performance.","marker":"[54]"},{"why":"Reports the lowest Dice (0.409) for object-detection-assisted segmentation, anchoring that category's mean.","marker":"[59]"},{"why":"Defines the U-Net encoder-decoder architecture on which most included studies build.","marker":"[6]"},{"why":"Defines the FCN convolution-only architecture used by several CNN methods in the review.","marker":"[22]"}],"fun_headline_variants":["CNN, ultrasound lead lymph node segmentation review","CNN best, ultrasound best in lymph node segmentation review","First systematic review: CNNs, ultrasound lead segmentation","Ultrasound, CNNs top 23-study lymph node segmentation review","Dice 0.836: CNNs lead in lymph node segmentation review"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The rankings assume that the accuracy scores reported by the 23 studies, which use different private datasets with different sizes, body sites, scanning protocols, and annotation quality, can be averaged together as if they were directly comparable.","fun_headline_variants_meta":{"raw":{"variants":["CNN, ultrasound lead lymph node segmentation review","CNN best, ultrasound best in lymph node segmentation review","First systematic review: CNNs, ultrasound lead segmentation","Ultrasound, CNNs top 23-study lymph node segmentation review","Dice 0.836: CNNs lead in lymph node segmentation review"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000566,"raw_usage":{"total_tokens":2688,"prompt_tokens":959,"completion_tokens":1729,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":575,"completion_tokens_details":{"reasoning_tokens":1646}},"tokens_in":575,"tokens_out":1729,"duration_ms":11028,"temperature":1.0,"reasoning_tokens":1646,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T22:46:55.545702+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the representative architectures on shared public lymph-node datasets with fixed training/test splits and identical preprocessing and evaluation software; if the ordering among CNN, encoder-decoder, transformer, and object-detection-assisted methods changes—or if CT matches or beats ultrasound—the paper's pooled mean-Dice rankings are an artifact of comparing incompatible studies.","supporting_citations":[{"cited_title":"A review on the use of deep learning for medical images segmentation,","cited_arxiv_id":null,"evidence_quote":"Supplies the systematic-review reporting standard that structures the search, screening, and inclusion flow."},{"cited_title":"The tumor target segmentation of nasopharyngeal cancer in ct images based on deep learning methods,","cited_arxiv_id":null,"evidence_quote":"Reports the highest CT Dice (0.935) with a modified UNet++ on the largest datasets, setting the review's top score."},{"cited_title":"Mediastinal lymph node detection and segmentation using deep learning,","cited_arxiv_id":null,"evidence_quote":"Gives the best PET/CT result (Dice 0.770) via DiSegNet, defining that modality's ceiling."},{"cited_title":"Efficientvit: Lightweight multi-scale attention for high-resolution dense prediction,","cited_arxiv_id":null,"evidence_quote":"Supplies the best MRI result (Dice 0.815) with a Mask R-CNN detection-assisted pipeline."},{"cited_title":"Deep learning-assisted ultrasonic diagnosis of cervical lymph node metastasis of thyroid cancer: a retrospective study of 3059 patients,","cited_arxiv_id":null,"evidence_quote":"Represents the only transformer-based method, DE-Net, used to estimate transformer performance."},{"cited_title":"Medvit: a robust vision transformer for generalized medical image classification,","cited_arxiv_id":null,"evidence_quote":"Reports the lowest Dice (0.409) for object-detection-assisted segmentation, anchoring that category's mean."},{"cited_title":"Imagenet classification with deep convolutional neural networks,","cited_arxiv_id":null,"evidence_quote":"Defines the FCN convolution-only architecture used by several CNN methods in the review."}],"review_version":1}