{"id":"9bd541db-8a2a-4ae6-b3a0-ad6c89e1140b","arxiv_id":"2507.19230","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Applying the ULS23 lesion segmentation model to longitudinal CT data causes a sharp drop in follow-up accuracy and lesion tracking, driven by the model's assumption that every lesion sits at the center of its input patch.","lead":"Researchers tested a popular AI tool that finds lesions in CT scans on pairs of scans from the same cancer patients over time. They found the tool often fails to find the same lesion in the follow-up scan, especially when the scan alignment is even slightly off.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Real-data causal claim is not established; the controlled displacement experiment, restricted to the 30 best-segmented lesions, does not demonstrate that VOI misalignment, rather than appearance change, drives the follow-up Dice drop.","rationale":"The reader's weakest_assumption is exactly the causal attribution concern: the follow-up Dice drop is attributed to VOI misalignment, but appearance and size changes are acknowledged confounds that are not controlled. I agree with that identification and add a specific technical observation that strengthens it: the controlled displacement experiment deliberately selects the 30 best-segmented lesions, which are also the most likely to be unaffected by appearance-based degradation, so the mechanism's transfer to the real-data cohort is not tested. A targeted stratification test (binning by registration error and comparing Dice change per bin, and re-running the analysis on the full lesion set) would settle the question. The paper's core observation—ULS23's sharp dependence on centered lesions—is well supported by the controlled experiment itself, so the verdict should remain CONDITIONAL rather than REJECT. The central caveat is real and addressable, and the paper honestly acknowledges it in the Discussion; my concern is that the claimed confirmation in Section 3.2 overstates what the experimental design can establish. I therefore agree with the reader's assessment and recommend no change to the verdict.","tokens_in":7162,"tokens_out":1702,"duration_ms":14554,"concrete_test":"Stratify the longitudinal experiment by lesion-level registration error. For each follow-up lesion, compute the Euclidean distance between propagated and true centroid (data already present in Fig. 3a) and bin lesions into <5, 5-10, 10-20, >20 mm. Plot the follow-up Dice relative to baseline Dice within each bin. If the Dice drop is small for low-error bins and large only for high-error bins, the causal claim is supported. Additionally, restrict both the real-data and displacement analyses to the same subset (e.g., all lesions, not only the 30 best) to test whether the selection artifact changes the conclusion. If the Dice drop appears equally in well-registered lesions with large appearance change, the confound dominates.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim rests on two results: (1) the follow-up Dice drop in the real longitudinal experiment (Fig. 3b) is primarily caused by VOI misalignment, and (2) the controlled displacement experiment (Fig. 5a) reveals the mechanism. The controlled experiment convincingly shows that shifts >20 mm cause near-total failure. However, the causal attribution in the real-data experiment is not established. The paper itself acknowledges (Discussion, paragraph 2) that changes in lesion morphology and appearance between scans are confounding factors, and the real-data analysis does not control for lesion size change, split/merge events, or new lesion appearance. Moreover, the controlled displacement experiment was performed only on the 30 lesions with the highest baseline/follow-up Dice (Section 2.3, 'Controlled VOI displacement analysis'), a selection that removes the very cases where appearance or size changes are most likely to degrade segmentation. Therefore, the claim that 'this controlled experiment confirms that the performance drop seen in real-world follow-up data (Figure 3) is caused by VOI misalignment' (end of Section 3.2) overreaches. The mechanism is real for large displacements, but its dominance in the real-data cohort is unverified; the real-data portion would also require knowing the distribution of registration errors among the lesions that actually fail. The sharp threshold behavior also does not, by itself, explain the gradual, long-tailed distribution of follow-up Dice scores; a direct comparison of registration error magnitude with per-lesion Dice change is missing.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper evaluates the ULS23 baseline lesion segmentation model on a public longitudinal CT dataset of melanoma patients. It compares segmentation Dice and lesion correspondence outcomes between baseline and follow-up scans, showing a statistically significant follow-up Dice drop and a long-tailed distribution of registration errors for propagated lesion centroids. To explain this drop, the authors run a controlled experiment in which 30 of the best-segmented lesions are artificially displaced from the VOI center; they observe a sharp threshold around 20 mm beyond which segmentation fails and correspondence breaks down. The paper concludes that inter-scan registration errors violate ULS23's centered-lesion assumption and that this is the cause of the real-data follow-up degradation, motivating a shift toward end-to-end longitudinal models.","tokens_in":7403,"tokens_out":4287,"duration_ms":41925,"significance":"If the causal claim were fully supported, the paper would be a valuable cautionary result for longitudinal lesion tracking: it identifies a concrete mechanism (center-of-VOI sensitivity) in a widely used open model and evaluates it on an open dataset, which makes the work reproducible. The controlled displacement experiment is a clean, self-contained probe of the model's behavior and credibly demonstrates that large displacements destroy segmentation. However, the real-data attribution to registration error is not directly tested, and the selection of the 30 best-segmented lesions limits extrapolation to the full cohort. The contribution is therefore a solid mechanism demonstration with an overreaching causal interpretation in its current form.","major_comments":[{"comment":"The sentence 'This controlled experiment confirms that the performance drop seen in real-world follow-up data (Figure 3) is caused by VOI misalignment' is not supported by the evidence presented. The real-data analysis in §3.1 never correlates per-lesion registration error with follow-up Dice, so it cannot separate VOI misalignment from the morphology and appearance confounds acknowledged in the Discussion. Because the displacement experiment uses only the 30 lesions with the highest Dice (Section 2.3), it excludes the cases where appearance changes are most likely to be severe and hence cannot estimate what fraction of the real-data drop is attributable to displacement. The authors should add a per-lesion or stratified analysis, such as plotting Dice change versus registration error or conditioning on lesions that actually failed, before claiming causation.","section":"§3.2 (last paragraph; Fig. 5)"},{"comment":"The two panels of Figure 5 are in tension with the proposed failure cascade. Figure 5a shows that Dice remains high for displacements below about 20 mm, while Figure 5b shows correct assignments falling below 50% by 10 mm. This means that for moderate displacements, the center-proximity correspondence rule fails even though segmentation quality is largely preserved, so the correspondence breakdown is not, in those cases, a consequence of segmentation failure. The narrative that 'registration error leads to segmentation failure, which in turn causes correspondence failure' is only valid for large displacements; the moderate-displacement regime is dominated by the assignment rule itself. Please reconcile the two panels or revise the mechanism claim accordingly.","section":"§3.2, Fig. 5a vs. Fig. 5b"},{"comment":"The Wilcoxon signed-rank test demonstrates a statistically significant Dice drop, but it does not control for lesion size change, split/merge events, or new lesion appearance, all of which are explicitly present in the dataset (Section 2.2). These factors can directly lower follow-up Dice independently of any registration error. The Discussion acknowledges this possibility, but the abstract and conclusion treat the registration-error connection as established. A stratified analysis with matched lesions of similar size change but different registration errors, or an adjustment for lesion volume change, is needed to support the causal interpretation.","section":"§3.1 / Fig. 3b"},{"comment":"The correspondence rule in Eq. (1), which assigns the lesion ID to the segmented component whose centroid is closest to the VOI center, is the authors' own evaluation choice rather than a component of the ULS23 model. The paper nevertheless uses the resulting failure rates (Fig. 4) as evidence that segmentation unreliability causes correspondence failure. Because an alternative matching rule (e.g., largest component, or a learned matcher) would produce different correspondence outcomes, the reported correspondence failure rates are partly an artifact of this rule. The authors should evaluate at least one alternative rule or explicitly discuss the sensitivity of their conclusions to this choice.","section":"§2.3, Eq. (1)"}],"minor_comments":[{"comment":"The text contains inconsistent spacing in '256× 256× 128 voxel'; please use consistent notation such as '256 × 256 × 128'.","section":"§2.3"},{"comment":"The figure captions use 'Fig.3' and 'Fig.4' without a space; please use 'Fig. 3' and 'Fig. 4' for consistency.","section":"Figures 3 and 4"},{"comment":"The definition of 'False Negative' says 'No component was segmented as s∗' but the earlier symbol definition uses s∗ for the assigned component; clarify whether a false negative means no component at all or no component selected by the assignment rule.","section":"§2.3"},{"comment":"The citation style '[10,5,9]' is unusual and the order is inconsistent with the reference numbering; please use a standard citation format and ensure the cited works include the relevant longitudinal methods.","section":"§5"},{"comment":"The criteria for selecting the '30 best-segmented lesions' are not fully specified; state whether this is based on baseline Dice, follow-up Dice, or an average, and report the range of those Dice values.","section":"§2.3"}],"recommendation":"major_revision","confidential_remarks":"The paper evaluates the authors' own ULS23 model, which creates a mild self-citation concern but is not disqualifying. The controlled displacement experiment is genuinely informative; the main weakness is the gap between the mechanism demonstrated there and the causal claim about the real-data cohort. If the authors add a direct per-lesion analysis linking registration error to Dice change, and temper the conclusion accordingly, the paper could become acceptable. As it stands, the real-data causal claim needs additional support before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a useful, clearly written evaluation of ULS23 that demonstrates a real failure mode — the model assumes its target lesion sits at the center of the input VOI, and its segmentation collapses when that assumption is violated by more than about 20 mm. The controlled displacement experiment is the heart of the paper and it is done well: same lesions, systematic shifts, clean threshold behavior, and a simple explanation grounded in the model's training. That part is worth publishing.\n\nThe new things are the public Longitudinal-CT evaluation and the sharp-threshold quantification, not the existence of the center bias, which was already stated in the ULS23 challenge paper. What the paper adds is showing how consequential that bias is in a realistic longitudinal pipeline, including the cascade into correspondence failures.\n\nThe soft spot is the causal attribution in the real-data experiment. The paper shows follow-up Dice drops and a registration-error distribution with a long tail, but it never directly connects the two per-lesion. The controlled experiment — the only evidence for the mechanism — was run on the 30 best-segmented lesions, which are exactly the cases where appearance change and size change are least likely to confound. The authors acknowledge this confound in the Discussion but do not control for it in the analysis. So the stronger claim that the real-data drop 'is caused by VOI misalignment' overreaches. The mechanism is real for large displacements; its dominance in practice is plausible but unverified. A per-lesion comparison of registration error versus Dice change would settle it.\n\nAlso, the use of true follow-up centroids when propagated coordinates were missing (a best-case scenario) slightly muddies the real-data comparison, though it doesn't change the direction.\n\nThe citation pattern is fine: they evaluate their own model and a dataset one of them co-authored, but the measurements are self-contained and the model/dataset are legitimate. I'd send this to peer review. A good referee can ask for the per-lesion correlation and a more careful caveat in the conclusion. It is not a paradigm-shift paper, but it is a solid empirical finding that the field should know about.","headline":"A solid, useful evaluation of ULS23's center assumption; the controlled displacement experiment carries the paper, but the real-data causal claim needs a direct per-lesion test.","tokens_in":7957,"tokens_out":1742,"would_cite":true,"duration_ms":16084,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The ULS23 segmentation model collapses when a lesion is displaced from the center of its input volume, and follow-up CT registration errors routinely cause that displacement, breaking lesion tracking.","keywords":["Longitudinal Analysis","Lesion Segmentation","Computed Tomography","Deep Learning","Registration Error","ULS23","Lesion Correspondence","Centered-lesion assumption"],"falsifier":"Compare Dice on follow-up lesions whose propagated centroid error is under 5 mm with the same lesions at baseline. If the drop persists when the lesion is effectively centered in the input crop, then appearance change, not misalignment, drives the real-world degradation—a result the controlled displacement experiment cannot provide because it reuses the same image content.","tokens_in":6944,"feed_emoji":"🩻","tokens_out":7256,"duration_ms":66606,"temperature":0.7,"pith_summary":"The paper aims to establish that ULS23, a general-purpose universal lesion segmentation model trained on single time-point CT scans, cannot be naively applied to longitudinal oncology workflows because it silently assumes the target lesion sits at the center of a fixed-size input volume. Using a public longitudinal melanoma CT dataset, the authors show that registration errors displace propagated lesion centroids in follow-up scans, and that this displacement produces a sharp drop in segmentation quality and a breakdown of the lesion correspondence step. A controlled experiment, in which the input volume is artificially shifted by known distances, isolates the mechanism: Dice remains high up to roughly 20 mm of displacement, then collapses to near zero beyond that threshold. The authors conclude that cascaded single-time-point tools compound errors and argue that robust tracking requires end-to-end temporal models.","feed_headline":"A 20 mm shift collapses CT lesion segmentation to near zero","feed_subtitle":"Follow-up scans inherit registration errors that push lesions out of the model's centered-crop comfort zone, breaking tracking.","key_machinery":"The load-bearing object is the ULS23 model, a 3D Residual Encoder UNet that disables default resampling and patching and processes fixed-size volumes of interest ($256\\times256\\times128$ voxels) with the assumption that a single target lesion is centered in each volume. This centered-lesion bias was reinforced during training on volumes cropped around lesion centroids and on non-exhaustively annotated data, which taught the model to treat peripheral lesions as background. The paper converts that implicit assumption into a measurable transfer function through a controlled displacement experiment, and pairs it with a center-proximity correspondence rule—each lesion ID is assigned to the segmented connected component closest to the volume center—which is the downstream mechanism that turns a segmentation miss into a tracking failure.","core_discovery":"On the paper's own terms, the central discovery is that ULS23's segmentation accuracy is governed by a centered-lesion assumption with a sharp tolerance: displacement of the lesion from the input-volume center by more than about 20 mm drives Dice to near zero, and even 10–15 mm displacements push the downstream center-proximity correspondence below 50% correct assignments. In the real longitudinal data, propagated centroids frequently exceed 10 mm of registration error, producing a statistically significant Dice decrease at follow-up relative to baseline ($p<0.001$) and a rise in clinically critical tracking errors—incorrect assignments plus false negatives—from 26.3% at baseline to 31.8% at follow-up. The paper interprets this as a fundamental limitation of applying single-time-point models to temporal analysis, since the failure cascade (registration error → off-center lesion → segmentation failure → correspondence failure) is built into the pipeline structure.","pith_inferences":["A practical extension the paper does not develop: the sharp 20 mm threshold could be turned into an automatic quality flag, so that any propagated centroid whose distance to the segmented component exceeds the threshold is routed to a human reader instead of trusted.","The same centered-crop failure mode likely applies to interactive click-based segmentation tools and to any model that crops a volume around a user-provided coordinate, not only to registration-based longitudinal pipelines.","The displacement experiment suggests a testable training fix: augmenting ULS23 with off-center volumes or conditioning on the centroid-to-lesion distance might flatten the threshold and restore robustness, although the paper argues the cascaded architecture would still remain fragile.","A direct comparison of follow-up Dice for lesions with minimal registration error versus lesions with large error would separate the geometric mechanism from treatment-induced appearance change; the paper's real-data analysis does not fully control for that confound."],"forward_implications":["ULS23 maintains high Dice for lesion-to-center displacements up to roughly 20 mm, then fails almost completely beyond that threshold in the controlled experiment.","In follow-up CT scans, registration-based propagated centroids with errors above 10 mm are common enough to produce a significant overall Dice drop relative to baseline.","Lesion correspondence by center proximity is only as reliable as the segmentation feeding it; when segmentation fails, incorrect assignments and false negatives rise to 31.8% of follow-up lesions.","Cascading independent stages (registration, volume-of-interest segmentation, correspondence) compounds errors, so robust longitudinal tracking should move toward integrated models that process temporal context jointly."],"supporting_citations":[{"why":"Supplies the ULS23 model, its training strategy, and the centered-lesion assumption under test.","marker":"[3]"},{"why":"Supplies the public longitudinal CT dataset with baseline and follow-up scans, manual annotations, and propagated centroids used in the experiments.","marker":"[7]"},{"why":"Supplies the automated configuration and architecture design approach on which ULS23's 3D UNet is based.","marker":"[5]"}],"fun_headline_variants":["Off-center lesions break CT segmentation and tracking","A 20 mm shift kills lesion segmentation accuracy","Longitudinal CT fails when lesions drift from center","Registration errors push lesions out of model's comfort zone","Centered-lesion assumption undermines longitudinal tracking"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise for the real-data portion is that the follow-up performance drop is caused primarily by misalignment of the input crop around the lesion, rather than by treatment-induced changes in lesion size and appearance between scans; the paper acknowledges that it does not control for this confound.","fun_headline_variants_meta":{"raw":{"variants":["Off-center lesions break CT segmentation and tracking","A 20 mm shift kills lesion segmentation accuracy","Longitudinal CT fails when lesions drift from center","Registration errors push lesions out of model's comfort zone","Centered-lesion assumption undermines longitudinal tracking"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000281,"raw_usage":{"total_tokens":1655,"prompt_tokens":926,"completion_tokens":729,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":542,"completion_tokens_details":{"reasoning_tokens":658}},"tokens_in":542,"tokens_out":729,"duration_ms":6819,"temperature":1.0,"reasoning_tokens":658,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T17:56:59.259595+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compare Dice on follow-up lesions whose propagated centroid error is under 5 mm with the same lesions at baseline. If the drop persists when the lesion is effectively centered in the input crop, then appearance change, not misalignment, drives the real-world degradation—a result the controlled displacement experiment cannot provide because it reuses the same image content.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the public longitudinal CT dataset with baseline and follow-up scans, manual annotations, and propagated centroids used in the experiments."}],"review_version":1}