{"id":"f7cc0589-08bc-4c1f-8b10-0193a022702c","arxiv_id":"2509.01433","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A frame-wise masked autoencoder with a temporal contrastive loss achieves 0.88 AUROC for binary EF classification on EchoNet-Dynamic, below the cited 0.93 AUROC of ECHO-VISION-FM.","lead":"This paper augments a masked autoencoder with a temporal contrastive loss for ultrasound videos, targeting ejection fraction estimation from echocardiography. It reports 0.88 AUROC on EchoNet-Dynamic for binary EF classification, but the claimed gain depends on an oracle frame-selection setting and smaller non-oracle temporal gains contradict the headline.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Temporal model's only win (AUROC 0.88) relies on oracle phase-aligned frames; without oracle it is 0.83, below the 0.86 frame-based end-to-end baseline, so the claimed real-time improvement is not established.","rationale":"The reader's weakest-assumption analysis correctly identifies the Oracle Setting as the load-bearing flaw. The paper's only headline-supported number, AUROC 0.88, is achieved exclusively under an oracle that aligns frames to systole/diastole, which contradicts the real-time framing of the abstract and is not applied to the baselines. The absence of the 'Temporal, End-to-End, Contrastive' non-oracle row makes it impossible to attribute the gain to the proposed temporal contrastive loss rather than to privileged frame selection. The comparison with ECHO-VISION-FM in Table 2 is also unfavorable on all metrics, but the internal Table 1 comparison is the primary evidence for the central claim. Given that the reader's REJECT verdict is based on this same confound, no verdict adjustment is needed. The requested test would settle whether the concern lands: if the no-oracle temporal-contrastive model does not beat the frame-based end-to-end baseline, or if frame-based models get a comparable oracle boost, the abstract's real-time claim should not stand.","tokens_in":6687,"tokens_out":2711,"duration_ms":30697,"concrete_test":"Run Table 1's 'End-to-End, Contrastive' temporal model on EchoNet-Dynamic using exactly the same 10 uniformly sampled 32×32 frames as the frame-based end-to-end row, with no systole/diastole oracle selection, same ViT-T backbone and training budget; report AUROC over 3 seeds with standard deviations. Also run the frame-based end-to-end model on the same oracle-selected frames used for the temporal oracle row. If the no-oracle temporal-contrastive AUROC does not exceed the frame-based end-to-end 0.86, or if the oracle-induced gain appears for the frame-based model too, then the claimed temporal advantage is an artifact of oracle frame selection rather than temporal contrastive learning.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that temporally consistent masking plus contrastive learning yields a substantial improvement in EF prediction for real-time ultrasound analysis. Internally, the only temporal model that beats the frame-based end-to-end baseline is 'End-to-End, Contrastive, Oracle' (AUROC 0.88 vs 0.86). This row depends on the Oracle Setting defined in Section 3.1: frames are perfectly aligned with key cardiac phases (systole/diastole) during both pretraining and inference. The same temporal model without oracle ('Temporal, End-to-End') scores 0.83, which is below the frame-based end-to-end baseline of 0.86. The oracle condition is not part of a real-time pipeline, and the frame-based baselines are not given the same oracle-selected frames. Moreover, Table 1 omits the decisive ablation: 'Temporal, End-to-End, Contrastive' without oracle. As reported, the contrastive loss is only evaluated together with oracle alignment, so the 0.06 AUROC jump from 'Temporal, End-to-End, Oracle' (0.82) to 'Temporal, End-to-End, Contrastive, Oracle' (0.88) cannot be separated from the oracle effect. Consequently, the abstract's 'substantial improvement in EF prediction accuracy' is not supported by a fair, real-time-compatible comparison.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a temporal masked-autoencoder framework for ultrasound video, combining frame-wise random masking with a margin-based temporal contrastive loss to learn temporally coherent representations. It evaluates the method on EchoNet-Dynamic for binary ejection-fraction classification (normal vs. reduced EF) at 32×32 resolution with 10 input frames, using ViT-Tiny. The central claim is that temporally-aware self-supervised pretraining yields a substantial improvement in EF prediction accuracy. The strongest reported result is an AUROC of 0.88 for the 'End-to-End, Contrastive, Oracle' configuration in Table 1, compared with 0.86 for the frame-based end-to-end baseline. However, this improvement is only achieved under the Oracle Setting (Section 3.1), which assumes perfect alignment of input frames with key cardiac phases during both pretraining and inference. The same temporal model without oracle achieves 0.83, below the frame-based end-to-end baseline. The paper omits the decisive non-oracle contrastive ablation, reports no error bars or significance tests, and Table 2 mislabels the configuration used in the state-of-the-art comparison. The abstract and introduction overstate the support for the method's claimed advantage.","tokens_in":7043,"tokens_out":6195,"duration_ms":67436,"significance":"If the central claim were established, the paper would offer a useful result: a low-resolution, low-parameter temporal self-supervised method competitive with much larger models on echocardiography EF classification, which could be relevant for real-time deployment. The self-supervised pretraining objective does not use EF labels, so there is no circularity in that respect. However, the reported experiments do not isolate the contribution of the temporal contrastive loss from the oracle frame-selection assumption, and the non-oracle temporal model is worse than the frame-based baseline. The manuscript therefore does not currently provide credible evidence for its stated contribution.","major_comments":[{"comment":"The only configuration in which the temporal model outperforms the frame-based end-to-end baseline is 'End-to-End, Contrastive, Oracle' (AUROC 0.88 vs. 0.86). This result depends on the Oracle Setting, which 'assumes optimal frame selection during pretraining and inference, where frames are perfectly aligned with key cardiac phases.' The same temporal model without oracle ('End-to-End') reaches only 0.83, below the frame-based end-to-end baseline of 0.86. Since oracle-aligned frames are not available in a real-time pipeline and the frame-based baselines are not given the same oracle selection, the claimed 'substantial improvement in EF prediction accuracy' is not supported by a real-time-compatible comparison.","section":"Section 3.1, Table 1"},{"comment":"The decisive ablation is missing: there is no 'Temporal, End-to-End, Contrastive' row without oracle. As reported, the improvement from 'End-to-End, Oracle' (0.82) to 'End-to-End, Contrastive, Oracle' (0.88) cannot be attributed to the contrastive loss, because the oracle condition is varied alongside it. The authors should report the non-oracle contrastive configuration and, for completeness, a frame-based contrastive oracle configuration, to separate the effects of contrastive pretraining and oracle frame selection.","section":"Table 1"},{"comment":"Table 2 labels the compared method as 'Ours (Temporal, End-to-end, Oracle)' with AUROC 0.88, but Table 1's corresponding row 'End-to-End, Oracle' has AUROC 0.82; the value 0.88 belongs to 'End-to-End, Contrastive, Oracle'. This mislabeling makes the state-of-the-art comparison inaccurate. Furthermore, the text describes the result as 'competitive performance' with ECHO-VISION-FM (AUROC 0.93), but a 0.88 vs. 0.93 gap, combined with very different pretraining data, model size, and resolution, is not a direct or strong comparison.","section":"Table 2"},{"comment":"No error bars, confidence intervals, or significance tests are reported for any configuration. On a dataset of ~10,000 videos, AUROC differences of 0.02–0.06 may be within run-to-run variability. The paper should specify the train/validation/test split, the number of random seeds, and the variance across runs before the relative ranking of models can be assessed.","section":"Table 1 and Section 3.2"}],"minor_comments":[{"comment":"The phrase 'substantial improvement in EF prediction accuracy' is not supported by the current evidence; please temper the claim or qualify it as applying only to the oracle setting.","section":"Abstract and Introduction"},{"comment":"The text states that frames are 'uniformly sampled over a one-second interval of the cardiac cycle' and also that the Oracle Setting assumes 'optimal frame selection... perfectly aligned with key cardiac phases.' These statements are in tension; please clarify how uniform sampling and oracle-aligned selection coexist.","section":"Section 3.1"},{"comment":"The notation |M| is used but M is not explicitly defined as the set of masked patches across all frames. Please define it precisely.","section":"Equation (2)"},{"comment":"The temporal-positional embedding notation Et = Epos(t) + Etime(t) is confusing because t appears on both sides; consider using separate indices for frame and spatial position.","section":"Section 2.1"},{"comment":"The paper mentions ViT-Tiny and ViT-Base backbones, but Table 1 only reports ViT-T results. Either provide ViT-Base results or remove the mention.","section":"Section 3.1"},{"comment":"The hyperparameters τp, τm, λ, and the masking ratio are not specified, and no sensitivity analysis is provided. Please report the chosen values and, ideally, ablations over them.","section":"Section 2.3"}],"recommendation":"reject","confidential_remarks":"The paper's central claim is not supported by the reported experiments because the only positive result requires an oracle frame-selection assumption that is unavailable in real-time use, the non-oracle temporal model is worse than the frame-based baseline, and the key ablation is missing. The mislabeled Table 2 further complicates the comparison. If the authors can provide a non-oracle contrastive ablation with error bars and show a robust improvement, a resubmission could be considered, but the current manuscript does not meet the bar for acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"I read the whole thing, including the oracle definition in Section 3.1. The honest take: the method is a reasonable adaptation of VideoMAE-style masking for low-resolution, 10-frame echo clips. The novelty is incremental—frame-wise random masking and a margin-based temporal contrastive loss are both known in spirit, and the combination is a legitimate tweak rather than a new mechanism. The work is clear about its small model and compute budget, and the writing is straightforward.\n\nWhat the paper does well: it tests a genuinely cheap setup (32x32, 10 frames, ViT-Tiny), and the idea of enforcing temporal coherence with a simple contrastive margin is sensible. The pretraining objective does not leak EF labels into the pretext task, so circularity burden is low. Credit for that.\n\nNow the soft spot, and it is load-bearing. Table 1 shows the only temporal model that beats the frame-based end-to-end baseline (0.86 AUROC) is the one with the Oracle setting: 0.88. Drop the oracle and the same temporal model scores 0.83—below the plain frame-based baseline. The oracle means frames are perfectly aligned to systole/diastole phases at both pretraining and inference, which is not a real-time condition and is not given to the baselines. The decisive ablation—contrastive loss without oracle—is missing, so you cannot separate the contribution of the contrastive loss from the oracle's effect. The 0.06 jump between 'Oracle' (0.82) and 'Contrastive, Oracle' (0.88) could be partly or largely the oracle itself. As reported, the abstract's 'substantial improvement' is not supported by a fair real-time-compatible comparison.\n\nThe comparison to ECHO-VISION-FM in Table 2 is honest in that they acknowledge lower performance across every metric, but the framing 'competitive performance' is generous given the gap is 0.05 AUROC with far fewer data and lower resolution—that part is fair to mention as a resource trade-off, not a win. No error bars, code, or data release further weaken confidence in the headline.\n\nBottom line: the work is serious and the direction is fine, but the current evidence does not establish the claimed temporal-model benefit for real-time use. The oracle design is a confound, not a minor caveat. It deserves a serious referee—the question of whether temporal contrastive pretraining helps at extremely low resolution is worth one careful look—but the paper needs revision: report the no-oracle contrastive ablation, give baselines the same oracle frames or remove the oracle setting, and temper the abstract. I would engage with a revised version.","headline":"The key result is a plausible incremental extension (frame-wise masking plus margin-based temporal contrastive loss), but the paper's headline improvement rests on an oracle frame-alignment setup that makes the central comparison unfair; the abstract overclaims.","tokens_in":7469,"tokens_out":670,"would_cite":false,"duration_ms":9258,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A masked autoencoder that masks differently per frame and adds a margin-based temporal contrastive loss reaches 0.88 AUROC for ejection-fraction classification on EchoNet-Dynamic, using an 8M-parameter model at 32×32 resolution.","keywords":["temporal representation learning","masked autoencoder","contrastive learning","ejection fraction","echocardiography","ultrasound video","self-supervised pretraining","EchoNet-Dynamic"],"falsifier":"Retrain the temporal model on EchoNet-Dynamic with the same 10-frame, 32×32 setup but no oracle and no phase information, fine-tune end-to-end, and compare AUROC with the 0.86 frame-based end-to-end baseline; if it does not beat 0.86, the claim that the temporal contrastive loss is the source of the improvement is falsified. Conversely, if a frame-based model given the same oracle-aligned frames matches or beats 0.88, the advantage is not temporal at all.","tokens_in":6587,"feed_emoji":"🫀","tokens_out":12794,"duration_ms":120985,"temperature":0.7,"pith_summary":"The paper claims that the temporal continuity of ultrasound video is a learnable signal that frame-by-frame models waste, and that a self-supervised pretraining recipe can harvest it: mask different spatial patches in each frame, reconstruct the clip, and add a margin-based contrastive loss that pulls temporally close frames together and pushes far-apart frames apart. On the EchoNet-Dynamic ejection-fraction benchmark this reaches 0.88 AUROC — above every frame-based baseline — using an 8M-parameter ViT at 32×32 resolution with 10 frames per clip. That is competitive with echocardiography foundation models trained on roughly 20 times more data at 224×224, which the authors take as evidence that explicit temporal coherence, not scale or resolution, is the lever for real-time ultrasound analysis. The best number requires oracle-aligned frames — a caveat registered in the paper's own tables, where the same temporal model without that alignment scores 0.83.","feed_headline":"Temporal pretraining lifts cardiac ultrasound to 0.88 AUROC","feed_subtitle":"An 8M-parameter model on 10 low-res frames closes in on echo models trained on 20x more data.","key_machinery":"The carrying mechanism is the margin-based temporal contrastive loss paired with frame-wise random masking. A frame's representation f_t is the average of all its patch tokens; the cosine distance d(t, Δt) to every other frame in the clip is computed, and the loss applies two thresholds — frames closer than the positive window τp are pulled together (minimize d²), frames beyond it are pushed apart until they clear the negative margin τm (penalty [τm − d]₊²). Because masking is applied independently per frame, the reconstruction objective cannot be satisfied by copying still images; the encoder must represent motion across frames. The total pretraining objective is ℒ_total = ℒ_rec + λℒ_contra","core_discovery":"The central claim is that a video-adapted masked autoencoder learns cardiac motion well enough that a deliberately small model performs like much larger systems. Pretraining combines a frame-wise masked reconstruction loss — each frame loses a different random set of patches that a decoder must rebuild — with a temporal contrastive loss computed on whole-frame representations, defined as the mean of a frame's patch tokens. The contrastive term measures cosine distance between frame pairs and applies two thresholds: pairs closer than τp are pulled together, pairs farther than τp are pushed to distance at least τm. With oracle-aligned frames, the fully configured model reaches 0.88 AUROC for b","pith_inferences":["If the oracle is the price of the win, the deployable version needs an automatic phase detector: adding a cheap systole/diastole aligner and measuring the AUROC gap to 0.88 is a direct testable extension.","The paper's own tables show the temporal model wins only with the oracle — it beats the frame-based end-to-end model 0.88 to 0.84 with alignment, but falls behind 0.83 to 0.86 without it — so the practical value of the method hinges on whether phase alignment can be obtained automatically at inference time.","A controlled ablation holding dataset, resolution, and backbone fixed, varying only the pretraining objective (plain MAE vs MAE plus contrastive), would isolate whether the margin-based loss rather than the video input format causes the improvement.","The contrastive loss treats time as locally smooth but never learns the heart's period; coupling it with an explicit periodicity or cycle-consistency term could remove the need for oracle alignment entirely."],"forward_implications":["The contrastive loss is the measured source of the gain: within the oracle-aligned temporal family, AUROC rises from 0.82 (end-to-end) to 0.88 (end-to-end plus contrastive pretraining).","Scale is no longer the only path: an 8M-parameter model at 32×32 with 10 frames lands at 0.88 AUROC, within 0.05 of ECHO-VISION-FM's 0.93, which uses ~98M parameters, 224×224 input, and 20× more pretraining videos.","Real-time deployment becomes plausible: the model's operating point fits bedside compute and latency budgets, the paper's stated motivation for the EF case study.","The recipe is portable: frame-wise masking plus a margin-based temporal contrastive objective transfers to other smoothly varying ultrasound modalities, including the fetal and vascular applications the paper names."],"supporting_citations":[{"why":"Supplies the EchoNet-Dynamic dataset and the ejection-fraction benchmark on which all experiments and comparisons are run.","marker":"(Ouyang et al., 2020)"},{"why":"The Masked Autoencoder architecture this method extends; source of the masked-reconstruction objective.","marker":"(He et al., 2022)"},{"why":"VideoMAE, the video-MAE predecessor whose tube-masking strategy the paper replaces with frame-wise random masking.","marker":"(Tong et al., 2022)"},{"why":"ECHO-VISION-FM, the state-of-the-art echocardiography foundation model used as the headline comparison in Table 2.","marker":"(Zhang et al., 2024)"},{"why":"EchoFM, prior work introducing a periodic contrastive loss for cardiac cycles, which positions the paper's temporal contrastive objective.","marker":"(Kim et al., 2024)"},{"why":"Introduces the Vision Transformer that the temporal backbone is built from.","marker":"(Dosovitskiy et al., 2020)"},{"why":"Supplies the ViT-Tiny and ViT-Base variants used as the encoders.","marker":"(Wu et al., 2022)"}],"fun_headline_variants":["8M-param model, 10 frames, 0.88 AUROC: temporal learning works","Temporal consistency beats frame-wise: small echo model hits 0.88 AUROC","0.88 AUROC from tiny echo model via temporal pretraining","Echo EF prediction: temporal learning with 8M params rivals 20x larger"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The paper's best result is reached only under its 'Oracle Setting' — frames assumed to be perfectly aligned with key cardiac phases such as systole and diastole at both pretraining and inference — and without that alignment the temporal model scores 0.83 AUROC, below the 0.86 of the frame-based end-to-end baseline.","fun_headline_variants_meta":{"raw":{"variants":["8M-param model, 10 frames, 0.88 AUROC: temporal learning works","Temporal consistency beats frame-wise: small echo model hits 0.88 AUROC","0.88 AUROC from tiny echo model via temporal pretraining","Echo EF prediction: temporal learning with 8M params rivals 20x larger"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000759,"raw_usage":{"total_tokens":3186,"prompt_tokens":702,"completion_tokens":2484,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":446,"completion_tokens_details":{"reasoning_tokens":2394}},"tokens_in":446,"tokens_out":2484,"duration_ms":18504,"temperature":1.0,"reasoning_tokens":2394,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T12:31:37.539622+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Retrain the temporal model on EchoNet-Dynamic with the same 10-frame, 32×32 setup but no oracle and no phase information, fine-tune end-to-end, and compare AUROC with the 0.86 frame-based end-to-end baseline; if it does not beat 0.86, the claim that the temporal contrastive loss is the source of the improvement is falsified. Conversely, if a frame-based model given the same oracle-aligned frames matches or beats 0.88, the advantage is not temporal at all.","supporting_citations":[{"cited_title":"Video MAE : Masked autoencoders are data-efficient learners for self-supervised video pre-training","cited_arxiv_id":null,"evidence_quote":"VideoMAE, the video-MAE predecessor whose tube-masking strategy the paper replaces with frame-wise random masking."},{"cited_title":"Echo-vision-fm: A pre-training and fine-tuning framework for echocardiogram video vision foundation model","cited_arxiv_id":null,"evidence_quote":"ECHO-VISION-FM, the state-of-the-art echocardiography foundation model used as the headline comparison in Table 2."}],"review_version":1}