{"id":"d26f1d7f-8c06-46e2-95e3-b6960d1c06e5","arxiv_id":"2508.10432","paper_version":1,"verdict":"UNVERDICTED","confidence":"UNKNOWN","novelty_score":4.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"The abstract and the body are two different papers, so the CRISP claims have no visible method, experiments, or derivation.","lead":"This submission's abstract describes CRISP, a continual video instance segmentation method with results on YouTube-VIS-2019/2021, but the full text provided is an unrelated paper on multimodal math reasoning, WE-MATH 2.0 (arXiv 2508.10433). The stated contribution cannot be assessed from the provided document.","discovery_kind":"unclear","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The supplied full text is arXiv:2508.10433 (WE-MATH 2.0), not 2508.10432; the CRISP abstract's claims (ARSP, three losses, YouTube-VIS results) have no supporting method or experiments in the document, so the central claim is unassessable.","rationale":"The reader's weakest_assumption is exactly the document-level mismatch between the CRISP abstract and the provided full text. My independent reading of the full text confirms it: the body is the WE-MATH 2.0 paper and contains none of CRISP's technical content. This is the dominant finding and makes the stated central claim unassessable. I found no additional technical flaw to evaluate because none of CRISP's components appear in the document. The proposed concrete test—retrieving the actual arXiv:2508.10432 and searching for CRISP-specific terms—would decisively distinguish a retrieval error from an abstract misrepresenting the paper. If the correct CRISP text is retrieved, the review should be restarted on that content; if the provided text is the actual submission, the verdict remains unverdictable. No scientific score can be attached to the CRISP claims based on the current document.","tokens_in":26180,"tokens_out":2269,"duration_ms":20468,"concrete_test":"Fetch the official arXiv PDF/HTML for identifier 2508.10432 and search for the strings 'CRISP', 'ARSP', 'YouTube-VIS', 'instance correlation loss', and 'semantic consistency loss'. If the retrieved content is identical to the provided WE-MATH 2.0 text (arXiv:2508.10433), the submission/retrieval mismatch is confirmed and the CRISP claim remains unverifiable. If the retrieved content instead matches the CRISP abstract and contains the claimed method and experiments, then the provided full text was simply the wrong attachment and the review should be re-run on the correct PDF.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that CRISP significantly outperforms existing continual segmentation methods on YouTube-VIS-2019 and YouTube-VIS-2021—requires at minimum a method section defining ARSP and the three claimed losses (instance correlation, semantic consistency, prompt initialization) and an experiments section reporting comparisons on those datasets. The provided full text is the WE-MATH 2.0 paper (arXiv:2508.10433) on multimodal mathematical reasoning: it introduces a MathBook knowledge hierarchy, GeoGebra-rendered datasets, and a two-stage RL pipeline. It never mentions CRISP, ARSP, video instance segmentation, YouTube-VIS, or any component of the abstract. Under the in-scope evidence rule, the submitted text is the only evidence available, and it provides zero support for the abstract's claims. This is not a critique of the method's internal correctness—no method is present—but a document-level failure: the body-abstract mismatch is the load-bearing fact. If the mismatch is a retrieval or submission error, the CRISP content is simply missing from this review; if not, the abstract misrepresents the paper. Either way, no technical evaluation of CRISP is possible from this document.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The submission is headed arXiv:2508.10432 and titled 'CRISP: Contrastive Residual Injection and Semantic Prompting for Continual Video Instance Segmentation'. The abstract promises a method (CRISP) with three components—an instance correlation loss, an adaptive residual semantic prompt (ARSP) learning framework with query-prompt matching, a contrastive semantic consistency loss, and an incremental prompt initialization strategy—and claims state-of-the-art performance on YouTube-VIS-2019 and YouTube-VIS-2021. However, the full text supplied for review is a different paper: 'WE-MATH 2.0: A Versatile MathBook System for Incentivizing Visual Mathematical Reasoning' (arXiv:2508.10433v1 [cs.AI]). That full text is about multimodal mathematical reasoning and never mentions CRISP, continual video instance segmentation, ARSP, the three claimed losses, or YouTube-VIS. The submitted document therefore contains no architecture, no equations, no experiments, and no ablations supporting the abstract's claims.","tokens_in":26356,"tokens_out":2894,"duration_ms":28788,"significance":"If CRISP were fully specified and its experimental claims supported, it would be a credible contribution to continual video instance segmentation, a task of active interest in the computer vision community. The claimed combination of instance-, category-, and task-wise mechanisms is potentially interesting. However, the manuscript under review contains none of the required evidence: there is no method section for CRISP, no definition of the losses, no definition of the ARSP pool, and no experiments on the claimed benchmarks. The abstract's GitHub link cannot compensate for the absence of the method and results in the body. Consequently, the significance of the work cannot be assessed from this document; there is also no machine-checked proof or reproducible code for CRISP to credit, since the body does not mention CRISP at all.","major_comments":[{"comment":"The full text is not the CRISP paper. Its title is 'WE-MATH 2.0: A Versatile MathBook System for Incentivizing Visual Mathematical Reasoning' and its header reads arXiv:2508.10433v1 [cs.AI], not arXiv:2508.10432. The body never mentions CRISP, ARSP, instance correlation loss, semantic consistency loss, continual video instance segmentation, or YouTube-VIS. Under the in-scope evidence rule, this full text is the only evidence in the manuscript. The abstract's central claim — that CRISP significantly outperforms existing continual segmentation methods on YouTube-VIS-2019 and YouTube-VIS-2021 — is therefore entirely unsupported by the submitted document.","section":"Abstract vs. Full Text (document-level mismatch)"},{"comment":"The abstract lists three load-bearing design components: instance correlation loss, adaptive residual semantic prompt (ARSP) learning with an adjustive query-prompt matching mechanism, and a semantic consistency loss based on contrastive learning, plus an initialization strategy for incremental prompts. None of these components is defined, formulated, or even named anywhere in the full text. There is no equation for any loss, no definition of the prompt pool or its learning rule, and no description of the matching mechanism. The method is therefore unassessable; there is no way to check for internal consistency, novelty, or correctness.","section":"§1 / Method (CRISP components)"},{"comment":"The claimed experiments on YouTube-VIS-2019 and YouTube-VIS-2021 do not appear in the submitted document. The full text reports experiments on MathVista, MathVision, We-Math, MathVerse, and MathBookEval for the WE-MATH 2.0 system. There are no continual video instance segmentation baselines, no evaluation metrics for VIS, no comparison tables involving CRISP, and no ablation of the three claimed losses or of the prompt initialization strategy. Thus the abstract's empirical claim of 'significantly outperforming existing continual segmentation methods' has no supporting data in the manuscript.","section":"§5 / Experiments"}],"minor_comments":[{"comment":"The title, author list, and arXiv identifier on the header do not match the abstract. The reference list likewise corresponds to the WE-MATH 2.0 paper, not to CRISP. Even the abstract's GitHub URL for CRISP cannot be cross-checked against the body because the body contains no code availability statement for CRISP.","section":"Title page and metadata"}],"recommendation":"reject","confidential_remarks":"This appears to be a submission or metadata error: the supplied full text is a completely different paper (WE-MATH 2.0) by different authors with no connection to the CRISP abstract. I did not evaluate the scientific content of the WE-MATH 2.0 paper, since it is not the manuscript under consideration. If this was an upload error, the correct CRISP manuscript should be submitted rather than reviewed in this form; as it stands, the submitted document cannot support any of the claims in the abstract."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, quick take on 2508.10432: the abstract and the full text are not the same paper. The abstract describes CRISP, a continual video instance segmentation method with three losses and YouTube-VIS results. The provided full text is WE-MATH 2.0 (arXiv 2508.10433), a multimodal math reasoning paper. Nothing in the body mentions CRISP, ARSP, video instance segmentation, or YouTube-VIS. Under any reasonable reading, this submission cannot be evaluated.\n\nWhat does the paper do well? Honestly, not much that I can assess. The abstract alone sketches a plausible combination of known continual-learning ideas—instance correlation loss, a learnable semantic prompt pool initialized incrementally, a contrastive semantic consistency loss. That is a reasonable recipe, but it is also a recipe of existing ingredients. The WE-MATH text that follows appears to be a serious dataset/benchmark paper, but it is a different paper; I am not crediting it to CRISP. The GitHub link in the abstract is not evidence of content.\n\nSoft spots: the load-bearing one is the body-abstract mismatch. There is no method section, no equations, no ablation, no YouTube-VIS experiments for CRISP. The abstract's claim of significant outperformance is therefore unsupported inside the document. Even taking the abstract at face value, the novelty is incremental—contrastive losses and prompt pools are standard in continual learning. But the bigger problem is you cannot check any of that because the content is missing. I see no internal contradiction in the abstract itself, but the document as a whole is incoherent: it is two different preprints stapled together.\n\nWho is this for? A reader wanting the CRISP method should go to the arXiv listing directly; this submission gives them nothing. A reader interested in WE-MATH 2.0 should read that paper, but it is not the submission. My recommendation: desk reject or return to the authors for a correct full text. Do not send to peer review. There is nothing for a referee to check.","headline":"The abstract and full text are two different papers; CRISP cannot be evaluated from this submission.","tokens_in":26953,"tokens_out":2637,"would_cite":false,"duration_ms":25780,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"CRISP claims a video segmentation method that learns new classes without forgetting old ones, yet the supplied body text is a different paper.","keywords":["continual video instance segmentation","catastrophic forgetting","contrastive learning","semantic prompting","adaptive residual semantic prompt","incremental learning","YouTube-VIS","instance correlation loss"],"falsifier":"Look inside the manuscript for CRISP's method section: if neither ARSP, the instance correlation loss, the semantic consistency loss, nor the YouTube-VIS-2019/2021 experimental tables appear anywhere, the central claim has no support in the submitted document. Alternatively, run CRISP's public code on YouTube-VIS with disjoint category tasks and compare old-class AP after each task to a standard fine-tuned baseline; a non-improvement there would refute the forgetting-avoidance claim.","tokens_in":25945,"feed_emoji":"🎬","tokens_out":4123,"duration_ms":43855,"temperature":0.7,"pith_summary":"The paper proposes CRISP (Contrastive Residual Injection and Semantic Prompting), a method for continual video instance segmentation that aims to learn new object categories without forgetting previously learned ones. The abstract claims CRISP significantly outperforms existing continual segmentation methods on YouTube-VIS-2019 and YouTube-VIS-2021, resolving instance-wise, category-wise, and task-wise confusion. The claimed mechanism pairs an instance correlation loss with an adaptive residual semantic prompt (ARSP) pool generated from category text, plus a contrastive semantic consistency loss and an incremental prompt initialization strategy. The full text supplied for this review is actually a mathematics-reasoning paper (WE-MATH 2.0), so none of CRISP's equations, ablations, or experimental tables appears in the body. Thus the abstract's claims are on record, but the document as provided cannot support or refute them.","feed_headline":"CRISP learns new video classes without forgetting old ones","feed_subtitle":"It links object queries to category-text prompts and claims top results on YouTube-VIS benchmarks.","key_machinery":"The load-bearing components are the three training-time mechanisms: the instance correlation loss (which treats the prior query space as an anchor and enforces current-task specificity), the adaptive residual semantic prompt (ARSP) pool (category text projected into learnable residual prompts and matched to object queries), and the contrastive semantic consistency loss (which pulls object queries and residual prompts into semantic coherence). An incremental prompt initialization strategy is the fourth component, meant to preserve inter-task query-space correlations. Together these are claimed to prevent forgetting at the instance, category, and task levels.","core_discovery":"The central claim, as stated in the abstract, is that CRISP outperforms existing continual segmentation methods on long-term continual video instance segmentation on YouTube-VIS-2019 and YouTube-VIS-2021 while avoiding catastrophic forgetting. Its designed contributions are: (1) an instance correlation loss that models tracking by aligning current task queries with the prior query space while sharpening current-task specificity; (2) ARSP, a learnable residual prompt pool generated from category text, with an adjustive query-prompt matching mechanism; (3) a contrastive semantic consistency loss linking object queries and residual prompts during incremental training; and (4) a prompt initializ","pith_inferences":["My reading: the document-level mismatch means the substantive contribution of this submission cannot be evaluated from the provided text; if the WE-MATH 2.0 body was attached by retrieval error, the CRISP evaluation requires the original manuscript.","If the abstract's design is taken at face value, the ARSP pool offers a transferable pattern: category text as a prior for residual prompts could apply to other incremental recognition settings, such as open-vocabulary detection or long-tail classification, not just video instance segmentation.","A testable extension would be to ablate the three losses separately against single-task continual baselines; the abstract does not report which loss carries the main forgetting reduction.","The claimed 'concise yet powerful' prompt initialization strategy would be worth testing as a standalone replay-free baseline for task-wise forgetting."],"forward_implications":["If CRISP works as claimed, continual video instance segmentation models can add new categories across tasks without a steep drop in old-category performance.","The ARSP query-prompt matching implies new categories can be incorporated through prompt-pool assignment rather than full model retraining.","The contrastive semantic consistency loss should keep object queries semantically stable as tasks accumulate, directly attacking catastrophic forgetting.","On the reported benchmarks, CRISP would set the current state of the art among continual segmentation methods for long-term settings.","The dual losses and prompt initialization target the three distinct confusion types, so each component is a separable intervention for diagnosis."],"supporting_citations":[],"fun_headline_variants":["CRISP outscores prior continual video segmentation methods on YouTube-VIS","CRISP: prompt-based continual video segmentation tops YouTube-VIS benchmarks","CRISP avoids catastrophic forgetting, sets new state-of-the-art on YouTube-VIS","CRISP links queries to text prompts, beating prior continual video segmentation","CRISP beats prior methods on continual video segmentation tasks"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The load-bearing premise is document-level: the body text must actually describe CRISP and its experiments for the claimed state-of-the-art result to be checkable; here the body is an unrelated mathematics-reasoning paper, so the abstract's claim stands unsupported by the supplied text.","fun_headline_variants_meta":{"raw":{"variants":["CRISP outscores prior continual video segmentation methods on YouTube-VIS","CRISP: prompt-based continual video segmentation tops YouTube-VIS benchmarks","CRISP avoids catastrophic forgetting, sets new state-of-the-art on YouTube-VIS","CRISP links queries to text prompts, beating prior continual video segmentation","CRISP beats prior methods on continual video segmentation tasks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000821,"raw_usage":{"total_tokens":3449,"prompt_tokens":780,"completion_tokens":2669,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":524,"completion_tokens_details":{"reasoning_tokens":2578}},"tokens_in":524,"tokens_out":2669,"duration_ms":20285,"temperature":1.0,"reasoning_tokens":2578,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T20:26:52.515641+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Look inside the manuscript for CRISP's method section: if neither ARSP, the instance correlation loss, the semantic consistency loss, nor the YouTube-VIS-2019/2021 experimental tables appear anywhere, the central claim has no support in the submitted document. Alternatively, run CRISP's public code on YouTube-VIS with disjoint category tasks and compare old-class AP after each task to a standard fine-tuned baseline; a non-improvement there would refute the forgetting-avoidance claim.","supporting_citations":[],"review_version":1}