{"id":"773a8b1b-f690-47c6-bec0-d9373ca39f7f","arxiv_id":"2508.12343","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":4.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"A feature-enhancement module trained jointly with YOLOv8m improves underwater object detection accuracy while keeping 46.5 FPS.","lead":"The authors present AquaFeat, a plug-in module that refines feature maps inside an object detector and is trained together with the detector. It reports higher precision and recall on underwater images at real-time speed, a promising step for marine monitoring and inspection.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The abstract reports accuracy gains but provides no ablation isolating AquaFeat from training recipe, backbone, or dataset; the causal claim is therefore unsupported.","rationale":"The reader's weakest assumption identified exactly the gap: the abstract attributes gains to the feature enhancement network without isolating it from joint-training recipe, datasets, or backbone. My stress-test confirms this is the most load-bearing point. The correctness of the central claim hinges on a controlled ablation; absent that, the paper is unverdictable. I agree with the reader's UNVERDICTED verdict and do not see grounds to move it. The concrete test I propose would settle whether the concern lands: an identical-everything ablation (or its absence). If the ablation exists and shows a clear improvement, the central claim would be supported; if not, the claim should be weakened to an engineering report.","tokens_in":628,"tokens_out":2957,"duration_ms":31442,"concrete_test":"If the full paper is available, locate or run the ablation that removes AquaFeat (identity pass-through) while keeping every other component identical: same YOLOv8m backbone, same data splits, same augmentations, same epochs, same optimizer hyperparameters, same loss weighting. Compare mAP@0.5 and mAP@[.5:.95]. If the delta is within typical run-to-run variance (~0.5–1.0 mAP point), or if no such ablation is reported, the central attribution claim should be treated as unverified.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that the multi-scale feature enhancement network, trained end-to-end with the detector loss, causes the reported state-of-the-art Precision and Recall. The abstract gives no experimental protocol: no dataset splits, no baseline training details, no comparison to YOLOv8m trained identically without AquaFeat. The reported numbers (P=0.877, R=0.624, mAP@0.5=0.677, mAP@[.5:.95]=0.421 at 46.5 FPS) are self-reported and could arise from longer training, different augmentation, a different optimizer schedule, or dataset selection. Without an ablation or controlled re-implementation, the improvement cannot be attributed to the feature enhancement mechanism. This is not an internal inconsistency, but an evidential gap in the argument as presented.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript, as submitted, consists of an abstract plus no visible full text. The abstract proposes AquaFeat, a plug-and-play multi-scale feature enhancement network trained end-to-end with the detector loss, intended to improve underwater object detection. Integrated with YOLOv8m, it reportedly achieves precision 0.877, recall 0.624, mAP@0.5 0.677, mAP@[0.5:0.95] 0.421, and 46.5 FPS on challenging underwater datasets, with state-of-the-art precision and recall. Because the full text is absent, the evaluation below is based entirely on the abstract and on the surrounding submission metadata.","tokens_in":828,"tokens_out":1897,"duration_ms":23560,"significance":"If substantiated, the contribution is potentially useful: a lightweight, plug-and-play enhancement module that is trained jointly with a detector could be a practical addition to underwater perception systems. The reported speed/accuracy trade-off is attractive. However, the significance cannot currently be assessed because no experimental protocol, ablations, comparisons, or reproducibility materials are provided. The central causal claim—that the feature enhancement module, rather than the training recipe, backbone, or dataset selection, produces the accuracy gains—is plausible but entirely unsupported by the abstract alone. The paper would merit serious consideration if the full experimental evidence backs the stated numbers.","major_comments":[{"comment":"The abstract attributes the reported accuracy gains to the multi-scale feature enhancement network trained with the detector loss, but provides no ablation isolating AquaFeat from the YOLOv8m backbone, the joint-training recipe, augmentation, optimizer schedule, or dataset selection. A controlled comparison of YOLOv8m trained identically with and without AquaFeat is load-bearing for the paper's central claim; without it, the measured improvements could come from other factors. This is an evidential gap, not an internal inconsistency, but it is the key missing support.","section":"Abstract"},{"comment":"The abstract reports precise numbers (P=0.877, R=0.624, mAP@0.5=0.677, mAP@[0.5:0.95]=0.421, 46.5 FPS) without any statement of which underwater datasets were used, how the train/val/test splits were created, the hardware, the number of runs, or error bars. Without this information the numbers are unverifiable, and the claim of 'state-of-the-art' precision/recall cannot be checked against existing benchmarks or reimplementations.","section":"Abstract"},{"comment":"The abstract does not name any baselines: no comparison to traditional image enhancement methods, learning-based enhancement models, or detection-only baselines is reported. The phrase 'state-of-the-art' therefore has no operational meaning in this manuscript. The authors should specify which published results they improve upon and provide a table with matched evaluation settings.","section":"Abstract"}],"minor_comments":[{"comment":"The term 'novel' should be supported by a brief positioning against existing joint enhancement-detection methods; otherwise it is a claim without evidence. Also, 'plug-and-play' is ambiguous: does the module add parameters, require retraining, or operate at inference only?","section":"Abstract"},{"comment":"The submitted manuscript contains no full text, references, figures, or tables. At minimum, a complete version with the experimental section, network architecture, and training details is required for review.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The abstract describes a plausible and potentially relevant contribution, but the submission currently contains no experimental details, ablations, or baselines. The central claim is unsupported by the available evidence, though it is not internally inconsistent. The manuscript can be made publishable if the authors provide a full experimental evaluation with controlled baselines and ablations. I am not recommending rejection because the missing information is additive and within the scope of a revision, but the current state is far from acceptable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is an abstract-only submission, and from what's in front of me, the central claim — that the multi-scale feature enhancement module, not the training recipe or backbone, causes the reported accuracy gains — is not supported by the evidence shown. The abstract gives numbers but no experimental protocol.\n\nWhat's genuinely plausible here: the idea of training a feature enhancement network end-to-end with the detector loss is a legitimate approach, and there's a real practical appeal in a plug-and-play module that slots into YOLOv8m and keeps 46.5 FPS. If the full paper shows that this module generalizes across underwater datasets, it would be a useful, incremental contribution to marine monitoring applications.\n\nThe soft spot is the gap between the headline and the evidence. No ablations, no dataset splits, no comparison to an identically trained YOLOv8m baseline, no error bars. The stress-test note is right: the reported precision and recall could come from longer training, different augmentation, or a different optimizer schedule. That's not a fatal flaw in the idea — it's a missing section in the paper. The precision/recall tradeoff (0.877 / 0.624) also looks like a high-precision operating point, and without a precision-recall curve or mAP breakdown, it's hard to tell if the gain is real.\n\nThe authors cite state-of-the-art performance, but that claim is unverifiable from the abstract alone. I'm not saying the work is wrong; I'm saying the evidence isn't there yet.\n\nWho is this for? Researchers in underwater vision who care about practical detection pipelines. If you're working on joint enhancement and detection, you'll want to read the full version when it's available.\n\nMy recommendation: don't desk-reject it. The idea is coherent and the application is real. Send it to peer review with the expectation that reviewers require a proper ablation isolating AquaFeat from the training recipe and a comparison to the base YOLOv8m. If those are in the full text, this becomes a solid incremental paper. If not, the state-of-the-art claim stays unsupported.","headline":"Abstract-only claim of SOTA underwater detection via a plug-and-play feature enhancement module; plausible idea but the causal claim is unsupported without ablations.","tokens_in":1260,"tokens_out":2279,"would_cite":false,"duration_ms":25358,"reading_group":"maybe","serious_thinker":"unclear","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper proposes AquaFeat, a plug-and-play feature-enhancement module that, trained end-to-end with the detector's loss, improves underwater object detection with YOLOv8m, reporting precision of 0.877, recall of 0.624, mAP@0.5 of 0.677,","keywords":["underwater object detection","feature enhancement","multi-scale network","YOLOv8","image enhancement","real-time detection","marine monitoring"],"falsifier":"Run YOLOv8m on the same underwater datasets with identical training settings but with the AquaFeat module ablated (or replaced by a fixed identity mapping). If precision and recall stay at 0.877 and 0.624, the module is not the cause.","tokens_in":611,"feed_emoji":"🌊","tokens_out":4104,"duration_ms":39614,"temperature":0.7,"pith_summary":"This paper proposes AquaFeat, a plug-and-play module that enhances image features inside an object detector rather than enhancing the images themselves. The module is a multi-scale feature enhancement network trained end-to-end with the detector's loss, so it learns to refine features that matter for detection. Integrated with YOLOv8m on underwater datasets, it reports state-of-the-art precision (0.877) and recall (0.624), with competitive mAP@0.5 (0.677) and mAP@[0.5:0.95] (0.421), at 46.5 FPS. The authors argue this offers an efficient alternative to conventional underwater image enhancement for real-time applications.","feed_headline":"Feature-enhancement module lifts underwater detection to 0.877","feed_subtitle":"Plug-and-play AquaFeat keeps real-time speed at 46.5 FPS while boosting recall to 0.624.","key_machinery":"AquaFeat: a multi-scale feature enhancement network inserted into YOLOv8m and trained end-to-end with the detector's loss. It acts on the feature maps, refining them to be more informative for detection, rather than on the input pixels.","core_discovery":"The core discovery is that task-driven feature enhancement—a network trained jointly with the detection loss to emphasize multi-scale features—can outperform traditional pre-processing image enhancement when paired with YOLOv8m. The reported numbers on challenging underwater datasets are precision 0.877, recall 0.624, mAP@0.5 0.677, mAP@[0.5:0.95] 0.421, at 46.5 FPS.","pith_inferences":["The reported gains are not yet isolated from the joint-training recipe; an ablation that removes only the enhancement module would tell how much of the improvement it causes.","Because the paper does not compare against state-of-the-art underwater enhancement detectors on identical backbones, the 'state-of-the-art' claim depends on the specific comparison set.","The module's task-driven nature might extend beyond underwater imagery to other degraded domains such as fog or low light, but that is untested.","If integrated with newer backbones, the detector might see further gains, but the module's relative contribution may shrink as backbones get stronger."],"forward_implications":["If correct, marine monitoring and infrastructure inspection can use real-time underwater detectors without separate image-enhancement preprocessing.","The plug-and-play design suggests the module could be attached to other detection architectures, potentially improving them as well.","Because it is trained with the detector loss, the enhancement is tailored to the task, possibly avoiding artifacts that generic enhancement introduces.","The reported 46.5 FPS indicates the module adds little computational cost, supporting deployment in real-time systems."],"supporting_citations":[],"fun_headline_variants":["Plug-and-play AquaFeat hits 0.877 underwater precision at 46.5 FPS","AquaFeat: task-driven feature enhancement sharpens underwater detection to 0.877","Underwater detection gets 0.877 precision with plug-and-play AquaFeat","Real-time AquaFeat lifts underwater detection precision to 0.877","AquaFeat plug-and-play lifts underwater detection precision to 0.877"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The experiments attribute the accuracy gains to the AquaFeat module, but without an ablation separating the module from the joint-training setup, the reported improvement could come from the training recipe or added capacity.","fun_headline_variants_meta":{"raw":{"variants":["Plug-and-play AquaFeat hits 0.877 underwater precision at 46.5 FPS","AquaFeat: task-driven feature enhancement sharpens underwater detection to 0.877","Underwater detection gets 0.877 precision with plug-and-play AquaFeat","Real-time AquaFeat lifts underwater detection precision to 0.877","AquaFeat plug-and-play lifts underwater detection precision to 0.877"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.002308,"raw_usage":{"total_tokens":8702,"prompt_tokens":666,"completion_tokens":8036,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":410,"completion_tokens_details":{"reasoning_tokens":7923}},"tokens_in":410,"tokens_out":8036,"duration_ms":60076,"temperature":1.0,"reasoning_tokens":7923,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T19:29:57.450753+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run YOLOv8m on the same underwater datasets with identical training settings but with the AquaFeat module ablated (or replaced by a fixed identity mapping). If precision and recall stay at 0.877 and 0.624, the module is not the cause.","supporting_citations":[],"review_version":1}