Pith. sign in

REVIEW 1 cited by

AquaticCLIP: A Vision-Language Foundation Model for Underwater Scene Analysis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.01785 v1 pith:O7CBR7O6 submitted 2025-02-03 cs.CV cs.AI

classification cs.CVcs.AI
keywords aquaticclipaquaticmodelunderwaterdatasetscenevision-languagecontrastive
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The preservation of aquatic biodiversity is critical in mitigating the effects of climate change. Aquatic scene understanding plays a pivotal role in aiding marine scientists in their decision-making processes. In this paper, we introduce AquaticCLIP, a novel contrastive language-image pre-training model tailored for aquatic scene understanding. AquaticCLIP presents a new unsupervised learning framework that aligns images and texts in aquatic environments, enabling tasks such as segmentation, classification, detection, and object counting. By leveraging our large-scale underwater image-text paired dataset without the need for ground-truth annotations, our model enriches existing vision-language models in the aquatic domain. For this purpose, we construct a 2 million underwater image-text paired dataset using heterogeneous resources, including YouTube, Netflix, NatGeo, etc. To fine-tune AquaticCLIP, we propose a prompt-guided vision encoder that progressively aggregates patch features via learnable prompts, while a vision-guided mechanism enhances the language encoder by incorporating visual context. The model is optimized through a contrastive pretraining loss to align visual and textual modalities. AquaticCLIP achieves notable performance improvements in zero-shot settings across multiple underwater computer vision tasks, outperforming existing methods in both robustness and interpretability. Our model sets a new benchmark for vision-language applications in underwater environments. The code and dataset for AquaticCLIP are publicly available on GitHub at xxx.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Energy Constrained Hierarchical Underwater Monitoring via Local Multi-Agent RAG

    cs.IR 2026-07 conditional novelty 6.0 of 10

    A master–satellite edge station pairs MAX78000/02 always-on visual/acoustic sentinels with selective Jetson multimodal RAG, local species ID, and multi-agent reporting to cut energy and uplink cost.

Pith tools