Skip to main navigation Skip to search Skip to main content

AquaticCLIP: A Vision-Language Foundation Model and Dataset for Underwater Scene Analysis

    • The University of Western Australia
    • Information Technology University

    Research output: Contribution to journalArticlepeer-review

    Abstract

    The preservation of aquatic biodiversity is critical in mitigating the effects of climate change. Aquatic scene understanding plays a pivotal role in aiding marine scientists in their decision-making processes. In this article, we introduce AquaticCLIP, a novel contrastive language-image pretraining (CLIP) model tailored for aquatic scene understanding. AquaticCLIP presents an underwater domain-specific learning framework that aligns images and texts in aquatic environments, enabling tasks such as segmentation, classification, detection, and object counting. By leveraging our large-scale underwater image-text paired dataset without the need for ground-truth (GT) annotations, our model enriches existing vision-language models (VLMs) in the aquatic domain. For this purpose, we construct a 2-million underwater image-text paired dataset using heterogeneous resources, including YouTube, Netflix, National Geographic (NatGeo), etc. To fine-tune AquaticCLIP, we propose a prompt-guided vision encoder (PGVE) that progressively aggregates patch features via learnable prompts, while a vision-guided mechanism enhances the language encoder by incorporating visual context. The model is optimized through a contrastive pretraining loss to align visual and textual modalities. AquaticCLIP achieves notable performance improvements in zero-shot settings across multiple underwater computer vision tasks, outperforming existing methods in both accuracy and robustness. Our model sets a new benchmark for vision-language applications in underwater environments.

    Original languageBritish English
    JournalIEEE Transactions on Neural Networks and Learning Systems
    DOIs
    StateAccepted/In press - 2026

    UN SDGs

    This output contributes to the following UN Sustainable Development Goals (SDGs)

    1. SDG 13 - Climate Action
      SDG 13 Climate Action

    Keywords

    • Object counting
    • object segmentation
    • underwater object detection
    • underwater scene analysis
    • vision language model

    Fingerprint

    Dive into the research topics of 'AquaticCLIP: A Vision-Language Foundation Model and Dataset for Underwater Scene Analysis'. Together they form a unique fingerprint.

    Cite this