Table of Contents
Fetching ...

MARIS: Marine Open-Vocabulary Instance Segmentation with Geometric Enhancement and Semantic Alignment

Bingyu Li, Feiyu Wang, Da Zhang, Zhiyuan Zhao, Junyu Gao, Xuelong Li

TL;DR

MARIS tackles the challenge of open-vocabulary underwater instance segmentation by introducing a dedicated fine-grained dataset and a two-branch framework. The Geometric Prior Enhancement Module (GPEM) leverages depth-derived geometric priors to stabilize features under underwater degradation, while the Semantic Alignment Injection Mechanism (SAIM) enriches language priors with underwater-aware prompts and template-adaptive selection. The combination, along with a CLIP-based visual-geometry fusion and Q-Former-based semantic bridging, yields state-of-the-art results in both in-domain and cross-domain settings, demonstrated on MARIS with 158 fine-grained categories across 9 super-classes. This work provides a robust benchmark and a principled approach for advancing open-vocabulary perception in challenging marine environments, with implications for biodiversity monitoring and autonomous underwater exploration.

Abstract

Most existing underwater instance segmentation approaches are constrained by close-vocabulary prediction, limiting their ability to recognize novel marine categories. To support evaluation, we introduce \textbf{MARIS} (\underline{Mar}ine Open-Vocabulary \underline{I}nstance \underline{S}egmentation), the first large-scale fine-grained benchmark for underwater Open-Vocabulary (OV) segmentation, featuring a limited set of seen categories and diverse unseen categories. Although OV segmentation has shown promise on natural images, our analysis reveals that transfer to underwater scenes suffers from severe visual degradation (e.g., color attenuation) and semantic misalignment caused by lack underwater class definitions. To address these issues, we propose a unified framework with two complementary components. The Geometric Prior Enhancement Module (\textbf{GPEM}) leverages stable part-level and structural cues to maintain object consistency under degraded visual conditions. The Semantic Alignment Injection Mechanism (\textbf{SAIM}) enriches language embeddings with domain-specific priors, mitigating semantic ambiguity and improving recognition of unseen categories. Experiments show that our framework consistently outperforms existing OV baselines both In-Domain and Cross-Domain setting on MARIS, establishing a strong foundation for future underwater perception research.

MARIS: Marine Open-Vocabulary Instance Segmentation with Geometric Enhancement and Semantic Alignment

TL;DR

MARIS tackles the challenge of open-vocabulary underwater instance segmentation by introducing a dedicated fine-grained dataset and a two-branch framework. The Geometric Prior Enhancement Module (GPEM) leverages depth-derived geometric priors to stabilize features under underwater degradation, while the Semantic Alignment Injection Mechanism (SAIM) enriches language priors with underwater-aware prompts and template-adaptive selection. The combination, along with a CLIP-based visual-geometry fusion and Q-Former-based semantic bridging, yields state-of-the-art results in both in-domain and cross-domain settings, demonstrated on MARIS with 158 fine-grained categories across 9 super-classes. This work provides a robust benchmark and a principled approach for advancing open-vocabulary perception in challenging marine environments, with implications for biodiversity monitoring and autonomous underwater exploration.

Abstract

Most existing underwater instance segmentation approaches are constrained by close-vocabulary prediction, limiting their ability to recognize novel marine categories. To support evaluation, we introduce \textbf{MARIS} (\underline{Mar}ine Open-Vocabulary \underline{I}nstance \underline{S}egmentation), the first large-scale fine-grained benchmark for underwater Open-Vocabulary (OV) segmentation, featuring a limited set of seen categories and diverse unseen categories. Although OV segmentation has shown promise on natural images, our analysis reveals that transfer to underwater scenes suffers from severe visual degradation (e.g., color attenuation) and semantic misalignment caused by lack underwater class definitions. To address these issues, we propose a unified framework with two complementary components. The Geometric Prior Enhancement Module (\textbf{GPEM}) leverages stable part-level and structural cues to maintain object consistency under degraded visual conditions. The Semantic Alignment Injection Mechanism (\textbf{SAIM}) enriches language embeddings with domain-specific priors, mitigating semantic ambiguity and improving recognition of unseen categories. Experiments show that our framework consistently outperforms existing OV baselines both In-Domain and Cross-Domain setting on MARIS, establishing a strong foundation for future underwater perception research.
Paper Structure (55 sections, 17 equations, 15 figures, 14 tables)

This paper contains 55 sections, 17 equations, 15 figures, 14 tables.

Figures (15)

  • Figure 1: The challenges of transferring OV instance segmentation to underwater scenarios in terms of (a) datasets and (b-c) methods, which have motivated the contributions of this study.
  • Figure 2: Visualization and analysis of the MARIS dataset. (a) Sample images from the MARIS dataset with object annotations. (b) Class split analysis, including Train Class, Insected Class, and OV Class. (c) Configuration of OV tasks, covering in-domain and cross-domain settings.
  • Figure 3: Overall framework of the proposed Method. The Geometric Prior Enhancement module strengthens structural representations via visual–geometric fusion and transformer-based query refinement. The Semantic Alignment Injection mechanism align category semantics with degraded underwater conditions.
  • Figure 4: Top-10 Best and Worst Classes: Comparison of in-domain and cross-domain AP, illustrating performance drops and gains with geometric-enhanced fusion.
  • Figure 5: Qualitative Results of visual information, geometric information, and their geometric-enhanced fusion, demonstrating clear improvements (viridis on the left and jet on the right).
  • ...and 10 more figures