Abstract
Deep-sea manganese nodules constitute a key seabed metal resource, and accurate segmentation of their spatial distribution in near-bottom imagery is critical for supporting autonomous mining-vehicle operations, path planning, and resource evaluation. However, achieving accurate segmentation remains challenging due to low contrast, complex backgrounds, and significant scale variation, and existing models still struggle to capture fine details and robust multi-scale cues. To address these issues, we present WEAR-Seg model, which establishes a unified spatial–frequency representation framework that enhances boundaries, textures, and multi-scale structures in complex underwater imagery. The model integrates three key components:WaveFuse, which uses wavelet decomposition and reconstruction to fuse frequency cues with spatial context to enhance structural details;AdaRep, which adaptively captures multi-scale features to highlight key regions and suppress background noise; andLiteConv, a lightweight convolutional module that improves small-object recognition and model deployability. We construct a dedicated manganese nodule dataset from Jiaolong submersible sea-trial imagery to support training and evaluation. Experimental results show that WEAR-Seg outperforms representative baselines across standard metrics. Ablation studies confirm thatWaveFuseandAdaRepcontribute most, especially for small targets and complex backgrounds, demonstrating the model’s effectiveness in accurate and fine-detail segmentation under challenging underwater conditions.
Showing the abstract — retrieve the full paper via the Exa API.