TransUTD: Underwater cross-domain collaborative spatial-temporal transformer detector.
Journal:
Neural networks : the official journal of the International Neural Network Society
Published Date:
Feb 10, 2026
Abstract
Research on underwater object detection has primarily focused on addressing degraded imagery. Single-frame feature refinement is inherently limited by restricted static spatial information, while joint image enhancement and detection paradigms encounter non-trivial challenges arising from irreversible artifacts and conflicting optimization objectives. In contrast, temporal information from video sequences offers a direct solution. Temporal semantic information enhances the feature representation of degraded underwater frames, whereas temporal positional cues furnish dynamic geometric associations that facilitate precise object localization. We propose Transformer Underwater Spatial-Temporal Cross-domain Collaborative Detection (TransUTD), reformulating underwater degraded feature representation as a temporal contextual modeling problem. By synergistic exploitation of spatial-temporal information, TransUTD naturally learns complementary features across frames to compensate for feature degradation in key frames, rather than relying on specific heuristic components. This simplifies the detection pipeline and eliminates hand-crafted modules. In our framework, the spatial-temporal fusion encoder aggregates multi-frame features to strengthen semantic representations in degraded images. The spatial-temporal query interaction refines localization in complex underwater scenes by correlating spatial-temporal geometric cues. Finally, the temporal hybrid collaborative decoder performs dense supervision through collaborative optimization of temporal positive queries. Concurrently, we construct UVID, the first underwater video object detection dataset. Experimental evaluations demonstrate that TransUTD achieves state-of-the-art performance, delivering AP improvements of 1.5% and 1.9% on the DUO and UVID datasets, respectively. Moreover, it attains near SOTA performance on ImageNetVID with AP50 of 86.0%. Our dataset and code are available at https://github.com/Anchor1566/TransUTD.
Authors
Keywords
No keywords available for this article.