Triple-domain Feature Learning with Frequency-aware Memory Enhancement for Moving Infrared Small Target Detection

Weiwei Duan; Luping Ji; Shengjia Chen; Sicheng Zhu; Mao Ye

Triple-domain Feature Learning with Frequency-aware Memory Enhancement for Moving Infrared Small Target Detection

Weiwei Duan, Luping Ji, Shengjia Chen, Sicheng Zhu, Mao Ye

TL;DR

This work tackles moving infrared small target detection by introducing Tridos, a triple-domain feature learning framework that fuses spatio-temporal information with frequency-domain features via a Fourier-based frequency-aware memory enhancement. The architecture comprises three branches—Memory-enhanced Spatial Relationship Module (MSRM), Temporal Dynamics Encoding Module (TDEM), and Local-global Frequency-aware Module (LGFM)—coupled with a Residual Compensation Unit (RCU) and a dual-view regression loss (L_dvr) to robustly detect tiny targets under challenging backgrounds. Through extensive experiments on DAUB, ITSDT-15K, and IRDST, Tridos achieves state-of-the-art performance, with ablations confirming the effectiveness of each component (MSRM, TDEM, LGFM, RCU) and the frequency-aware fusion strategy. While the approach yields notable accuracy gains, it incurs higher computational cost, motivating future work on efficient, lightweight implementations for real-time deployment in infrared target detection scenarios.

Abstract

As a sub-field of object detection, moving infrared small target detection presents significant challenges due to tiny target sizes and low contrast against backgrounds. Currently-existing methods primarily rely on the features extracted only from spatio-temporal domain. Frequency domain has hardly been concerned yet, although it has been widely applied in image processing. To extend feature source domains and enhance feature representation, we propose a new Triple-domain Strategy (Tridos) with the frequency-aware memory enhancement on spatio-temporal domain for infrared small target detection. In this scheme, it effectively detaches and enhances frequency features by a local-global frequency-aware module with Fourier transform. Inspired by human visual system, our memory enhancement is designed to capture the spatial relations of infrared targets among video frames. Furthermore, it encodes temporal dynamics motion features via differential learning and residual enhancing. Additionally, we further design a residual compensation to reconcile possible cross-domain feature mismatches. To our best knowledge, proposed Tridos is the first work to explore infrared target feature learning comprehensively in spatio-temporal-frequency domains. The extensive experiments on three datasets (i.e., DAUB, ITSDT-15K and IRDST) validate that our triple-domain infrared feature learning scheme could often be obviously superior to state-of-the-art ones. Source codes are available at https://github.com/UESTC-nnLab/Tridos.

Triple-domain Feature Learning with Frequency-aware Memory Enhancement for Moving Infrared Small Target Detection

TL;DR

Abstract

Paper Structure (30 sections, 21 equations, 11 figures, 12 tables)

This paper contains 30 sections, 21 equations, 11 figures, 12 tables.

Introduction
Related Work
Single-frame Infrared Small Target Detection
Multi-frame Infrared Small Target Detection
METHODOLOGY
Overall Architecture
Memory-enhanced Spatial Relationship Module
Temporal Dynamics Encoding Module
Local-global Frequency-aware Module
Residual Compensation Unit
Dual View Regression Loss
EXPERIMENTS
Datasets and Evaluation Metrics
Implementation Details
Comparisons With Other Methods
...and 15 more sections

Figures (11)

Figure 1: The comparison between existing spatio-temporal scheme and our triple-domain learning scheme. Our scheme extracts features in spatio-temporal-frequency domains.
Figure 2: Overview of proposed framework Tridos. $\textbf{I}_\textbf{c}$ is a group of video frames for Tridos. It consists of a backbone and three primary branches, i.e., a Memory-enhanced Spatial Relationship extraction branch, a Temporal Dynamic Encoding branch, and a Local-Global Frequency-aware branch. The features extracted by the first two branches are fused together to generate $\boldsymbol{F_{st}}$. The two outputs of the third branch, $\boldsymbol{F_{lf}}$ and $\boldsymbol{F_{gf}}$ are fused with $\boldsymbol{F_{st}}$ by RCU to generate $\boldsymbol{F_{stf_1}}$ and $\boldsymbol{F_{stf_2}}$, respectively. Finally, $\boldsymbol{F_{stf_1}}$ and $\boldsymbol{F_{stf_2}}$ are further compensated by RCU to obtain the refined features $\boldsymbol{F_{st}}$ for designed detection head.
Figure 3: The workflow of Tridos. It contains the calculation process of input $\textbf{I}_\textbf{c}$ and the feature flow of target detection.
Figure 4: The details of our proposed TDEM, with time window $T = 5$. $\text{ResB}_1$ and $\text{ResB}_2$ are two residual blocks for enhancing differential information.
Figure 5: The visualization comparisons of 14 methods on DAUB, with 21/63.bmp. GT is ground truth. Red and blue boxes represent detected targets and amplified detection regions, respectively. Yellow circles denote false alarms.
...and 6 more figures

Triple-domain Feature Learning with Frequency-aware Memory Enhancement for Moving Infrared Small Target Detection

TL;DR

Abstract

Triple-domain Feature Learning with Frequency-aware Memory Enhancement for Moving Infrared Small Target Detection

Authors

TL;DR

Abstract

Table of Contents

Figures (11)