상세 보기
Lightweight transformer-driven multi-scale trapezoidal attention network for saliency detection
- Usman, Muhammad Talha;
- Khan, Habib;
- Rida, Imad;
- Koo, JaKeoung
WEB OF SCIENCE
75SCOPUS
86초록
State-of-the-art (SOTA) methods in salient object detection (SOD) often struggle to balance computational efficiency with performance. They encounter challenges in capturing complex features due to insufficient attention mechanisms and ineffective integration of multi-scale features. While deep convolutional neural networks (CNNs) provide lightweight architectures, they lack an understanding of global contexts. Transformers excel at capturing global context, but they require considerable resources, limiting their use on constrained platforms. To overcome these limitations, we present a novel saliency detection (SD) network that utilizes a pyramid vision transformer backbone to efficiently extract multi-scale features. The intermediate multi-scale features are refined through contextual feature refinement blocks (CFRBs) with dilated convolutions, which capture rich contextual information at each scale and enhance feature representation. Furthermore, these features are passed on to our newly proposed trapezoidal attention module (TAM), which integrates effective attention blocks. The Adaptive Spatial Coordinate Attention (ASCA) blocks are employed to highlight the importance of spatial locations in the initial two feature maps. As the high-level refined features require channel adjustments, we introduce Compact Channel Gate (CCG) blocks that adaptively recalibrate channel-wise feature responses. Additionally, the medium-scale two feature maps are processed through Feature-Aware Multi-Head Attention (FAMHA) blocks to capture long-range dependencies and global context. The extracted features are progressively upsampled and concatenated to create a high resolution saliency map, which is used for the final predictions. Extensive empirical analysis on six benchmark SD datasets demonstrates the effectiveness of our network, surpassing over 26 SOTA SD methods while maintaining a lightweight architecture. Code, qualitative results, and trained models will be available at: https://github.com/TalhaUsman-ZERO/TRSNet.
키워드
- 제목
- Lightweight transformer-driven multi-scale trapezoidal attention network for saliency detection
- 저자
- Usman, Muhammad Talha; Khan, Habib; Rida, Imad; Koo, JaKeoung
- 발행일
- 2025-09
- 유형
- Article
- 권
- 155