HMPFormer: Hierarchical vision transformer with multi-perspective feature learning for precise polyp segmentation

  • Usman, Muhammad Talha
  • Khan, Habib
  • Khan, Haseeb
  • Rida, Imad
  • Zhu, Xianxun
  • ... Koo, JaKeoung
Citations

WEB OF SCIENCE

8
Citations

SCOPUS

9

초록

Accurate polyp segmentation (PS) in endoscopic images is critical for early diagnosis and treatment of colorectal cancer, yet remains challenging due to significant variation in polyp morphology and complex boundaries. Existing approaches are limited by their reliance on uniform feature extraction mechanisms and basic multi-scale fusion strategies. Furthermore, single-stream attention mechanisms struggle to handle the diverse morphological characteristics and subtle boundary details of polyps in colonoscopy images. To address these limitations, we introduce a novel Hierarchical Vision Transformer with Multi-Perspective Feature Learning (MPFL) framework called HMPFormer for precise PS. Our framework integrates effective feature learning with attention-guided refinement, allowing robust feature representation and accurate boundary delineation across multiple scales. The proposed architecture leverages a Pyramid Vision Transformer v2 (PVTv2) backbone enhanced with a MPFL module, which enriches feature diversity through the impactful package of parallel dynamic, deformable, and standard convolutions. A Cross-Scale Context Fusion (CSCF) module establishes bidirectional cross-resolution interactions, enabling adaptive aggregation of multi-level semantic and spatial information. For precise mask generation, the progressive decoder equipped with Hierarchical Complementary Attention Module (HCAM) employs reverse refinement with residual attention, emphasizing overlooked regions through complementary masking. Extensive experiments on five benchmark datasets (Kvasir-SEG, CVC-ClinicDB, CVC-ColonDB, ETIS-LaribPolypDB, and Endoscene) demonstrate HMPFormer's superior performance over state-of-the-art approaches. Code, predcictions and trained models will be made publicly available at: https://github.com/TalhaUsman-ZERO/HMPFormer.

키워드

Endoscopic imagingPolyp segmentationCross fusionMulti-perspective learningHierarchical transformerATTENTION
제목
HMPFormer: Hierarchical vision transformer with multi-perspective feature learning for precise polyp segmentation
저자
Usman, Muhammad TalhaKhan, HabibKhan, HaseebRida, ImadZhu, XianxunKoo, JaKeoung
DOI
10.1016/j.imavis.2025.105777
발행일
2025-12
유형
Article
저널명
Image and Vision Computing
164