Multimodal Fusion of Endoscopic and Histopathological Images for Lesion Detection Using Hybrid Deep Learning

  • Sahu, Premananda
  • Bharany, Salil
  • Singh, Jaibir
  • Mohamed, Heba G.
  • Rehman, Ateeq Ur
Citations

WEB OF SCIENCE

0
Citations

SCOPUS

0

초록

Introduction The classification of endoscopic lesions and the early diagnosis of gastrointestinal (GI) disorders rely heavily on precise lesion identification. However, subtle lesion appearances and considerable interobserver variability among clinicians often hinder accurate diagnosis. These challenges emphasize the need for robust, automated diagnostic systems that can integrate diverse image features to improve diagnostic accuracy and reliability.Materials and Methods To address these limitations, a novel deep learning framework, the Spatio-Histological Fusion Network (SHF-Net), was developed to enhance lesion detection by integrating multiple feature-extraction approaches. CNNs were used to capture local spatial characteristics from endoscopic images, ViTs to model long-range dependencies, and GNNs to incorporate structural histopathological information. A multi-channel attention mechanism was employed to fuse these features effectively. The model was trained on the HyperKvasir dataset, which contains 111,079 labeled endoscopic images and corresponding histopathological data across five lesion categories-normal, polyp, ulcer, cancer, and erosion. Preprocessing steps included stain normalization, and Generative Adversarial Networks (GANs) were used to overcome data scarcity. The training process combined cross-entropy and topological loss functions, optimized with AdamW using transfer learning.Results SHF-Net achieved exceptional classification performance with an accuracy of 98.84%, precision of 98.57%, recall of 98.21%, and an F1-score of 98.46% across the five lesion categories. Qualitative analysis using attention heatmaps showed that the model consistently focused on clinically significant lesion regions in both endoscopic and histopathological images, confirming its interpretability and clinical relevance.Discussion The integration of CNNs, ViTs, and GNNs in SHF-Net enabled the model to capture complementary spatial, contextual, and structural features, effectively addressing the challenges posed by subtle lesion patterns and interobserver variability. The multi-channel attention mechanism enhanced interpretability, an important factor for clinical adoption. These results demonstrate the advantages of combining multi-modal data and diverse network architectures for accurate and reliable GI lesion classification.Conclusion SHF-Net represents a significant advancement in automated GI diagnosis through the integration of spatial and histological imaging modalities. Its high accuracy, strong interpretability, and multi-modal design make it suitable for real-time deployment in clinical workflows. Future research will aim to extend SHF-Net to additional imaging modalities, validate it across diverse clinical datasets, and ensure its scalability for broader healthcare applications.

키워드

Multimodal fusionEndoscopic histopathological integrationAttention fusionSHF-NetCNNViTsGNNHyperKvasirGastrointestinal lesion classification
제목
Multimodal Fusion of Endoscopic and Histopathological Images for Lesion Detection Using Hybrid Deep Learning
저자
Sahu, PremanandaBharany, SalilSingh, JaibirMohamed, Heba G.Rehman, Ateeq Ur
DOI
10.2174/0115734056437477260427145329
발행일
2026-05
유형
Article
저널명
Current Medical Imaging Reviews
22