Transformer Based on Multi-Domain Feature Fusion for AI-Generated Image Detection

Citations

WEB OF SCIENCE

1
Citations

SCOPUS

1

초록

With the rapid advancement of Generative Adversarial Networks (GANs), diffusion models, and other deep generative techniques, AI-generated images have achieved unprecedented levels of visual realism, posing severe challenges to the authenticity, security, and credibility of digital content. This paper proposes a novel hybrid transformer model that integrates spatial and frequency domains. It leverages CLIP to extract semantic inconsistencies in the image's spatial domain while employing wavelet transforms to capture multi-scale frequency anomalies in AI-generated images. After cross-domain feature fusion, global modeling is performed within the Swin-Transformer architecture, enabling robust authenticity detection of AI-generated images. Extensive experiments demonstrate that our detector maintains high accuracy across diverse datasets.

키워드

AI-generated image detectionforged forensicsfrequency domain analysissemantic analysis
제목
Transformer Based on Multi-Domain Feature Fusion for AI-Generated Image Detection
저자
Man, QiaoyueCho, Young-Im
DOI
10.3390/electronics15030716
발행일
2026-02
유형
Article
저널명
ELECTRONICS
15
3