Unveiling malicious PDF behavior: Interpretable classification and profiling malicious PDF using TabNet

Citations

WEB OF SCIENCE

0
Citations

SCOPUS

0

초록

With the inevitable growth of information digitization, Portable Document Format (PDF) has become one of the most popular exploited file formats for document exchange among various applications and platforms. Consequently, PDF files have become an attractive target for attackers to infect and deliver malicious codes to users. Despite the efficacy and success of machine learning classifiers in detecting malicious PDF files, they require tedious feature engineering and have some limitations. Additionally, one of the main reasons for resistance to using deep learning models is their lack of interpretability. To address these challenges, this study proposes using the TabNet model for malicious PDF detection, offering global and local interpretability while maintaining high or competitive detection performance. The Optuna optimization framework is employed to further enhance the model's capabilities. The proposed approach is evaluated on the real-world Evasive-PDFMal2022 dataset and demonstrates state-of-the-art performance compared to baseline methods.

키워드

Malicious PDFInterpretabilityProfilingDeep learningOF-THE-ARTNETWORK
제목
Unveiling malicious PDF behavior: Interpretable classification and profiling malicious PDF using TabNet
저자
Roudsari, Arousha HaghighianLashkari, Arash HabibiLoh, Woong-Kee
DOI
10.1016/j.jisa.2026.104487
발행일
2026-07
유형
Article
저널명
Journal of Information Security and Applications
100