Skip to main navigation Skip to search Skip to main content

MalGPT: A Generative Explainable Model for Malware Binaries

  • Mohd Saqib
  • , Benjamin C.M. Fung*
  • , Steven H.H. Ding
  • , Philippe Charland
  • *Corresponding author for this work

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

1 Scopus citations

Abstract

Explaining malware binaries poses significant challenges, as existing approaches often focus on surface-level features, dynamic behaviors, or assembly code analysis. While these models highlight features contributing to classification, they remain inaccessible to non-experts. In this work, we propose MalGPT, a multi-model and transformer-based approach that generates human-readable explanations of malware binaries in natural language. We manually analyzed malware binaries from different malware families, including benign files, using various tools to create a ground truth dataset with high-level explanations. As per the literature, this is the first contribution of a malware dataset paired with natural language explanations, along with a high-level explanatory model developed for the cybersecurity community. Our approach includes complex feature engineering, followed by a novel architecture, Cross-Hierarchical Attention Network (CHAiN), which learns relationships not only within individual features, but across different feature sets in a multi-model architecture. We developed a Generative Pretrained Transformer (GPT)-style architecture optimized for multi-modal malware binary analysis, designed to seamlessly integrate heterogeneous features, such as numeric data, printable strings, and graph-based representations of assembly code. The architecture aligns syntactic structures with semantic context, to transform encoded multi-modal inputs into coherent and precise explanations. This innovative approach enhances compatibility with diverse data modalities, providing robust and interpretable insights into malware behavior, while enabling detailed and contextually accurate textual explanations. In future work, we aim to scale this approach with larger datasets, enhancing its capacity to explain emerging malware variants and address different cybersecurity landscapes, such as malicious apps or network viruses, ultimately contributing to risk mitigation.

Original languageEnglish
Title of host publicationMachine Learning and Knowledge Discovery in Databases. Research Track - European Conference, ECML PKDD 2025, Proceedings
EditorsRita P. Ribeiro, Bernhard Pfahringer, Nathalie Japkowicz, Pedro Larrañaga, Alípio M. Jorge, Carlos Soares, Pedro H. Abreu, João Gama
PublisherSpringer Science and Business Media Deutschland GmbH
Pages130-148
Number of pages19
ISBN (Print)9783032060778
DOIs
StatePublished - 4 Oct 2025
Externally publishedYes
EventEuropean Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases, ECML PKDD 2025 - Porto, Portugal
Duration: 15 Sep 202519 Sep 2025

Publication series

NameLecture Notes in Computer Science
Volume16016 LNCS
ISSN (Print)0302-9743
ISSN (Electronic)1611-3349

Conference

ConferenceEuropean Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases, ECML PKDD 2025
Country/TerritoryPortugal
CityPorto
Period15/09/2519/09/25

Bibliographical note

Publisher Copyright:
© The Author(s), under exclusive license to Springer Nature Switzerland AG 2026.

Keywords

  • Explainable AI
  • GPT Architecture
  • Large Language Model
  • Malware Analysis

ASJC Scopus subject areas

  • Theoretical Computer Science
  • General Computer Science

Fingerprint

Dive into the research topics of 'MalGPT: A Generative Explainable Model for Malware Binaries'. Together they form a unique fingerprint.

Cite this