Abstract
Explaining malware binaries poses significant challenges, as existing approaches often focus on surface-level features, dynamic behaviors, or assembly code analysis. While these models highlight features contributing to classification, they remain inaccessible to non-experts. In this work, we propose MalGPT, a multi-model and transformer-based approach that generates human-readable explanations of malware binaries in natural language. We manually analyzed malware binaries from different malware families, including benign files, using various tools to create a ground truth dataset with high-level explanations. As per the literature, this is the first contribution of a malware dataset paired with natural language explanations, along with a high-level explanatory model developed for the cybersecurity community. Our approach includes complex feature engineering, followed by a novel architecture, Cross-Hierarchical Attention Network (CHAiN), which learns relationships not only within individual features, but across different feature sets in a multi-model architecture. We developed a Generative Pretrained Transformer (GPT)-style architecture optimized for multi-modal malware binary analysis, designed to seamlessly integrate heterogeneous features, such as numeric data, printable strings, and graph-based representations of assembly code. The architecture aligns syntactic structures with semantic context, to transform encoded multi-modal inputs into coherent and precise explanations. This innovative approach enhances compatibility with diverse data modalities, providing robust and interpretable insights into malware behavior, while enabling detailed and contextually accurate textual explanations. In future work, we aim to scale this approach with larger datasets, enhancing its capacity to explain emerging malware variants and address different cybersecurity landscapes, such as malicious apps or network viruses, ultimately contributing to risk mitigation.
| Original language | English |
|---|---|
| Title of host publication | Machine Learning and Knowledge Discovery in Databases. Research Track - European Conference, ECML PKDD 2025, Proceedings |
| Editors | Rita P. Ribeiro, Bernhard Pfahringer, Nathalie Japkowicz, Pedro Larrañaga, Alípio M. Jorge, Carlos Soares, Pedro H. Abreu, João Gama |
| Publisher | Springer Science and Business Media Deutschland GmbH |
| Pages | 130-148 |
| Number of pages | 19 |
| ISBN (Print) | 9783032060778 |
| DOIs | |
| State | Published - 4 Oct 2025 |
| Externally published | Yes |
| Event | European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases, ECML PKDD 2025 - Porto, Portugal Duration: 15 Sep 2025 → 19 Sep 2025 |
Publication series
| Name | Lecture Notes in Computer Science |
|---|---|
| Volume | 16016 LNCS |
| ISSN (Print) | 0302-9743 |
| ISSN (Electronic) | 1611-3349 |
Conference
| Conference | European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases, ECML PKDD 2025 |
|---|---|
| Country/Territory | Portugal |
| City | Porto |
| Period | 15/09/25 → 19/09/25 |
Bibliographical note
Publisher Copyright:© The Author(s), under exclusive license to Springer Nature Switzerland AG 2026.
Keywords
- Explainable AI
- GPT Architecture
- Large Language Model
- Malware Analysis
ASJC Scopus subject areas
- Theoretical Computer Science
- General Computer Science
Fingerprint
Dive into the research topics of 'MalGPT: A Generative Explainable Model for Malware Binaries'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver