Skip to main navigation Skip to search Skip to main content

Constructing an Arabic Deepfake Detection Dataset from Government Sources and LLM-Generated Text

  • Najla A. Aljaloud*
  • , Amal Sunba
  • , Tarek Helmy
  • *Corresponding author for this work

Research output: Contribution to journalConference articlepeer-review

Abstract

The increasing use of large language models (LLMs) such as ChatGPT and DeepSeek poses new challenges for detecting AI-generated (deepfake) Arabic news. This paper constructs a high-quality dataset of authentic governmental Arabic news and corresponding deepfake articles generated using advanced LLMs. A BERT model fine-tuned on this dataset achieved an accuracy of 84.3% on an unseen test set. ross-domain assessment indicated some generalization, achieving 77.3% accuracy on X (Twitter) but exhibiting reduced performance in other domains. These results emphasize the capabilities and constraints of transformer-based methods for identifying Arabic deepfake news.

Original languageEnglish
Pages (from-to)385-392
Number of pages8
JournalProcedia Computer Science
Volume275
DOIs
StatePublished - 2026
Event7th International Conference on AI in Computational Linguistics, ACLing 2025 - Hybrid, Dubai, United Arab Emirates
Duration: 6 Dec 20257 Dec 2025

Bibliographical note

Publisher Copyright:
© 2025 The Authors. Published by Elsevier B.V.

Keywords

  • Arabic Deepfake Detection
  • BERT
  • Cross-Domain Validation
  • Dataset Construction
  • Large Language Models (LLMs)

ASJC Scopus subject areas

  • General Computer Science

Fingerprint

Dive into the research topics of 'Constructing an Arabic Deepfake Detection Dataset from Government Sources and LLM-Generated Text'. Together they form a unique fingerprint.

Cite this