Abstract
Lip reading, the process of understanding spoken words through visual observation of lip movements, holds significant promise for improving accessibility in autonomous transportation systems. This research introduces a visual Arabic lip-reading model aimed at facilitating inclusive access to in-vehicle information for people with hearing challenges. As transportation technologies evolve toward greater autonomy and intelligence, the exchange between passengers and vehicles must go beyond sound-based communication. The proposed framework allows users to engage with the vehicle through lip gestures, enabling natural interaction for accessing real-time updates in their native language. This study addresses this gap by presenting an empirical investigation into Arabic visual lip reading, emphasizing the challenges unique to this language. To advance research in this area, we constructed a comprehensive dataset designed to capture the diversity of Arabic speakers, accounting for variations in pronunciation, dialect, and speaking styles. Furthermore, we explored the efficacy of combining 3D Convolutional Neural Networks (3DCNN) with Bidirectional Long Short-Term Memory (Bi-LSTM) networks to perform visual speech recognition tasks in Arabic. The results from our experiments reveal the complexity of processing Arabic speech, particularly as the number of target classes increases. These findings highlight the need for continued innovation in model architecture and dataset development to overcome the unique challenges posed by Arabic lip reading, paving the way for more inclusive and robust visual speech recognition systems.
| Original language | English |
|---|---|
| Pages (from-to) | 141-147 |
| Number of pages | 7 |
| Journal | Transportation Research Procedia |
| Volume | 97 |
| DOIs | |
| State | Published - 2026 |
| Event | 13th International Conference on Transport Survey Methods, 2026 - Danang, Viet Nam Duration: 30 Mar 2025 → 4 Apr 2025 |
Bibliographical note
Publisher Copyright:Copyright © 2026. Published by Elsevier B.V.
Keywords
- Arabic dataset
- Bi-LSTM
- CNN
- Deep learning
- Lip reading
ASJC Scopus subject areas
- Transportation
Fingerprint
Dive into the research topics of 'Empirical Study of Arabic Visual Lip Reading for Human–Vehicle Interaction in Autonomous Mobility'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver