Skip to main navigation Skip to search Skip to main content

Attentional Multimodal Speech Enhancement for Voice User Interface in Consumer Electronics

  • Nasir Saleem*
  • , Sami Bourouis
  • , Muhammad Irfan
  • , Kia Dashtipour
  • , Hela Elmannai
  • , Ahmed Al-Dubai
  • , Tughrul Arslan
  • , Amir Hussain
  • *Corresponding author for this work

Research output: Contribution to journalArticlepeer-review

7 Scopus citations

Abstract

Voice User Interfaces (VUI) in consumer electronics often operate in noisy environments where speech quality degrades significantly. We propose AV-Net, a lightweight audiovisual speech enhancement model that addresses this challenge through novel cross-attentional feature fusion. Our approach dynamically integrates audio and visual modalities using a computationally efficient architecture combining a convolutional encoder-decoder (5.2M parameters) with a P3D-ResNet18 video encoder. The key innovation is a cross-attention mechanism that learns inter-modal correlations while preserving modality-specific features, outperforming conventional fusion methods. Evaluated on TCD-TIMIT and AVSE3 datasets under challenging conditions (SNR \leq -5dB), AV-Net achieves a PESQ of 2.56 (vs. 1.26 baseline) and STOI improvement of 22%, while maintaining real-time performance (RTF=0.11). The model demonstrates strong generalization to unseen speakers and diverse noise types, making it particularly suitable for resource-constrained edge devices in healthcare, automotive, and smart home applications where robust speech interaction is critical.

Original languageEnglish
Pages (from-to)10123-10133
Number of pages11
JournalIEEE Transactions on Consumer Electronics
Volume71
Issue number4
DOIs
StatePublished - 2025
Externally publishedYes

Bibliographical note

Publisher Copyright:
© 1975-2011 IEEE.

Keywords

  • Attention process
  • audiovisual speech enhancement
  • multimodal feature integration
  • voice user interface (VUI)

ASJC Scopus subject areas

  • Media Technology
  • Electrical and Electronic Engineering

Fingerprint

Dive into the research topics of 'Attentional Multimodal Speech Enhancement for Voice User Interface in Consumer Electronics'. Together they form a unique fingerprint.

Cite this