Abstract
Voice User Interfaces (VUI) in consumer electronics often operate in noisy environments where speech quality degrades significantly. We propose AV-Net, a lightweight audiovisual speech enhancement model that addresses this challenge through novel cross-attentional feature fusion. Our approach dynamically integrates audio and visual modalities using a computationally efficient architecture combining a convolutional encoder-decoder (5.2M parameters) with a P3D-ResNet18 video encoder. The key innovation is a cross-attention mechanism that learns inter-modal correlations while preserving modality-specific features, outperforming conventional fusion methods. Evaluated on TCD-TIMIT and AVSE3 datasets under challenging conditions (SNR \leq -5dB), AV-Net achieves a PESQ of 2.56 (vs. 1.26 baseline) and STOI improvement of 22%, while maintaining real-time performance (RTF=0.11). The model demonstrates strong generalization to unseen speakers and diverse noise types, making it particularly suitable for resource-constrained edge devices in healthcare, automotive, and smart home applications where robust speech interaction is critical.
| Original language | English |
|---|---|
| Pages (from-to) | 10123-10133 |
| Number of pages | 11 |
| Journal | IEEE Transactions on Consumer Electronics |
| Volume | 71 |
| Issue number | 4 |
| DOIs | |
| State | Published - 2025 |
| Externally published | Yes |
Bibliographical note
Publisher Copyright:© 1975-2011 IEEE.
Keywords
- Attention process
- audiovisual speech enhancement
- multimodal feature integration
- voice user interface (VUI)
ASJC Scopus subject areas
- Media Technology
- Electrical and Electronic Engineering
Fingerprint
Dive into the research topics of 'Attentional Multimodal Speech Enhancement for Voice User Interface in Consumer Electronics'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver