Abstract
Human speech perception naturally integrates visual and auditory cues, with lip movements providing critical disambiguation in noisy environments where audio signals are degraded (SNR ≤ −5 dB). While existing audio-visual speech enhancement (AVSE) models use this multimodal synergy, they often fail to exploit phoneme-specific viseme correlations. We propose VG-MCA-AVSENet, a cognitively inspired architecture that introduces a viseme-gated multilayer cross-attention mechanism within a convolutional encoder-decoder framework. The model dynamically weights visual features using auditory context and learned viseme importance (prioritizing lip closures for plosives), while maintaining efficiency through depthwise separable convolutions and MobileNetV3-based visual encoding. Our AVSE system comprises an audio encoder, visual encoder with temporal modeling, separation module with 3-layer viseme-aware attention, and neural vocoder decoder. Evaluated on GRID-CHiME, NTCD-TIMIT, and LRS3 datasets, the model achieves 72.58% STOI (+12.28% over baseline), 2.56 PESQ, and 7.80 dB SI-SDR in extreme noise conditions (SNR≤ -6 dB). The viseme gating specifically improves plosive intelligibility by 7% with a few additional parameters, showing that explicit phoneme-viseme modeling significantly outperforms conventional approaches.
| Original language | English |
|---|---|
| Pages (from-to) | 469-481 |
| Number of pages | 13 |
| Journal | IEEE Transactions on Audio, Speech and Language Processing |
| Volume | 34 |
| DOIs | |
| State | Published - 2026 |
| Externally published | Yes |
Bibliographical note
Publisher Copyright:© 2025 IEEE.
Keywords
- Viseme gate attention
- audiovisual speech enhancement
- cross-attention
- multimodal feature fusion
ASJC Scopus subject areas
- Computer Science (miscellaneous)
- Electrical and Electronic Engineering
- Computational Mathematics
- Acoustics and Ultrasonics
Fingerprint
Dive into the research topics of 'Viseme-Gated Multilayer Cross-Attentional Feature Fusion for Cognitively-Inspired Multimodal Speech Enhancement'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver