Abstract
Audio-visual speech enhancement (AVSE) is a sub-field of machine learning that aims to improve speech quality from noisy audio signals, leveraging visual cues to guide information flow and enhance learning. In this context, the COG-MHEAR program supports an early competition, namely the AVSE Challenge (AVSEC), to advance the development of new solutions and AVSE methods while fostering the community. This paper introduces the RecognAVSE-V2, an improved and more efficient version of RecognAVSE proposed at AVSEC 2024, that implements a Time-Synced Cross-Attention Mechanism that models the temporal dependencies between both modalities. Its architecture consists of a video encoder for spatiotemporal feature learning, an audio encoder operating on STFT representations, and a time-synchronized cross-Attention module that aligns audio and video features. Experimental results conducted over AVSE Challenge and CHiME3 datasets show RecognAVSE-V2 is capable of outperforming the baseline and its prior version in some cases while consisting of a model with reduced complexity, 20 times smaller, whose inference is 26% faster, and training demands a quarter of the total time.
| Original language | English |
|---|---|
| Title of host publication | Proceedings of the IEEE International Conference on Complex Systems and Intelligent Computing, ICCSIC 2026 |
| Editors | Rahma Fourati, Habib M. Kammoun, Joao Paulo Papa, M. Tanveer |
| Publisher | Institute of Electrical and Electronics Engineers Inc. |
| Pages | 25-32 |
| Number of pages | 8 |
| ISBN (Electronic) | 9798319531162 |
| DOIs | |
| State | Published - 2026 |
| Event | 2026 IEEE International Conference on Complex Systems and Intelligent Computing, ICCSIC 2026 - Tabarka, Tunisia Duration: 26 Mar 2026 → 29 Mar 2026 |
Publication series
| Name | Proceedings of the IEEE International Conference on Complex Systems and Intelligent Computing, ICCSIC 2026 |
|---|
Conference
| Conference | 2026 IEEE International Conference on Complex Systems and Intelligent Computing, ICCSIC 2026 |
|---|---|
| Country/Territory | Tunisia |
| City | Tabarka |
| Period | 26/03/26 → 29/03/26 |
Bibliographical note
Publisher Copyright:© 2026 IEEE.
Keywords
- Audio-Video Temporal Alignment
- Audio-Visual Speech Enhancement
- Cross-Attention Mechanism
- Feature Fusion
- Lip-Speech Synchronization
ASJC Scopus subject areas
- Artificial Intelligence
- Computational Theory and Mathematics
- Hardware and Architecture
Fingerprint
Dive into the research topics of 'RecognAVSE-V2: An Improved Cross-Attention-Based Audio-Visual Speech Enhancement Approach'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver