Skip to main navigation Skip to search Skip to main content

RecognAVSE-V2: An Improved Cross-Attention-Based Audio-Visual Speech Enhancement Approach

  • Leandro A. Passos*
  • , Joao R. Manesco
  • , Rahma Fourati
  • , Joao P. Papa
  • , Amir Hussain
  • *Corresponding author for this work

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Audio-visual speech enhancement (AVSE) is a sub-field of machine learning that aims to improve speech quality from noisy audio signals, leveraging visual cues to guide information flow and enhance learning. In this context, the COG-MHEAR program supports an early competition, namely the AVSE Challenge (AVSEC), to advance the development of new solutions and AVSE methods while fostering the community. This paper introduces the RecognAVSE-V2, an improved and more efficient version of RecognAVSE proposed at AVSEC 2024, that implements a Time-Synced Cross-Attention Mechanism that models the temporal dependencies between both modalities. Its architecture consists of a video encoder for spatiotemporal feature learning, an audio encoder operating on STFT representations, and a time-synchronized cross-Attention module that aligns audio and video features. Experimental results conducted over AVSE Challenge and CHiME3 datasets show RecognAVSE-V2 is capable of outperforming the baseline and its prior version in some cases while consisting of a model with reduced complexity, 20 times smaller, whose inference is 26% faster, and training demands a quarter of the total time.

Original languageEnglish
Title of host publicationProceedings of the IEEE International Conference on Complex Systems and Intelligent Computing, ICCSIC 2026
EditorsRahma Fourati, Habib M. Kammoun, Joao Paulo Papa, M. Tanveer
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages25-32
Number of pages8
ISBN (Electronic)9798319531162
DOIs
StatePublished - 2026
Event2026 IEEE International Conference on Complex Systems and Intelligent Computing, ICCSIC 2026 - Tabarka, Tunisia
Duration: 26 Mar 202629 Mar 2026

Publication series

NameProceedings of the IEEE International Conference on Complex Systems and Intelligent Computing, ICCSIC 2026

Conference

Conference2026 IEEE International Conference on Complex Systems and Intelligent Computing, ICCSIC 2026
Country/TerritoryTunisia
CityTabarka
Period26/03/2629/03/26

Bibliographical note

Publisher Copyright:
© 2026 IEEE.

Keywords

  • Audio-Video Temporal Alignment
  • Audio-Visual Speech Enhancement
  • Cross-Attention Mechanism
  • Feature Fusion
  • Lip-Speech Synchronization

ASJC Scopus subject areas

  • Artificial Intelligence
  • Computational Theory and Mathematics
  • Hardware and Architecture

Fingerprint

Dive into the research topics of 'RecognAVSE-V2: An Improved Cross-Attention-Based Audio-Visual Speech Enhancement Approach'. Together they form a unique fingerprint.

Cite this