Skip to main navigation Skip to search Skip to main content

Mandarin Electrolaryngeal Speech Voice Conversion with Speech Encoder Loss Learning and Seq2seq Modeling

  • Ming Chi Yen*
  • , Chia Hua Wu
  • , Shu Wei Tsai
  • , Jyh Shing Roger Jang
  • , Yu Tsao
  • , Amir Hussain
  • , Hsin Min Wang
  • *Corresponding author for this work

Research output: Contribution to journalArticlepeer-review

1 Scopus citations

Abstract

Electrolaryngeal (EL) speech utilizes excitation signals generated by an electrolarynx instead of human vocal vibrations. In daily communication, EL speech is less natural and more difficult to understand than natural (NL) speech due to mechanical vibration noise and fixed pitch. Different methods have been proposed to improve the quality and intelligibility of EL speech, but limited training data and atypical acoustic characteristics pose challenges. Voice conversion (VC) is one popular method, and the task is called EL speech VC (ELVC). Sequence-to-sequence (seq2seq) modeling with pretraining strategies has been proposed for ELVC. However, seq2seq ELVC still faces the problem of incomplete and missing phonemes. Furthermore, although previous work has evaluated simulated EL (sEL) speech produced by healthy speakers using electrolarynxes, the effectiveness of seq2seq ELVC on patient EL (pEL) speech has not been studied. In this article, we propose three approaches to address the issues of ELVC implementation. First, we utilize sEL speech in the pretraining stage to close the gap between pEL speech and NL speech. Second, we adopt a speech encoder loss to solve the problem of incomplete and missing phonemes. Third, we introduce waveform similarity overlap-and-add to augment pEL training speech. We conduct systematic experiments on pEL speech to evaluate our approaches. Ablation studies show that incorporating our approaches improves the converted speech in both objective and subjective evaluations compared to the baseline model.

Original languageEnglish
Pages (from-to)22-28
Number of pages7
JournalIEEE Internet of Things Magazine
Volume8
Issue number4
DOIs
StatePublished - 2025
Externally publishedYes

Bibliographical note

Publisher Copyright:
© 2018 IEEE.

ASJC Scopus subject areas

  • Software
  • Computer Networks and Communications
  • Computer Science Applications
  • Hardware and Architecture
  • Information Systems

Fingerprint

Dive into the research topics of 'Mandarin Electrolaryngeal Speech Voice Conversion with Speech Encoder Loss Learning and Seq2seq Modeling'. Together they form a unique fingerprint.

Cite this