Skip to main navigation Skip to search Skip to main content

Direct speech-to-speech translation with a sequence-to-sequence model

  • Ye Jia
  • , Ron J. Weiss
  • , Fadi Biadsy
  • , Wolfgang Macherey
  • , Melvin Johnson
  • , Zhifeng Chen
  • , Yonghui Wu

Research output: Contribution to journalConference articlepeer-review

162 Scopus citations

Abstract

We present an attention-based sequence-to-sequence neural network which can directly translate speech from one language into speech in another language, without relying on an intermediate text representation. The network is trained end-to-end, learning to map speech spectrograms into target spectrograms in another language, corresponding to the translated content (in a different canonical voice). We further demonstrate the ability to synthesize translated speech using the voice of the source speaker. We conduct experiments on two Spanish-to-English speech translation datasets, and find that the proposed model slightly underperforms a baseline cascade of a direct speech-to-text translation model and a text-to-speech synthesis model, demonstrating the feasibility of the approach on this very challenging task.

Original languageEnglish
Pages (from-to)1123-1127
Number of pages5
JournalProceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH
Volume2019-September
DOIs
StatePublished - 1 Jan 2019
Externally publishedYes
Event20th Annual Conference of the International Speech Communication Association: Crossroads of Speech and Language, INTERSPEECH 2019 - Graz, Austria
Duration: 15 Sep 201919 Sep 2019

Keywords

  • Attention
  • End-to-end model
  • Sequence-to-sequence model
  • Speech-to-speech translation
  • Voice transfer

ASJC Scopus subject areas

  • Language and Linguistics
  • Human-Computer Interaction
  • Signal Processing
  • Software
  • Modeling and Simulation

Fingerprint

Dive into the research topics of 'Direct speech-to-speech translation with a sequence-to-sequence model'. Together they form a unique fingerprint.

Cite this