Fetching the paper…
Reading the bibliography…
Expressive speech-to-speech translation (S2ST) aims to transfer prosodic attributes of source speech to target speech while maintaining translation accuracy.
“JANUS: a speech-to-speech translation system using connectionist and symbolic processing strategies,”
A. McNair, A. Jain, A. Waibel, J. Tebelskis, A. Hauptmann, and H. Saito, · 1991
Earlier work this paper cites.
“Prosody Generation for Speech-to-Speech Translation,”
P. D. Aguero, J. Adell, and A. Bonafonte, · 2006
Earlier work this paper cites.
“The Blizzard Challenge 2013,”
S. King and V. Karaiskos, · 2013
Earlier work this paper cites.
“Appraisal Theories of Emotion: State of the Art and Future Development,”
A. Moors, P. C Ellsworth, K. R Scherer, and N. H Frijda, · 2013
Earlier work this paper cites.
“Mixed Emotions and Coping: The Benefits of Secondary Emotions,”
A. Braniecka, E. Trzebinska, A. Dowgiert, and A. Wytykowska, · 2014
Earlier work this paper cites.
“The case for mixed emotions,”
J. T Larsen and A P. McGraw, · 2014
Earlier work this paper cites.
“WORLD: A Vocoder-Based High-Quality Speech Synthesis System for Real-Time Applications,”
M. Morise, F. Yokomori, and K. Ozawa, · 2016
Earlier work this paper cites.
“Preserving Word-Level Emphasis in Speech-to-Speech Translation,”
Q. T. Do, T. Toda, G. Neubig, S. Sakti, and S. Nakamura, · 2017
Earlier work this paper cites.
“Montreal Forced Aligner: Trainable Text-Speech Alignment Using Kaldi,”
M. McAuliffe, M. Socolof, S. Mihuc, M. Wagner, and M. Sonderegger, · 2017
Earlier work this paper cites.
“Subjective evaluation of speech quality with a crowdsourcing approach,” 2018
ITU-T Recommendation P.808, · 2018
Earlier work this paper cites.
“Bilingual Prosodic Dataset Compilation for Spoken Language Translation,”
A. Öktem, M. Farrús, and A. Bonafonte, · 2018
Earlier work this paper cites.
“Sequence-to-Sequence Models for Emphasis Speech Translation,”
Q. T. Do, S. Sakti, and S. Nakamura, · 2018
Earlier work this paper cites.
“Towards End-to-End Prosody Transfer for Expressive Speech Synthesis with Tacotron,”
RJ Skerry-Ryan, E. Battenberg, Y. Xiao, Y. Wang, D. Stanton, J. Shor, R. Weiss, R. Clark, and R. A. Saurous, · 2018
Cited alongside, same era.
“Style Tokens: Unsupervised Style Modeling, Control and Transfer in End-to-End Speech Synthesis,”
Y. Wang, D. Stanton, Y. Zhang, RJ-Skerry Ryan, E. Battenberg, J. Shor, Y. Xiao, Y. Jia, F. Ren, and R. A. Saurous, · 2018
Cited alongside, same era.
“Direct Speech-to-Speech Translation with a Sequence-to-Sequence Model,”
Y. Jia, R. J. Weiss, F. Biadsy, W. Macherey, M. Johnson, Z. Chen, and Y. Wu, · 2019
Cited alongside, same era.
“Learning Latent Representations for Style Control and Transfer in End-to-end Speech Synthesis,”
Y.-J. Zhang, S. Pan, L. He, and Z.-H. Ling, · 2019
Cited alongside, same era.
“Hierarchical Generative Modeling for Controllable Speech Synthesis,”
W.-N. Hsu, Y. Zhang, R. Weiss, H. Zen, Y. Wu, Y. Cao, and Y. Wang, · 2019
Cited alongside, same era.
“Meta-StyleSpeech : Multi-Speaker Adaptive Text-to-Speech Generation,”
D. Min, D. B. Lee, E. Yang, and S. J. Hwang, · 2021
Later among the works it cites.
“FastSpeech 2: Fast and High-Quality End-to-End Text to Speech,”
Y. Ren, C. Hu, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, · 2021
Later among the works it cites.
“PortaSpeech: Portable and High-Quality Generative Text-to-Speech,”
Y. Ren, J. Liu, and Z. Zhao, · 2021
Later among the works it cites.
“fairseq sˆ2: A scalable and integrable speech synthesis toolkit,”
C. Wang, W.-N. Hsu, Y. Adi, A. Polyak, A. Lee, P.-J. Chen, J. Gu, and J. Pino, · 2021
Later among the works it cites.
“VoxPopuli: A large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation,”
C. Wang, M. Riviere, A. Lee, A. Wu, C. Talnikar, D. Haziza, M. Williamson, J. Pino, and E. Dupoux, · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Robust and Fine-grained Prosody Control of End-to-end Speech Synthesis,”
Y. Lee and T. Kim, · 2019
Cited alongside, same era.
“Fine-Grained Robust Prosody Transfer for Single-Speaker Neural Text-To-Speech,”
V. Klimkov, S. Ronanki, J. Rohnke, and T. Drugman, · 2019
Cited alongside, same era.
“fairseq: A Fast, Extensible Toolkit for Sequence Modeling,”
M. Ott, S. Edunov, A. Baevski, A. Fan, S. Gross, N. Ng, D. Grangier, and M. Auli, · 2019
Cited alongside, same era.
“Fairseq S2T: Fast speech-to-text modeling with fairseq,”
C. Wang, Y. Tang, X. Ma, A. Wu, D. Okhonko, and J. Pino, · 2020
Cited alongside, same era.
“HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis,”
J. Kong, J. Kim, and J. Bae, · 2020
Cited alongside, same era.
“Real Time Speech Enhancement in the Waveform Domain,”
A. Défossez, G. Synnaeve, and Y. Adi, · 2020
Cited alongside, same era.
Z.-Y. Dou and G. Neubig, · 2021
Later among the works it cites.
“Multimodal and Multilingual Embeddings for Large-Scale Speech Mining,”
P.-A. Duquenne, H. Gong, and H. Schwenk, · 2021
Later among the works it cites.
“Translatotron 2: High-quality direct speech-to-speech translation with voice preservation,”
Y. Jia, M. T. Ramanovich, T. Remez, and R. Pomerantz, · 2022
Later among the works it cites.
“Direct Speech-to-Speech Translation With Discrete Units,”
A. Lee, P.-J. Chen, C. Wang, J. Gu, S. Popuri, X. Ma, A. Polyak, Y. Adi, Q. He, Y. Tang, J. Pino, and W.-N. Hsu, · 2022
Later among the works it cites.
“Textless Speech-to-Speech Translation on Real Data,”
A. Lee, H. Gong, P.-A. Duquenne, H. Schwenk, P.-J. Chen, C. Wang, S. Popuri, Y. Adi, J. Pino, J. Gu, and W.-N. Hsu, · 2022
Later among the works it cites.
“Face-Dubbing++: Lip-Synchronous, Voice Preserving Translation of Videos,”
A. Waibel, M. Behr, F. I. Eyiokur, D. Yaman, T.-N. Nguyen, C. Mullov, M. A. Demirtas, A. Kantarcı, S. Constantin, and H. K. Ekenel, · 2022
Later among the works it cites.