Fetching the paper…
Reading the bibliography…
This letter presents an incremental text-to-speech (TTS) method that performs synthesis in small linguistic units while maintaining the naturalness of output speech.
“Binary codes capable of correcting deletions, insertions and reversals,”
V. Levenshtein, · 1966
Earlier work this paper cites.
“An HMM-based speech synthesis system applied to English,”
K. Tokuda, H. Zen, and A. W. Black, · 2002
Earlier work this paper cites.
“Statistical parametric speech synthesis,”
H. Zen, K. Tokuda, and A. Black, · 2009
Earlier work this paper cites.
“Amazon’s Mechanical Turk: A new source of inexpensive, yet high-quality, data?,”
M. Buhrmester, T. Kwang, and S. D. Gosling, · 2011
Earlier work this paper cites.
“Real-time incremental speech-to-speech translation of dialogs,”
S. Bangalore, V. K. Rangarajan Sridhar, P. Kolan, L. Golipour, and A. Jimenez, · 2012
Earlier work this paper cites.
“Statistical parametric speech synthesis using deep neural networks,”
H. Zen, A. Senior, and M. Schuster, · 2013
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
D. Kingma and B. Jimmy, · 2014
Earlier work this paper cites.
“Librispeech: An ASR corpus based on public domain audio books,”
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, · 2015
Earlier work this paper cites.
“Mean opinion score (MOS) revisited: methods and applications, limitations and alternatives,”
R. Streijl, S. Winkler, and D. Hands, · 2016
Earlier work this paper cites.
“Tacotron: Towards end-to-end speech synthesis,”
Y. Wang, R. J. Skerry-Ryan, D. Stanton, Y. Wu, Ron J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengio, Q. Le, Y. Agiomyrgiannakis, R. Clark, and R. A. Saurous, · 2017
Cited alongside, same era.
“The LJ Speech Dataset,” https://keithito.com/LJ-Speech-Dataset/
K. Ito and L. Johnson, · 2017
Cited alongside, same era.
“Gentle: A robust yet lenient forced aligner built on Kaldi,” https://lowerquality.com/gentle/
R. M. Ochshorn and M. Hawkins., · 2017
Cited alongside, same era.
“Joint CTC-attention based end-to-end speech recognition using multi-task learning,”
S. Kim, T. Hori, and S. Watanabe, · 2017
Cited alongside, same era.
“Natural TTS synthesis by conditioning WaveNet on mel spectrogram predictions,”
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerrv-Ryan, R. A. Saurous, Y. Agiomvrgiannakis, and Y. Wu, · 2018
Cited alongside, same era.
“Language models are unsupervised multitask learners,” https://openai.com/blog/better-language-models/
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever, · 2019
Later among the works it cites.
“Neural speech synthesis with transformer network,”
N. Li, S. Liu, Y. Liu, S. Zhao, and M. Liu, · 2019
Later among the works it cites.
“Neural iTTS: Toward synthesizing speech in real-time with end-to-end neural text-to-speech framework,”
T. Yanagita, S. Sakti, and S. Nakamura, · 2019
Later among the works it cites.
“Simultaneous speech-to-speech translation system with neural incremental ASR, MT, and TTS,”
K. Sudoh, T. Kano, S. Novitasari, T. Yanagita, S. Sakti, and S. Nakamura, · 2020
Closest in time.
“Incremental text-to-speech synthesis with prefix-to-prefix framework,”
M. Ma, B. Zheng, K. Liu, R. Zheng, H. Liu, K. Peng, K. Church, and L. Huang, · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“WaveGlow: A flow-based generative network for speech synthesis,”
R. Prenger, R. Valle, and B. Catanzaro, · 2018
Cited alongside, same era.
“Style Tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis,”
Y. Wang, D. Stanton, Y. Zhang, R. Skerry-Ryan, E. Battenberg, J. Shor, Y. Xiao, F. Ren, Y. Jia, and R. A. Saurous, · 2018
Cited alongside, same era.
“Hierarchical neural story generation,”
A. Fan, M. Lewis, and Y. Dauphin, · 2018
Cited alongside, same era.
“ESPnet: End-to-end speech processing toolkit,”
S. Watanabe, T. Hori, S. Karita, T. Hayashi, J. Nishitoba, Y. Unno, N. E. Y. Soplin, J. Heymann, M. Wiesner, N. Chen, A. Renduchintala, and T. Ochiai, · 2018
Cited alongside, same era.
Y. Cong, R. Zhang, and J. Luan, · 2020
Closest in time.
“What the future brings: Investigating the impact of lookahead for incremental neural TTS,”
B. Stephenson, L. Besacier, L. Girin, and T. Hueber, · 2020
Closest in time.
“Incremental text to speech for neural sequence-to-sequence models using reinforcement learning,”
D. S. R. Mohan, R. Lenain, L. Foglianti, T. H. Teh, M. Staib, and A. Torresquintero, · 2020
Closest in time.