Fetching the paper…
Reading the bibliography…
This paper describes Tacotron 2, a neural network architecture for speech synthesis directly from text.
“Comparison of parametric representations for monosyllabic word recognition in continuously spoken sentences,”
S. Davis and P. Mermelstein, · 1980
Earlier work this paper cites.
“Signal estimation from modified short-time Fourier transform,”
D. W. Griffin and J. S. Lim, · 1984
Earlier work this paper cites.
“Mixture density networks,”
C. M. Bishop, · 1994
Earlier work this paper cites.
“Unit selection in a concatenative speech synthesis system using a large speech database,”
A. J. Hunt and A. W. Black, · 1996
Earlier work this paper cites.
“Automatically clustering similar units for unit selection in speech synthesis,”
A. W. Black and P. Taylor, · 1997
Earlier work this paper cites.
“Bidirectional recurrent neural networks,”
M. Schuster and K. K. Paliwal, · 1997
Earlier work this paper cites.
“Long short-term memory,”
S. Hochreiter and J. Schmidhuber, · 1997
Earlier work this paper cites.
On supervised learning from sequential data with applications for speech recognition
M. Schuster, · 1999
Earlier work this paper cites.
“Speech parameter generation algorithms for HMM-based speech synthesis,”
K. Tokuda, T. Yoshimura, T. Masuko, T. Kobayashi, and T. Kitamura, · 2000
Earlier work this paper cites.
Text-to-Speech Synthesis
P. Taylor, · 2009
Earlier work this paper cites.
“Statistical parametric speech synthesis,”
H. Zen, K. Tokuda, and A. W. Black, · 2009
Earlier work this paper cites.
“Statistical parametric speech synthesis using deep neural networks,”
H. Zen, A. Senior, and M. Schuster, · 2013
Cited alongside, same era.
“Speech synthesis based on hidden Markov models,”
K. Tokuda, Y. Nankaku, T. Toda, H. Zen, J. Yamagishi, and K. Oura, · 2013
Cited alongside, same era.
“Sequence to sequence learning with neural networks.,”
I. Sutskever, O. Vinyals, and Q. V. Le, · 2014
Cited alongside, same era.
“Dropout: a simple way to prevent neural networks from overfitting.,”
N. Srivastava, G. E. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, · 2014
Cited alongside, same era.
“Batch normalization: Accelerating deep network training by reducing internal covariate shift,”
S. Ioffe and C. Szegedy, · 2015
Cited alongside, same era.
“Attention-based models for speech recognition,”
J. K. Chorowski, D. Bahdanau, D. Serdyuk, K. Cho, and Y. Bengio, · 2015
“Fast, compact, and high quality LSTM-RNN based statistical parametric speech synthesizers for mobile devices,”
H. Zen, Y. Agiomyrgiannakis, N. Egberts, F. Henderson, and P. Szczepaniak, · 2016
Later among the works it cites.
“Deep voice: Real-time neural text-to-speech,”
S. Ö. Arik, M. Chrzanowski, A. Coates, G. Diamos, A. Gibiansky, Y. Kang, X. Li, J. Miller, J. Raiman, S. Sengupta, and M. Shoeybi, · 2017
Closest in time.
“Deep voice 2: Multi-speaker neural text-to-speech,”
S. Ö. Arik, G. F. Diamos, A. Gibiansky, J. Miller, K. Peng, W. Ping, J. Raiman, and Y. Zhou, · 2017
Closest in time.
“Deep voice 3: 2000-speaker neural text-to-speech,”
W. Ping, K. Peng, A. Gibiansky, S. Ö. Arik, A. Kannan, S. Narang, J. Raiman, and J. Miller, · 2017
Closest in time.
“Tacotron: Towards end-to-end speech synthesis,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Neural machine translation by jointly learning to align and translate,”
D. Bahdanau, K. Cho, and Y. Bengio, · 2015
Cited alongside, same era.
“Adam: A method for stochastic optimization,”
D. P. Kingma and J. Ba, · 2015
Cited alongside, same era.
“WaveNet: A generative model for raw audio,”
A. van den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. W. Senior, and K. Kavukcuoglu, · 2016
Cited alongside, same era.
“Recent advances in Google real-time HMM-driven unit selection synthesizer,”
X. Gonzalvo, S. Tazari, C.-a. Chan, M. Becker, A. Gutkin, and H. Silen, · 2016
Cited alongside, same era.
Y. Wang, R. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengio, Q. Le, Y. Agiomyrgiannakis, R. Clark, and R. A. Saurous, · 2017
Closest in time.
“Speaker-dependent WaveNet vocoder,”
A. Tamamori, T. Hayashi, K. Kobayashi, K. Takeda, and T. Toda, · 2017
Closest in time.
“Char2Wav: End-to-end speech synthesis,”
J. Sotelo, S. Mehri, K. Kumar, J. F. Santos, K. Kastner, A. Courville, and Y. Bengio, · 2017
Closest in time.
“Zoneout: Regularizing RNNs by randomly preserving hidden activations,”
D. Krueger, T. Maharaj, J. Kramár, M. Pezeshki, N. Ballas, N. R. Ke, A. Goyal, Y. Bengio, H. Larochelle, A. Courville, et al., · 2017
Closest in time.
“PixelCNN++: Improving the PixelCNN with discretized logistic mixture likelihood and other modifications,”
T. Salimans, A. Karpathy, X. Chen, and D. P. Kingma, · 2017
Closest in time.
“Parallel WaveNet: Fast High-Fidelity Speech Synthesis,”
A. van den Oord, Y. Li, I. Babuschkin, K. Simonyan, O. Vinyals, K. Kavukcuoglu, G. van den Driessche, E. Lockhart, L. C. Cobo, F. Stimberg, N. Casagrande, D. Grewe, S. Noury, S. Dieleman, E. Elsen, N. Kalchbrenner, H. Zen, A. Graves, H. King, T. Walters, D. Belov, and D. Hassabis, · 2017
Closest in time.