Fetching the paper…
Reading the bibliography…
We describe a sequence-to-sequence neural network which directly generates speech waveforms from text inputs.
“Signal estimation from modified short-time Fourier transform,”
D. Griffin and J. Lim, · 1984
Earlier work this paper cites.
“Mel-cepstral distance measure for objective speech quality assessment,”
R. Kubichek, · 1993
Earlier work this paper cites.
“Using dynamic time warping to find patterns in time series,”
D. J. Berndt and J. Clifford, · 1994
Earlier work this paper cites.
“RNADE: The real-valued neural autoregressive density-estimator,”
B. Uria, I. Murray, and H. Larochelle, · 2013
Earlier work this paper cites.
“Variational inference with normalizing flows,”
D. Rezende and S. Mohamed, · 2015
Earlier work this paper cites.
“Attention-based models for speech recognition,”
J. K. Chorowski, D. Bahdanau, D. Serdyuk, K. Cho, and Y. Bengio, · 2015
Earlier work this paper cites.
“NICE: Non-linear independent components estimation,”
L. Dinh, D. Krueger, and Y. Bengio, · 2015
Earlier work this paper cites.
“WaveNet: A generative model for raw audio,”
A. van den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, et al., · 2016
Earlier work this paper cites.
“Tacotron: Towards end-to-end speech synthesis,”
Y. Wang, R. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, et al., · 2017
Earlier work this paper cites.
“Char2Wav: End-to-end speech synthesis,”
J. Sotelo, S. Mehri, K. Kumar, J. F. Santos, K. Kastner, A. Courville, and Y. Bengio, · 2017
Earlier work this paper cites.
“Density estimation using Real NVP,”
L. Dinh and S. Bengio, · 2017
Earlier work this paper cites.
“Attention is all you need,”
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, · 2017
Earlier work this paper cites.
“The LJ Speech Dataset,” https://keithito.com/LJ-Speech-Dataset/ , 2017
K. Ito, · 2017
Earlier work this paper cites.
“Natural TTS synthesis by conditioning WaveNet on mel spectrogram predictions,”
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, et al., · 2018
Cited alongside, same era.
“Efficient neural audio synthesis,”
N. Kalchbrenner, E. Elsen, K. Simonyan, S. Noury, N. Casagrande, E. Lockhart, F. Stimberg, A. van den Oord, et al., · 2018
Cited alongside, same era.
“Parallel WaveNet: Fast High-Fidelity Speech Synthesis,”
A. van den Oord, Y. Li, I. Babuschkin, K. Simonyan, O. Vinyals, K. Kavukcuoglu, G. v. d. Driessche, E. Lockhart, L. C. Cobo, et al., · 2018
Cited alongside, same era.
“Fast spectrogram inversion using multi-head convolutional neural networks,”
S. Ö. Arık, H. Jun, and G. Diamos, · 2018
Cited alongside, same era.
“Glow: Generative flow with invertible 1x1 convolutions,”
D. P. Kingma and P. Dhariwal, · 2018
Cited alongside, same era.
“Towards end-to-end prosody transfer for expressive speech synthesis with tacotron,”
“Hierarchical generative modeling for controllable speech synthesis,”
W. N. Hsu, Y. Zhang, R. J. Weiss, H. Zen, Y. Wu, Y. Wang, Y. Cao, Y. Jia, Z. Chen, J. Shen, P. Nguyen, and R. Pang, · 2019
Later among the works it cites.
“WaveFlow: A compact flow-based model for raw audio,”
W. Ping, K. Peng, K. Zhao, and Z. Song, · 2020
Closest in time.
“High Fidelity Speech Synthesis with Adversarial Networks,”
M. Bińkowski, J. Donahue, S. Dieleman, A. Clark, E. Elsen, N. Casagrande, L. C. Cobo, and K. Simonyan, · 2020
Closest in time.
“A Spectral Energy Distance for Parallel Speech Synthesis,”
A. A. Gritsenko, T. Salimans, R. v. d. Berg, J. Snoek, and N. Kalchbrenner, · 2020
Closest in time.
“Flow-TTS: A non-autoregressive network for text to speech based on flow,”
C. Miao, S. Liang, M. Chen, J. Ma, S. Wang, and J. Xiao, · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
R. Skerry-Ryan, E. Battenberg, Y. Xiao, Y. Wang, D. Stanton, J. Shor, R. J. Weiss, R. Clark, and R. A. Saurous, · 2018
Cited alongside, same era.
“ClariNet: Parallel Wave Generation in End-to-End Text-to-Speech,”
W. Ping, K. Peng, and J. Chen, · 2019
Cited alongside, same era.
“FloWaveNet: A generative flow for raw audio,”
S. Kim, S.-G. Lee, J. Song, J. Kim, and S. Yoon, · 2019
Cited alongside, same era.
“WaveGlow: A Flow-based Generative Network for Speech Synthesis,”
R. Prenger, R. Valle, and B. Catanzaro, · 2019
Cited alongside, same era.
“MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis,”
K. Kumar, R. Kumar, T. de Boissiere, L. Gestin, W. Z. Teoh, J. Sotelo, A. de Brebisson, Y. Bengio, et al., · 2019
Cited alongside, same era.
“Neural source-filter-based waveform model for statistical parametric speech synthesis,”
X. Wang, S. Takaki, and J. Yamagishi, · 2019
Cited alongside, same era.
“FastSpeech: Fast, robust and controllable text to speech,”
Y. Ren, Y. Ruan, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, · 2019
Cited alongside, same era.
“Flowtron: an autoregressive flow-based generative network for text-to-speech synthesis,”
R. Valle, K. Shih, R. Prenger, and B. Catanzaro, · 2020
Closest in time.
“Glow-TTS: A generative flow for text-to-speech via monotonic alignment search,”
J. Kim, S. Kim, J. Kong, and S. Yoon, · 2020
Closest in time.
“Fastspeech 2: Fast and high-quality end-to-end text-to-speech,”
Y. Ren, C. Hu, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, · 2020
Closest in time.
“End-to-end adversarial text-to-speech,”
J. Donahue, S. Dieleman, M. Bińkowski, E. Elsen, and K. Simonyan, · 2020
Closest in time.
“Location-relative attention mechanisms for robust long-form speech synthesis,”
E. Battenberg, R. Skerry-Ryan, S. Mariooryad, D. Stanton, D. Kao, et al., · 2020
Closest in time.
“Generating diverse and natural text-to-speech samples using a quantized fine-grained VAE and auto-regressive prosody prior,”
G. Sun, Y. Zhang, R. J. Weiss, Y. Cao, H. Zen, A. Rosenberg, B. Ramabhadran, and Y. Wu, · 2020
Closest in time.
“WaveGrad: Estimating gradients for waveform generation,”
N. Chen, Y. Zhang, H. Zen, R. J. Weiss, M. Norouzi, and W. Chan, · 2021
Closest in time.