Fetching the paper…
Reading the bibliography…
This paper proposes VARA-TTS, a non-autoregressive (non-AR) text-to-speech (TTS) model using a very deep Variational Autoencoder (VDVAE) with Residual Attention mechanism, which refines the textual-to-acoustic alignment layer-wisely.
Deep voice 3: 2000-speaker neural text-to-speech
Ping, W., Peng, K., Gibiansky, A., Arik, S. O., Kannan, A., Narang, S., Raiman, J., and Miller, J · 2000
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M · 2013
Earlier work this paper cites.
Stochastic backpropagation and approximate inference in deep generative models
Rezende, D. J., Mohamed, S., and Wierstra, D · 2014
Earlier work this paper cites.
Attention-based models for speech recognition
Chorowski, J. K., Bahdanau, D., Serdyuk, D., Cho, K., and Bengio, Y · 2015
Earlier work this paper cites.
A recurrent latent variable model for sequential data
Chung, J., Kastner, K., Dinh, L., Goel, K., Courville, A. C., and Bengio, Y · 2015
Earlier work this paper cites.
Infogan: Interpretable representation learning by information maximizing generative adversarial nets
Chen, X., Duan, Y., Houthooft, R., Schulman, J., Sutskever, I., and Abbeel, P · 2016
Earlier work this paper cites.
Gaussian error linear units (gelus)
Hendrycks, D. and Gimpel, K · 2016
Earlier work this paper cites.
Wavenet: A generative model for raw audio
Oord, A. v. d., Dieleman, S., Zen, H., Simonyan, K., Vinyals, O., Graves, A., Kalchbrenner, N., Senior, A., and Kavukcuoglu, K · 2016
Earlier work this paper cites.
Ladder variational autoencoders
Sønderby, C. K., Raiko, T., Maaløe, L., Sønderby, S. K., and Winther, O · 2016
Earlier work this paper cites.
Conditional image generation with pixelcnn decoders
Van den Oord, A., Kalchbrenner, N., Espeholt, L., Vinyals, O., Graves, A., et al · 2016
Earlier work this paper cites.
Deep voice: Real-time neural text-to-speech
Arik, S. O., Chrzanowski, M., Coates, A., Diamos, G., Gibiansky, A., Kang, Y., Li, X., Miller, J., Ng, A., Raiman, J., et al · 2017
Earlier work this paper cites.
Deep voice 2: Multi-speaker neural text-to-speech
Gibiansky, A., Arik, S., Diamos, G., Miller, J., Peng, K., Ping, W., Raiman, J., and Zhou, Y · 2017
Earlier work this paper cites.
The lj speech dataset
Ito, K. and Johnson, L · 2017
Earlier work this paper cites.
Neural discrete representation learning
Van Den Oord, A., Vinyals, O., et al · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Earlier work this paper cites.
Tacotron: Towards end-to-end speech synthesis
Wang, Y., Skerry-Ryan, R., Stanton, D., Wu, Y., Weiss, R. J., Jaitly, N., Yang, Z., Xiao, Y., Chen, Z., Bengio, S., et al · 2017
Cited alongside, same era.
Fixing a broken ELBO
Alemi, A. A., Poole, B., Fischer, I., Dillon, J. V., Saurous, R. A., and Murphy, K · 2018
Cited alongside, same era.
Understanding disentangling in β \beta -vae
Burgess, C. P., Higgins, I., Pal, A., Matthey, L., Watters, N., Desjardins, G., and Lerchner, A · 2018
Cited alongside, same era.
Fftnet: A real-time speaker-dependent neural vocoder
Jin, Z., Finkelstein, A., Mysore, G. J., and Lu, J · 2018
Cited alongside, same era.
Efficient neural audio synthesis
Kalchbrenner, N., Elsen, E., Simonyan, K., Noury, S., Casagrande, N., Lockhart, E., Stimberg, F., Oord, A. v. d., Dieleman, S., and Kavukcuoglu, K · 2018
Cited alongside, same era.
Libritts: A corpus derived from librispeech for text-to-speech
Zen, H., Dang, V., Clark, R., Zhang, Y., Weiss, R. J., Jia, Y., Chen, Z., and Wu, Y · 2019
Later among the works it cites.
Location-relative attention mechanisms for robust long-form speech synthesis
Battenberg, E., Skerry-Ryan, R., Mariooryad, S., Stanton, D., Kao, D., Shannon, M., and Bagby, T · 2020
Later among the works it cites.
Very deep vaes generalize autoregressive models and can outperform them on images
Child, R · 2020
Later among the works it cites.
End-to-end adversarial text-to-speech
Donahue, J., Dieleman, S., Bińkowski, M., Elsen, E., and Simonyan, K · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Parallel wavenet: Fast high-fidelity speech synthesis
Oord, A., Li, Y., Babuschkin, I., Simonyan, K., Vinyals, O., Kavukcuoglu, K., Driessche, G., Lockhart, E., Cobo, L., Stimberg, F., et al · 2018
Cited alongside, same era.
Film: Visual reasoning with a general conditioning layer
Perez, E., Strub, F., De Vries, H., Dumoulin, V., and Courville, A · 2018
Cited alongside, same era.
Rezende, D. J. and Viola, F · 2018
Cited alongside, same era.
Natural tts synthesis by conditioning wavenet on mel spectrogram predictions
Shen, J., Pang, R., Weiss, R. J., Schuster, M., Jaitly, N., Yang, Z., Chen, Z., Zhang, Y., Wang, Y., Skerrv-Ryan, R., et al · 2018
Cited alongside, same era.
Efficiently trainable text-to-speech system based on deep convolutional networks with guided attention
Tachibana, H., Uenoyama, K., and Aihara, S · 2018
Cited alongside, same era.
Forward attention in sequence-to-sequence acoustic modeling for speech synthesis
Zhang, J.-X., Ling, Z.-H., and Dai, L.-R · 2018
Cited alongside, same era.
Robust sequence-to-sequence acoustic modeling with stepwise monotonic attention for neural tts
He, M., Deng, Y., and He, L · 2019
Cited alongside, same era.
He, R., Ravula, A., Kanagal, B., and Ainslie, J · 2020
Later among the works it cites.
Glow-tts: A generative flow for text-to-speech via monotonic alignment search
Kim, J., Kim, S., Kong, J., and Yoon, S · 2020
Later among the works it cites.
Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis
Kong, J., Kim, J., and Bae, J · 2020
Later among the works it cites.
Flow-tts: A non-autoregressive network for text to speech based on flow
Miao, C., Liang, S., Chen, M., Ma, J., Wang, S., and Xiao, J · 2020
Later among the works it cites.
Non-autoregressive neural text-to-speech
Peng, K., Ping, W., Song, Z., and Zhao, K · 2020
Later among the works it cites.
Fastspeech 2: Fast and high-quality end-to-end text-to-speech
Ren, Y., Hu, C., Qin, T., Zhao, S., Zhao, Z., and Liu, T.-Y · 2020
Later among the works it cites.
Flowtron: an autoregressive flow-based generative network for text-to-speech synthesis
Valle, R., Shih, K., Prenger, R., and Catanzaro, B · 2020
Later among the works it cites.
Wave-tacotron: Spectrogram-free end-to-end text-to-speech synthesis
Weiss, R. J., Skerry-Ryan, R., Battenberg, E., Mariooryad, S., and Kingma, D. P · 2020
Later among the works it cites.
Parallel wavegan: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram
Yamamoto, R., Song, E., and Kim, J.-M · 2020
Later among the works it cites.
Bidirectional variational inference for non-autoregressive text-to-speech
Lee, Y., Shin, J., and Jung, K · 2021
Closest in time.