Fetching the paper…
Reading the bibliography…
In this work, we propose ParaNet, a non-autoregressive seq2seq model that converts text to spectrogram.
Text-to-Speech Synthesis
Taylor, P · 2009
Earlier work this paper cites.
CrowdMOS: An approach for crowdsourcing mean opinion score studies
Ribeiro, F., Florêncio, D., Zhang, C., and Seltzer, M · 2011
Earlier work this paper cites.
Auto-encoding variational Bayes
Kingma, D. P. and Welling, M · 2014
Earlier work this paper cites.
Stochastic backpropagation and approximate inference in deep generative models
Rezende, D. J., Mohamed, S., and Wierstra, D · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y · 2015
Earlier work this paper cites.
Scheduled sampling for sequence prediction with recurrent neural networks
Bengio, S., Vinyals, O., Jaitly, N., and Shazeer, N · 2015
Earlier work this paper cites.
Attention-based models for speech recognition
Chorowski, J. K., Bahdanau, D., Serdyuk, D., Cho, K., and Bengio, Y · 2015
Earlier work this paper cites.
A recurrent latent variable model for sequential data
Chung, J., Kastner, K., Dinh, L., Goel, K., Courville, A. C., and Bengio, Y · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Variational inference with normalizing flows
Rezende, D. J. and Mohamed, S · 2015
Earlier work this paper cites.
Generating sentences from a continuous space
Bowman, S. R., Vilnis, L., Vinyals, O., Dai, A., Jozefowicz, R., and Bengio, S · 2016
Earlier work this paper cites.
Improving variational inference with inverse autoregressive flow
Kingma, D. P., Salimans, T., Jozefowicz, R., Chen, X., Sutskever, I., and Welling, M · 2016
Earlier work this paper cites.
WaveNet: A generative model for raw audio
van den Oord, A., Dieleman, S., Zen, H., Simonyan, K., Vinyals, O., Graves, A., Kalchbrenner, N., Senior, A., and Kavukcuoglu, K · 2016
Earlier work this paper cites.
Density estimation using real NVP
Dinh, L., Sohl-Dickstein, J., and Bengio, S · 2017
Earlier work this paper cites.
Char2wav: End-to-end speech synthesis
Sotelo, J., Mehri, S., Kumar, K., Santos, J. F., Kastner, K., Courville, A., and Bengio, Y · 2017
Cited alongside, same era.
Neural discrete representation learning
van den Oord, A., Vinyals, O., and Kavukcuoglu, K · 2017
Cited alongside, same era.
Tacotron: Towards end-to-end speech synthesis
Wang, Y., Skerry-Ryan, R., Stanton, D., Wu, Y., Weiss, R. J., Jaitly, N., Yang, Z., Xiao, Y., Chen, Z., Bengio, S., Le, Q., Agiomyrgiannakis, Y., Clark, R., and Saurous, R. A · 2017
Cited alongside, same era.
Neural voice cloning with a few samples
Arik, S. O., Chen, J., Peng, K., Ping, W., and Zhou, Y · 2018
Cited alongside, same era.
Non-autoregressive neural machine translation
Gu, J., Bradbury, J., Xiong, C., Li, V. O., and Socher, R · 2018
Cited alongside, same era.
Transfer learning from speaker verification to multispeaker text-to-speech synthesis
Jia, Y., Zhang, Y., Weiss, R. J., Wang, Q., Shen, J., Ren, F., Chen, Z., Nguyen, P., Pang, R., and Moreno, I. L · 2018
Fast spectrogram inversion using multi-head convolutional neural networks
Arık, S. Ö., Jun, H., and Diamos, G · 2019
Closest in time.
Sample efficient adaptive text-to-speech
Chen, Y., Assael, Y., Shillingford, B., Budden, D., Reed, S., Zen, H., Wang, Q., Cobo, L. C., Trask, A., Laurie, B., et al · 2019
Closest in time.
Hierarchical generative modeling for controllable speech synthesis
Hsu, W.-N., Zhang, Y., Weiss, R. J., Zen, H., Wu, Y., Wang, Y., Cao, Y., Jia, Y., Chen, Z., Shen, J., et al · 2019
Closest in time.
Probability distillation: A caveat and alternatives
Huang, C.-W., Ahmed, F., Kumar, K., Lacoste, A., and Courville, A · 2019
Closest in time.
FloWaveNet: A generative flow for raw audio
Kim, S., Lee, S.-g., Song, J., and Yoon, S · 2019
Closest in time.
Melgan: Generative adversarial networks for conditional waveform synthesis
Kumar, K., Kumar, R., de Boissiere, T., Gestin, L., Teoh, W. Z., Sotelo, J., de Brébisson, A., Bengio, Y., and Courville, A. C · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Fast decoding in sequence models using discrete latent variables
Kaiser, Ł., Roy, A., Vaswani, A., Pamar, N., Bengio, S., Uszkoreit, J., and Shazeer, N · 2018
Cited alongside, same era.
Efficient neural audio synthesis
Kalchbrenner, N., Elsen, E., Simonyan, K., Noury, S., Casagrande, N., Lockhart, E., Stimberg, F., Oord, A. v. d., Dieleman, S., and Kavukcuoglu, K · 2018
Cited alongside, same era.
Glow: Generative flow with invertible 1x1 convolutions
Kingma, D. P. and Dhariwal, P · 2018
Cited alongside, same era.
Deterministic non-autoregressive neural sequence modeling by iterative refinement
Lee, J., Mansimov, E., and Cho, K · 2018
Cited alongside, same era.
Fitting new speakers based on a short untranscribed sample
Nachmani, E., Polyak, A., Taigman, Y., and Wolf, L · 2018
Cited alongside, same era.
Natural TTS synthesis by conditioning WaveNet on mel spectrogram predictions
Shen, J., Pang, R., Weiss, R. J., Schuster, M., Jaitly, N., Yang, Z., Chen, Z., Zhang, Y., Wang, Y., Skerry-Ryan, R., et al · 2018
Cited alongside, same era.
Closest in time.
Neural speech synthesis with transformer network
Li, N., Liu, S., Liu, Y., Zhao, S., Liu, M., and Zhou, M · 2019
Closest in time.
WaveGlow: A flow-based generative network for speech synthesis
Prenger, R., Valle, R., and Catanzaro, B · 2019
Closest in time.
Fastspeech: Fast, robust and controllable text to speech
Ren, Y., Ruan, Y., Tan, X., Qin, T., Zhao, S., Zhao, Z., and Liu, T.-Y · 2019
Closest in time.
Neural source-filter-based waveform model for statistical parametric speech synthesis
Wang, X., Takaki, S., and Yamagishi, J · 2019
Closest in time.
Yamamoto, R., Song, E., and Kim, J.-M · 2019
Closest in time.
High fidelity speech synthesis with adversarial networks
Bińkowski, M., Donahue, J., Dieleman, S., Clark, A., Elsen, E., Casagrande, N., Cobo, L. C., and Simonyan, K · 2020
Closest in time.
WaveFlow: A compact flow-based model for raw audio
Ping, W., Peng, K., Zhao, K., and Song, Z · 2020
Closest in time.