Fetching the paper…
Reading the bibliography…
In this work, we propose WaveFlow, a small-footprint generative flow for raw audio, which is directly trained with maximum likelihood.
Yamamoto, R., Song, E., and Kim, J.-M · 1904
Earlier work this paper cites.
Yamamoto, R., Song, E., and Kim, J.-M · 1910
Earlier work this paper cites.
CrowdMOS: An approach for crowdsourcing mean opinion score studies
Ribeiro, F., Florêncio, D., Zhang, C., and Seltzer, M · 2011
Earlier work this paper cites.
NICE: Non-linear independent components estimation
Dinh, L., Krueger, D., and Bengio, Y · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Unsupervised representation learning with deep convolutional generative adversarial networks
Radford, A., Metz, L., and Chintala, S · 2015
Earlier work this paper cites.
Variational inference with normalizing flows
Rezende, D. J. and Mohamed, S · 2015
Earlier work this paper cites.
Multi-scale context aggregation by dilated convolutions
Yu, F. and Koltun, V · 2015
Earlier work this paper cites.
Improving variational inference with inverse autoregressive flow
Kingma, D. P., Salimans, T., Jozefowicz, R., Chen, X., Sutskever, I., and Welling, M · 2016
Earlier work this paper cites.
Fast wavenet generation algorithm
Paine, T. L., Khorrami, P., Chang, S., Zhang, Y., Ramachandran, P., Hasegawa-Johnson, M. A., and Huang, T. S · 2016
Earlier work this paper cites.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
Salimans, T. and Kingma, D. P · 2016
Earlier work this paper cites.
WaveNet: A generative model for raw audio
van den Oord, A., Dieleman, S., Zen, H., Simonyan, K., Vinyals, O., Graves, A., Kalchbrenner, N., Senior, A., and Kavukcuoglu, K · 2016
Earlier work this paper cites.
Conditional image generation with pixelcnn decoders
Van den Oord, A., Kalchbrenner, N., Espeholt, L., Vinyals, O., Graves, A., et al · 2016
Earlier work this paper cites.
Density estimation using Real NVP
Dinh, L., Sohl-Dickstein, J., and Bengio, S · 2017
Earlier work this paper cites.
The LJ speech dataset
Ito, K · 2017
Earlier work this paper cites.
SampleRNN: An unconditional end-to-end neural audio generation model
Mehri, S., Kumar, K., Gulrajani, I., Kumar, R., Jain, S., Sotelo, J., Courville, A., and Bengio, Y · 2017
Cited alongside, same era.
Masked autoregressive flow for density estimation
Papamakarios, G., Pavlakou, T., and Murray, I · 2017
Cited alongside, same era.
Char2wav: End-to-end speech synthesis
Sotelo, J., Mehri, S., Kumar, K., Santos, J. F., Kastner, K., Courville, A., and Bengio, Y · 2017
Cited alongside, same era.
Tacotron: Towards end-to-end speech synthesis
Wang, Y., Skerry-Ryan, R., Stanton, D., Wu, Y., Weiss, R. J., Jaitly, N., Yang, Z., Xiao, Y., Chen, Z., Bengio, S., Le, Q., Agiomyrgiannakis, Y., Clark, R., and Saurous, R. A · 2017
Cited alongside, same era.
Sylvester normalizing flows for variational inference
Berg, R. v. d., Hasenclever, L., Tomczak, J. M., and Welling, M · 2018
Cited alongside, same era.
VoiceLoop: Voice fitting and synthesis via a phonological loop
Taigman, Y., Wolf, L., Polyak, A., and Nachmani, E · 2018
Later among the works it cites.
Parallel WaveNet: Fast high-fidelity speech synthesis
van den Oord, A., Li, Y., Babuschkin, I., Simonyan, K., Vinyals, O., Kavukcuoglu, K., Driessche, G. v. d., Lockhart, E., Cobo, L. C., Stimberg, F., et al · 2018
Later among the works it cites.
High fidelity speech synthesis with adversarial networks
Bińkowski, M., Donahue, J., Dieleman, S., Clark, A., Elsen, E., Casagrande, N., Cobo, L. C., and Simonyan, K · 2019
Closest in time.
Ho, J., Chen, X., Srinivas, A., Duan, Y., and Abbeel, P · 2019
Closest in time.
Emerging convolutions for generative normalizing flows
Hoogeboom, E., Berg, R. v. d., and Welling, M · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Brock, A., Donahue, J., and Simonyan, K · 2018
Cited alongside, same era.
The challenge of realistic music generation: modelling raw audio at scale
Dieleman, S., van den Oord, A., and Simonyan, K · 2018
Cited alongside, same era.
Donahue, C., McAuley, J., and Puckette, M · 2018
Cited alongside, same era.
Huang, C.-W., Krueger, D., Lacoste, A., and Courville, A · 2018
Cited alongside, same era.
Efficient neural audio synthesis
Kalchbrenner, N., Elsen, E., Simonyan, K., Noury, S., Casagrande, N., Lockhart, E., Stimberg, F., Oord, A. v. d., Dieleman, S., and Kavukcuoglu, K · 2018
Cited alongside, same era.
Glow: Generative flow with invertible 1x1 convolutions
Kingma, D. P. and Dhariwal, P · 2018
Cited alongside, same era.
Generating high fidelity images with subscale pixel networks and multidimensional upscaling
Menick, J. and Kalchbrenner, N · 2018
Cited alongside, same era.
Closest in time.
FloWaveNet: A generative flow for raw audio
Kim, S., Lee, S.-g., Song, J., and Yoon, S · 2019
Closest in time.
Melgan: Generative adversarial networks for conditional waveform synthesis
Kumar, K., Kumar, R., de Boissiere, T., Gestin, L., Teoh, W. Z., Sotelo, J., de Brébisson, A., Bengio, Y., and Courville, A. C · 2019
Closest in time.
Neural speech synthesis with transformer network
Li, N., Liu, S., Liu, Y., Zhao, S., Liu, M., and Zhou, M · 2019
Closest in time.
Parallel neural text-to-speech
Peng, K., Ping, W., Song, Z., and Zhao, K · 2019
Closest in time.
ClariNet: Parallel wave generation in end-to-end text-to-speech
Ping, W., Peng, K., and Chen, J · 2019
Closest in time.
WaveGlow: A flow-based generative network for speech synthesis
Prenger, R., Valle, R., and Catanzaro, B · 2019
Closest in time.
Fastspeech: Fast, robust and controllable text to speech
Ren, Y., Ruan, Y., Tan, X., Qin, T., Zhao, S., Zhao, Z., and Liu, T.-Y · 2019
Closest in time.
Blow: a single-scale hyperconditioned flow for non-parallel raw-audio voice conversion
Serrà, J., Pascual, S., and Segura, C · 2019
Closest in time.
Discrete flows: Invertible generative models of discrete data
Tran, D., Vafa, K., Agrawal, K. K., Dinh, L., and Poole, B · 2019
Closest in time.
Neural source-filter-based waveform model for statistical parametric speech synthesis
Wang, X., Takaki, S., and Yamagishi, J · 2019
Closest in time.