Fetching the paper…
Reading the bibliography…
In this work, we propose DiffWave, a versatile diffusion probabilistic model for conditional and unconditional waveform generation.
CrowdMOS: An approach for crowdsourcing mean opinion score studies
Flávio Ribeiro, Dinei Florêncio, Cha Zhang, and Michael Seltzer · 2011
Earlier work this paper cites.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Auto-encoding variational Bayes
Diederik P Kingma and Max Welling · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Jascha Sohl-Dickstein, Eric A Weiss, Niru Maheswaranathan, and Surya Ganguli · 2015
Earlier work this paper cites.
Persistent rnns: Stashing recurrent weights on-chip
Greg Diamos, Shubho Sengupta, Bryan Catanzaro, Mike Chrzanowski, Adam Coates, Erich Elsen, Jesse Engel, Awni Hannun, and Sanjeev Satheesh · 2016
Earlier work this paper cites.
Improved techniques for training gans
Tim Salimans, Ian Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen · 2016
Earlier work this paper cites.
WaveNet: A generative model for raw audio
Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
Variational walkback: Learning a transition operator as a stochastic recurrent net
Anirudh Goyal Alias Parth Goyal, Nan Rosemary Ke, Surya Ganguli, and Yoshua Bengio · 2017
Earlier work this paper cites.
Deligan: Generative adversarial networks for diverse and limited data
Swaminathan Gurumurthy, Ravi Kiran Sarvadevabhatla, and R Venkatesh Babu · 2017
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter · 2017
Earlier work this paper cites.
The LJ speech dataset
Keith Ito · 2017
Earlier work this paper cites.
SampleRNN: An unconditional end-to-end neural audio generation model
Soroush Mehri, Kundan Kumar, Ishaan Gulrajani, Rithesh Kumar, Shubham Jain, Jose Sotelo, Aaron Courville, and Yoshua Bengio · 2017
Earlier work this paper cites.
Char2wav: End-to-end speech synthesis
Jose Sotelo, Soroush Mehri, Kundan Kumar, Joao Felipe Santos, Kyle Kastner, Aaron Courville, and Yoshua Bengio · 2017
Earlier work this paper cites.
Neural discrete representation learning
Aaron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Tacotron: Towards end-to-end speech synthesis
Yuxuan Wang, RJ Skerry-Ryan, Daisy Stanton, Yonghui Wu, Ron J Weiss, Navdeep Jaitly, Zongheng Yang, Ying Xiao, Zhifeng Chen, Samy Bengio, Quoc Le, Yannis Agiomyrgiannakis, Rob Clark, and Rif A. Saurous · 2017
Earlier work this paper cites.
Aggregated residual transformations for deep neural networks
Saining Xie, Ross Girshick, Piotr Dollár, Zhuowen Tu, and Kaiming He · 2017
Earlier work this paper cites.
Convolutional neural networks for Google speech commands data set with PyTorch , 2017
Yuan Xu and Erdene-Ochir Tuguldur · 2017
Cited alongside, same era.
Activation maximization generative adversarial nets
Zhiming Zhou, Han Cai, Shu Rong, Yuxuan Song, Kan Ren, Weinan Zhang, Yong Yu, and Jun Wang · 2017
Cited alongside, same era.
Large scale GAN training for high fidelity natural image synthesis
Andrew Brock, Jeff Donahue, and Karen Simonyan · 2018
Cited alongside, same era.
Efficient neural audio synthesis
Nal Kalchbrenner, Erich Elsen, Karen Simonyan, Seb Noury, Norman Casagrande, Edward Lockhart, Florian Stimberg, Aaron van den Oord, Sander Dieleman, and Koray Kavukcuoglu · 2018
Cited alongside, same era.
Chae Young Lee, Anoop Toffy, Gue Jun Jung, and Woo-Jin Han · 2018
WaveGlow: A flow-based generative network for speech synthesis
Ryan Prenger, Rafael Valle, and Bryan Catanzaro · 2019
Later among the works it cites.
Fastspeech: Fast, robust and controllable text to speech
Yi Ren, Yangjun Ruan, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu · 2019
Later among the works it cites.
Generative modeling by estimating gradients of the data distribution
Yang Song and Stefano Ermon · 2019
Later among the works it cites.
Melnet: A generative model for audio in the frequency domain
Sean Vasquez and Mike Lewis · 2019
Later among the works it cites.
Neural source-filter-based waveform model for statistical parametric speech synthesis
Xin Wang, Shinji Takaki, and Junichi Yamagishi · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
End-to-end music source separation: is it possible in the waveform domain?
Francesc Lluís, Jordi Pons, and Xavier Serra · 2018
Cited alongside, same era.
Deep Voice 3: Scaling text-to-speech with convolutional sequence learning
Wei Ping, Kainan Peng, Andrew Gibiansky, Sercan O Arik, Ajay Kannan, Sharan Narang, Jonathan Raiman, and John Miller · 2018
Cited alongside, same era.
A wavenet for speech denoising
Dario Rethage, Jordi Pons, and Xavier Serra · 2018
Cited alongside, same era.
On gans and gmms
Eitan Richardson and Yair Weiss · 2018
Cited alongside, same era.
Natural TTS synthesis by conditioning WaveNet on mel spectrogram predictions
Jonathan Shen, Ruoming Pang, Ron J Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, RJ Skerry-Ryan, et al · 2018
Cited alongside, same era.
VoiceLoop: Voice fitting and synthesis via a phonological loop
Yaniv Taigman, Lior Wolf, Adam Polyak, and Eliya Nachmani · 2018
Cited alongside, same era.
Parallel WaveNet: Fast high-fidelity speech synthesis
Aaron van den Oord, Yazhe Li, Igor Babuschkin, Karen Simonyan, Oriol Vinyals, Koray Kavukcuoglu, George van den Driessche, Edward Lockhart, Luis C Cobo, Florian Stimberg, et al · 2018
Cited alongside, same era.
A neural vocoder with hierarchical generation of amplitude and phase spectra for statistical parametric speech synthesis
Yang Ai and Zhen-Hua Ling · 2020
Closest in time.
High fidelity speech synthesis with adversarial networks
Mikołaj Bińkowski, Jeff Donahue, Sander Dieleman, Aidan Clark, Erich Elsen, Norman Casagrande, Luis C Cobo, and Karen Simonyan · 2020
Closest in time.
WaveGrad: Estimating gradients for waveform generation
Nanxin Chen, Yu Zhang, Heiga Zen, Ron J Weiss, Mohammad Norouzi, and William Chan · 2020
Closest in time.
End-to-end adversarial text-to-speech
Jeff Donahue, Sander Dieleman, Mikołaj Bińkowski, Erich Elsen, and Karen Simonyan · 2020
Closest in time.
Ddsp: Differentiable digital signal processing
Jesse Engel, Lamtharn Hantrakul, Chenjie Gu, and Adam Roberts · 2020
Closest in time.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Closest in time.
Conditional spoken digit generation with stylegan
Kasperi Palkama, Lauri Juvela, and Alexander Ilin · 2020
Closest in time.
Non-autoregressive neural text-to-speech
Kainan Peng, Wei Ping, Zhao Song, and Kexin Zhao · 2020
Closest in time.
WaveFlow: A compact flow-based model for raw audio
Wei Ping, Kainan Peng, Kexin Zhao, and Zhao Song · 2020
Closest in time.
Improved techniques for training score-based generative models
Yang Song and Stefano Ermon · 2020
Closest in time.
Flowtron: an autoregressive flow-based generative network for text-to-speech synthesis
Rafael Valle, Kevin Shih, Ryan Prenger, and Bryan Catanzaro · 2020
Closest in time.
Parallel wavegan: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram
Ryuichi Yamamoto, Eunwoo Song, and Jae-Min Kim · 2020
Closest in time.