Fetching the paper…
Reading the bibliography…
Conditional waveform synthesis models learn a distribution of audio waveforms given conditioning such as text, mel-spectrograms, or MIDI.
Tentative standards for sound level meters
RG McCurdy · 1936
Earlier work this paper cites.
Acoustic theory of speech production
Gunnar Fant · 1970
Earlier work this paper cites.
Digital processing of speech signals
Lawrence R Rabiner and Ronald W Schafer · 1978
Earlier work this paper cites.
Signal estimation from modified short-time fourier transform
Daniel Griffin and Jae Lim · 1984
Earlier work this paper cites.
Yin, a fundamental frequency estimator for speech and music
Alain De Cheveigné and Hideki Kawahara · 2002
Earlier work this paper cites.
Estimation of non-normalized statistical models by score matching
Aapo Hyvärinen and Peter Dayan · 2005
Earlier work this paper cites.
Mixing secrets for the small studio
Mike Senior · 2011
Earlier work this paper cites.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
Kyunghyun Cho, Bart Van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Earlier work this paper cites.
Nice: Non-linear independent components estimation
Laurent Dinh, David Krueger, and Yoshua Bengio · 2014
Earlier work this paper cites.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Can we automatically transform speech recorded on common consumer devices in real-world environments into professional production quality speech?—a dataset, insights, and challenges
Gautham J Mysore · 2014
Earlier work this paper cites.
Deep speech 2: End-to-end speech recognition in english and mandarin
Dario Amodei, Sundaram Ananthanarayanan, Rishita Anubhai, Jingliang Bai, Eric Battenberg, Carl Case, Jared Casper, Bryan Catanzaro, Qiang Cheng, Guoliang Chen, et al · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Autoencoding beyond pixels using a learned similarity metric
Anders Boesen Lindbo Larsen, Søren Kaae Sønderby, Hugo Larochelle, and Ole Winther · 2016
Earlier work this paper cites.
Samplernn: An unconditional end-to-end neural audio generation model
Soroush Mehri, Kundan Kumar, Ishaan Gulrajani, Rithesh Kumar, Shubham Jain, Jose Sotelo, Aaron Courville, and Yoshua Bengio · 2016
Earlier work this paper cites.
World: a vocoder-based high-quality speech synthesis system for real-time applications
Masanori Morise, Fumiya Yokomori, and Kenji Ozawa · 2016
Earlier work this paper cites.
Wavenet: A generative model for raw audio
Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
Tim Salimans and Durk P Kingma · 2016
Cited alongside, same era.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter · 2017
Cited alongside, same era.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Cited alongside, same era.
Block-sparse recurrent neural networks
Sharan Narang, Eric Undersander, and Gregory Diamos · 2017
Cited alongside, same era.
Enabling factorized piano music modeling and generation with the maestro dataset
Curtis Hawthorne, Andriy Stasyuk, Adam Roberts, Ian Simon, Cheng-Zhi Anna Huang, Sander Dieleman, Erich Elsen, Jesse Engel, and Douglas Eck · 2018
Neural source-filter-based waveform model for statistical parametric speech synthesis
Xin Wang, Shinji Takaki, and Junichi Yamagishi · 2019
Later among the works it cites.
Cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit (version 0.92)
Junichi Yamagishi, Christophe Veaux, Kirsten MacDonald, et al · 2019
Later among the works it cites.
Wavegrad: Estimating gradients for waveform generation
Nanxin Chen, Yu Zhang, Heiga Zen, Ron J Weiss, Mohammad Norouzi, and William Chan · 2020
Later among the works it cites.
Jukebox: A generative model for music
Prafulla Dhariwal, Heewoo Jun, Christine Payne, Jong Wook Kim, Alec Radford, and Ilya Sutskever · 2020
Later among the works it cites.
End-to-end adversarial text-to-speech
Jeff Donahue, Sander Dieleman, Mikołaj Bińkowski, Erich Elsen, and Karen Simonyan · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Fftnet: A real-time speaker-dependent neural vocoder
Zeyu Jin, Adam Finkelstein, Gautham J Mysore, and Jingwan Lu · 2018
Cited alongside, same era.
Efficient neural audio synthesis
Nal Kalchbrenner, Erich Elsen, Karen Simonyan, Seb Noury, Norman Casagrande, Edward Lockhart, Florian Stimberg, Aaron Oord, Sander Dieleman, and Koray Kavukcuoglu · 2018
Cited alongside, same era.
Crepe: A convolutional representation for pitch estimation
Jong Wook Kim, Justin Salamon, Peter Li, and Juan Pablo Bello · 2018
Cited alongside, same era.
Glow: Generative flow with invertible 1x1 convolutions
Diederik P Kingma and Prafulla Dhariwal · 2018
Cited alongside, same era.
Spectral normalization for generative adversarial networks
Takeru Miyato, Toshiki Kataoka, Masanori Koyama, and Yuichi Yoshida · 2018
Cited alongside, same era.
High fidelity speech synthesis with adversarial networks
Mikołaj Bińkowski, Jeff Donahue, Sander Dieleman, Aidan Clark, Erich Elsen, Norman Casagrande, Luis C Cobo, and Karen Simonyan · 2019
Cited alongside, same era.
Generating long sequences with sparse transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever · 2019
Cited alongside, same era.
Later among the works it cites.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Later among the works it cites.
Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae · 2020
Later among the works it cites.
torchcrepe, 6 2020
Max Morrison · 2020
Later among the works it cites.
Controllable neural prosody synthesis
Max Morrison, Zeyu Jin, Justin Salamon, Nicholas J Bryan, and Gautham J Mysore · 2020
Later among the works it cites.
Waveflow: A compact flow-based model for raw audio
Wei Ping, Kainan Peng, Kexin Zhao, and Zhao Song · 2020
Later among the works it cites.
Unsupervised speech decomposition via triple information bottleneck
Kaizhi Qian, Yang Zhang, Shiyu Chang, Mark Hasegawa-Johnson, and David Cox · 2020
Later among the works it cites.
You only need adversarial supervision for semantic image synthesis
Edgar Schönfeld, Vadim Sushko, Dan Zhang, Juergen Gall, Bernt Schiele, and Anna Khoreva · 2020
Later among the works it cites.
To prune, or not to prune: exploring the efficacy of pruning for model compression
Michael Zhu and Suyog Gupta · 2020
Later among the works it cites.
Sam Bond-Taylor, Adam Leach, Yang Long, and Chris G Willcocks · 2021
Closest in time.
Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech
Jaehyeon Kim, Jungil Kong, and Juhee Son · 2021
Closest in time.
Segmented DAPS (Device and Produced Speech) Dataset, May 2021
Max Morrison, Zeyu Jin, Nicholas J. Bryan, Juan-Pablo Caceres, and Bryan Pardo · 2021
Closest in time.
Wave-tacotron: Spectrogram-free end-to-end text-to-speech synthesis
Ron J Weiss, RJ Skerry-Ryan, Eric Battenberg, Soroosh Mariooryad, and Diederik P Kingma · 2021
Closest in time.