KERMIT: Generative Insertion-Based Modeling for Sequences
Original
William Chan, Nikita Kitaev, Kelvin Guu, Mitchell Stern, and Jakob Uszkoreit · 1906
Earlier work this paper cites.
An Empirical Study of Generation Order for Machine Translation
Original
William Chan, Mitchell Stern, Jamie Kiros, and Jakob Uszkoreit · 1910
Earlier work this paper cites.
Mel-Cepstral Distance Measure for Objective Speech Quality Assessment
Robert Kubichek · 1993
Earlier work this paper cites.
Estimation of Non-Normalized Statistical Models by Score Matching
Aapo Hyvärinen · 2005
Earlier work this paper cites.
Multi-band MelGAN: Faster Waveform Generation for High-Quality Text-to-Speech
Original
Geng Yang, Shan Yang, Kai Liu, Peng Fang, Wei Chen, and Lei Xie · 2005
Earlier work this paper cites.
VocGAN: A High-Fidelity Real-time Vocoder with a Hierarchically-nested Adversarial Network
Original
Jinhyeok Yang, Junmo Lee, Youngik Kim, Hoonyoung Cho, and Injung Kim · 2007
Earlier work this paper cites.
Reducing F0 Frame Error of F0 Tracking Algorithms under Noisy Conditions with an Unvoiced/Voiced Classification Frontend
Wei Chu and Abeer Alwan · 2009
Earlier work this paper cites.
A Connection Between Score Matching and Denoising Autoencoders
Pascal Vincent · 2011
Earlier work this paper cites.
Exact Solutions to the Nonlinear Dynamics of Learning in Deep Linear Neural Networks
Andrew M. Saxe, James L. McClelland, and Surya Ganguli · 2014
Earlier work this paper cites.
Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Deep Unsupervised Learning using Nonequilibrium Thermodynamics
Jascha Sohl-Dickstein, Eric A. Weiss, Niru Maheswaranathan, and Surya Ganguli · 2015
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
WaveNet: A Generative Model for Raw Audio
Original
Aäron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
The LJ Speech Dataset, 2017
Keith Ito and Linda Johnson · 2017
Earlier work this paper cites.
SampleRNN: An Unconditional End-to-End Neural Audio Generation Model
Soroush Mehri, Kundan Kumar, Ishaan Gulrajani, Rithesh Kumar, Shubham Jain, Jose Sotelo, Aaron Courville, and Yoshua Bengio · 2017
Earlier work this paper cites.
PixelCNN++: Improving the PixelCNN with Discretized Logistic Mixture Likelihood and Other Modifications
Tim Salimans, Andrej Karpathy, Xi Chen, and Diederik P. Kingma · 2017
Earlier work this paper cites.
Char2Wav: End-to-End Speech Synthesis
Jose Sotelo, Soroush Mehri, Kundan Kumar, Joao Felipe Santos, Kyle Kastner, Aaron C. Courville, and Yoshua Bengio · 2017
Earlier work this paper cites.
Attention Is All You Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Tacotron: Towards End-to-End Speech Synthesis
Yuxuan Wang, RJ Skerry-Ryan, Daisy Stanton, Yonghui Wu, Ron J. Weiss, Navdeep Jaitly, Zongheng Yang, Ying Xiao, Zhifeng Chen, Samy Bengio, Quoc Le, Yannis Agiomyrgiannakis, Rob Clark, and Rif A. Saurous · 2017
Earlier work this paper cites.
Adversarial Audio Synthesis
Original
Chris Donahue, Julian McAuley, and Miller Puckette · 2018
Earlier work this paper cites.
Feature-wise Transformations
Vincent Dumoulin, Ethan Perez, Nathan Schucher, Florian Strub, Harm de Vries, Aaron Courville, and Yoshua Bengio · 2018
Earlier work this paper cites.
Efficient Neural Audio Synthesis
Nal Kalchbrenner, Erich Elsen, Karen Simonyan, Seb Noury, Norman Casagrande, Edward Lockhart, Florian Stimberg, Aäron van den Oord, Sander Dieleman, and Koray Kavukcuoglu · 2018
Earlier work this paper cites.