Fetching the paper…
Reading the bibliography…
Capturing high-level structure in audio waveforms is challenging because a single second of audio spans tens of thousands of timesteps.
Signal estimation from modified short-time fourier transform
Daniel Griffin and Jae Lim · 1984
Earlier work this paper cites.
Mixture density networks
Christopher M Bishop · 1994
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Multi-dimensional recurrent neural networks
Alex Graves, Santiago Fernández, and Jürgen Schmidhuber · 2007
Earlier work this paper cites.
Offline handwriting recognition with multidimensional recurrent neural networks
Alex Graves and Jürgen Schmidhuber · 2009
Earlier work this paper cites.
The blizzard challenge 2011, 2011
Simon King · 2011
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman and Geoffrey Hinton · 2012
Earlier work this paper cites.
Generating sequences with recurrent neural networks
Alex Graves · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2013
Earlier work this paper cites.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Inversion of auditory spectrograms, traditional spectrograms, and other envelope representations
Rémi Decorsière, Peter L Søndergaard, Ewen N MacDonald, and Torsten Dau · 2015
Earlier work this paper cites.
Nal Kalchbrenner, Ivo Danihelka, and Alex Graves · 2015
Earlier work this paper cites.
Generative image modeling using spatial lstms
Lucas Theis and Matthias Bethge · 2015
Earlier work this paper cites.
Renet: A recurrent neural network based alternative to convolutional networks
Francesco Visin, Kyle Kastner, Kyunghyun Cho, Matteo Matteucci, Aaron Courville, and Yoshua Bengio · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Exploring multidimensional lstms for large vocabulary asr
Jinyu Li, Abdelrahman Mohamed, Geoffrey Zweig, and Yifan Gong · 2016
Cited alongside, same era.
Samplernn: An unconditional end-to-end neural audio generation model
Soroush Mehri, Kundan Kumar, Ishaan Gulrajani, Rithesh Kumar, Shubham Jain, Jose Sotelo, Aaron Courville, and Yoshua Bengio · 2016
Cited alongside, same era.
Modeling time-frequency patterns with lstm vs. convolutional architectures for lvcsr tasks
Tara N Sainath and Bo Li · 2016
Cited alongside, same era.
Deep voice: Real-time neural text-to-speech
Sercan O Arik, Mike Chrzanowski, Adam Coates, Gregory Diamos, Andrew Gibiansky, Yongguo Kang, Xian Li, John Miller, Andrew Ng, Jonathan Raiman, et al · 2017
Cited alongside, same era.
Gansynth: Adversarial neural audio synthesis
Jesse Engel, Kumar Krishna Agrawal, Shuo Chen, Ishaan Gulrajani, Chris Donahue, and Adam Roberts · 2018
Later among the works it cites.
Enabling factorized piano music modeling and generation with the maestro dataset
Curtis Hawthorne, Andriy Stasyuk, Adam Roberts, Ian Simon, Cheng-Zhi Anna Huang, Sander Dieleman, Erich Elsen, Jesse Engel, and Douglas Eck · 2018
Later among the works it cites.
Ted-lium 3: twice as much data and corpus repartition for experiments on speaker adaptation
François Hernandez, Vincent Nguyen, Sahar Ghannay, Natalia Tomashenko, and Yannick Esteve · 2018
Later among the works it cites.
Hierarchical generative modeling for controllable speech synthesis
Wei-Ning Hsu, Yu Zhang, Ron J Weiss, Heiga Zen, Yonghui Wu, Yuxuan Wang, Yuan Cao, Ye Jia, Zhifeng Chen, Jonathan Shen, et al · 2018
Later among the works it cites.
Music transformer: Generating music with long-term structure
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xi Chen, Nikhil Mishra, Mostafa Rohaninejad, and Pieter Abbeel · 2017
Cited alongside, same era.
Pixel recursive super resolution
Ryan Dahl, Mohammad Norouzi, and Jonathon Shlens · 2017
Cited alongside, same era.
Progressive growing of gans for improved quality, stability, and variation
Tero Karras, Timo Aila, Samuli Laine, and Jaakko Lehtinen · 2017
Cited alongside, same era.
Parallel multiscale autoregressive density estimation
Scott Reed, Aäron van den Oord, Nal Kalchbrenner, Sergio Gómez Colmenarejo, Ziyu Wang, Dan Belov, and Nando de Freitas · 2017
Cited alongside, same era.
Tim Salimans, Andrej Karpathy, Xi Chen, and Diederik P Kingma · 2017
Cited alongside, same era.
Char2wav: End-to-end speech synthesis
Jose Sotelo, Soroush Mehri, Kundan Kumar, Joao Felipe Santos, Kyle Kastner, Aaron Courville, and Yoshua Bengio · 2017
Cited alongside, same era.
Tacotron: Towards end-to-end speech synthesis
Yuxuan Wang, RJ Skerry-Ryan, Daisy Stanton, Yonghui Wu, Ron J Weiss, Navdeep Jaitly, Zongheng Yang, Ying Xiao, Zhifeng Chen, Samy Bengio, et al · 2017
Cited alongside, same era.
Expressive speech synthesis via modeling expressions with variational autoencoder
Kei Akuzawa, Yusuke Iwasawa, and Yutaka Matsuo · 2018
Cited alongside, same era.
Cheng-Zhi Anna Huang, Ashish Vaswani, Jakob Uszkoreit, Noam Shazeer, Curtis Hawthorne, Andrew M Dai, Matthew D Hoffman, and Douglas Eck · 2018
Later among the works it cites.
Efficient neural audio synthesis
Nal Kalchbrenner, Erich Elsen, Karen Simonyan, Seb Noury, Norman Casagrande, Edward Lockhart, Florian Stimberg, Aaron van den Oord, Sander Dieleman, and Koray Kavukcuoglu · 2018
Later among the works it cites.
Glow: Generative flow with invertible 1x1 convolutions
Durk P Kingma and Prafulla Dhariwal · 2018
Later among the works it cites.
Conditioning deep generative raw audio models for structured automatic music
Rachel Manzelli, Vijay Thakkar, Ali Siahkamari, and Brian Kulis · 2018
Later among the works it cites.
Generating high fidelity images with subscale pixel networks and multidimensional upscaling
Jacob Menick and Nal Kalchbrenner · 2018
Later among the works it cites.
Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Łukasz Kaiser, Noam Shazeer, and Alexander Ku · 2018
Later among the works it cites.
Clarinet: Parallel wave generation in end-to-end text-to-speech
Wei Ping, Kainan Peng, and Jitong Chen · 2018
Later among the works it cites.
Natural tts synthesis by conditioning wavenet on mel spectrogram predictions
Jonathan Shen, Ruoming Pang, Ron J Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, Rj Skerrv-Ryan, et al · 2018
Later among the works it cites.
Voiceloop: Voice fitting and synthesis via a phonological loop
Yaniv Taigman, Lior Wolf, Adam Polyak, and Eliya Nachmani · 2018
Later among the works it cites.
Fast spectrogram inversion using multi-head convolutional neural networks
Sercan Ö Arık, Heewoo Jun, and Gregory Diamos · 2019
Closest in time.
Generating long sequences with sparse transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever · 2019
Closest in time.
Waveglow: A flow-based generative network for speech synthesis
Ryan Prenger, Rafael Valle, and Bryan Catanzaro · 2019
Closest in time.