Fetching the paper…
Reading the bibliography…
Most generative models of audio directly generate samples in one of two domains: time or frequency.
Speech analysis/synthesis based on a sinusoidal representation
Robert McAulay and Thomas Quatieri · 1986
Earlier work this paper cites.
Multilayer feedforward networks are universal approximators
Kurt Hornik, Maxwell Stinchcombe, and Halbert White · 1989
Earlier work this paper cites.
Backpropagation applied to handwritten zip code recognition
Y. LeCun, B. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. Hubbard, and L. D. Jackel · 1989
Earlier work this paper cites.
Spectral modeling synthesis: A sound analysis/synthesis system based on a deterministic plus stochastic decomposition
Xavier Serra and Julius Smith · 1990
Earlier work this paper cites.
Backpropagation through time: what it does and how to do it
P. J. Werbos · 1990
Earlier work this paper cites.
Gradient-based learning algorithms for recurrent connectionist networks
Ronald J Williams and David Zipser · 1990
Earlier work this paper cites.
Timbre morphing of sounds with unequal numbers of features
Edwin Tellman, Lippold Haken, and Bryan Holloway · 1995
Earlier work this paper cites.
Robust multipitch estimation for the analysis and manipulation of polyphonic musical signals
Anssi Klapuri, Tuomas Virtanen, and Jan-Markus Holm · 2000
Earlier work this paper cites.
Hiln-the mpeg-4 parametric audio coding tools
Heiko Purnhagen and Nikolaus Meine · 2000
Earlier work this paper cites.
Multiresolution spectrotemporal analysis of complex sounds
Taishih Chi, Powen Ru, and Shihab A Shamma · 2005
Earlier work this paper cites.
Feature-based synthesis: Mapping acoustic and perceptual features onto synthesis parameters
Matthew D Hoffman and Perry R Cook · 2006
Earlier work this paper cites.
Analysis, synthesis, and perception of musical sounds
James W Beauchamp · 2007
Earlier work this paper cites.
Speech dereverberation
Patrick A Naylor and Nikolay D Gaubitch · 2010
Earlier work this paper cites.
Physical audio signal processing: For virtual musical instruments and audio effects
Julius Orion Smith · 2010
Earlier work this paper cites.
Processing of natural sounds in human auditory cortex: tonotopy, spectral tuning, and relation to voice sensitivity
Michelle Moerel, Federico De Martino, and Elia Formisano · 2012
Cited alongside, same era.
Speech analysis synthesis and perception , volume 3
James L Flanagan · 2013
Cited alongside, same era.
Active learning of intuitive control knobs for synthesizers using gaussian processes
Cheng-Zhi Anna Huang, David Duvenaud, Kenneth C Arnold, Brenton Partridge, Josiah W Oberholtzer, and Krzysztof Z Gajos · 2014
Cited alongside, same era.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V Le · 2014
Cited alongside, same era.
Neural processing of natural sounds
Frédéric E Theunissen and Julie E Elie · 2014
Cited alongside, same era.
Generating images with perceptual similarity metrics based on deep networks
Alexey Dosovitskiy and Thomas Brox · 2016
Tacotron: Towards end-to-end speech synthesis
Yuxuan Wang, RJ Skerry-Ryan, Daisy Stanton, Yonghui Wu, Ron J Weiss, Navdeep Jaitly, Zongheng Yang, Ying Xiao, Zhifeng Chen, Samy Bengio, et al · 2017
Later among the works it cites.
Fast spectrogram inversion using multi-head convolutional neural networks
S. O. Arik, H. Jun, and G. Diamos · 2018
Later among the works it cites.
Sing: Symbol-to-instrument neural generator
Alexandre Defossez, Neil Zeghidour, Nicolas Usunier, Leon Bottou, and Francis Bach · 2018
Later among the works it cites.
David Ha and Jürgen Schmidhuber · 2018
Later among the works it cites.
Efficient neural audio synthesis
Nal Kalchbrenner, Erich Elsen, Karen Simonyan, Seb Noury, Norman Casagrande, Edward Lockhart, Florian Stimberg, Aaron van den Oord, Sander Dieleman, and Koray Kavukcuoglu · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Samplernn: An unconditional end-to-end neural audio generation model
Soroush Mehri, Kundan Kumar, Ishaan Gulrajani, Rithesh Kumar, Shubham Jain, Jose Sotelo, Aaron Courville, and Yoshua Bengio · 2016
Cited alongside, same era.
World: a vocoder-based high-quality speech synthesis system for real-time applications
Masanori Morise, Fumiya Yokomori, and Kenji Ozawa · 2016
Cited alongside, same era.
Wavenet: A generative model for raw audio
Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
A neural parametric singing synthesizer modeling timbre and expression from natural songs
Merlijn Blaauw and Jordi Bonada · 2017
Cited alongside, same era.
Neural audio synthesis of musical notes with WaveNet autoencoders
Jesse Engel, Cinjon Resnick, Adam Roberts, Sander Dieleman, Douglas Eck, Karen Simonyan, and Mohammad Norouzi · 2017
Cited alongside, same era.
Crepe: A convolutional representation for pitch estimation
Jong Wook Kim, Justin Salamon, Peter Li, and Juan Pablo Bello · 2018
Later among the works it cites.
Challenging common assumptions in the unsupervised learning of disentangled representations
Francesco Locatello, Stefan Bauer, Mario Lucic, Sylvain Gelly, Bernhard Schölkopf, and Olivier Bachem · 2018
Later among the works it cites.
Wgansing: A multi-voice singing voice synthesizer based on the wasserstein-gan
Pritish Chandna, Merlijn Blaauw, Jordi Bonada, and Emilia Gomez · 2019
Later among the works it cites.
Adversarial audio synthesis
Chris Donahue, Julian McAuley, and Miller Puckette · 2019
Later among the works it cites.
GANSynth: Adversarial neural audio synthesis
Jesse Engel, Kumar Krishna Agrawal, Shuo Chen, Ishaan Gulrajani, Chris Donahue, and Adam Roberts · 2019
Later among the works it cites.
Universal audio synthesizer control with normalizing flows
Philippe Esling, Naotake Masuda, Adrien Bardet, Romeo Despres, et al · 2019
Later among the works it cites.
Fast and flexible neural audio synthesis
Lamtharn (Hanoi) Hantrakul, Jesse Engel, Adam Roberts, and Chenjie Gu · 2019
Later among the works it cites.
Neural source-filter waveform models for statistical parametric speech synthesis
Xin Wang, Shinji Takaki, and Junichi Yamagishi · 2019
Later among the works it cites.