Fetching the paper…
Reading the bibliography…
Deep generative models applied to audio have improved by a large margin the state-of-the-art in many speech and music related tasks.
Neural source-filter waveform models for statistical parametric speech synthesis
Xin Wang, Shinji Takaki, and Junichi Yamagishi · 1904
Earlier work this paper cites.
MelNet: A Generative Model for Audio in the Frequency Domain
Sean Vasquez and Mike Lewis · 1906
Earlier work this paper cites.
DurIAN: Duration Informed Attention Network For Multimodal Synthesis
Chengzhu Yu, Heng Lu, Na Hu, Meng Yu, Chao Weng, Kun Xu, Peng Liu, Deyi Tuo, Shiyin Kang, Guangzhi Lei, Dan Su, and Dong Yu · 1909
Earlier work this paper cites.
MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis
Kundan Kumar, Rithesh Kumar, Thibault de Boissiere, Lucas Gestin, Wei Zhen Teoh, Jose Sotelo, Alexandre de Brebisson, Yoshua Bengio, and Aaron Courville · 1910
Earlier work this paper cites.
Near-Perfect-Reconstruction Pseudo-QMF Banks
Truong Q. Nguyen · 1994
Earlier work this paper cites.
A new algorithm for designing prototype filters for M-band Pseudo QMF banks, 1996
M. Rossi, Jin-Yun Zhang, and W. Steenaart · 1996
Earlier work this paper cites.
A kaiser window approach for the design of prototype filters of cosine modulated filterbanks
Yuan Pei Lin and P. P. Vaidyanathan · 1998
Earlier work this paper cites.
Jukebox: A Generative Model for Music
Prafulla Dhariwal, Heewoo Jun, Christine Payne, Jong Wook Kim, Alec Radford, and Ilya Sutskever · 2005
Earlier work this paper cites.
Multi-band MelGAN: Faster Waveform Generation for High-Quality Text-to-Speech
Geng Yang, Shan Yang, Kai Liu, Peng Fang, Wei Chen, and Lei Xie · 2005
Earlier work this paper cites.
Generative Adversarial Networks
Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P. Kingma and Max Welling · 2014
Cited alongside, same era.
A Recurrent Latent Variable Model for Sequential Data
Junyoung Chung, Kyle Kastner, Laurent Dinh, Kratarth Goel, Aaron Courville, and Yoshua Bengio · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Lei Ba · 2015
Cited alongside, same era.
Autoencoding beyond pixels using a learned similarity metric
Anders Boesen Lindbo Larsen, Søren Kaae Sønderby, Hugo Larochelle, and Ole Winther · 2015
Cited alongside, same era.
Sequential Neural Models with Stochastic Layers
Marco Fraccaro, Søren Kaae Sønderby, Ulrich Paquet, and Ole Winther · 2016
Samplernn: An unconditional end-to-end neural audio generation model
Soroush Mehri, Kundan Kumar, Ishaan Gulrajani, Rithesh Kumar, Shubham Jain, Jose Sotelo, Aaron Courville, and Yoshua Bengio · 2017
Later among the works it cites.
SING: Symbol-to-Instrument Neural Generator
Alexandre Défossez, Neil Zeghidour, Nicolas Usunier, Léon Bottou, and Francis Bach · 2018
Later among the works it cites.
Generative timbre spaces: regularizing variational auto-encoders with perceptual metrics
Philippe Esling, Axel Chemla-Romeu-Santos, and Adrien Bitton · 2018
Later among the works it cites.
A Universal Music Translation Network
Noam Mor, Lior Wolf, Adam Polyak, and Yaniv Taigman · 2018
Later among the works it cites.
ClariNet: Parallel Wave Generation in End-to-End Text-to-Speech
Wei Ping, Kainan Peng, and Jitong Chen · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
beta-VAE: Learning Basic Visual Concepts with a Constrained Variational Framework, 11 2016
Irina Higgins, Loic Matthey, Arka Pal, Christopher Burgess, Xavier Glorot, Matthew Botvinick, Shakir Mohamed, and Alexander Lerchner · 2016
Cited alongside, same era.
WaveNet: A Generative Model for Raw Audio
Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
Neural Audio Synthesis of Musical Notes with WaveNet Autoencoders, 7 2017
Jesse Engel, Cinjon Resnick, Adam Roberts, Sander Dieleman, Mohammad Norouzi, Douglas Eck, and Karen Simonyan · 2017
Cited alongside, same era.
Fader Networks: Manipulating Images by Sliding Attributes
Guillaume Lample, Neil Zeghidour, Nicolas Usunier, Antoine Bordes, Ludovic Denoyer, and Marc’aurelio Ranzato · 2017
Cited alongside, same era.
Later among the works it cites.
DDSP: Differentiable Digital Signal Processing
Jesse Engel, Lamtharn Hantrakul, Chenjie Gu, and Adam Roberts · 2019
Later among the works it cites.
Waveglow: A Flow-based Generative Network for Speech Synthesis
Ryan Prenger, Rafael Valle, and Bryan Catanzaro · 2019
Later among the works it cites.
CSTR VCTK Corpus: English Multi-speaker Corpus for CSTR Voice Cloning Toolkit, 2019
Junichi Yamagishi, Christophe Veaux, and Kirsten MacDonald · 2019
Later among the works it cites.
Parallel Wavegan: A Fast Waveform Generation Model Based on Generative Adversarial Networks with Multi-Resolution Spectrogram
Ryuichi Yamamoto, Eunwoo Song, and Jae Min Kim · 2020
Later among the works it cites.