Fetching the paper…
Reading the bibliography…
A number of recent advances in neural audio synthesis rely on upsampling layers, which can introduce undesired artifacts.
“Performance measurement in blind audio source separation,”
Emmanuel Vincent, Rémi Gribonval, and Cédric Févotte, · 2006
Earlier work this paper cites.
Mathematics of the discrete Fourier transform (DFT): with audio applications
Julius Orion Smith, · 2007
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
Diederik P Kingma and Jimmy Ba, · 2014
Earlier work this paper cites.
“Understanding deep image representations by inverting them,”
Aravindh Mahendran and Andrea Vedaldi, · 2015
Earlier work this paper cites.
“Inceptionism: Going deeper into neural networks,”
Alexander Mordvintsev, Christopher Olah, and Mike Tyka, · 2015
Earlier work this paper cites.
“librosa: Audio and music signal analysis in python,”
Brian McFee, Colin Raffel, Dawen Liang, Daniel PW Ellis, Matt McVicar, Eric Battenberg, and Oriol Nieto, · 2015
Earlier work this paper cites.
“Wavenet: A generative model for raw audio,”
Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu, · 2016
Earlier work this paper cites.
“Deconvolution and checkerboard artifacts,”
Augustus Odena, Vincent Dumoulin, and Chris Olah, · 2016
Earlier work this paper cites.
“Synthesizing the preferred inputs for neurons in neural networks via deep generator networks,”
Anh Nguyen, Alexey Dosovitskiy, Jason Yosinski, Thomas Brox, and Jeff Clune, · 2016
Earlier work this paper cites.
“Real-time single image and video super-resolution using an efficient sub-pixel convolutional neural network,”
Wenzhe Shi, Jose Caballero, Ferenc Huszár, Johannes Totz, Andrew P Aitken, Rob Bishop, Daniel Rueckert, and Zehan Wang, · 2016
Earlier work this paper cites.
“Segan: Speech enhancement generative adversarial network,”
Santiago Pascual, Antonio Bonafonte, and Joan Serrà, · 2017
Earlier work this paper cites.
“Audio super resolution using neural networks,”
Volodymyr Kuleshov, S Zayd Enam, and Stefano Ermon, · 2017
Cited alongside, same era.
“Feature visualization,”
Chris Olah, Alexander Mordvintsev, and Ludwig Schubert, · 2017
Cited alongside, same era.
“Checkerboard artifact free sub-pixel convolution: A note on sub-pixel convolution, resize convolution and convolution resize,”
Andrew Aitken, Christian Ledig, Lucas Theis, Jose Caballero, Zehan Wang, and Wenzhe Shi, · 2017
Cited alongside, same era.
“Timbre analysis of music audio signals with convolutional neural networks,”
Jordi Pons, Olga Slizovskaia, Rong Gong, Emilia Gómez, and Xavier Serra, · 2017
Cited alongside, same era.
“Improving music source separation based on deep neural networks through data augmentation and network blending,”
Stefan Uhlich, Marcello Porcu, Franck Giron, Michael Enenkl, Thomas Kemp, Naoya Takahashi, and Yuki Mitsufuji, · 2017
Cited alongside, same era.
“Melgan: Generative adversarial networks for conditional waveform synthesis,”
Kundan Kumar, Rithesh Kumar, Thibault de Boissiere, Lucas Gestin, Wei Zhen Teoh, Jose Sotelo, Alexandre de Brébisson, Yoshua Bengio, and Aaron C Courville, · 2019
Later among the works it cites.
“Waveglow: A flow-based generative network for speech synthesis,”
Ryan Prenger, Rafael Valle, and Bryan Catanzaro, · 2019
Later among the works it cites.
“Blow: a single-scale hyperconditioned flow for non-parallel raw-audio voice conversion,”
Joan Serrà, Santiago Pascual, and Carlos Segura Perales, · 2019
Later among the works it cites.
“High fidelity speech synthesis with adversarial networks,”
Mikołaj Bińkowski, Jeff Donahue, Sander Dieleman, Aidan Clark, Erich Elsen, Norman Casagrande, Luis C Cobo, and Karen Simonyan, · 2019
Later among the works it cites.
“Speech denoising with deep feature losses,”
Francois G Germain, Qifeng Chen, and Vladlen Koltun, · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zafar Rafii, Antoine Liutkus, Fabian-Robert Stöter, Stylianos Ioannis Mimilakis, and Rachel Bittner, · 2017
Cited alongside, same era.
“Wave-u-net: A multi-scale neural network for end-to-end audio source separation,”
Daniel Stoller, Sebastian Ewert, and Simon Dixon, · 2018
Cited alongside, same era.
“A wavenet for speech denoising,”
Dario Rethage, Jordi Pons, and Xavier Serra, · 2018
Cited alongside, same era.
“Multi-target voice conversion without parallel data by adversarially learning disentangled audio representations,”
Ju-chieh Chou, Cheng-chieh Yeh, Hung-yi Lee, and Lin-shan Lee, · 2018
Cited alongside, same era.
“Adversarial audio synthesis,”
Chris Donahue, Julian McAuley, and Miller Puckette, · 2019
Cited alongside, same era.
“Music source separation in the waveform domain,”
Alexandre Défossez, Nicolas Usunier, Léon Bottou, and Francis Bach, · 2019
Cited alongside, same era.
Ritwik Giri, Umut Isik, and Arvindh Krishnaswamy, · 2019
Later among the works it cites.
“Open-unmix - a reference implementation for music source separation,”
Fabian-Robert Stöter, S. Uhlich, A. Liutkus, and Y. Mitsufuji, · 2019
Later among the works it cites.
“Sams-net: A sliced attention-based neural network for music source separation,”
Haowen Hou Ming Li Tingle Li, Jiawei Chen, · 2019
Later among the works it cites.
“Densely connected neural network with dilated convolutions for real-time speech enhancement in the time domain,”
Ashutosh Pandey and DeLiang Wang, · 2020
Closest in time.
“A spectral energy distance for parallel speech synthesis,”
Alexey A Gritsenko, Tim Salimans, Rianne van den Berg, Jasper Snoek, and Nal Kalchbrenner, · 2020
Closest in time.