Fetching the paper…
Reading the bibliography…
Data-driven models for audio source separation such as U-Net or Wave-U-Net are usually models dedicated to and specifically trained for a single task, e.g.
Performance measurement in blind audio source separation
E. Vincent, R. Gribonval, and C. Févotte · 2006
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
mir_eval: a transparent implementation of common mir metrics
C. Raffel, B. Mcfee, E. J. Humphrey, J. Salamon, O. Nieto, D. Liang, and D. P. W. Ellis · 2014
Earlier work this paper cites.
Joint optimization of masks and deep recurrent neural networks for monaural source separation
Po-Sen Huang, Minje Kim, Mark Hasegawa-Johnson, and Paris Smaragdis · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
O. Ronneberger, P. Fischer, and T. Brox · 2015
Earlier work this paper cites.
Wavenet: A generative model for raw audio
A. van den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. W. Senior, and K. Kavukcuoglu · 2016
Earlier work this paper cites.
Singing voice separation with deep u-net convolutional networks
N. Montecchio R. Bittner A. Kumar T. Weyde A. Jansson, E. J. Humphrey · 2017
Earlier work this paper cites.
Monoaural audio source separation using deep convolutional neural networks
P. Chandna, M. Miron, J. Janer, and E. Gómez · 2017
Earlier work this paper cites.
Modulating early visual processing by language
H. de Vries, F. Strub, J. Mary, H. Larochelle, O. Pietquin, and A. C. Courville · 2017
Cited alongside, same era.
Dynamic layer normalization for adaptive neural acoustic modeling in speech recognition
T. Kim, I. Song, and Y. Bengio · 2017
Cited alongside, same era.
Impact of phase estimation on single-channel speech separation based on time-frequency masking
F. Mayer, D. Williamson, P. Mowlaee, and D. Wang · 2017
Cited alongside, same era.
The MUSDB18 corpus for music separation, 2017
Zafar Rafii, Antoine Liutkus, Fabian-Robert Stöter, Stylianos Ioannis Mimilakis, and Rachel Bittner · 2017
Cited alongside, same era.
Parallel wavenet: Fast high-fidelity speech synthesis
A. van den Oord, Y. Li, I. Babuschkin, K. Simonyan, O. Vinyals, K. Kavukcuoglu, G. van den Driessche, E. Lockhart, L. C. Cobo, F. Stimberg, N. Casagrande, D. Grewe, S. Noury, S. Dieleman, E. Elsen, N. Kalchbrenner, H. Zen, A. Graves, H. King, T. Walters, D. Belov, and D. Hassabis · 2017
Semi-blind source separation with multichannel variational autoencoder
H. Kameoka, Li Li, S. Inoue, and S. Makino · 2018
Later among the works it cites.
Film: Visual reasoning with a general conditioning layer
E. Perez, F. Strub, H. de Vries, V. Dumoulin, and A. C. Courville · 2018
Later among the works it cites.
An Overview of Lead and Accompaniment Separation in Music
Z. Rafii, A. Liutkus, F.-R. Stöter, S. Ioannis Mimilakis, D. Fitzgerald, and B. Pardo · 2018
Later among the works it cites.
Natural TTS synthesis by conditioning wavenet on mel spectrogram predictions
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. J. Skerry-Ryan, R. A. Saurous, Y. Agiomyrgiannakis, and Y. Wu · 2018
Later among the works it cites.
Wave-u-net: A multi-scale neural network for end-to-end audio source separation
D. Stoller, S. Ewert, and S. Dixon · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Li-Chia Yang, Szu-Yu Chou, and Yi-Hsuan Yang · 2017
Cited alongside, same era.
Feature-wise transformations
V. Dumoulin, E. Perez, N. Schucher, F. Strub, H. de Vries, Aaron Courville, and Y. Bengio · 2018
Cited alongside, same era.
Onsets and frames: Dual-objective piano transcription
C. Hawthorne, E. Elsen, J. Song, A. Roberts, I. Simon, C. Raffel, J. Engel, S. Oore, and D. Eck · 2018
Cited alongside, same era.
An improved relative self-attention mechanism for transformer with application to music generation
C. A. Huang, A. Vaswani, J. Uszkoreit, N. Shazeer, C. Hawthorne, A. M. Dai, M. D. Hoffman, and D. Eck · 2018
Cited alongside, same era.
Visual reasoning with multi-hop feature modulation
F. Strub, M. Seurin, E. Perez, H. de Vries, J. Mary, P. Preux, A. C. Courville, and O. Pietquin · 2018
Later among the works it cites.
Between-class learning for image classification
Y. Tokozume, Y. Ushiku, and T. Harada · 2018
Later among the works it cites.
Improving singing voice separation using deep u-net and wave-u-net with data augmentation
Alice Cohen-Hadria, Axel Roebel, and Geoffroy Peeters · 2019
Closest in time.