Fetching the paper…
Reading the bibliography…
In recent years, rapid progress has been made on the problem of single-channel sound separation using supervised training of deep neural networks.
Image method for efficiently simulating small-room acoustics
J. B. Allen and D. A. Berkley · 1979
Earlier work this paper cites.
Using confidence intervals in within-subject designs
G. R. Loftus and M. E. Masson · 1994
Earlier work this paper cites.
One microphone source separation
S. T. Roweis · 2001
Earlier work this paper cites.
Audio-visual sound separation via hidden markov models
J. R. Hershey and M. Casey · 2002
Earlier work this paper cites.
Non-negative matrix factor deconvolution; extraction of multiple sound sources from monophonic inputs
P. Smaragdis · 2004
Earlier work this paper cites.
Super-human multi-talker speech recognition: The IBM 2006 speech separation challenge system
T. Kristjansson, J. Hershey, P. Olsen, S. Rennie, and R. Gopinath · 2006
Earlier work this paper cites.
Single-channel speech separation using sparse non-negative matrix factorization
M. N. Schmidt and R. K. Olsson · 2006
Earlier work this paper cites.
Speech separation using speaker-adapted eigenvoice speech models
R. J. Weiss and D. P. W. Ellis · 2010
Earlier work this paper cites.
Deep learning for monaural speech separation
P.-S. Huang, M. Kim, M. Hasegawa-Johnson, and P. Smaragdis · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Earlier work this paper cites.
LibriSpeech: an ASR corpus based on public domain audio books
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur · 2015
Earlier work this paper cites.
Speech enhancement with LSTM recurrent neural networks and its application to noise-robust ASR
F. Weninger, H. Erdogan, S. Watanabe, E. Vincent, J. Le Roux, J. R. Hershey, and B. Schuller · 2015
Earlier work this paper cites.
Domain separation networks
K. Bousmalis, G. Trigeorgis, N. Silberman, D. Krishnan, and D. Erhan · 2016
Earlier work this paper cites.
Domain-adversarial training of neural networks
Y. Ganin, E. Ustinova, H. Ajakan, P. Germain, H. Larochelle, F. Laviolette, M. Marchand, and V. Lempitsky · 2016
Earlier work this paper cites.
Deep clustering: Discriminative embeddings for segmentation and separation
J. R. Hershey, Z. Chen, J. Le Roux, and S. Watanabe · 2016
Earlier work this paper cites.
Single-channel multi-speaker separation using deep clustering
Y. Isik, J. L. Roux, Z. Chen, S. Watanabe, and J. R. Hershey · 2016
Earlier work this paper cites.
YFCC100M: The new data in multimedia research
B. Thomee, D. A. Shamma, G. Friedland, B. Elizalde, K. Ni, D. Poland, D. Borth, and L.-J. Li · 2016
Earlier work this paper cites.
Unsupervised pixel-level domain adaptation with generative adversarial networks
K. Bousmalis, N. Silberman, D. Dohan, D. Erhan, and D. Krishnan · 2017
Cited alongside, same era.
Freesound datasets: a platform for the creation of open audio datasets
E. Fonseca, J. Pons Puig, X. Favory, F. Font Corbera, D. Bogdanov, A. Ferraro, S. Oramas, A. Porter, and X. Serra · 2017
Cited alongside, same era.
Audio set: An ontology and human-labeled dataset for audio events
J. F. Gemmeke, D. P. W. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter · 2017
Cited alongside, same era.
Adversarial discriminative domain adaptation
E. Tzeng, J. Hoffman, K. Saenko, and T. Darrell · 2017
Cited alongside, same era.
Permutation invariant training of deep models for speaker-independent multi-talker speech separation
D. Yu, M. Kolbæk, Z.-H. Tan, and J. Jensen · 2017
Cited alongside, same era.
Noise2Noise: Learning image restoration without clean data
Conv-TasNet: Surpassing ideal time–frequency magnitude masking for speech separation
Y. Luo and N. Mesgarani · 2019
Later among the works it cites.
Cutting music source separation some slakh: A dataset to study the impact of training data quality and quantity
E. Manilow, G. Wichern, P. Seetharaman, and J. Le Roux · 2019
Later among the works it cites.
Finding strength in weakness: Learning to separate sounds with weak supervision
F. Pishdadian, G. Wichern, and J. Le Roux · 2019
Later among the works it cites.
Bootstrapping single-channel source separation via unsupervised spatial clustering on stereo mixtures
P. Seetharaman, G. Wichern, J. Le Roux, and B. Pardo · 2019
Later among the works it cites.
Unsupervised deep clustering for source separation: Direct learning from mixtures using spatial information
E. Tzinis, S. Venkataramani, and P. Smaragdis · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Lehtinen, J. Munkberg, J. Hasselgren, S. Laine, T. Karras, M. Aittala, and T. Aila · 2018
Cited alongside, same era.
Building corpora for single-channel speech separation across multiple domains
M. Maciejewski, G. Sell, L. P. Garcia-Perera, S. Watanabe, and S. Khudanpur · 2018
Cited alongside, same era.
Multi-channel deep clustering: Discriminative spectral and spatial embeddings for speaker-independent speech separation
Z.-Q. Wang, J. Le Roux, and J. R. Hershey · 2018
Cited alongside, same era.
mixup: Beyond empirical risk minimization
H. Zhang, M. Cisse, Y. N. Dauphin, and D. Lopez-Paz · 2018
Cited alongside, same era.
Self-supervised deep learning-based speech denoising
N. Alamdari, A. Azarang, and N. Kehtarnavaz · 2019
Cited alongside, same era.
ReMixMatch: Semi-supervised learning with distribution alignment and augmentation anchoring
D. Berthelot, N. Carlini, E. D. Cubuk, A. Kurakin, K. Sohn, H. Zhang, and C. Raffel · 2019
Cited alongside, same era.
MixMatch: A holistic approach to semi-supervised learning
D. Berthelot, N. Carlini, I. Goodfellow, N. Papernot, A. Oliver, and C. A. Raffel · 2019
Cited alongside, same era.
Differentiable consistency constraints for improved deep speech enhancement
S. Wisdom, J. R. Hershey, K. Wilson, J. Thorpe, M. Chinen, B. Patton, and R. A. Saurous · 2019
Later among the works it cites.
https://universal-sound-separation.github.io/unsupervised_sound_separation
2020
Closest in time.
LibriMix: An open-source dataset for generalizable speech separation
J. Cosentino, M. Pariente, S. Cornell, A. Deleforge, and E. Vincent · 2020
Closest in time.
FSD50k: an open dataset of human-labeled sound events
E. Fonseca, X. Favory, J. Pons, F. Font, and X. Serra · 2020
Closest in time.
Mixup-breakdown: a consistency training method for improving generalization of speech separation models
M. W. Lam, J. Wang, D. Su, and D. Yu · 2020
Closest in time.
Dual-path RNN: efficient long sequence modeling for time-domain single-channel speech separation
Y. Luo, Z. Chen, and T. Yoshioka · 2020
Closest in time.
Voice separation with an unknown number of multiple speakers
E. Nachmani, Y. Adi, and L. Wolf · 2020
Closest in time.
Filterbank design for end-to-end speech separation
M. Pariente, S. Cornell, A. Deleforge, and E. Vincent · 2020
Closest in time.
Improving universal sound separation using sound classification
E. Tzinis, S. Wisdom, J. R. Hershey, A. Jansen, and D. P. W. Ellis · 2020
Closest in time.
Free Universal Sound Separation (FUSS) dataset
S. Wisdom, H. Erdogan, D. P. W. Ellis, and J. R. Hershey · 2020
Closest in time.
What’s all the FUSS about free universal sound separation data?
S. Wisdom, H. Erdogan, D. P. W. Ellis, R. Serizel, N. Turpault, E. Fonseca, J. Salamon, P. Seetharaman, and J. R. Hershey · 2020
Closest in time.
Wavesplit: End-to-end speech separation by speaker clustering
N. Zeghidour and D. Grangier · 2020
Closest in time.