Fetching the paper…
Reading the bibliography…
Universal source separation (USS) is a fundamental research task for computational auditory scene analysis, which aims to separate mono recordings into individual source tracks.
S. R. Quackenbush, T. P. Barnwell, and M. A. Clements, Objective Measures of Speech Quality . Prentice Hall, 1988
1988
Earlier work this paper cites.
G. J. Brown and M. Cooke, “Computational Auditory Scene Analysis,” Computer Speech & Language , vol. 8, no. 4, pp. 297–336, 1994
1994
Earlier work this paper cites.
D. D. Lee and H. S. Seung, “Algorithms for non-negative matrix factorization,” in Neural Information Processing Systems (NeurIPS) , 2000
2000
Earlier work this paper cites.
ITU-T, “Perceptual evaluation of speech quality (PESQ): An objective method for end-to-end speech quality assessment of narrow-band telephone networks and speech codecs,” Rec. ITU-T P. 862 , 2001
2001
Earlier work this paper cites.
S. Haykin and Z. Chen, “The cocktail party problem,” Neural Computation , vol. 17, no. 9, pp. 1875–1902, 2005
2005
Earlier work this paper cites.
D. Wang and G. J. Brown, Computational auditory scene analysis: Principles, algorithms, and applications . Wiley-IEEE press, 2006
2006
Earlier work this paper cites.
E. Vincent, R. Gribonval, and C. Févotte, “Performance measurement in blind audio source separation,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 14, no. 4, pp. 1462–1469, 2006
2006
Earlier work this paper cites.
P. C. Loizou, Speech Enhancement: Theory and Practice . CRC press, 2007
2007
Earlier work this paper cites.
T. Heittola, A. Mesaros, T. Virtanen, and A. Eronen, “Sound event detection in multisource environments using source separation,” in CHiME Workshop on Machine Listening in Multisource Environments (CHiME 2011) , 2011, pp. 36–40
2011
Earlier work this paper cites.
A. Narayanan and D. Wang, “Ideal ratio mask estimation using deep neural networks for robust speech recognition,” in IEEE International Conference on Acoustics, Speech and Signal Processing , 2013, pp. 7092–7096
2013
Earlier work this paper cites.
C. Veaux, J. Yamagishi, and S. King, “The Voice Bank Corpus: Design, collection and data analysis of a large regional accent speech database,” in International Conference Oriental COCOSDA with Conference on Asian Spoken Language Research and Evaluation (O-COCOSDA/CASLRE) , 2013
2013
Earlier work this paper cites.
J. Thiemann, N. Ito, and E. Vincent, “DEMAND: a collection of multi-channel recordings of acoustic noise in diverse environments,” in Proceedings of Meetings on Acoustics , 2013
2013
Earlier work this paper cites.
Y. Xu, J. Du, L.-R. Dai, and C.-H. Lee, “A regression approach to speech enhancement based on deep neural networks,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 23, no. 1, pp. 7–19, 2014
2014
Earlier work this paper cites.
J. Salamon, C. Jacoby, and J. P. Bello, “A dataset and taxonomy for urban sound research,” in Proceedings of the 22nd ACM international conference on Multimedia , 2014, pp. 1041–1044
2014
Earlier work this paper cites.
D. S. Williamson, Y. Wang, and D. Wang, “Complex ratio masking for monaural speech separation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 24, no. 3, pp. 483–492, 2015
2015
Earlier work this paper cites.
P.-S. Huang, M. Kim, M. Hasegawa-Johnson, and P. Smaragdis, “Joint optimization of masks and deep recurrent neural networks for monaural source separation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 23, no. 12, pp. 2136–2147, 2015
2015
Earlier work this paper cites.
S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in International Conference on Machine Learning , 2015, pp. 448–456
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in International Conference on Learning Representations (ICLR) , 2015
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Identity mappings in deep residual networks,” in European Conference on Computer Vision . Springer, 2016, pp. 630–645
2016
Earlier work this paper cites.
J. F. Gemmeke, D. P. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio set: An ontology and human-labeled dataset for audio events,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2017, pp. 776–780
2017
Earlier work this paper cites.
Q. Kong, Y. Xu, W. Wang, and M. D. Plumbley, “A joint detection-classification model for audio tagging of weakly labelled data,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2017, pp. 641–645
2017
Earlier work this paper cites.
S. Uhlich, M. Porcu, F. Giron, M. Enenkl, T. Kemp, N. Takahashi, and Y. Mitsufuji, “Improving music source separation based on deep neural networks through data augmentation and network blending,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2017, pp. 261–265
2017
Earlier work this paper cites.
P. Chandna, M. Miron, J. Janer, and E. Gómez, “Monoaural audio source separation using deep convolutional neural networks,” in International Conference on Latent Variable Analysis and Signal Separation . Springer, 2017, pp. 258–266
2017
Earlier work this paper cites.
A. Jansson, E. Humphrey, N. Montecchio, R. Bittner, A. Kumar, and T. Weyde, “Singing voice separation with deep U-Net convolutional networks,” in International Society for Music Information Retrieval (ISMIR) , 2017
2017
Earlier work this paper cites.
Z. Rafii, A. Liutkus, F.-R. Stöter, S. I. Mimilakis, and R. Bittner, “The MUSDB18 corpus for music separation,” Dec. 2017. [Online]. Available: https://doi.org/10.5281/zenodo.1117372
2017
Cited alongside, same era.
S. Hershey, S. Chaudhuri, D. P. Ellis, J. F. Gemmeke, A. Jansen, R. C. Moore, M. Plakal, D. Platt, R. A. Saurous, B. Seybold et al. , “CNN architectures for large-scale audio classification,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2017, pp. 131–135
2017
Cited alongside, same era.
F.-R. Stöter, A. Liutkus, and N. Ito, “The 2018 signal separation evaluation campaign,” in International Conference on Latent Variable Analysis and Signal Separation . Springer, 2018, pp. 293–305
2018
Cited alongside, same era.
D. Stoller, S. Ewert, and S. Dixon, “Wave-U-Net: A multi-scale neural network for end-to-end audio source separation,” in International Society for Music Information Retrieval (ISMIR) , 2018
2018
Cited alongside, same era.
E. Tzinis, Z. Wang, and P. Smaragdis, “Sudo rm-rf: Efficient networks for universal audio source separation,” in IEEE International Workshop on Machine Learning for Signal Processing (MLSP) , 2020
2020
Later among the works it cites.
E. Tzinis, S. Wisdom, J. R. Hershey, A. Jansen, and D. P. Ellis, “Improving universal sound separation using sound classification,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2020, pp. 96–100
2020
Later among the works it cites.
F. Pishdadian, G. Wichern, and J. Le Roux, “Learning to separate sounds from weakly labeled scenes,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2020, pp. 91–95
2020
Later among the works it cites.
Q. Kong, Y. Cao, T. Iqbal, Y. Wang, W. Wang, and M. D. Plumbley, “PANNs: Large-scale pretrained audio neural networks for audio pattern recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 28, pp. 2880–2894, 2020
2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
Y. Wang, “Polyphonic sound event detection with weak labeling,” PhD thesis , 2018
2018
Cited alongside, same era.
Y. Luo and N. Mesgarani, “Tasnet: time-domain audio separation network for real-time, single-channel speech separation,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2018, pp. 696–700
2018
Cited alongside, same era.
N. Takahashi, N. Goswami, and Y. Mitsufuji, “MMDenseLSTM: An efficient combination of convolutional and recurrent neural networks for audio source separation,” in IEEE International Workshop on Acoustic Signal Enhancement (IWAENC) , 2018, pp. 106–110
2018
Cited alongside, same era.
E. Fonseca, M. Plakal, F. Font, D. P. Ellis, X. Favory, J. Pons, and X. Serra, “General-purpose tagging of Freesound audio with Audioset labels: Task description, dataset, and baseline,” in Proceedings of the Detection and Classification of Acoustic Scenes and Events 2018 Workshop (DCASE) , 2018
2018
Cited alongside, same era.
E. Perez, F. Strub, H. De Vries, V. Dumoulin, and A. Courville, “FiLM: Visual reasoning with a general conditioning layer,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 32, no. 1, 2018
2018
Cited alongside, same era.
I. Kavalerov, S. Wisdom, H. Erdogan, B. Patton, K. Wilson, J. Le Roux, and J. R. Hershey, “Universal sound separation,” in IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA) , 2019, pp. 175–179
2019
Cited alongside, same era.
P. Seetharaman, G. Wichern, S. Venkataramani, and J. Le Roux, “Class-conditional embeddings for music source separation,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2019, pp. 301–305
2019
Cited alongside, same era.
Later among the works it cites.
Q. Kong, Y. Wang, X. Song, Y. Cao, W. Wang, and M. D. Plumbley, “Source separation with weakly labelled data: An approach to computational auditory scene analysis,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2020, pp. 101–105
2020
Later among the works it cites.
Y. Hu, Y. Liu, S. Lv, M. Xing, S. Zhang, Y. Fu, J. Wu, B. Zhang, and L. Xie, “DCCRN: Deep complex convolution recurrent network for phase-aware speech enhancement,” in INTERSPEECH , 2020
2020
Later among the works it cites.
R. Hennequin, A. Khlif, F. Voituret, and M. Moussallam, “Spleeter: a fast and efficient music source separation tool with pre-trained models,” Journal of Open Source Software , vol. 5, no. 50, p. 2154, 2020
2020
Later among the works it cites.
E. Fonseca, X. Favory, J. Pons, F. Font, and X. Serra, “FSD50k: an open dataset of human-labeled sound events,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , 2020
2020
Later among the works it cites.
Y. Luo, Z. Chen, C. Han, C. Li, T. Zhou, and N. Mesgarani, “Rethinking the separation layers in speech separation networks,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 1–5
2021
Later among the works it cites.
B. Gfeller, D. Roblek, and M. Tagliasacchi, “One-shot conditional audio filtering of arbitrary sounds,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2021, pp. 501–505
2021
Later among the works it cites.
W. Choi, M. Kim, J. Chung, and S. Jung, “LaSAFT: Latent Source Attentive Frequency Transformation for Conditioned Source Separation,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2021, pp. 171–175
2021
Later among the works it cites.
E. Tzinis, Z. Wang, X. Jiang, and P. Smaragdis, “Compute and memory efficient universal sound source separation,” Journal of Signal Processing Systems , pp. 245––259, 2021
2021
Later among the works it cites.
Y. Gong, Y.-A. Chung, and J. Glass, “PSLA: Improving audio tagging with pretraining, sampling, labeling, and aggregation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 29, pp. 3292–3306, 2021
2021
Later among the works it cites.
Q. Kong, Y. Cao, H. Liu, K. Choi, and Y. Wang, “Decoupling magnitude and phase estimation with deep resunet for music source separation,” in International Society for Music Information Retrieval (ISMIR) , 2021
2021
Later among the works it cites.
A. Défossez, “Hybrid spectrogram and waveform source separation,” in Proceedings of the ISMIR 2021 Workshop on Music Source Separation , 2021
2021
Later among the works it cites.
S. Wisdom, H. Erdogan, D. P. Ellis, R. Serizel, N. Turpault, E. Fonseca, J. Salamon, P. Seetharaman, and J. R. Hershey, “What’s all the fuss about free universal sound separation data?” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2021, pp. 186–190
2021
Later among the works it cites.
Y. Gong, Y.-A. Chung, and J. Glass, “AST: Audio spectrogram transformer,” in INTERSPEECH , 2021
2021
Later among the works it cites.
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , 2021, pp. 10 012–10 022
2021
Later among the works it cites.
X. Liu, H. Liu, Q. Kong, X. Mei, J. Zhao, Q. Huang, M. D. Plumbley, and W. Wang, “Separate what you describe: Language-queried audio source separation,” in INTERSPEECH , 2022
2022
Later among the works it cites.
K. Chen, X. Du, B. Zhu, Z. Ma, T. Berg-Kirkpatrick, and S. Dubnov, “HTS-AT: A hierarchical token-semantic audio transformer for sound classification and detection,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2022, pp. 646–650
2022
Later among the works it cites.
K. Chen, X. Du, B. Zhu, Z. Ma, T. Berg-Kirkpatrick, and S. Dubnov, “Zero-shot audio source separation through query-based learning from weakly-labeled data,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 36, no. 4, 2022, pp. 4441–4449
2022
Later among the works it cites.
2022
Later among the works it cites.
M. Delcroix, J. B. Vázquez, T. Ochiai, K. Kinoshita, Y. Ohishi, and S. Araki, “SoundBeam: Target sound extraction conditioned on sound-class labels and enrollment clues for increased performance and continuous learning,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , 2022
2022
Later among the works it cites.
S. Rouard, F. Massa, and A. Défossez, “Hybrid transformers for music source separation,” in International Conference on Acoustics, Speech, and Signal Processing (ICASSP) , 2023
2023
Closest in time.