Fetching the paper…
Reading the bibliography…
We propose a method of separating a desired sound source from a single-channel mixture, based on either a textual description or a short audio sample of the target source.
F. Weninger, H. Erdogan, S. Watanabe, E. Vincent, J. L. Roux, J. R. Hershey, and B. Schuller, “Speech enhancement with LSTM recurrent neural networks and its application to noise-robust ASR,” in Proc. LVA/ICA . Springer, 2015, pp. 91–99
2015
Earlier work this paper cites.
L. Le Magoarou, A. Ozerov, and N. Q. Duong, “Text-informed audio source separation. example-based approach using non-negative matrix partial co-factorization,” Journal of Signal Processing Systems , vol. 79, no. 2, pp. 117–131, 2015
2015
Earlier work this paper cites.
D. Clevert, T. Unterthiner, and S. Hochreiter, “Fast and accurate deep network learning by exponential linear units (ELUs),” in Proc. ICLR , Y. Bengio and Y. LeCun, Eds., 2016
2016
Earlier work this paper cites.
S. Hershey, S. Chaudhuri, D. P. Ellis, J. F. Gemmeke, A. Jansen, R. C. Moore, M. Plakal, D. Platt, R. A. Saurous, B. Seybold et al. , “CNN architectures for large-scale audio classification,” in Proc. ICASSP , 2017, pp. 131–135
2017
Earlier work this paper cites.
J. F. Gemmeke, D. P. W. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “AudioSet: An ontology and human-labeled dataset for audio events,” in Proc. ICASSP , New Orleans, LA, 2017
2017
Earlier work this paper cites.
S. Hershey, S. Chaudhuri, D. P. W. Ellis, J. F. Gemmeke, A. Jansen, R. C. Moore, M. Plakal, D. Platt, R. A. Saurous, B. Seybold, M. Slaney, R. J. Weiss, and K. Wilson, “CNN architectures for large-scale audio classification,” in Proc. ICASSP , 2017, pp. 131–135
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
A. Ephrat, I. Mosseri, O. Lang, T. Dekel, K. Wilson, A. Hassidim, W. T. Freeman, and M. Rubinstein, “Looking to listen at the cocktail party: A speaker-independent audio-visual model for speech separation,” ACM Trans. Graph. , vol. 37, no. 4, Jul. 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
E. Perez, F. Strub, H. de Vries, V. Dumoulin, and A. C. Courville, “FiLM: Visual reasoning with a general conditioning layer,” in Proc. AAAI , 2018, pp. 3942–3951
2018
Earlier work this paper cites.
Y. Wu and K. He, “Group normalization,” in ECCV 2018, Munich, Germany, September 8-14, 2018, Proceedings, Part XIII , ser. LNCS, vol. 11217. Springer, 2018, pp. 3–19
2018
Earlier work this paper cites.
2019
Earlier work this paper cites.
I. Kavalerov, S. Wisdom, H. Erdogan, B. Patton, K. W. Wilson, J. L. Roux, and J. R. Hershey, “Universal sound separation,” in Proc. WASPAA , 2019, pp. 175–179
2019
Earlier work this paper cites.
R. Gao and K. Grauman, “Co-separating sounds of visual objects,” in Proc. CVPR , 2019, pp. 3879–3888
2019
Cited alongside, same era.
X. Xu, B. Dai, and D. Lin, “Recursive visual sound separation using minus-plus net,” in Proc. ICCV , 2019, pp. 882–891
2019
Cited alongside, same era.
Q. Wang, H. Muckenhirn, K. Wilson, P. Sridhar, Z. Wu, J. Hershey, R. A. Saurous, R. J. Weiss, Y. Jia, and I. L. Moreno, “Voicefilter: Targeted voice separation by speaker-conditioned spectrogram masking,” in Proc. Interspeech , 2019
2019
Cited alongside, same era.
S. Birnbaum, V. Kuleshov, S. Z. Enam, P. W. Koh, and S. Ermon, “Temporal FiLM: Capturing long-range sequence dependencies with feature-wise modulations,” in Advances in Neural Information Processing Systems , 2019, pp. 10 287–10 298
2019
Cited alongside, same era.
C. D. Kim, B. Kim, H. Lee, and G. Kim, “Audiocaps: Generating captions for audios in the wild,” in NAACL-HLT , 2019
K. Drossos, S. Lipping, and T. Virtanen, “Clotho: An audio captioning dataset,” in Proc. ICASSP , 2020, pp. 736–740
2020
Later among the works it cites.
2021
Later among the works it cites.
A. Li, W. Liu, X. Luo, C. Zheng, and X. Li, “ICASSP 2021 deep noise suppression challenge: Decoupling magnitude and phase optimization with a two-stage deep network,” in Proc. ICASSP , 2021, pp. 6628–6632
2021
Later among the works it cites.
E. Tzinis, S. Wisdom, A. Jansen, S. Hershey, T. Remez, D. P. Ellis, and J. R. Hershey, “Into the wild with AudioScope: Unsupervised audio-visual separation of on-screen sounds,” in Proc. ICLR , 2021
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
J. L. Roux, S. Wisdom, H. Erdogan, and J. R. Hershey, “SDR - half-baked or well done?” in Proc. ICASSP , 2019, pp. 626–630
2019
Cited alongside, same era.
E. Tzinis, S. Wisdom, J. R. Hershey, A. Jansen, and D. P. W. Ellis, “Improving universal sound separation using sound classification,” in Proc. ICASSP , 2020, pp. 96–100
2020
Cited alongside, same era.
F. Pishdadian, G. Wichern, and J. Le Roux, “Finding strength in weakness: Learning to separate sounds with weak supervision,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 28, pp. 2386–2399, 2020
2020
Cited alongside, same era.
Q. Kong, Y. Wang, X. Song, Y. Cao, W. Wang, and M. D. Plumbley, “Source separation with weakly labelled data: an approach to computational auditory scene analysis,” in Proc. ICASSP , 2020, pp. 101–105
2020
Cited alongside, same era.
S. Wisdom, E. Tzinis, H. Erdogan, R. Weiss, K. Wilson, and J. Hershey, “Unsupervised sound separation using mixture invariant training,” in Advances in Neural Information Processing Systems , 2020
2020
Cited alongside, same era.
M. Tagliasacchi, Y. Li, K. Misiunas, and D. Roblek, “SEANet: A multi-modal speech enhancement network,” in Proc. Interspeech , 2020, pp. 1126–1130
2020
Cited alongside, same era.
2020
Cited alongside, same era.
B. Gfeller, D. Roblek, and M. Tagliasacchi, “One-shot conditional audio filtering of arbitrary sounds,” in Proc. ICASSP , 2021, pp. 501–505
2021
Later among the works it cites.
D. Michelsanti, Z.-H. Tan, S.-X. Zhang, Y. Xu, M. Yu, D. Yu, and J. Jensen, “An overview of deep-learning-based audio-visual speech enhancement and separation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
K. Schulze-Forster, C. S. J. Doire, G. Richard, and R. Badeau, “Phoneme level lyrics alignment and text-informed singing voice separation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 29, pp. 2382–2395, 2021
2021
Later among the works it cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever, “Learning transferable visual models from natural language supervision,” in Proc. ICML , 2021, pp. 8748–8763
2021
Later among the works it cites.
C. Jia, Y. Yang, Y. Xia, Y.-T. Chen, Z. Parekh, H. Pham, Q. Le, Y.-H. Sung, Z. Li, and T. Duerig, “Scaling up visual and vision-language representation learning with noisy text supervision,” in Proc. ICLR , 2021, pp. 4904–4916
2021
Later among the works it cites.
2021
Later among the works it cites.
S. Wisdom, H. Erdogan, D. P. W. Ellis, R. Serizel, N. Turpault, E. Fonseca, J. Salamon, P. Seetharaman, and J. R. Hershey, “What’s all the FUSS about free universal sound separation data?” in Proc. ICASSP , 2021
2021
Later among the works it cites.
E. Fonseca, X. Favory, J. Pons, F. Font, and X. Serra, “Fsd50k: an open dataset of human-labeled sound events,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 30, pp. 829–852, 2022
2022
Closest in time.