Fetching the paper…
Reading the bibliography…
Universal sound separation (USS) aims to extract arbitrary types of sounds from real-world recordings.
L. van der Maaten and G. Hinton, “Visualizing data using t-SNE,” J. Mach. Learn. Res. , vol. 9, no. 86, pp. 2579–2605, 2008
2008
Earlier work this paper cites.
F. Font, G. Roma, and X. Serra, “Freesound technical demo,” in Proc. 21st ACM Int. Conf. Multimedia (ACM-MM) , 2013, pp. 411–412
2013
Earlier work this paper cites.
O. Ronneberger, P. Fischer, and T. Brox, “U-Net: Convolutional networks for biomedical image segmentation,” in Proc. Med. Imag. Comp. and Comp.-Assist. Interv. (MICCAI) , 2015, pp. 234–241
2015
Earlier work this paper cites.
K. J. Piczak, “ESC: Dataset for environmental sound classification,” in Proc. 23rd ACM Conf. Multimedia (ACM-MM) , 2015, pp. 1015–1018
2015
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Proc. Adv. Neural Inf. Process. Syst. (NeurIPS) , vol. 30, 2017
2017
Earlier work this paper cites.
D. Yu, M. Kolbæk, Z.-H. Tan, and J. Jensen, “Permutation invariant training of deep models for speaker-independent multi-talker speech separation,” in Proc. IEEE Int. Conf. Acoustics Speech Signal Process. (ICASSP) , 2017, pp. 241–245
2017
Earlier work this paper cites.
J. F. Gemmeke, D. P. W. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio Set: An ontology and human-labeled dataset for audio events,” in Proc. IEEE Int. Conf. Acoustics Speech Signal Process. (ICASSP) , 2017, pp. 776–780
2017
Earlier work this paper cites.
D. Wang and J. Chen, “Supervised speech separation based on deep learning: An overview,” IEEE/ACM Trans. Audio, Speech, Lang. Process. , vol. 26, no. 10, pp. 1702–1726, 2018
2018
Earlier work this paper cites.
Z. Rafii, A. Liutkus, F.-R. Stöter, S. I. Mimilakis, D. FitzGerald, and B. Pardo, “An overview of lead and accompaniment separation in music,” IEEE/ACM Trans. Audio, Speech, Lang. Process. , vol. 26, no. 8, pp. 1307–1335, 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
H. Zhao, C. Gan, A. Rouditchenko, C. Vondrick, J. McDermott, and A. Torralba, “The sound of pixels,” in Proc. Eur. Conf. Comput. Vis. (ECCV) , 2018, pp. 570–586
2018
Earlier work this paper cites.
E. Perez, F. Strub, H. De Vries, V. Dumoulin, and A. Courville, “FiLM: Visual reasoning with a general conditioning layer,” in Proc. AAAI Conf. Artif. Intell. , 2018, pp. 3942–3951
2018
Earlier work this paper cites.
E. Fonseca, M. Plakal, F. Font, D. Ellis, X. Favory, J. Pons, and X. Serra, “General-purpose tagging of freesound audio with audioset labels: Task description, dataset and baseline,” in Proc. Detect. Classif. Acoust. Scenes Events (DCASE) , 2018, pp. 69–73
2018
Earlier work this paper cites.
Y. Luo and N. Mesgarani, “Conv-TasNet: Surpassing ideal time–frequency magnitude masking for speech separation,” IEEE/ACM Trans. Audio, Speech, Lang. Process. , vol. 27, no. 8, pp. 1256–1266, 2019
2019
Earlier work this paper cites.
I. Kavalerov, S. Wisdom, H. Erdogan, B. Patton, K. Wilson, J. Le Roux, and J. R. Hershey, “Universal sound separation,” in IEEE Workshop on Appl. Signal Process. Audio Acoust. (WASPAA) , 2019, pp. 175–179
2019
Earlier work this paper cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in Proc. Conf. North Amer. Chap. Assoc. Comput. Linguist. - Human Lang. Tech. (NAACL-HLT) , 2019, pp. 4171–4186
2019
Earlier work this paper cites.
J. L. Roux, S. Wisdom, H. Erdogan, and J. R. Hershey, “SDR – half-baked or well done?” in Proc. IEEE Int. Conf. Acoustics Speech Signal Process. (ICASSP) , 2019, pp. 626–630
2019
Earlier work this paper cites.
C. D. Kim, B. Kim, H. Lee, and G. Kim, “AudioCaps: Generating captions for audios in the wild,” in Proc. Conf. North Amer. Chap. Assoc. Comput. Linguist. - Human Lang. Tech. (NAACL-HLT) , 2019, pp. 119–132
2019
Earlier work this paper cites.
D. S. Park, W. Chan, Y. Zhang, C.-C. Chiu, B. Zoph, E. D. Cubuk, and Q. V. Le, “SpecAugment: A simple data augmentation method for automatic speech recognition,” in Proc. INTERSPEECH , 2019, pp. 2613–2617
2019
Earlier work this paper cites.
I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” in Proc. Int. Conf. Learn. Represent. (ICLR) , 2019
2019
Cited alongside, same era.
S. Wisdom, E. Tzinis, H. Erdogan, R. Weiss, K. Wilson, and J. Hershey, “Unsupervised sound separation using mixture invariant training,” Proc. Adv. Neural Inf. Process. Syst. (NeurIPS) , vol. 33, pp. 3846–3857, 2020
2020
Cited alongside, same era.
T. Ochiai, M. Delcroix, Y. Koizumi, H. Ito, K. Kinoshita, and S. Araki, “Listen to what you want: Neural network-based universal sound selector,” in Proc. INTERSPEECH , 2020, pp. 1441–1445
2020
Cited alongside, same era.
A. Pandey and D. Wang, “Dense CNN with self-attention for time-domain speech enhancement,” IEEE/ACM Trans. Audio, Speech, Lang. Process. , vol. 29, pp. 1270–1279, 2021
2021
Cited alongside, same era.
B. Veluri, J. Chan, M. Itani, T. Chen, T. Yoshioka, and S. Gollakota, “Real-time target sound extraction,” in Proc. IEEE Int. Conf. Acoustics Speech Signal Process. (ICASSP) , 2023
2023
Later among the works it cites.
M. Delcroix, J. B. Vázquez, T. Ochiai, K. Kinoshita, Y. Ohishi, and S. Araki, “SoundBeam: Target sound extraction conditioned on sound-class labels and enrollment clues for increased performance and continuous learning,” IEEE/ACM Trans. Audio, Speech, Lang. Process. , vol. 31, pp. 121–136, 2023
2023
Later among the works it cites.
H.-W. Dong, N. Takahashi, Y. Mitsufuji, J. McAuley, and T. Berg-Kirkpatrick, “CLIPSep: Learning text-queried sound separation with noisy unlabeled videos,” in Proc. Int. Conf. Learn. Represent. (ICLR) , 2023
2023
Later among the works it cites.
2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Q. Kong, Y. Cao, H. Liu, K. Choi, and Y. Wang, “Decoupling magnitude and phase estimation with deep ResUNet for music source separation.” in Proc. Int. Soc. Music Inf. Retr. (ISMIR) , 2021, pp. 342–349
2021
Cited alongside, same era.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in Proc. Int. Conf. Mach. Learn. (ICML) , 2021, pp. 8748–8763
2021
Cited alongside, same era.
C. Subakan, M. Ravanelli, S. Cornell, M. Bronzi, and J. Zhong, “Attention is all you need in speech separation,” in Proc. IEEE Int. Conf. Acoustics Speech Signal Process. (ICASSP) , 2021, pp. 21–25
2021
Cited alongside, same era.
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin Transformer: Hierarchical vision transformer using shifted windows,” in Proc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV) , 2021, pp. 10 012–10 022
2021
Cited alongside, same era.
K. Kilgour, B. Gfeller, Q. Huang, A. Jansen, S. Wisdom, and M. Tagliasacchi, “Text-Driven Separation of Arbitrary Sounds,” in Proc. INTERSPEECH , 2022, pp. 5403–5407
2022
Cited alongside, same era.
K. Chen*, X. Du*, B. Zhu, Z. Ma, T. Berg-Kirkpatrick, and S. Dubnov, “Zero-shot audio source separation via query-based learning from weakly-labeled data,” in Proc. AAAI Conf. Artif. Intell. , 2022, pp. 4441–4449
2022
Cited alongside, same era.
X. Liu, H. Liu, Q. Kong, X. Mei, J. Zhao, Q. Huang, M. D. Plumbley, and W. Wang, “Separate what you describe: Language-queried audio source separation,” in Proc. INTERSPEECH , 2022, pp. 1801–1805
2022
Cited alongside, same era.
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen, “LoRA: Low-rank adaptation of large language models,” in Proc. Int. Conf. Learn. Represent. (ICLR) , 2022
2022
Cited alongside, same era.
Later among the works it cites.
Y. Wu, K. Chen, T. Zhang, Y. Hui, T. Berg-Kirkpatrick, and S. Dubnov, “Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,” in Proc. IEEE Int. Conf. Acoustics Speech Signal Process. (ICASSP) , 2023
2023
Later among the works it cites.
C. Li, Y. Qian, Z. Chen, D. Wang, T. Yoshioka, S. Liu, Y. Qian, and M. Zeng, “Target sound extraction with variable cross-modality clues,” in Proc. IEEE Int. Conf. Acoustics Speech Signal Process. (ICASSP) , 2023, pp. 1–5
2023
Later among the works it cites.
2023
Later among the works it cites.
H. Liu, Z. Chen, Y. Yuan, X. Mei, X. Liu, D. Mandic, W. Wang, and M. D. Plumbley, “AudioLDM: Text-to-audio generation with latent diffusion models,” Proc. Int. Conf. Mach. Learn. (ICML) , pp. 21 450–21 474, 2023
2023
Later among the works it cites.
H. Ma, Z. Peng, M. Shao, J. Li, and J. Liu, “Extending Whisper with prompt tuning to target-speaker ASR,” in Proc. IEEE Int. Conf. Acoustics Speech Signal Process. (ICASSP) , 2024, pp. 12 516–12 520
2024
Closest in time.
Y. Liu, X. Liu, Y. Zhao, Y. Wang, R. Xia, P. Tain, and Y. Wang, “Audio prompt tuning for universal sound separation,” in Proc. IEEE Int. Conf. Acoustics Speech Signal Process. (ICASSP) , 2024, pp. 1446–1450
2024
Closest in time.
Y. Wang, H. Chen, D. Yang, J. Yu, C. Weng, Z. Wu, and H. Meng, “Consistent and relevant: Rethink the query embedding in general sound separation,” in Proc. IEEE Int. Conf. Acoustics Speech Signal Process. (ICASSP) . IEEE, 2024, pp. 961–965
2024
Closest in time.
2024
Closest in time.
H. Liu, Y. Yuan, X. Liu, X. Mei, Q. Kong, Q. Tian, Y. Wang, W. Wang, Y. Wang, and M. D. Plumbley, “Audioldm 2: Learning holistic audio generation with self-supervised pretraining,” IEEE/ACM Tran. Audio, Speech, Lang. Process. , vol. 32, pp. 2871–2883, 2024
2024
Closest in time.
Y. Yuan, H. Liu, X. Liu, Q. Huang, M. D. Plumbley, and W. Wang, “Retrieval-augmented text-to-audio generation,” in Proc. IEEE Int. Conf. Acoustics Speech Signal Process. (ICASSP) . IEEE, 2024, pp. 581–585
2024
Closest in time.
J. Kim, J. Jung, J. Lee, and S. H. Woo, “Enclap: Combining neural audio codec and audio-text joint embedding for automated audio captioning,” in Proc. IEEE Int. Conf. Acoustics Speech Signal Process. (ICASSP) . IEEE, 2024, pp. 6735–6739
2024
Closest in time.
S. Deshmukh, B. Elizalde, D. Emmanouilidou, B. Raj, R. Singh, and H. Wang, “Training audio captioning models without audio,” pp. 371–375, 2024
2024
Closest in time.
Y. Zhang, E. Sui, and S. Yeung-Levy, “Connect, collapse, corrupt: Learning cross-modal tasks with uni-modal data,” in Proc. Int. Conf. Learn. Represent. (ICLR) , 2024
2024
Closest in time.