Fetching the paper…
Reading the bibliography…
The goal of text-queried target sound extraction (TSE) is to extract from a mixture a sound source specified with a natural-language caption.
K. J. Piczak, “ESC: Dataset for environmental sound classification,” in Proc. ACM Multimedia , 2015, pp. 1015–1018
2015
Earlier work this paper cites.
J. R. Hershey, Z. Chen, J. Le Roux, and S. Watanabe, “Deep clustering: Discriminative embeddings for segmentation and separation,” in Proc. ICASSP , 2016, pp. 31–35
2016
Earlier work this paper cites.
D. Yu, M. Kolbæk, Z. H. Tan, and J. Jensen, “Permutation invariant training of deep models for speaker-independent multi-talker speech separation,” in Proc. ICASSP , 2017, pp. 241–245
2017
Earlier work this paper cites.
J. F. Gemmeke, D. P. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio Set: An ontology and human-labeled dataset for audio events,” in Proc. ICASSP , 2017, pp. 776–780
2017
Earlier work this paper cites.
H. Zhao, C. Gan, A. Rouditchenko, C. Vondrick, J. McDermott, and A. Torralba, “The sound of pixels,” in Proc. ECCV , 2018, pp. 570–586
2018
Earlier work this paper cites.
E. Perez, F. Strub, H. De Vries, V. Dumoulin, and A. Courville, “FiLM: Visual reasoning with a general conditioning layer,” in Proc. AAAI , vol. 32, no. 1, 2018
2018
Earlier work this paper cites.
I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” in Proc. ICLR , 2018
2018
Earlier work this paper cites.
Y. Luo and N. Mesgarani, “Conv-TasNet: Surpassing ideal time-frequency magnitude masking for speech separation,” IEEE/ACM Trans. Audio, Speech, Lang. Process. , vol. 27, no. 8, pp. 1256–1266, 2019
2019
Earlier work this paper cites.
Q. Wang, H. Muckenhirn, K. Wilson, P. Sridhar, Z. Wu, J. R. Hershey, R. A. Saurous, R. J. Weiss, Y. Jia, and I. L. Moreno, “VoiceFilter: Targeted voice separation by speaker-conditioned spectrogram masking,” in Proc. Interspeech , 2019, pp. 2728–2732
2019
Earlier work this paper cites.
K. Žmolíková, M. Delcroix, K. Kinoshita, T. Ochiai, T. Nakatani, L. Burget, and J. Černockỳ, “Speakerbeam: Speaker aware neural network for target speaker extraction in speech mixtures,” IEEE Journal of Selected Topics in Signal Processing , vol. 13, no. 4, pp. 800–814, 2019
2019
Earlier work this paper cites.
C. D. Kim, B. Kim, H. Lee, and G. Kim, “AudioCaps: Generating captions for audios in the wild,” in Proc. NAACL , 2019, pp. 119–132
2019
Earlier work this paper cites.
J. Le Roux, S. Wisdom, H. Erdogan, and J. R. Hershey, “SDR–half-baked or well done?” in Proc. ICASSP , 2019, pp. 626–630
2019
Earlier work this paper cites.
Y. Luo, Z. Chen, and T. Yoshioka, “Dual-Path RNN: Efficient long sequence modeling for time-domain single-channel speech separation,” in Proc. ICASSP , 2020
2020
Cited alongside, same era.
T. Ochiai, M. Delcroix, Y. Koizumi, H. Ito, K. Kinoshita, and S. Araki, “Listen to what you want: Neural network-based universal sound selector,” in Proc. Interspeech , 2020, pp. 1441–1445
2020
Cited alongside, same era.
H. Chen, W. Xie, A. Vedaldi, and A. Zisserman, “VGGSound: A large-scale audio-visual dataset,” in Proc. ICASSP , 2020, pp. 721–725
2020
Cited alongside, same era.
K. Drossos, S. Lipping, and T. Virtanen, “Clotho: An audio captioning dataset,” in Proc. ICASSP , 2020, pp. 736–740
2020
Cited alongside, same era.
A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu et al. , “Conformer: Convolution-augmented transformer for speech recognition,” in Proc. Interspeech , 2020, pp. 5036–5040
2023
Later among the works it cites.
H.-W. Dong, N. Takahashi, Y. Mitsufuji, J. McAuley, and T. Berg-Kirkpatrick, “CLIPSep: Learning text-queried sound separation with noisy unlabeled videos,” in Proc. ICLR , 2023
2023
Later among the works it cites.
B. Elizalde, S. Deshmukh, M. Al Ismail, and H. Wang, “CLAP: Learning audio concepts from natural language supervision,” in Proc. ICASSP , 2023, pp. 1–5
2023
Later among the works it cites.
Y. Wu, K. Chen, T. Zhang, Y. Hui, T. Berg-Kirkpatrick, and S. Dubnov, “Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,” in Proc. ICASSP , 2023, pp. 1–5
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
C. Subakan, M. Ravanelli, S. Cornell, M. Bronzi, and J. Zhong, “Attention is all you need in speech separation,” in Proc. ICASSP , 2021
2021
Cited alongside, same era.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in Proc. ICML , 2021, pp. 8748–8763
2021
Cited alongside, same era.
F. Wang and H. Liu, “Understanding the behaviour of contrastive loss,” in Proc. CVPR , 2021, pp. 2495–2504
2021
Cited alongside, same era.
X. Liu, H. Liu, Q. Kong, X. Mei, J. Zhao, Q. Huang, M. D. Plumbley, and W. Wang, “Separate what you describe: Language-queried audio source separation,” in Proc. Interspeech , 2022, pp. 1801–1805
2022
Cited alongside, same era.
K. Kilgour, B. Gfeller, Q. Huang, A. Jansen, S. Wisdom, and M. Tagliasacchi, “Text-driven separation of arbitrary sounds,” in Proc. Interspeech , 2022, pp. 5403–5407
2022
Cited alongside, same era.
V. W. Liang, Y. Zhang, Y. Kwon, S. Yeung, and J. Y. Zou, “Mind the gap: Understanding the modality gap in multi-modal contrastive representation learning,” in Proc. NeurIPS , vol. 35, 2022, pp. 17 612–17 625
2022
Cited alongside, same era.
2022
Cited alongside, same era.
H. Liu, Z. Chen, Y. Yuan, X. Mei, X. Liu, D. Mandic, W. Wang, and M. D. Plumbley, “AudioLDM: text-to-audio generation with latent diffusion models,” in Proc. ICML , 2023, pp. 21 450–21 474
2023
Later among the works it cites.
X. Liu, Q. Kong, Y. Zhao, H. Liu, Y. Yuan, Y. Liu, R. Xia, Y. Wang, M. D. Plumbley, and W. Wang, “Separate anything you describe,” 2023. [Online]. Available: https://github.com/Audio-AGI/AudioSep
2023
Later among the works it cites.
K. Saijo and T. Ogawa, “Remixing-based unsupervised source separation from scratch,” in Proc. Interspeech , 2023, pp. 1678–1682
2023
Later among the works it cites.
2024
Closest in time.
2024
Closest in time.
B. Elizalde, S. Deshmukh, and H. Wang, “Natural language supervision for general-purpose audio representations,” in Proc. ICASSP , 2024, pp. 336–340
2024
Closest in time.
S. Deshmukh, B. Elizalde, D. Emmanouilidou, B. Raj, R. Singh, and H. Wang, “Training audio captioning models without audio,” in ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2024, pp. 371–375
2024
Closest in time.