Fetching the paper…
Reading the bibliography…
Multi-modal contrastive learning techniques in the audio-text domain have quickly become a highly active area of research.
“A dataset and taxonomy for urban sound research,”
Justin Salamon, Christopher Jacoby, and Juan Pablo Bello, · 2014
Earlier work this paper cites.
“Esc: Dataset for environmental sound classification,”
Karol J Piczak, · 2015
Earlier work this paper cites.
“Representation learning with contrastive predictive coding,”
Aaron van den Oord, Yazhe Li, and Oriol Vinyals, · 2018
Earlier work this paper cites.
“Localizing moments in video with temporal language,”
Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan Russell, · 2018
Earlier work this paper cites.
“Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks,”
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee, · 2019
Earlier work this paper cites.
“Lxmert: Learning cross-modality encoder representations from transformers,”
Hao Tan and Mohit Bansal, · 2019
Earlier work this paper cites.
“Bert: Pre-training of deep bidirectional transformers for language understanding,”
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova, · 2019
Earlier work this paper cites.
“Audiocaps: Generating captions for audios in the wild,”
Chris Dongjoo Kim, Byeongchang Kim, Hyunmin Lee, and Gunhee Kim, · 2019
Earlier work this paper cites.
“Roberta: A robustly optimized bert pretraining approach,”
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov, · 2019
Earlier work this paper cites.
“Sound event detection in domestic environments with weakly labeled data and soundscape synthesis,”
Nicolas Turpault, Romain Serizel, Ankit Parag Shah, and Justin Salamon, · 2019
Earlier work this paper cites.
“Panns: Large-scale pretrained audio neural networks for audio pattern recognition,”
Qiuqiang Kong, Yin Cao, Turab Iqbal, Yuxuan Wang, Wenwu Wang, and Mark D Plumbley, · 2020
Earlier work this paper cites.
“Clotho: An audio captioning dataset,”
Konstantinos Drossos, Samuel Lipping, and Tuomas Virtanen, · 2020
Cited alongside, same era.
“Beyond accuracy: Behavioral testing of nlp models with checklist,”
Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh, · 2020
Cited alongside, same era.
“Learning transferable visual models from natural language supervision,”
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, et al., · 2021
Cited alongside, same era.
“Localizing visual sounds the hard way,”
Honglie Chen, Weidi Xie, Triantafyllos Afouras, Arsha Nagrani, Andrea Vedaldi, and Andrew Zisserman, · 2021
Cited alongside, same era.
“Out of order: How important is the sequential order of words in a sentence in natural language understanding tasks?,”
Thang M Pham, Trung Bui, Long Mai, and Anh Nguyen, · 2021
Cited alongside, same era.
“Audioclip: Extending clip to image, text and audio,”
Andrey Guzhov, Federico Raue, Jörn Hees, and Andreas Dengel, · 2022
Later among the works it cites.
“Language-based audio retrieval task in dcase 2022 challenge,”
Huang Xie, Samuel Lipping, and Tuomas Virtanen, · 2022
Later among the works it cites.
“Contrastive audio-language learning for music,”
Ilaria Manco, Emmanouil Benetos, Elio Quinton, and György Fazekas, · 2022
Later among the works it cites.
“Connecting the dots between audio and text without parallel data through visual knowledge transfer,”
Yanpeng Zhao, Jack Hessel, Youngjae Yu, Ximing Lu, Rowan Zellers, and Yejin Choi, · 2022
Later among the works it cites.
“Audio-text retrieval in context,”
Siyu Lou, Xuenan Xu, Mengyue Wu, and Kai Yu, · 2022
Later among the works it cites.
“On metric learning for audio-text cross-modal retrieval,”
Xinhao Mei, Xubo Liu, Jianyuan Sun, Mark D Plumbley, and Wenwu Wang, · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Irene Martín-Morató and Annamaria Mesaros, · 2021
Cited alongside, same era.
“Fsd50k: an open dataset of human-labeled sound events,”
Eduardo Fonseca, Xavier Favory, Jordi Pons, Frederic Font, and Xavier Serra, · 2021
Cited alongside, same era.
“Lit: Zero-shot transfer with locked-image text tuning,”
Xiaohua Zhai, Xiao Wang, Basil Mustafa, Andreas Steiner, Daniel Keysers, Alexander Kolesnikov, and Lucas Beyer, · 2022
Cited alongside, same era.
“How to listen? rethinking visual sound localization,”
Ho-Hsiang Wu, Magdalena Fuentes, Prem Seetharaman, and Juan Pablo Bello, · 2022
Cited alongside, same era.
“It’s time for artistic correspondence in music and video,”
Dídac Surís, Carl Vondrick, Bryan Russell, and Justin Salamon, · 2022
Cited alongside, same era.
“Wav2clip: Learning robust audio representations from clip,”
Ho-Hsiang Wu, Prem Seetharaman, Kundan Kumar, and Juan Pablo Bello, · 2022
Cited alongside, same era.
Later among the works it cites.
“Audio retrieval with natural language queries: A benchmark study,”
A Sophia Koepke, Andreea-Maria Oncescu, Joao Henriques, Zeynep Akata, and Samuel Albanie, · 2022
Later among the works it cites.
“Clap: Learning audio concepts from natural language supervision,”
Benjamin Elizalde, Soham Deshmukh, Mahmoud Al Ismail, and Huaming Wang, · 2022
Later among the works it cites.
“The SJTU system for DCASE2022 challenge task 6: Audio captioning with audio-text retrieval pre-training,”
Xuenan Xu, Zeyu Xie, Mengyue Wu, and Kai Yu, · 2022
Later among the works it cites.
“Language-based audio retrieval with pre-trained models,”
Xinhao Mei, Xubo Liu, Haohe Liu, Jianyuan Sun, Mark D. Plumbley, and Wenwu Wang, · 2022
Later among the works it cites.