Fetching the paper…
Reading the bibliography…
Contrastive language-audio pretraining~(CLAP) has been developed to align the representations of audio and language, achieving remarkable performance in retrieval and classification tasks.
“A dataset and taxonomy for urban sound research,”
J. Salamon, C. Jacoby, and J. P. Bello, · 2014
Earlier work this paper cites.
“ESC: dataset for environmental sound classification,”
Karol J Piczak, · 2015
Earlier work this paper cites.
“Roberta: A robustly optimized bert pretraining approach,”
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov, · 2019
Earlier work this paper cites.
“AudioCaps: Generating captions for audios in the wild,”
Chris Dongjoo Kim, Byeongchang Kim, Hyunmin Lee, and Gunhee Kim, · 2019
Earlier work this paper cites.
“Clotho: An audio captioning dataset,”
Konstantinos Drossos, Samuel Lipping, and Tuomas Virtanen, · 2020
Earlier work this paper cites.
“VGGSound: A large-scale audio-visual dataset,”
Honglie Chen, Weidi Xie, Andrea Vedaldi, and Andrew Zisserman, · 2020
Earlier work this paper cites.
“Learning transferable visual models from natural language supervision,”
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al., · 2021
Earlier work this paper cites.
“CLIPcap: CLIP prefix for image captioning,”
Ron Mokady, Amir Hertz, and Amit H Bermano, · 2021
Earlier work this paper cites.
“Zero-shot image-to-text generation for visual-semantic arithmetic,”
Yoad Tewel, Yoav Shalev, Idan Schwartz, and Lior Wolf, · 2021
Earlier work this paper cites.
“Masked language modeling and the distributional hypothesis: Order word matters pre-training for little,”
Koustuv Sinha, Robin Jia, Dieuwke Hupkes, Joelle Pineau, Adina Williams, and Douwe Kiela, · 2021
Earlier work this paper cites.
“Audio retrieval with natural language queries,”
Andreea Maria Oncescu, A Koepke, João F Henriques, Zeynep Akata, and Samuel Albanie, · 2021
Earlier work this paper cites.
“Conditional prompt learning for vision-language models,”
Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu, · 2022
Earlier work this paper cites.
“Learning to prompt for vision-language models,”
Kaiyang Zhou, Jingkang Yang, Chenchange Loy, and Ziwei Liu, · 2022
Cited alongside, same era.
“Effective conditioned and composed image retrieval combining clip-based features,”
Alberto Baldrati, Marco Bertini, Tiberio Uricchio, and Alberto Del Bimbo, · 2022
Cited alongside, same era.
“EDIFFI: Text-to-image diffusion models with an ensemble of expert denoisers,”
Yogesh Balaji, Seungjun Nah, Xun Huang, Arash Vahdat, Jiaming Song, Karsten Kreis, Miika Aittala, Timo Aila, Samuli Laine, Bryan Catanzaro, et al., · 2022
Cited alongside, same era.
“GLIDE: Towards photorealistic image generation and editing with text-guided diffusion models,”
Alexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob Mcgrew, Ilya Sutskever, and Mark Chen, · 2022
Cited alongside, same era.
“Photorealistic text-to-image diffusion models with deep language understanding,”
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, et al., · 2022
“Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,”
Yusong Wu, Ke Chen, Tianyu Zhang, Yuchen Hui, Taylor Berg-Kirkpatrick, and Shlomo Dubnov, · 2023
Later among the works it cites.
“AudioLDM: Text-to-Audio generation with latent diffusion models,”
Haohe Liu, Zehua Chen, Yi Yuan, Xinhao Mei, Xubo Liu, Danilo Mandic, Wenwu Wang, and Mark D. Plumbley, · 2023
Later among the works it cites.
“Retrieval-augmented text-to-audio generation,”
Yi Yuan, Haohe Liu, Xubo Liu, Qiushi Huang, Mark D Plumbley, and Wenwu Wang, · 2023
Later among the works it cites.
“CREPE: Can vision-language foundation models reason compositionally?,”
Zixian Ma, Jerry Hong, Mustafa Omer Gul, Mona Gandhi, Irena Gao, and Ranjay Krishna, · 2023
Later among the works it cites.
“Enhance Temporal Relations in Audio Captioning with Sound Event Detection,”
Zeyu Xie, Xuenan Xu, Mengyue Wu, and Kai Yu, · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“HTS-AT: A hierarchical token-semantic audio transformer for sound classification and detection,”
Ke Chen, Xingjian Du, Bilei Zhu, Zejun Ma, Taylor Berg-Kirkpatrick, and Shlomo Dubnov, · 2022
Cited alongside, same era.
“Audio retrieval with Wavtext5K and CLAP training,”
Soham Deshmukh, Benjamin Elizalde, and Huaming Wang, · 2022
Cited alongside, same era.
“On metric learning for audio-text cross-modal retrieval,”
Xinhao Mei, Xubo Liu, Jianyuan Sun, Mark D Plumbley, and Wenwu Wang, · 2022
Cited alongside, same era.
“Wav2clip: Learning robust audio representations from clip,”
Ho Hsiang Wu, Prem Seetharaman, Kundan Kumar, and Juan Pablo Bello, · 2022
Cited alongside, same era.
“AudioCLIP: Extending CLIP to image, text and audio,”
Andrey Guzhov, Federico Raue, Jörn Hees, and Andreas Dengel, · 2022
Cited alongside, same era.
“Clip for all things zero-shot sketch-based image retrieval, fine-grained or not,”
Aneeshan Sain, Ayan Kumar Bhunia, Pinaki Nath Chowdhury, Subhadeep Koley, Tao Xiang, and Yi-Zhe Song, · 2023
Cited alongside, same era.
“CLAP learning audio concepts from natural language supervision,”
Benjamin Elizalde, Soham Deshmukh, Mahmoud Al Ismail, and Huaming Wang, · 2023
Later among the works it cites.
“Improving text-audio retrieval by text-aware attention pooling and prior matrix revised loss,”
Yifei Xin, Dongchao Yang, and Yuexian Zou, · 2023
Later among the works it cites.
“Teaching CLIP to count to ten,”
Roni Paiss, Ariel Ephrat, Omer Tov, Shiran Zada, Inbar Mosseri, Michal Irani, and Tali Dekel, · 2023
Later among the works it cites.
“A large-scale dataset for audio-language representation learning,”
Luoyi Sun, Xuenan Xu, Mengyue Wu, and Weidi Xie, · 2023
Later among the works it cites.
“Audio-text models do not yet leverage natural language,”
Ho Hsiang Wu, Oriol Nieto, Juan Pablo Bello, and Justin Salomon, · 2023
Later among the works it cites.
“Blat: Bootstrapping language-audio pre-training based on audioset tag-guided synthetic data,”
Xuenan Xu, Zhiling Zhang, Zelin Zhou, Pingyue Zhang, Zeyu Xie, Mengyue Wu, and Kenny Q Zhu, · 2023
Later among the works it cites.