Fetching the paper…
Reading the bibliography…
Zero-shot audio captioning aims at automatically generating descriptive textual captions for audio content without prior training for this task.
Bleu: a method and for automatic and evaluation of machine and translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Semantic-audio retrieval
Malcolm Slaney · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
S. Banerjee and A. Lavie · 2005
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
Ramakrishna Vedantam, C. L. Zitnick, and D. Parikh · 2015
Earlier work this paper cites.
Spice: Semantic propositional image caption evaluation
Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould · 2016
Earlier work this paper cites.
Automated audio captioning with recurrent neural networks
Konstantinos Drossos, Sharath Adavanne, and Tuomas Virtanen · 2017
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events
Jort F. Gemmeke, Daniel P. W. Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R. Channing Moore, Manoj Plakal, and Marvin Ritter · 2017
Earlier work this paper cites.
Improved image captioning via policy gradient optimization of spider
Siqi Liu, Zhenhai Zhu, Ning Ye, Sergio Guadarrama, and Kevin Murphy · 2017
Earlier work this paper cites.
Audiocaps: Generating captions for audios in the wild
Chris Dongjoo Kim, Byeongchang Kim, Hyunmin Lee, and Gunhee Kim · 2019
Earlier work this paper cites.
Audio caption: Listen and tell
Mengyue Wu, Heinrich Dinkel, and Kai Yu · 2019
Earlier work this paper cites.
Clotho: An audio captioning dataset
Konstantinos Drossos, Samuel Lipping, and Tuomas Virtanen · 2020
Earlier work this paper cites.
Audio captioning based on combined audio and semantic embeddings
Ayşegül Özkaya Eren and Mustafa Sert · 2020
Cited alongside, same era.
Yuma Koizumi, Yasunori Ohishi, Daisuke Niizumi, Daiki Takeuchi, and Masahiro Yasuda · 2020
Cited alongside, same era.
Automated audio captioning by fine-tuning bart with audioset tags
Félix Gontier, Romain Serizel, and Christophe Cerisara · 2021
Cited alongside, same era.
Audioclip: Extending clip to image, text and audio
Andrey Guzhov, Federico Raue, Jörn Hees, and Andreas R. Dengel · 2021
Cited alongside, same era.
Cl4ac: A contrastive loss for audio captioning
Xubo Liu, Qiushi Huang, Xinhao Mei, Tom Ko, H Lilian Tang, Mark D Plumbley, and Wenwu Wang · 2021
Cited alongside, same era.
Yusong Wu, K. Chen, Tianyu Zhang, Yuchen Hui, Taylor Berg-Kirkpatrick, and Shlomo Dubnov · 2022
Later among the works it cites.
DCASE2022 challenge task 6: Automated audio captioning, 2022
Huang Xie, Felix Gontier, Samuel Lipping, Konstantinos Drossos, Tuomas Virtanen, and Romain Serizel · 2022
Later among the works it cites.
Opt: Open pre-trained transformer language models
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al · 2022
Later among the works it cites.
Hyu submission for the dcase 2023 task 6a: automated audio captioning model using al-mixgen and synonyms substitution
Jae-Heung Cho, Yoon-Ah Park, Jaewon Kim, and Joon-Hyuk Chang · 2023
Closest in time.
Imagebind one embedding space to bind them all
Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, and Ishan Misra · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xinhao Mei, Xubo Liu, Qiushi Huang, Mark D Plumbley, and Wenwu Wang · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Cited alongside, same era.
Investigating local and global information for automated audio captioning with transfer learning
Xuenan Xu, Heinrich Dinkel, Mengyue Wu, Zeyu Xie, and Kai Yu · 2021
Cited alongside, same era.
Audio retrieval with natural language queries: A benchmark study
A. Sophia Koepke, Andreea-Maria Oncescu, Joao Henriques, Zeynep Akata, and Samuel Albanie · 2022
Cited alongside, same era.
Audio-text retrieval in context
Siyu Lou, Xuenan Xu, Mengyue Wu, and Kai Yu · 2022
Cited alongside, same era.
Language models can see: Plugging visual controls in text generation
Yixuan Su, Tian Lan, Yahui Liu, Fangyu Liu, Dani Yogatama, Yan Wang, Lingpeng Kong, and Nigel Collier · 2022
Cited alongside, same era.
Zero-shot image captioning by anchor-augmented vision-language space alignment
Junyan Wang, Yi Zhang, Ming Yan, Ji Chao Zhang, and Jitao Sang · 2022
Cited alongside, same era.
Closest in time.
Exploring train and test-time augmentations for audio-language learning
Eungbeom Kim, Jinhee Kim, Yoori Oh, Kyungsu Kim, Minju Park, Jaeheon Sim, Jinwoo Lee, and Kyogu Lee · 2023
Closest in time.
Decap: Decoding CLIP latents for zero-shot captioning via text-only training
Wei Li, Linchao Zhu, Longyin Wen, and Yi Yang · 2023
Closest in time.
Zero-shot translation of attention patterns in vqa models to natural language
Leonard Salewski, A. Sophia Koepke, Hendrik P.A. Lensch, and Zeynep Akata · 2023
Closest in time.
Zero-shot audio captioning via audibility guidance
Tal Shaharabany, Ariel Shaulov, and Lior Wolf · 2023
Closest in time.
Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation
Yusong Wu, Ke Chen, Tianyu Zhang, Yuchen Hui, Taylor Berg-Kirkpatrick, and Shlomo Dubnov · 2023
Closest in time.
Socratic models: Composing zero-shot multimodal reasoning with language
Andy Zeng, Maria Attarian, brian ichter, Krzysztof Marcin Choromanski, Adrian Wong, Stefan Welker, Federico Tombari, Aveek Purohit, Michael S Ryoo, Vikas Sindhwani, Johnny Lee, Vincent Vanhoucke, and Pete Florence · 2023
Closest in time.