Fetching the paper…
Reading the bibliography…
The task of audio captioning is similar in essence to tasks such as image and video captioning.
“Bleu: a method for automatic evaluation of machine translation,”
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu, · 2002
Earlier work this paper cites.
“Rouge: A package for automatic evaluation of summaries,”
Chin-Yew Lin, · 2004
Earlier work this paper cites.
“Meteor: An automatic metric for mt evaluation with improved correlation with human judgments,”
Satanjeev Banerjee and Alon Lavie, · 2005
Earlier work this paper cites.
“Cider: Consensus-based image description evaluation,”
Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh, · 2015
Earlier work this paper cites.
“Representation learning with contrastive predictive coding,”
Aaron van den Oord, Yazhe Li, and Oriol Vinyals, · 2018
Earlier work this paper cites.
“Bert: Pre-training of deep bidirectional transformers for language understanding,”
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova, · 2018
Earlier work this paper cites.
“Why is there so much more research on vision than on any other sensory modality?,”
Fabian Hutmacher, · 2019
Earlier work this paper cites.
“Audiocaps: Generating captions for audios in the wild,”
Chris Dongjoo Kim, Byeongchang Kim, Hyunmin Lee, and Gunhee Kim, · 2019
Earlier work this paper cites.
“Language models are unsupervised multitask learners,”
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al., · 2019
Earlier work this paper cites.
“Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter,”
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf, · 2019
Cited alongside, same era.
“Plug and play language models: A simple approach to controlled text generation,”
Sumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung, Eric Frank, Piero Molino, Jason Yosinski, and Rosanne Liu, · 2020
Cited alongside, same era.
“Audio captioning transformer,”
Xinhao Mei, Xubo Liu, Qiushi Huang, Mark D Plumbley, and Wenwu Wang, · 2021
Cited alongside, same era.
“Learning transferable visual models from natural language supervision,”
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al., · 2021
Cited alongside, same era.
“Diffusion models beat gans on image synthesis,”
Prafulla Dhariwal and Alexander Nichol, · 2021
Cited alongside, same era.
“Robust speech recognition via large-scale weak supervision,”
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever, · 2023
Closest in time.
“Valor: Vision-audio-language omni-perception pretraining model and dataset,”
Sihan Chen, Xingjian He, Longteng Guo, Xinxin Zhu, Weining Wang, Jinhui Tang, and Jing Liu, · 2023
Closest in time.
“Pengi: An audio language model for audio tasks,”
Soham Deshmukh, Benjamin Elizalde, Rita Singh, and Huaming Wang, · 2023
Closest in time.
“Disclip: Open-vocabulary referring expression generation,”
Lior Bracha, Eitan Shaar, Aviv Shamsian, Ethan Fetaya, and Gal Chechik, · 2023
Closest in time.
“Zero-shot video captioning with evolving pseudo-tokens,”
Yoad Tewel, Yoav Shalev, Roy Nadler, Idan Schwartz, and Lior Wolf, · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Leveraging pre-trained bert for audio captioning,”
Xubo Liu, Xinhao Mei, Qiushi Huang, Jianyuan Sun, Jinzheng Zhao, Haohe Liu, Mark D Plumbley, Volkan Kilic, and Wenwu Wang, · 2022
Cited alongside, same era.
“Language models can see: Plugging visual controls in text generation,”
Yixuan Su, Tian Lan, Yahui Liu, Fangyu Liu, Dani Yogatama, Yan Wang, Lingpeng Kong, and Nigel Collier, · 2022
Cited alongside, same era.
“Zerocap: Zero-shot image-to-text generation for visual-semantic arithmetic,”
Yoad Tewel, Yoav Shalev, Idan Schwartz, and Lior Wolf, · 2022
Cited alongside, same era.
“A whisper transformer for audio captioning trained with synthetic captions and transfer learning,”
Marek Kadlčík, Adam Hájek, Jürgen Kieslich, and Radosław Winiecki, · 2023
Cited alongside, same era.
Changhao Shi, Haomiao Ni, Kai Li, Shaobo Han, Mingfu Liang, and Martin Renqiang Min, · 2023
Closest in time.
“Diffusion self-guidance for controllable image generation,”
Dave Epstein, Allan Jabri, Ben Poole, Alexei A Efros, and Aleksander Holynski, · 2023
Closest in time.
“Imagebind: One embedding space to bind them all,”
Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, and Ishan Misra, · 2023
Closest in time.