Fetching the paper…
Reading the bibliography…
We present RECAP (REtrieval-Augmented Audio CAPtioning), a novel and effective audio captioning system that generates captions conditioned on an input audio and other captions similar to the audio retrieved from a datastore.
“Semantic parsing on freebase from question-answer pairs,”
Jonathan Berant, Andrew Chou, Roy Frostig, and Percy Liang, · 2013
Earlier work this paper cites.
“Audiocaps: Generating captions for audios in the wild,”
Chris Dongjoo Kim, Byeongchang Kim, Hyunmin Lee, and Gunhee Kim, · 2019
Earlier work this paper cites.
“Natural questions: a benchmark for question answering research,”
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al., · 2019
Earlier work this paper cites.
“Clotho: An audio captioning dataset,”
Konstantinos Drossos, Samuel Lipping, and Tuomas Virtanen, · 2020
Earlier work this paper cites.
Yuma Koizumi, Yasunori Ohishi, Daisuke Niizumi, Daiki Takeuchi, and Masahiro Yasuda, · 2020
Earlier work this paper cites.
“Retrieval-augmented generation for knowledge-intensive nlp tasks,”
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al., · 2020
Earlier work this paper cites.
“Audio captioning based on combined audio and semantic embeddings,”
Ayşegül Özkaya Eren and Mustafa Sert, · 2020
Earlier work this paper cites.
“Audio captioning based on transformer and pre-trained cnn.,”
Kun Chen, Yusong Wu, Ziyue Wang, Xuan Zhang, Fudong Nian, Shengchen Li, and Xi Shao, · 2020
Cited alongside, same era.
“BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension,”
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer, · 2020
Cited alongside, same era.
“Automated audio captioning by fine-tuning bart with audioset tags,”
Félix Gontier, Romain Serizel, and Christophe Cerisara, · 2021
Cited alongside, same era.
“Investigating local and global information for automated audio captioning with transfer learning,”
Xuenan Xu, Heinrich Dinkel, Mengyue Wu, Zeyu Xie, and Kai Yu, · 2021
Cited alongside, same era.
“Audio retrieval with natural language queries: A benchmark study,”
A Sophia Koepke, Andreea-Maria Oncescu, Joao Henriques, Zeynep Akata, and Samuel Albanie, · 2022
Cited alongside, same era.
“A survey on retrieval-augmented text generation,”
Huayang Li, Yixuan Su, Deng Cai, Yan Wang, and Lemao Liu, · 2022
Later among the works it cites.
“Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,”
Yusong Wu, Ke Chen, Tianyu Zhang, Yuchen Hui, Taylor Berg-Kirkpatrick, and Shlomo Dubnov, · 2023
Closest in time.
“Prefix tuning for automated audio captioning,”
Minkyu Kim, Kim Sung-Bin, and Tae-Hyun Oh, · 2023
Closest in time.
“A whisper transformer for audio captioning trained with synthetic captions and transfer learning,”
Marek Kadlčík, Adam Hájek, Jürgen Kieslich, and Radosław Winiecki, · 2023
Closest in time.
“Retrieval-augmented image captioning,”
Rita Ramos, Desmond Elliott, and Bruno Martins, · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Automated audio captioning using transfer learning and reconstruction latent space similarity regularization,”
Andrew Koh, Xue Fuzhao, and Chng Eng Siong, · 2022
Cited alongside, same era.
“Few-shot learning with retrieval augmented language models,”
Izacard et al., · 2022
Cited alongside, same era.
“Beats-based audio captioning model with instructor embedding supervision and chatgpt mix-up,”
Wu et al.,
Cited in the paper.
“Leveraging multi-task training and image retrieval with clap for audio captioning,”
Haoran Sun, Zhiyong Yan, Yongqing Wang, Heinrich Dinkel, Junbo Zhang, and Yujun Wang,
Cited in the paper.
“Label-refined sequential training with noisy data for automated audio captioning,”
Jaeheon Sim, Eungbeom Kim, and Kyogu Lee,
Cited in the paper.
“Audio captioning transformer,”
Xinhao Mei, Xubo Liu, Qiushi Huang, Mark D. Plumbley, and Wenwu Wang,
Cited in the paper.
Closest in time.
“Smallcap: Lightweight image captioning prompted with retrieval augmentation,”
Rita Ramos, Bruno Martins, Desmond Elliott, and Yova Kementchedjhieva, · 2023
Closest in time.