Fetching the paper…
Reading the bibliography…
There has been significant research on developing pretrained transformer architectures for multimodal-to-text generation tasks.
“Object hallucination in image captioning,”
Anna Rohrbach, Lisa Anne Hendricks, Kaylee Burns, Trevor Darrell, and Kate Saenko, · 2018
Earlier work this paper cites.
“Sentence-BERT: Sentence embeddings using Siamese BERT-networks,”
Nils Reimers and Iryna Gurevych, · 2019
Earlier work this paper cites.
“AudioCaps: Generating captions for audios in the wild,”
Chris Dongjoo Kim, Byeongchang Kim, Hyunmin Lee, and Gunhee Kim, · 2019
Earlier work this paper cites.
“Clotho: an audio captioning dataset,”
Konstantinos Drossos, Samuel Lipping, and Tuomas Virtanen, · 2020
Earlier work this paper cites.
“Best-First Beam Search,”
Clara Meister, Tim Vieira, and Ryan Cotterell, · 2020
Earlier work this paper cites.
“Improving the performance of automated audio captioning via integrating the acoustic and semantic information,”
Zhongjie Ye, Helin Wang, Dongchao Yang, and Yuexian Zou, · 2021
Earlier work this paper cites.
“Compressing Large-Scale Transformer-Based Models: A Case Study on BERT,”
Prakhar Ganesh, Yao Chen, Xin Lou, Mohammad Ali Khan, Yin Yang, Hassan Sajjad, Preslav Nakov, Deming Chen, and Marianne Winslett, · 2021
Earlier work this paper cites.
“Diversity and bias in audio captioning datasets,”
Irene Martin Morato and Annamaria Mesaros, · 2021
Cited alongside, same era.
“AST: Audio Spectrogram Transformer,”
Yuan Gong, Yu-An Chung, and James Glass, · 2021
Cited alongside, same era.
“An encoder-decoder based audio captioning system with transfer and reinforcement learning,”
Xinhao Mei, Qiushi Huang, Xubo Liu, Gengyun Chen, Jingqian Wu, Yusong Wu, Jinzheng Zhao, Shengchen Li, Tom Ko, H Lilian Tang, Xingkun Shao, MarkD . Plumbley, and Wenwu Wang, · 2021
Cited alongside, same era.
“iCNN-Transformer: An improved CNN-Transformer with Channel-spatial Attention and Keyword Prediction for Automated Audio Captioning,”
Kun Chen, Jun Wang, Feng Deng, and Xiaorui Wang, · 2022
Cited alongside, same era.
“Correcting diverse factual errors in abstractive summarization via post-editing and language model infilling,”
Vidhisha Balachandran, Hannaneh Hajishirzi, William Cohen, and Yulia Tsvetkov, · 2022
Cited alongside, same era.
“Thinking hallucination for video captioning,”
Nasib Ullah and Partha Pratim Mohanta, · 2022
Later among the works it cites.
“Survey of hallucination in natural language generation,”
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung, · 2023
Closest in time.
Xinhao Mei, Chutong Meng, Haohe Liu, Qiuqiang Kong, Tom Ko, Chengqi Zhao, Mark D. Plumbley, Yuexian Zou, and Wenwu Wang, · 2023
Closest in time.
“Plausible may not be faithful: Probing object hallucination in vision-language pre-training,”
Wenliang Dai, Zihan Liu, Ziwei Ji, Dan Su, and Pascale Fung, · 2023
Closest in time.
“Judging llm-as-a-judge with mt-bench and chatbot arena,” 2023
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric. P Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica, · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Improved beam search for hallucination mitigation in abstractive summarization,”
Arvind Krishna Sridhar and Erik Visser, · 2022
Cited alongside, same era.
“Don’t say what you don’t know: Improving the consistency of abstractive summarization by constraining beam search,”
Daniel King, Zejiang Shen, Nishant Subramani, Daniel S. Weld, Iz Beltagy, and Doug Downey, · 2022
Cited alongside, same era.
“Llama 2: Open foundation and fine-tuned chat models,”
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al., · 2023
Closest in time.
“Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,”
Yusong Wu*, Ke Chen*, Tianyu Zhang*, Yuchen Hui*, Taylor Berg-Kirkpatrick, and Shlomo Dubnov, · 2023
Closest in time.