Fetching the paper…
Reading the bibliography…
Image Captioning is a traditional vision-and-language task that aims to generate the language description of an image.
“Bleu: a method for automatic evaluation of machine translation,”
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu, · 2002
Earlier work this paper cites.
“METEOR: An automatic metric for MT evaluation with improved correlation with human judgments,”
Satanjeev Banerjee and Alon Lavie, · 2005
Earlier work this paper cites.
“Microsoft coco: Common objects in context,” 2015
Tsung-Yi Lin, Michael Maire, Serge Belongie, Lubomir Bourdev, Ross Girshick, James Hays, Pietro Perona, Deva Ramanan, C. Lawrence Zitnick, and Piotr Dollár, · 2015
Earlier work this paper cites.
“Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models,”
Bryan A. Plummer, Liwei Wang, Chris M. Cervantes, Juan C. Caicedo, Julia Hockenmaier, and Svetlana Lazebnik, · 2015
Earlier work this paper cites.
“Deep visual-semantic alignments for generating image descriptions,”
Andrej Karpathy and Fei-Fei Li, · 2015
Earlier work this paper cites.
“Cider: Consensus-based image description evaluation,”
Ramakrishna Vedantam, C. Lawrence Zitnick, and Devi Parikh, · 2015
Earlier work this paper cites.
“Visual genome: Connecting language and vision using crowdsourced dense image annotations,”
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A. Shamma, Michael S. Bernstein, and Li Fei-Fei, · 2016
Earlier work this paper cites.
“Spice: Semantic propositional image caption evaluation,” 2016
Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould, · 2016
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Cited alongside, same era.
“Parameter-efficient transfer learning for NLP,”
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin de Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly, · 2019
Cited alongside, same era.
“Language models are unsupervised multitask learners,”
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever, · 2019
Cited alongside, same era.
“Nocaps: novel object captioning at scale,”
Harsh Agrawal, Karan Desai, Yufei Wang, Xinlei Chen, Rishabh Jain, Mark Johnson, Dhruv Batra, Devi Parikh, Stefan Lee, and Peter Anderson, · 2019
Cited alongside, same era.
“Decoupled weight decay regularization,”
Ilya Loshchilov and Frank Hutter, · 2019
Cited alongside, same era.
“Oscar: Object-semantics aligned pre-training for vision-language tasks,”
Xiujun Li, Xi Yin, Chunyuan Li, Pengchuan Zhang, Xiaowei Hu, Lei Zhang, Lijuan Wang, Houdong Hu, Li Dong, Furu Wei, Yejin Choi, and Jianfeng Gao, · 2020
“Xgpt: Cross-modal generative pre-training for image captioning,”
Qiaolin Xia, Haoyang Huang, Nan Duan, Dongdong Zhang, Lei Ji, Zhifang Sui, Edward Cui, Taroon Bharti, and Ming Zhou, · 2021
Later among the works it cites.
“Clipcap: Clip prefix for image captioning,” 2021
Ron Mokady, Amir Hertz, and Amit H. Bermano, · 2021
Later among the works it cites.
“Learning transferable visual models from natural language supervision,”
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever, · 2021
Later among the works it cites.
“Prefix-tuning: Optimizing continuous prompts for generation,”
Xiang Lisa Li and Percy Liang, · 2021
Later among the works it cites.
Xiao Liu, Yanan Zheng, Zhengxiao Du, Ming Ding, Yujie Qian, Zhilin Yang, and Jie Tang, · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Unified vision-language pre-training for image captioning and vqa,”
Luowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu, Jason Corso, and Jianfeng Gao, · 2020
Cited alongside, same era.
“Momentum contrast for unsupervised visual representation learning,”
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross B. Girshick, · 2020
Cited alongside, same era.
“Unifying vision-and-language tasks via text generation,”
Jaemin Cho, Jie Lei, Hao Tan, and Mohit Bansal, · 2021
Later among the works it cites.
“Unitab: Unifying text and box outputs for grounded vision-language modeling,”
Zhengyuan Yang, Zhe Gan, Jianfeng Wang, Xiaowei Hu, Faisal Ahmed, Zicheng Liu, Yumao Lu, and Lijuan Wang, · 2022
Closest in time.
“SimVLM: Simple visual language model pretraining with weak supervision,”
Zirui Wang, Jiahui Yu, Adams Wei Yu, Zihang Dai, Yulia Tsvetkov, and Yuan Cao, · 2022
Closest in time.