Fetching the paper…
Reading the bibliography…
Image Captioning is a fundamental task to join vision and language, concerning about cross-modal understanding and text generation.
Visual Entailment: A Novel Task for Fine-Grained Image Understanding
Ning Xie, Farley Lai, Derek Doran, and Asim Kadav. 2019 · 1901
Earlier work this paper cites.
Attention on Attention for Image Captioning
Lun Huang, Wenmin Wang, Jie Chen, and Xiao-Yong Wei. 2019 · 1908
Earlier work this paper cites.
Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training
Gen Li, Nan Duan, Yuejian Fang, Ming Gong, Daxin Jiang, and Ming Zhou. 2019a · 1908
Earlier work this paper cites.
VisualBERT: A Simple and Performant Baseline for Vision and Language
Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang. 2019b · 1908
Earlier work this paper cites.
LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Hao Tan and Mohit Bansal. 2019 · 1908
Earlier work this paper cites.
UNITER: UNiversal Image-TExt Representation Learning
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. 2020 · 1909
Earlier work this paper cites.
Unified Vision-Language Pre-Training for Image Captioning and VQA
Luowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu, Jason J. Corso, and Jianfeng Gao. 2019b · 1909
Earlier work this paper cites.
Meshed-Memory Transformer for Image Captioning
Marcella Cornia, Matteo Stefanini, Lorenzo Baraldi, and Rita Cucchiara. 2020 · 1912
Earlier work this paper cites.
Unifying Vision-and-Language Tasks via Text Generation. In Proceedings of the 38th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 139) , Marina Meila and Tong Zhang (Eds.). PMLR, 1931–1942
Jaemin Cho, Jie Lei, Hao Tan, and Mohit Bansal. 2021 · 1942
Earlier work this paper cites.
Policy Gradient Methods for Reinforcement Learning with Function Approximation. In Advances in Neural Information Processing Systems , S. Solla, T. Leen, and K. Müller (Eds.), Vol. 12. MIT Press
Richard S Sutton, David McAllester, Satinder Singh, and Yishay Mansour. 1999 · 1999
Earlier work this paper cites.
Bleu: a Method for Automatic Evaluation of Machine Translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics . Association for Computational Linguistics, Philadelphia, Pennsylvania, USA, 311–318
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
X-Linear Attention Networks for Image Captioning
Yingwei Pan, Ting Yao, Yehao Li, and Tao Mei. 2020b · 2003
Earlier work this paper cites.
Oscar: Object-Semantics Aligned Pre-training for Vision-Language Tasks
Xiujun Li, Xi Yin, Chunyuan Li, Pengchuan Zhang, Xiaowei Hu, Lei Zhang, Lijuan Wang, Houdong Hu, Li Dong, Furu Wei, Yejin Choi, and Jianfeng Gao. 2020 · 2004
Earlier work this paper cites.
METEOR: An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments. In Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization . Association for Computational Linguistics, Ann Arbor, Michigan, 65–72
Satanjeev Banerjee and Alon Lavie. 2005 · 2005
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database. In CVPR09
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei. 2009 · 2009
Earlier work this paper cites.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021 · 2010
Earlier work this paper cites.
Incorporating BERT into Parallel Sequence Decoding with Adapters
Junliang Guo, Zhirui Zhang, Linli Xu, Hao-Ran Wei, Boxing Chen, and Enhong Chen. 2020 · 2010
Earlier work this paper cites.
Connecting modalities: Semi-supervised segmentation and annotation of images using unaligned text corpora. In 2010 IEEE Computer Society Conference on Computer Vision and Pattern Recognition . 966–973
Richard Socher and Li Fei-Fei. 2010 · 2010
Earlier work this paper cites.
I2T: Image Parsing to Text Description
Benjamin Z. Yao, Xiong Yang, Liang Lin, Mun Wai Lee, and Song-Chun Zhu. 2010 · 2010
Earlier work this paper cites.
Shuming Ma, Jian Yang, Haoyang Huang, Zewen Chi, Li Dong, Dongdong Zhang, Hany Hassan Awadalla, Alexandre Muzio, Akiko Eriguchi, Saksham Singhal, Xia Song, Arul Menezes, and Furu Wei. 2020 · 2012
Earlier work this paper cites.
Learning a Recurrent Visual Representation for Image Caption Generation
Xinlei Chen and C. Lawrence Zitnick. 2014 · 2014
Earlier work this paper cites.
Sequence to Sequence Learning with Neural Networks. In NIPS
Ilya Sutskever, Oriol Vinyals, and Quoc V. Le. 2014 · 2014
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2015 · 2015
Earlier work this paper cites.
Deep Visual-Semantic Alignments for Generating Image Descriptions
Andrej Karpathy and Li Fei-Fei. 2015 · 2015
Earlier work this paper cites.
Microsoft COCO: Common Objects in Context
Tsung-Yi Lin, Michael Maire, Serge Belongie, Lubomir Bourdev, Ross Girshick, James Hays, Pietro Perona, Deva Ramanan, C. Lawrence Zitnick, and Piotr Dollár. 2015 · 2015
Earlier work this paper cites.
Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks. In Advances in Neural Information Processing Systems , C. Cortes, N. Lawrence, D. Lee, M. Sugiyama, and R. Garnett (Eds.), Vol. 28. Curran Associates, Inc
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015 · 2015
Cited alongside, same era.
CIDEr: Consensus-based Image Description Evaluation
Ramakrishna Vedantam, C. Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Cited alongside, same era.
VQA: Visual Question Answering
Aishwarya Agrawal, Jiasen Lu, Stanislaw Antol, Margaret Mitchell, C. Lawrence Zitnick, Dhruv Batra, and Devi Parikh. 2016 · 2016
Cited alongside, same era.
SPICE: Semantic Propositional Image Caption Evaluation
Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould. 2016 · 2016
Cited alongside, same era.
VL-BERT: Pre-training of Generic Visual-Linguistic Representations. In International Conference on Learning Representations
Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai. 2020 · 2020
Later among the works it cites.
MiniVLM: A Smaller and Faster Vision-Language Model
Jianfeng Wang, Xiaowei Hu, Pengchuan Zhang, Xiujun Li, Lijuan Wang, L. Zhang, Jianfeng Gao, and Zicheng Liu. 2020 · 2020
Later among the works it cites.
BEiT: BERT Pre-Training of Image Transformers
Hangbo Bao, Li Dong, and Furu Wei. 2021 · 2021
Later among the works it cites.
David Bau, Alex Andonian, Audrey Cui, YeonHwan Park, Ali Jahanian, Aude Oliva, and Antonio Torralba. 2021 · 2021
Later among the works it cites.
An Empirical Study of Training End-to-End Vision-and-Language Transformers
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bryan A. Plummer, Liwei Wang, Chris M. Cervantes, Juan C. Caicedo, Julia Hockenmaier, and Svetlana Lazebnik. 2016 · 2016
Cited alongside, same era.
SCA-CNN: Spatial and Channel-Wise Attention in Convolutional Networks for Image Captioning
Long Chen, Hanwang Zhang, Jun Xiao, Liqiang Nie, Jian Shao, Wei Liu, and Tat-Seng Chua. 2017 · 2017
Cited alongside, same era.
Knowing When to Look: Adaptive Attention via A Visual Sentinel for Image Captioning
Jiasen Lu, Caiming Xiong, Devi Parikh, and Richard Socher. 2017 · 2017
Cited alongside, same era.
Self-Critical Sequence Training for Image Captioning
Steven J. Rennie, Etienne Marcheret, Youssef Mroueh, Jerret Ross, and Vaibhava Goel. 2017b · 2017
Cited alongside, same era.
Attention is All you Need. In Advances in Neural Information Processing Systems , I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett (Eds.), Vol. 30. Curran Associates, Inc
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang. 2018 · 2018
Cited alongside, same era.
Jiasen Lu, Jianwei Yang, Dhruv Batra, and Devi Parikh. 2018 · 2018
Cited alongside, same era.
Continual Lifelong Learning with Neural Networks: A Review
German Ignacio Parisi, Ronald Kemker, Jose L. Part, Christopher Kanan, and Stefan Wermter. 2018 · 2018
Cited alongside, same era.
Zi-Yi Dou, Yichong Xu, Zhe Gan, Jianfeng Wang, Shuohang Wang, Lijuan Wang, Chenguang Zhu, Pengchuan Zhang, Lu Yuan, Nanyun Peng, Zicheng Liu, and Michael Zeng. 2021 · 2021
Later among the works it cites.
Injecting Semantic Concepts into End-to-End Image Captioning
Zhiyuan Fang, Jianfeng Wang, Xiaowei Hu, Lin Liang, Zhe Gan, Lijuan Wang, Yezhou Yang, and Zicheng Liu. 2021 · 2021
Later among the works it cites.
StyleGAN-NADA: CLIP-Guided Domain Adaptation of Image Generators
Rinon Gal, Or Patashnik, Haggai Maron, Gal Chechik, and Daniel Cohen-Or. 2021 · 2021
Later among the works it cites.
Masked Autoencoders Are Scalable Vision Learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick. 2021 · 2021
Later among the works it cites.
Prakhar Kaushik, Alex Gain, Adam Kortylewski, and Alan Loddon Yuille. 2021 · 2021
Later among the works it cites.
ViLT: Vision-and-Language Transformer Without Convolution or Region Supervision
Wonjae Kim, Bokyung Son, and Ildoo Kim. 2021 · 2021
Later among the works it cites.
Tianyi Liu, Zuxuan Wu, Wenhan Xiong, Jingjing Chen, and Yu-Gang Jiang. 2021 · 2021
Later among the works it cites.
ClipCap: CLIP Prefix for Image Captioning
Ron Mokady, Amir Hertz, and Amit H. Bermano. 2021 · 2021
Later among the works it cites.
Learning Transferable Visual Models From Natural Language Supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021 · 2021
Later among the works it cites.
Multilingual Translation via Grafting Pre-trained Language Models. In Findings of the Association for Computational Linguistics: EMNLP 2021 . Association for Computational Linguistics, Punta Cana, Dominican Republic, 2735–2747
Zewei Sun, Mingxuan Wang, and Lei Li. 2021 · 2021
Later among the works it cites.
XGPT: Cross-modal Generative Pre-Training for Image Captioning. In NLPCC (1) . 786–797
Qiaolin Xia, Haoyang Huang, Nan Duan, Dongdong Zhang, Lei Ji, Zhifang Sui, Edward Cui, Taroon Bharti, and Ming Zhou. 2021 · 2021
Later among the works it cites.
E2E-VLP: End-to-End Vision-Language Pre-training Enhanced by Visual Learning. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) . Association for Computational Linguistics, Online, 503–513
Haiyang Xu, Ming Yan, Chenliang Li, Bin Bi, Songfang Huang, Wenming Xiao, and Fei Huang. 2021 · 2021
Later among the works it cites.
Research on Automatic News Text Summarization Technology Based on GPT2 Model
Zhenmin Yang, Yonghao Dong, Jiange Deng, Baocheng Sha, and Tao Xu. 2021a · 2021
Later among the works it cites.
VOLO: Vision Outlooker for Visual Recognition
Li Yuan, Qibin Hou, Zihang Jiang, Jiashi Feng, and Shuicheng Yan. 2021 · 2021
Later among the works it cites.
Learning to Prompt for Vision-Language Models
Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu. 2021b · 2021
Later among the works it cites.
LAFITE: Towards Language-Free Training for Text-to-Image Generation
Yufan Zhou, Ruiyi Zhang, Changyou Chen, Chunyuan Li, Chris Tensmeyer, Tong Yu, Jiuxiang Gu, Jinhui Xu, and Tong Sun. 2021c · 2021
Later among the works it cites.
I-Tuning: Tuning Language Models with Image for Caption Generation
Ziyang Luo, Yadong Xi, Rongsheng Zhang, and Jing Ma. 2022 · 2022
Closest in time.
Object recognition datasets and challenges: A review
Aria Salari, Abtin Djavadifar, Xiangrui Liu, and Homayoun Najjaran. 2022 · 2022
Closest in time.
From Show to Tell: A Survey on Deep Learning-based Image Captioning
Matteo Stefanini, Marcella Cornia, Lorenzo Baraldi, Silvia Cascianelli, Giuseppe Fiameni, and Rita Cucchiara. 2022 · 2022
Closest in time.
Image BERT Pre-training with Online Tokenizer. In International Conference on Learning Representations
Jinghao Zhou, Chen Wei, Huiyu Wang, Wei Shen, Cihang Xie, Alan Yuille, and Tao Kong. 2022 · 2022
Closest in time.