Fetching the paper…
Reading the bibliography…
CNN-LSTM based architectures have played an important role in image captioning, but limited by the training efficiency and expression ability, researchers began to explore the CNN-Transformer based models and achieved great success.
Long Short-Term Memory
Hochreiter, S.; and Schmidhuber, J. 1997 · 1997
Earlier work this paper cites.
Bleu: a Method for Automatic Evaluation of Machine Translation
Papineni, K.; Roukos, S.; Ward, T.; and Zhu, W. 2002 · 2002
Earlier work this paper cites.
ROUGE: A Package for Automatic Evaluation of Summaries
Lin, C.-Y. 2004 · 2004
Earlier work this paper cites.
METEOR: An Automatic Metric for MT Evaluation with High Levels of Correlation with Human Judgments
Lavie, A.; and Agarwal, A. 2007 · 2007
Earlier work this paper cites.
Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation
Cho, K.; van Merrienboer, B.; Gülçehre, Ç.; Bahdanau, D.; Bougares, F.; Schwenk, H.; and Bengio, Y. 2014 · 2014
Earlier work this paper cites.
Microsoft COCO: Common Objects in Context
Lin, T.; Maire, M.; Belongie, S. J.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; and Zitnick, C. L. 2014 · 2014
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Kingma, D. P.; and Ba, J. 2015 · 2015
Earlier work this paper cites.
Very Deep Convolutional Networks for Large-Scale Image Recognition
Simonyan, K.; and Zisserman, A. 2015 · 2015
Earlier work this paper cites.
CIDEr: Consensus-based image description evaluation
Vedantam, R.; Zitnick, C. L.; and Parikh, D. 2015 · 2015
Earlier work this paper cites.
Show and tell: A neural image caption generator
Vinyals, O.; Toshev, A.; Bengio, S.; and Erhan, D. 2015 · 2015
Earlier work this paper cites.
SPICE: Semantic Propositional Image Caption Evaluation
Anderson, P.; Fernando, B.; Johnson, M.; and Gould, S. 2016 · 2016
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Earlier work this paper cites.
Deep Visual-Semantic Alignments for Generating Image Descriptions
Karpathy, A.; and Fei-Fei, L. 2017 · 2017
Cited alongside, same era.
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations
Krishna, R.; Zhu, Y.; Groth, O.; Johnson, J.; Hata, K.; Kravitz, J.; Chen, S.; Kalantidis, Y.; Li, L.; Shamma, D. A.; Bernstein, M. S.; and Fei-Fei, L. 2017 · 2017
Cited alongside, same era.
Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
Ren, S.; He, K.; Girshick, R. B.; and Sun, J. 2017 · 2017
Cited alongside, same era.
Self-Critical Sequence Training for Image Captioning
Rennie, S. J.; Marcheret, E.; Mroueh, Y.; Ross, J.; and Goel, V. 2017 · 2017
Cited alongside, same era.
Attention is All you Need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L.; and Polosukhin, I. 2017 · 2017
Cited alongside, same era.
Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering
Hierarchical Attention Network for Image Captioning
Wang, W.; Chen, Z.; and Hu, H. 2019 · 2019
Later among the works it cites.
Meshed-Memory Transformer for Image Captioning
Cornia, M.; Stefanini, M.; Baraldi, L.; and Cucchiara, R. 2020 · 2020
Later among the works it cites.
In Defense of Grid Features for Visual Question Answering
Jiang, H.; Misra, I.; Rohrbach, M.; Learned-Miller, E. G.; and Chen, X. 2020 · 2020
Later among the works it cites.
X-Linear Attention Networks for Image Captioning
Pan, Y.; Yao, T.; Li, Y.; and Mei, T. 2020 · 2020
Later among the works it cites.
ActBERT: Learning Global-Local Video-Text Representations
Zhu, L.; and Yang, Y. 2020 · 2020
Later among the works it cites.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; Uszkoreit, J.; and Houlsby, N. 2021 · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Anderson, P.; He, X.; Buehler, C.; Teney, D.; Johnson, M.; Gould, S.; and Zhang, L. 2018 · 2018
Cited alongside, same era.
Recurrent Fusion Network for Image Captioning
Jiang, W.; Ma, L.; Jiang, Y.; Liu, W.; and Zhang, T. 2018 · 2018
Cited alongside, same era.
Exploring Visual Relationship for Image Captioning
Yao, T.; Pan, Y.; Li, Y.; and Mei, T. 2018 · 2018
Cited alongside, same era.
Image Captioning: Transforming Objects into Words
Herdade, S.; Kappeler, A.; Boakye, K.; and Soares, J. 2019 · 2019
Cited alongside, same era.
Attention on Attention for Image Captioning
Huang, L.; Wang, W.; Chen, J.; and Wei, X. 2019 · 2019
Cited alongside, same era.
Entangled Transformer for Image Captioning
Li, G.; Zhu, L.; Liu, P.; and Yang, Y. 2019 · 2019
Cited alongside, same era.
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks
Lu, J.; Batra, D.; Parikh, D.; and Lee, S. 2019 · 2019
Cited alongside, same era.
Later among the works it cites.
Improving Image Captioning by Leveraging Intra- and Inter-layer Global Representation in Transformer Network
Ji, J.; Luo, Y.; Sun, X.; Chen, F.; Luo, G.; Wu, Y.; Gao, Y.; and Ji, R. 2021 · 2021
Later among the works it cites.
Swin Transformer: Hierarchical Vision Transformer using Shifted Windows
Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; and Guo, B. 2021 · 2021
Later among the works it cites.
Dual-level Collaborative Transformer for Image Captioning
Luo, Y.; Ji, J.; Sun, X.; Cao, L.; Wu, Y.; Huang, F.; Lin, C.; and Ji, R. 2021 · 2021
Later among the works it cites.
Learning Transferable Visual Models From Natural Language Supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; Krueger, G.; and Sutskever, I. 2021 · 2021
Later among the works it cites.
RSTNet: Captioning With Adaptive Attention on Visual and Non-Visual Words
Zhang, X.; Sun, X.; Luo, Y.; Ji, J.; Zhou, Y.; Wu, Y.; Huang, F.; and Ji, R. 2021 · 2021
Later among the works it cites.
Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
Xu, K.; Ba, J.; Kiros, R.; Cho, K.; Courville, A. C.; Salakhutdinov, R.; Zemel, R. S.; and Bengio, Y. 2015 · 2057
Closest in time.