Fetching the paper…
Reading the bibliography…
Image Captioning is a task that combines computer vision and natural language processing, where it aims to generate descriptive legends for images.
Understanding inverse document frequency: on theoretical arguments for idf
S. Robertson · 2004
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
A. Karpathy and F. Li · 2014
Earlier work this paper cites.
Microsoft COCO: common objects in context
T. Lin, M. Maire, S. J. Belongie, L. D. Bourdev, R. B. Girshick, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Earlier work this paper cites.
Image retrieval using scene graphs
J. Johnson, R. Krishna, M. Stark, L.-J. Li, D. A. Shamma, M. S. Bernstein, and L. Fei-Fei · 2015
Earlier work this paper cites.
Faster R-CNN: towards real-time object detection with region proposal networks
S. Ren, K. He, R. B. Girshick, and J. Sun · 2015
Earlier work this paper cites.
Generating semantically precise scene graphs from textual descriptions for improved image retrieval
S. Schuster, R. Krishna, A. Chang, L. Fei-Fei, and C. D. Manning · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
R. Vedantam, C. L. Zitnick, and D. Parikh · 2015
Earlier work this paper cites.
Guided open vocabulary image captioning with constrained beam search
P. Anderson, B. Fernando, M. Johnson, and S. Gould · 2016
Earlier work this paper cites.
Spice: Semantic propositional image caption evaluation
P. Anderson, B. Fernando, M. Johnson, and S. Gould · 2016
Cited alongside, same era.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L. Li, D. A. Shamma, M. S. Bernstein, and F. Li · 2016
Cited alongside, same era.
Self-critical sequence training for image captioning
S. J. Rennie, E. Marcheret, Y. Mroueh, J. Ross, and V. Goel · 2016
Cited alongside, same era.
Bottom-up and top-down attention for image captioning and VQA
P. Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, and L. Zhang · 2017
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
C. Finn, P. Abbeel, and S. Levine · 2017
Deep contextualized word representations
M. E. Peters, M. Neumann, M. Iyyer, M. Gardner, C. Clark, K. Lee, and L. Zettlemoyer · 2018
Later among the works it cites.
Improving image captioning with conditional generative adversarial nets
C. Chen, S. Mu, W. Xiao, Z. Ye, L. Wu, and Q. Ju · 2019
Later among the works it cites.
Meta learning for image captioning
N. Li, Z. Chen, and S. Liu · 2019
Later among the works it cites.
A systematic literature review on image captioning
R. Staniūtė and D. Šešok · 2019
Later among the works it cites.
Image captioning using deep learning: A systematic literature review
M. Chohan, A. Khan, M. Saleem, S. Hassan, A. Ghafoor, and M. Khan · 2020
Later among the works it cites.
Oscar: Object-semantics aligned pre-training for vision-language tasks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Openimages: A public dataset for large-scale multi-label and multi-class image classification
I. Krasin, T. Duerig, N. Alldrin, V. Ferrari, S. Abu-El-Haija, A. Kuznetsova, H. Rom, J. Uijlings, S. Popov, A. Veit, S. Belongie, V. Gomes, A. Gupta, C. Sun, G. Chechik, D. Cai, Z. Feng, D. Narayanan, and K. Murphy · 2017
Cited alongside, same era.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Cited alongside, same era.
nocaps: novel object captioning at scale
H. Agrawal, K. Desai, Y. Wang, X. Chen, R. Jain, M. Johnson, D. Batra, D. Parikh, S. Lee, and P. Anderson · 2018
Cited alongside, same era.
A comprehensive survey of deep learning for image captioning, 2018
M. Z. Hossain, F. Sohel, M. F. Shiratuddin, and H. Laga · 2018
Cited alongside, same era.
Objects and attention: The state of the art
B. J. Scholl
Cited in the paper.
X. Li, X. Yin, C. Li, P. Zhang, X. Hu, L. Zhang, L. Wang, H. Hu, L. Dong, F. Wei, et al · 2020
Later among the works it cites.
A review on automatic image captioning techniques
K. C. Nithya and V. V. Kumar · 2020
Later among the works it cites.
Vivo: Visual vocabulary pre-training for novel object captioning, 2021
X. Hu, X. Yin, K. Lin, L. Wang, L. Zhang, J. Gao, and Z. Liu · 2021
Closest in time.