Fetching the paper…
Reading the bibliography…
Image captioning models have achieved impressive results on datasets containing limited visual concepts and large amounts of paired image-caption training data.
Maximum likelihood from incomplete data via the EM algorithm
A. P. Dempster, N. M. Laird, and D. B. Rubin · 1977
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W. Zhu · 2002
Earlier work this paper cites.
Rouge: a package for automatic evaluation of summaries
C. Lin · 2004
Earlier work this paper cites.
Meteor: An automatic metric for MT evaluation with high levels of correlation with human judgments
A. Lavie and A. Agarwal · 2007
Earlier work this paper cites.
Visualizing data using t-SNE
L. van der Maaten and G. Hinton · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Large-scale machine learning with stochastic gradient descent
L. Bottou · 2010
Earlier work this paper cites.
Im2text: Describing images using 1 million captioned photographs
V. Ordonez, G. Kulkarni, and T. L. Berg · 2011
Earlier work this paper cites.
Framing image description as a ranking task: Data, models and evaluation metrics
M. Hodosh, P. Young, and J. Hockenmaier · 2013
Earlier work this paper cites.
Babytalk: Understanding and generating simple image descriptions
G. Kulkarni, V. Premraj, V. Ordonez, S. Dhar, S. Li, Y. Choi, A. C. Berg, and T. L. Berg · 2013
Earlier work this paper cites.
Dependency-based word embeddings
O. Levy and Y. Goldberg · 2014
Earlier work this paper cites.
GloVe: Global Vectors for Word Representation
J. Pennington, R. Socher, and C. D. Manning · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
P. Young, A. Lai, M. Hodosh, and J. Hockenmaier · 2014
Earlier work this paper cites.
Microsoft COCO captions: Data collection and evaluation server
X. Chen, H. Fang, T.-Y. Lin, R. Vedantam, S. Gupta, P. Dollár, and C. L. Zitnick · 2015
Earlier work this paper cites.
Microsoft COCO Captions: Data Collection and Evaluation Server
X. Chen, T.-Y. L. Hao Fang, R. Vedantam, S. Gupta, P. Dollar, and C. L. Zitnick · 2015
Earlier work this paper cites.
Long-term recurrent convolutional networks for visual recognition and description
J. Donahue, L. A. Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell · 2015
Earlier work this paper cites.
From captions to visual concepts and back
H. Fang, S. Gupta, F. N. Iandola, R. Srivastava, L. Deng, P. Dollar, J. Gao, X. He, M. Mitchell, J. C. Platt, C. L. Zitnick, and G. Zweig · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
A. Karpathy and L. Fei-Fei · 2015
Cited alongside, same era.
Unifying visual-semantic embeddings with multimodal neural language models
R. Kiros, R. Salakhutdinov, and R. S. Zemel · 2015
Cited alongside, same era.
Faster R-CNN: Towards real-time object detection with region proposal networks
S. Ren, K. He, R. Girshick, and J. Sun · 2015
Cited alongside, same era.
CIDEr: Consensus-based image description evaluation
R. Vedantam, C. L. Zitnick, and D. Parikh · 2015
Cited alongside, same era.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2015
Cited alongside, same era.
Show, attend and tell: Neural image caption generation with visual attention
Improved image captioning via policy gradient optimization of SPIDEr
S. Liu, Z. Zhu, N. Ye, S. Guadarrama, and K. Murphy · 2017
Later among the works it cites.
Knowing when to look: Adaptive attention via a visual sentinel for image captioning
J. Lu, C. Xiong, D. Parikh, and R. Socher · 2017
Later among the works it cites.
Extreme clicking for efficient object annotation
D. P. Papadopoulos, J. R. Uijlings, F. Keller, and V. Ferrari · 2017
Later among the works it cites.
Self-critical sequence training for image captioning
S. J. Rennie, E. Marcheret, Y. Mroueh, J. Ross, and V. Goel · 2017
Later among the works it cites.
Captioning Images with Diverse Objects
S. Venugopalan, L. A. Hendricks, M. Rohrbach, R. J. Mooney, T. Darrell, and K. Saenko · 2017
Later among the works it cites.
Incorporating copying mechanism in image captioning for learning novel objects
T. Yao, Y. Pan, Y. Li, and T. Mei · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
K. Xu, J. Ba, R. Kiros, K. Cho, A. C. Courville, R. Salakhutdinov, R. S. Zemel, and Y. Bengio · 2015
Cited alongside, same era.
SPICE: Semantic Propositional Image Caption Evaluation
P. Anderson, B. Fernando, M. Johnson, and S. Gould · 2016
Cited alongside, same era.
Deep Compositional Captioning: Describing Novel Object Categories without Paired Training Data
L. A. Hendricks, S. Venugopalan, M. Rohrbach, R. Mooney, K. Saenko, and T. Darrell · 2016
Cited alongside, same era.
Deep compositional captioning: Describing novel object categories without paired training data
L. A. Hendricks, S. Venugopalan, M. Rohrbach, R. J. Mooney, K. Saenko, and T. Darrell · 2016
Cited alongside, same era.
Training and evaluating multimodal word embeddings with large-scale web annotated images
J. Mao, J. Xu, K. Jing, and A. L. Yuille · 2016
Cited alongside, same era.
We don’t need no bounding-boxes: Training object class detectors using only human verification
D. P. Papadopoulos, J. R. Uijlings, F. Keller, and V. Ferrari · 2016
Cited alongside, same era.
Rich Image Captioning in the Wild
K. Tran, X. He, L. Zhang, J. Sun, C. Carapcea, C. Thrasher, C. Buehler, and C. Sienkiewicz · 2016
Cited alongside, same era.
Later among the works it cites.
Partially-supervised image captioning
P. Anderson, S. Gould, and M. Johnson · 2018
Closest in time.
Bottom-up and top-down attention for image captioning and visual question answering
P. Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, and L. Zhang · 2018
Closest in time.
Jointly predicting predicates and arguments in neural semantic role labeling
L. He, K. Lee, O. Levy, and L. Zettlemoyer · 2018
Closest in time.
Neural baby talk
J. Lu, J. Yang, D. Batra, and D. Parikh · 2018
Closest in time.
Deep contextualized word representations
M. Peters, M. Neumann, M. Iyyer, M. Gardner, C. Clark, K. Lee, and L. Zettlemoyer · 2018
Closest in time.
Do CIFAR-10 classifiers generalize to CIFAR-10?
B. Recht, R. Roelofs, L. Schmidt, and V. Shankar · 2018
Closest in time.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
P. Sharma, N. Ding, S. Goodman, and R. Soricut · 2018
Closest in time.
Decoupled novel object captioner
Y. Wu, L. Zhu, L. Jiang, and Y. Yang · 2018
Closest in time.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, and S. R. Bowman · 2019
Closest in time.
Evalai: Towards better evaluation systems for ai agents
D. Yadav, R. Jain, H. Agrawal, P. Chattopadhyay, T. Singh, A. Jain, S. B. Singh, S. Lee, and D. Batra · 2019
Closest in time.