Fetching the paper…
Reading the bibliography…
Recent captioning models are limited in their ability to scale and describe concepts unseen in paired image-text corpora.
Corpus-guided sentence generation of natural images
Y. Yang, C. L. Teo, H. Daumé III, and Y. Aloimonos · 2011
Earlier work this paper cites.
Midge: Generating image descriptions from computer vision detections
M. Mitchell, J. Dodge, A. Goyal, K. Yamaguchi, K. Stratos, X. Han, A. Mensch, A. C. Berg, T. L. Berg, and H. D. III · 2012
Earlier work this paper cites.
LSTM neural networks for language modeling
M. Sundermeyer, R. Schlüter, and H. Ney · 2012
Earlier work this paper cites.
Devise: A deep visual-semantic embedding model
A. Frome, G. S. Corrado, J. Shlens, S. Bengio, J. Dean, T. Mikolov, et al · 2013
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean · 2013
Earlier work this paper cites.
Zero-shot learning by convex combination of semantic embeddings
M. Norouzi, T. Mikolov, S. Bengio, Y. Singer, J. Shlens, A. Frome, G. S. Corrado, and J. Dean · 2013
Earlier work this paper cites.
Meteor universal: Language specific translation evaluation for any target language
M. Denkowski and A. Lavie · 2014
Earlier work this paper cites.
Multimodal neural language models
R. Kiros, R. Salakhutdinov, and R. Zemel · 2014
Earlier work this paper cites.
Treetalk: Composition and compression of trees for image descriptions
P. Kuznetsova, V. Ordonez, T. L. Berg, U. C. Hill, and Y. Choi · 2014
Earlier work this paper cites.
Is this a wampimuk? cross-modal mapping between distributional semantics and the visual world
A. Lazaridou, E. Bruni, and M. Baroni · 2014
Cited alongside, same era.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Cited alongside, same era.
The stanford corenlp natural language processing toolkit
C. Manning, M. Surdeanu, J. Bauer, J. Finkel, S. J. Bethard, and D. McClosky · 2014
Cited alongside, same era.
Glove: Global vectors for word representation
J. Pennington, R. Socher, and C. D. Manning · 2014
Cited alongside, same era.
ILSVRC, 2014
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
On using monolingual corpora in neural machine translation
C. Gulcehre, O. Firat, K. Xu, K. Cho, L. Barrault, H. Lin, F. Bougares, H. Schwenk, and Y. Bengio · 2015
Later among the works it cites.
Deep visual-semantic alignments for generating image descriptions
A. Karpathy and L. Fei-Fei · 2015
Later among the works it cites.
Unifying visual-semantic embeddings with multimodal neural language models
R. Kiros, R. Salakhutdinov, and R. S. Zemel · 2015
Later among the works it cites.
Deep captioning with multimodal recurrent neural networks (m-rnn)
J. Mao, W. Xu, Y. Yang, J. Wang, Z. Huang, and A. Yuille · 2015
Later among the works it cites.
Learning like a child: Fast novel visual concept learning from sentence descriptions of images
J. Mao, W. Xu, Y. Yang, J. Wang, Z. Huang, and A. L. Yuille · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
K. Simonyan and A. Zisserman · 2014
Cited alongside, same era.
Grounded compositional semantics for finding and describing images with sentences
R. Socher, A. Karpathy, Q. V. Le, C. D. Manning, and A. Y. Ng · 2014
Cited alongside, same era.
Long-term recurrent convolutional networks for visual recognition and description
J. Donahue, L. A. Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell · 2015
Cited alongside, same era.
From captions to visual concepts and back
H. Fang, S. Gupta, F. N. Iandola, R. Srivastava, L. Deng, P. Dollár, J. Gao, X. He, M. Mitchell, J. C. Platt, C. L. Zitnick, and G. Zweig · 2015
Cited alongside, same era.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2015
Later among the works it cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Closest in time.
Deep compositional captioning: Describing novel object categories without paired training data
L. A. Hendricks, S. Venugopalan, M. Rohrbach, R. Mooney, K. Saenko, and T. Darrell · 2016
Closest in time.
Improving LSTM-based video description with linguistic knowledge mined from text
S. Venugopalan, L. A. Hendricks, R. Mooney, and K. Saenko · 2016
Closest in time.