Fetching the paper…
Reading the bibliography…
This paper presents a novel approach for automatically generating image descriptions: visual detectors, language models, and multimodal similarity models learnt directly from a dataset of image captions.
Trigger-based language models: A maximum entropy approach
R. Lau, R. Rosenfeld, and S. Roukos · 1993
Earlier work this paper cites.
A maximum entropy approach to natural language processing
A. L. Berger, S. A. D. Pietra, and V. J. D. Pietra · 1996
Earlier work this paper cites.
A framework for multiple-instance learning
O. Maron and T. Lozano-Pérez · 1998
Earlier work this paper cites.
Trainable methods for surface natural language generation
A. Ratnaparkhi · 2000
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu · 2002
Earlier work this paper cites.
Trainable approaches to surface natural language generation and their application to conversational dialog systems
A. Ratnaparkhi · 2002
Earlier work this paper cites.
Minimum error rate training in statistical machine translation
F. J. Och · 2003
Earlier work this paper cites.
Automatic evaluation of machine translation quality using longest common subsequence and skip-bigram statistics
C.-Y. Lin and F. J. Och · 2004
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments
S. Banerjee and A. Lavie · 2005
Earlier work this paper cites.
Multiple instance boosting for object detection
C. Zhang, J. C. Platt, and P. A. Viola · 2005
Earlier work this paper cites.
Three new graphical models for statistical language modelling
A. Mnih and G. Hinton · 2007
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Toward an architecture for never-ending language learning
A. Carlson, J. Betteridge, B. Kisiel, B. Settles, E. R. Hruschka Jr, and T. M. Mitchell · 2010
Earlier work this paper cites.
The PASCAL visual object classes (VOC) challenge
M. Everingham, L. Van Gool, C. K. I. Williams, J. Winn, and A. Zisserman · 2010
Earlier work this paper cites.
Every picture tells a story: Generating sentences from images
A. Farhadi, M. Hejrati, M. A. Sadeghi, P. Young, C. Rashtchian, J. Hockenmaier, and D. Forsyth · 2010
Earlier work this paper cites.
Collecting image annotations using Amazon’s mechanical turk
C. Rashtchian, P. Young, M. Hodosh, and J. Hockenmaier · 2010
Earlier work this paper cites.
I2T: Image parsing to text description
B. Z. Yao, X. Yang, L. Lin, M. W. Lee, and S.-C. Zhu · 2010
Earlier work this paper cites.
Baby talk: Understanding and generating simple image descriptions
G. Kulkarni, V. Premraj, S. Dhar, S. Li, Y. Choi, A. C. Berg, and T. L. Berg · 2011
Cited alongside, same era.
Composing simple image descriptions using web-scale n-grams
S. Li, G. Kulkarni, T. L. Berg, A. C. Berg, and Y. Choi · 2011
Cited alongside, same era.
Strategies for training large scale neural network language models
T. Mikolov, A. Deoras, D. Povey, L. Burget, and J. Cernocky · 2011
Cited alongside, same era.
Im2text: Describing images using 1 million captioned photographs
V. Ordonez, G. Kulkarni, and T. L. Berg · 2011
Cited alongside, same era.
Corpus-guided sentence generation of natural images
Y. Yang, C. L. Teo, H. Daumé III, and Y. Aloimonos · 2011
Cited alongside, same era.
ImageNet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Rich feature hierarchies for accurate object detection and semantic segmentation
R. Girshick, J. Donahue, T. Darrell, and J. Malik · 2014
Closest in time.
Using k-poselets for detecting people and localizing their keypoints
G. Gkioxari, B. Hariharan, R. Girshick, and J. Malik · 2014
Closest in time.
Caffe: Convolutional architecture for fast feature embedding
Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell · 2014
Closest in time.
Deep fragment embeddings for bidirectional image sentence mapping
A. Karpathy, A. Joulin, and L. Fei-Fei · 2014
Closest in time.
Microsoft COCO: Common objects in context
T. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Collective generation of natural image descriptions
P. Kuznetsova, V. Ordonez, A. C. Berg, T. L. Berg, and Y. Choi · 2012
Cited alongside, same era.
Midge: Generating image descriptions from computer vision detections
M. Mitchell, X. Han, J. Dodge, A. Mensch, A. Goyal, A. Berg, K. Yamaguchi, T. Berg, K. Stratos, and H. Daumé III · 2012
Cited alongside, same era.
A fast and simple algorithm for training neural probabilistic language models
A. Mnih and Y. W. Teh · 2012
Cited alongside, same era.
Neil: Extracting visual knowledge from web data
X. Chen, A. Shrivastava, and A. Gupta · 2013
Cited alongside, same era.
Framing image description as a ranking task: Data, models and evaluation metrics
M. Hodosh, P. Young, and J. Hockenmaier · 2013
Cited alongside, same era.
Learning deep structured semantic models for web search using clickthrough data
P. Huang, X. He, J. Gao, L. Deng, A. Acero, and L. Heck · 2013
Cited alongside, same era.
J. Mao, W. Xu, Y. Yang, J. Wang, and A. L. Yuille · 2014
Closest in time.
A latent semantic model with convolutional-pooling structure for information retrieval
Y. Shen, X. He, J. Gao, L. Deng, and G. Mesnil · 2014
Closest in time.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Closest in time.
Cider: Consensus-based image description evaluation
R. Vedantam, C. L. Zitnick, and D. Parikh · 2014
Closest in time.
Edge boxes: Locating object proposals from edges
C. L. Zitnick and P. Dollár · 2014
Closest in time.
Microsoft coco captions: Data collection and evaluation server
X. Chen, H. Fang, T. Lin, R. Vedantam, S. Gupta, P. Dollár, and C. L. Zitnick · 2015
Closest in time.
Mind’s eye: A recurrent visual representation for image caption generation
X. Chen and C. L. Zitnick · 2015
Closest in time.
Long-term recurrent convolutional networks for visual recognition and description
J. Donahue, L. A. Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell · 2015
Closest in time.
Deep visual-semantic alignments for generating image descriptions
A. Karpathy and L. Fei-Fei · 2015
Closest in time.
R. Lebret, P. O. Pinheiro, and R. Collobert · 2015
Closest in time.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2015
Closest in time.
Show, attend and tell: Neural image caption generation with visual attention
K. Xu, J. Ba, R. Kiros, A. Courville, R. Salakhutdinov, R. Zemel, and Y. Bengio · 2015
Closest in time.