Fetching the paper…
Reading the bibliography…
We introduce the dense captioning task, which requires a computer vision system to both localize and describe salient regions in images in natural language.
Generalization of backpropagation with application to a recurrent gas market model
P. J. Werbos · 1988
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Matching words and pictures
K. Barnard, P. Duygulu, D. Forsyth, N. De Freitas, D. M. Blei, and M. I. Jordan · 2003
Earlier work this paper cites.
A neural probabilistic language model
Y. Bengio, R. Ducharme, P. Vincent, and C. Janvin · 2003
Earlier work this paper cites.
The pascal visual object classes (voc) challenge
M. Everingham, L. Van Gool, C. K. Williams, J. Winn, and A. Zisserman · 2010
Earlier work this paper cites.
Every picture tells a story: Generating sentences from images
A. Farhadi, M. Hejrati, M. A. Sadeghi, P. Young, C. Rashtchian, J. Hockenmaier, and D. Forsyth · 2010
Earlier work this paper cites.
Recurrent neural network based language model
T. Mikolov, M. Karafiát, L. Burget, J. Cernockỳ, and S. Khudanpur · 2010
Earlier work this paper cites.
Connecting modalities: Semi-supervised segmentation and annotation of images using unaligned text corpora
R. Socher and L. Fei-Fei · 2010
Earlier work this paper cites.
Torch7: A matlab-like environment for machine learning
R. Collobert, K. Kavukcuoglu, and C. Farabet · 2011
Earlier work this paper cites.
Learning cross-modality similarity for multinomial data
Y. Jia, M. Salzmann, and T. Darrell · 2011
Earlier work this paper cites.
Baby talk: Understanding and generating simple image descriptions
G. Kulkarni, V. Premraj, S. Dhar, S. Li, Y. Choi, A. C. Berg, and T. L. Berg · 2011
Earlier work this paper cites.
Generating text with recurrent neural networks
I. Sutskever, J. Martens, and G. E. Hinton · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Generating sequences with recurrent neural networks
A. Graves · 2013
Earlier work this paper cites.
Generalizing image captions for image-text parallel corpus
P. Kuznetsova, V. Ordonez, A. C. Berg, T. L. Berg, and Y. Choi · 2013
Earlier work this paper cites.
Visualizing and understanding convolutional neural networks
M. D. Zeiler and R. Fergus · 2013
Earlier work this paper cites.
Learning a recurrent visual representation for image caption generation
X. Chen and C. L. Zitnick · 2014
Earlier work this paper cites.
Meteor universal: Language specific translation evaluation for any target language
M. Denkowski and A. Lavie · 2014
Cited alongside, same era.
Long-term recurrent convolutional networks for visual recognition and description
J. Donahue, L. A. Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell · 2014
Cited alongside, same era.
Scalable object detection using deep neural networks
D. Erhan, C. Szegedy, A. Toshev, and D. Anguelov · 2014
Cited alongside, same era.
From captions to visual concepts and back
H. Fang, S. Gupta, F. Iandola, R. Srivastava, L. Deng, P. Dollár, J. Gao, X. He, M. Mitchell, J. Platt, et al · 2014
Cited alongside, same era.
Rich feature hierarchies for accurate object detection and semantic segmentation
R. Girshick, J. Donahue, T. Darrell, and J. Malik · 2014
Cited alongside, same era.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
P. Young, A. Lai, M. Hodosh, and J. Hockenmaier · 2014
Later among the works it cites.
Edge boxes: Locating object proposals from edges
C. L. Zitnick and P. Dollár · 2014
Later among the works it cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
C. M. C. J. C. C. J. H. Bryan A. Plummer, Liwei Wang and S. Lazebni · 2015
Closest in time.
Microsoft coco captions: Data collection and evaluation server
X. Chen, H. Fang, T.-Y. Lin, R. Vedantam, S. Gupta, P. Dollar, and C. L. Zitnick · 2015
Closest in time.
Describing multimedia content using attention-based encoder-decoder networks
K. Cho, A. C. Courville, and Y. Bengio · 2015
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep visual-semantic alignments for generating image descriptions
A. Karpathy and L. Fei-Fei · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2014
Cited alongside, same era.
Unifying visual-semantic embeddings with multimodal neural language models
R. Kiros, R. Salakhutdinov, and R. S. Zemel · 2014
Cited alongside, same era.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Cited alongside, same era.
Explain images with multimodal recurrent neural networks
J. Mao, W. Xu, Y. Yang, J. Wang, and A. L. Yuille · 2014
Cited alongside, same era.
Overfeat: Integrated recognition, localization and detection using convolutional networks
P. Sermanet, D. Eigen, X. Zhang, M. Mathieu, R. Fergus, and Y. LeCun · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Cited alongside, same era.
R. Girshick · 2015
Closest in time.
Draw: A recurrent neural network for image generation
K. Gregor, I. Danihelka, A. Graves, and D. Wierstra · 2015
Closest in time.
Spatial pyramid pooling in deep convolutional networks for visual recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Closest in time.
M. Jaderberg, K. Simonyan, A. Zisserman, and K. Kavukcuoglu · 2015
Closest in time.
Visualizing and understanding recurrent networks
A. Karpathy, J. Johnson, and L. Fei-Fei · 2015
Closest in time.
Fully convolutional networks for semantic segmentation
J. Long, E. Shelhamer, and T. Darrell · 2015
Closest in time.
Large scale retrieval and generation of image descriptions
V. Ordonez, X. Han, P. Kuznetsova, G. Kulkarni, M. Mitchell, K. Yamaguchi, K. Stratos, A. Goyal, J. Dodge, A. Mensch, et al · 2015
Closest in time.
You only look once: Unified, real-time object detection
J. Redmon, S. Divvala, R. Girshick, and A. Farhadi · 2015
Closest in time.
Faster r-cnn: Towards real-time object detection with region proposal networks
S. Ren, K. He, R. Girshick, and J. Sun · 2015
Closest in time.
ImageNet Large Scale Visual Recognition Challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei · 2015
Closest in time.
The new data and new challenges in multimedia research
B. Thomee, D. A. Shamma, G. Friedland, B. Elizalde, K. Ni, D. Poland, D. Borth, and L.-J. Li · 2015
Closest in time.
Show, attend and tell: Neural image caption generation with visual attention
K. Xu, J. Ba, R. Kiros, A. Courville, R. Salakhutdinov, R. Zemel, and Y. Bengio · 2015
Closest in time.