Fetching the paper…
Reading the bibliography…
Generating natural language descriptions for in-the-wild videos is a challenging task.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu · 2002
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
S. Banerjee and A. Lavie · 2005
Earlier work this paper cites.
Beyond bags of features: Spatial pyramid matching for recognizing natural scene categories
S. Lazebnik, C. Schmid, and J. Ponce · 2006
Earlier work this paper cites.
Object bank: A high-level image representation for scene classification & semantic feature sparsification
L.-J. Li, H. Su, L. Fei-Fei, and E. P. Xing · 2010
Earlier work this paper cites.
Collecting highly parallel data for paraphrase evaluation
D. L. Chen and W. B. Dolan · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Youtube2text: Recognizing and describing arbitrary activities using semantic hierarchies and zero-shot recognition
S. Guadarrama, N. Krishnamoorthy, G. Malkarnenkar, S. Venugopalan, R. Mooney, T. Darrell, and K. S. nko · 2013
Earlier work this paper cites.
Translating video content to natural language descriptions
M. Rohrbach, W. Qiu, I. Titov, S. Thater, M. Pinkal, and B. Schiele · 2013
Earlier work this paper cites.
Comparing automatic evaluation measures for image description
D. Elliott and F. Keller · 2014
Cited alongside, same era.
From captions to visual concepts and back
H. Fang, S. Gupta, F. Iandola, R. Srivastava, L. Deng, P. Dollár, J. Gao, X. He, M. Mitchell, J. Platt, et al · 2014
Cited alongside, same era.
Rich feature hierarchies for accurate object detection and semantic segmentation
R. Girshick, J. Donahue, T. Darrell, and J. Malik · 2014
Cited alongside, same era.
Caffe: Convolutional architecture for fast feature embedding
Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell · 2014
Cited alongside, same era.
Fully convolutional neural networks for crowd segmentation
K. Kang and X. Wang · 2014
Cited alongside, same era.
Fully convolutional multi-class multiple instance learning
D. Pathak, E. Shelhamer, J. Long, and T. Darrell · 2014
Later among the works it cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Later among the works it cites.
Going deeper with convolutions
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2014
Later among the works it cites.
Integrating language and vision to generate natural language descriptions of videos in the wild
J. Thomason, S. Venugopalan, S. Guadarrama, K. Saenko, and R. Mooney · 2014
Later among the works it cites.
Translating videos to natural language using deep recurrent neural networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep visual-semantic alignments for generating image descriptions
A. Karpathy and L. Fei-Fei · 2014
Cited alongside, same era.
Fully convolutional networks for semantic segmentation
J. Long, E. Shelhamer, and T. Darrell · 2014
Cited alongside, same era.
Explain images with multimodal recurrent neural networks
J. Mao, W. Xu, Y. Yang, J. Wang, and A. L. Yuille · 2014
Cited alongside, same era.
Weakly supervised object recognition with convolutional neural networks
M. Oquab, L. Bottou, I. Laptev, and J. Sivic · 2014
Cited alongside, same era.
S. Venugopalan, H. Xu, J. Donahue, M. Rohrbach, R. J. Mooney, and K. Saenko · 2014
Later among the works it cites.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2014
Later among the works it cites.
Sequence to sequence - video to text
S. Venugopalan, M. Rohrbach, J. Donahue, R. Mooney, T. Darrell, and K. Saenko · 2015
Closest in time.
Jointly modeling deep video and compositional text to bridge vision and language in a unified framework
R. Xu, C. Xiong, W. Chen, and J. J. Corso · 2015
Closest in time.
Describing videos by exploiting temporal structure
L. Yao, A. Torabi, K. Cho, N. Ballas, C. Pal, H. Larochelle, and A. Courville · 2015
Closest in time.