Fetching the paper…
Reading the bibliography…
Recent progress in using recurrent neural networks (RNNs) for image description has motivated the exploration of their application for video description.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Natural language description of human activities from video images based on concept hierarchy of actions
A. Kojima, T. Tamura, and K. Fukunaga · 2002
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu · 2002
Earlier work this paper cites.
Human detection using oriented histograms of flow and appearance
N. Dalal, B. Triggs, and C. Schmid · 2006
Earlier work this paper cites.
Evaluation of local spatio-temporal features for action recognition
H. Wang, M. M. Ullah, A. Kläser, I. Laptev, and C. Schmid · 2009
Earlier work this paper cites.
Theano: a CPU and GPU math expression compiler
J. Bergstra, O. Breuleux, F. Bastien, P. Lamblin, R. Pascanu, G. Desjardins, J. Turian, D. Warde-Farley, and Y. Bengio · 2010
Earlier work this paper cites.
Youtube scale, large vocabulary video annotation
N. Morsillo, G. Mann, and C. Pal · 2010
Earlier work this paper cites.
Convolutional learning of spatio-temporal features
G. W. Taylor, R. Fergus, Y. LeCun, and C. Bregler · 2010
Earlier work this paper cites.
Collecting highly parallel data for paraphrase evaluation
D. L. Chen and W. B. Dolan · 2011
Earlier work this paper cites.
HMDB: a large video database for human motion recognition
H. Kuehne, H. Jhuang, E. Garrote, T. Poggio, and T. Serre · 2011
Earlier work this paper cites.
Video in sentences out
A. Barbu, A. Bridge, Z. Burchill, D. Coroian, S. Dickinson, S. Fidler, A. Michaux, S. Mussman, S. Narayanaswamy, D. Salvi, et al · 2012
Earlier work this paper cites.
Theano: new features and speed improvements
F. Bastien, P. Lamblin, R. Pascanu, J. Bergstra, I. J. Goodfellow, A. Bergeron, N. Bouchard, and Y. Bengio · 2012
Earlier work this paper cites.
Random search for hyper-parameter optimization
J. Bergstra and Y. Bengio · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
UCF101: A dataset of 101 human actions classes from videos in the wild
K. Soomro, A. R. Zamir, and M. Shah · 2012
Earlier work this paper cites.
Learning latent temporal structure for complex event detection
K. Tang, L. Fei-Fei, and D. Koller · 2012
Earlier work this paper cites.
ADADELTA: an adaptive learning rate method
M. D. Zeiler · 2012
Cited alongside, same era.
Temporal localization of actions with actoms
A. Gaidon, Z. Harchaoui, and C. Schmid · 2013
Cited alongside, same era.
Framing image description as a ranking task: Data, models and evaluation metrics
M. Hodosh, P. Young, and J. Hockenmaier · 2013
Cited alongside, same era.
3d convolutional neural networks for human action recognition
S. Ji, W. Xu, M. Yang, and K. Yu · 2013
Cited alongside, same era.
Translating video content to natural language descriptions
M. Rohrbach, W. Qiu, I. Titov, S. Thater, M. Pinkal, and B. Schiele · 2013
Cited alongside, same era.
Weakly supervised action labeling in videos under ordering constraints
P. Bojanowski, R. Lajugie, F. Bach, I. Laptev, J. Ponce, C. Schmid, and J. Sivic · 2014
Cited alongside, same era.
Sequence to sequence learning with neural networks
I. Sutskever, O. Vinyals, and Q. V. V. Le · 2014
Later among the works it cites.
Integrating language and vision to generate natural language descriptions of videos in the wild
J. Thomason, S. Venugopalan, S. Guadarrama, K. Saenko, and R. Mooney · 2014
Later among the works it cites.
C3D: Generic features for video analysis
D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri · 2014
Later among the works it cites.
Recurrent neural network regularization
W. Zaremba, I. Sutskever, and O. Vinyals · 2014
Later among the works it cites.
Neural machine translation by jointly learning to align and translate
D. Bahdanau, K. Cho, and Y. Bengio · 2015
Closest in time.
Microsoft coco captions: Data collection and evaluation server
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning phrase representations using RNN encoder-decoder for statistical machine translation
K. Cho, B. van Merrienboer, C. Gulcehre, F. Bougares, H. Schwenk, and Y. Bengio · 2014
Cited alongside, same era.
Meteor universal: Language specific translation evaluation for any target language
M. Denkowski and A. Lavie · 2014
Cited alongside, same era.
Caffe: Convolutional architecture for fast feature embedding
Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell · 2014
Cited alongside, same era.
Deep visual-semantic alignments for generating image descriptions
A. Karpathy and L. Fei-Fei · 2014
Cited alongside, same era.
Large-scale video classification with convolutional neural networks
A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei · 2014
Cited alongside, same era.
Unifying visual-semantic embeddings with multimodal neural language models
R. Kiros, R. Salakhutdinov, and R. S. Zemel · 2014
Cited alongside, same era.
X. Chen, H. Fang, T.-Y. Lin, R. Vedantam, S. Gupta, P. Dollar, and C. L. Zitnick · 2015
Closest in time.
Long-term recurrent convolutional networks for visual recognition and description
J. Donahue, L. A. Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell · 2015
Closest in time.
A dataset for movie description
A. Rohrbach, M. Rohrbach, N. Tandon, and B. Schiele · 2015
Closest in time.
Unsupervised learning of video representations using lstms
N. Srivastava, E. Mansimov, and R. Salakhutdinov · 2015
Closest in time.
Going deeper with convolutions
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2015
Closest in time.
Using descriptive video services to create a large data source for video annotation research
A. Torabi, C. Pal, H. Larochelle, and A. Courville · 2015
Closest in time.
CIDEr: Consensus-based image description evaluation
R. Vedantam, C. L. Zitnick, and D. Parikh · 2015
Closest in time.
Translating videos to natural language using deep recurrent neural networks
S. Venugopalan, H. Xu, J. Donahue, M. Rohrbach, R. Mooney, and K. Saenko · 2015
Closest in time.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2015
Closest in time.
Show, attend and tell: Neural image caption generation with visual attention
K. Xu, J. Ba, R. Kiros, , K. Cho, A. Courville, R. Salakhutdinov, R. Zemel, and Y. Bengio · 2015
Closest in time.