Fetching the paper…
Reading the bibliography…
We present an approach that exploits hierarchical Recurrent Neural Networks (RNNs) to tackle the video captioning problem, i.e., generating one or multiple sentences to describe a realistic video.
Finding structure in time
J. L. Elman · 1990
Earlier work this paper cites.
Backpropagation through time: what does it do and how to do it
P. Werbos · 1990
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Bidirectional recurrent neural networks
M. Schuster and K. Paliwal · 1997
Earlier work this paper cites.
Efficient backprop
Y. LeCun, L. Bottou, G. Orr, and K. Müller · 1998
Earlier work this paper cites.
Natural language description of human activities from video images based on concept hierarchy of actions
A. Kojima, T. Tamura, and K. Fukunaga · 2002
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W. jing Zhu · 2002
Earlier work this paper cites.
Two-frame motion estimation based on polynomial expansion
G. Farnebäck · 2003
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments
S. Banerjee and A. Lavie · 2005
Earlier work this paper cites.
SAVE: A framework for semantic annotation of visual events
M. W. Lee, A. Hakeem, N. Haering, and S.-C. Zhu · 2008
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
V. Nair and G. E. Hinton · 2010
Earlier work this paper cites.
Collecting highly parallel data for paraphrase evaluation
D. L. Chen and W. B. Dolan · 2011
Earlier work this paper cites.
Human focused video description
M. U. G. Khan, L. Zhang, and Y. Gotoh · 2011
Earlier work this paper cites.
Towards coherent natural language description of video streams
M. U. G. Khan, L. Zhang, and Y. Gotoh · 2011
Earlier work this paper cites.
Action Recognition by Dense Trajectories
H. Wang, A. Kläser, C. Schmid, and C.-L. Liu · 2011
Earlier work this paper cites.
Video in sentences out
A. Barbu, A. Bridge, Z. Burchill, D. Coroian, S. Dickinson, S. Fidler, A. Michaux, S. Mussman, N. Siddharth, D. Salvi, L. Schmidt, J. Shangguan, J. M. Siskind, J. Waggoner, S. Wang, J. Wei, Y. Yin, and Z. Zhang · 2012
Earlier work this paper cites.
Automated textual descriptions for a wide range of video events with 48 human actions
P. Hanckmann, K. Schutte, and G. J. Burghouts · 2012
Earlier work this paper cites.
Aggregating local image descriptors into compact codes
H. Jegou, F. Perronnin, M. Douze, J. Sánchez, P. Perez, and C. Schmid · 2012
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
T. Tieleman and G. Hinton · 2012
Earlier work this paper cites.
A thousand frames in just a few words: Lingual description of videos through latent topics and sparse object stitching
P. Das, C. Xu, R. F. Doell, and J. J. Corso · 2013
Cited alongside, same era.
Youtube2text: Recognizing and describing arbitrary activities using semantic hierarchies and zero-shot recognition
S. Guadarrama, N. Krishnamoorthy, G. Malkarnenkar, S. Venugopalan, T. D. R. Mooney, and K. Saenko · 2013
Cited alongside, same era.
Recurrent continuous translation models
N. Kalchbrenner and P. Blunsom · 2013
Cited alongside, same era.
Generating natural-language video descriptions using text-mined knowledge
N. Krishnamoorthy, G. Malkarnenkar, R. J. Mooney, K. Saenko, and S. Guadarrama · 2013
Cited alongside, same era.
Translating video content to natural language descriptions
M. Rohrbach, W. Qiu, I. Titov, S. Thater, M. Pinkal, and B. Schiele · 2013
Cited alongside, same era.
A hierarchical neural autoencoder for paragraphs and documents
J. Li, M. Luong, and D. Jurafsky · 2015
Closest in time.
Hierarchical recurrent neural network for document modeling
R. Lin, S. Liu, M. Yang, M. Li, M. Zhou, and S. Li · 2015
Closest in time.
Deep captioning with multimodal recurrent neural networks (m-rnn)
J. Mao, W. Xu, Y. Yang, J. Wang, Z. Huang, and A. Yuille · 2015
Closest in time.
Learning like a child: Fast novel visual concept learning from sentence descriptions of images
J. Mao, W. Xu, Y. Yang, J. Wang, Z. Huang, and A. L. Yuille · 2015
Closest in time.
Jointly modeling embedding and translation to bridge video and language
Y. Pan, T. Mei, T. Yao, H. Li, and Y. Rui · 2015
Closest in time.
Expressing an image stream with a sequence of natural sentences
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Bordes, S. Chopra, and J. Weston · 2014
Cited alongside, same era.
Learning phrase representations using RNN encoder-decoder for statistical machine translation
K. Cho, B. van Merrienboer, Ç. Gülçehre, F. Bougares, H. Schwenk, and Y. Bengio · 2014
Cited alongside, same era.
Large-scale video classification with convolutional neural networks
A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei · 2014
Cited alongside, same era.
Unifying visual-semantic embeddings with multimodal neural language models
R. Kiros, R. Salakhutdinov, and R. S. Zemel · 2014
Cited alongside, same era.
Coherent multi-sentence video description with variable level of detail
A. Rohrbach, M. Rohrbach, W. Qiu, A. Friedrich, M. Pinkal, and B. Schiele · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Cited alongside, same era.
Dropout: A simple way to prevent neural networks from overfitting
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov · 2014
Cited alongside, same era.
C. C. Park and G. Kim · 2015
Closest in time.
Sequence level training with recurrent neural networks
M. Ranzato, S. Chopra, M. Auli, and W. Zaremba · 2015
Closest in time.
Imagenet large scale visual recognition challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei · 2015
Closest in time.
C3D: generic features for video analysis
D. Tran, L. D. Bourdev, R. Fergus, L. Torresani, and M. Paluri · 2015
Closest in time.
Cider: Consensus-based image description evaluation
R. Vedantam, C. L. Zitnick, and D. Parikh · 2015
Closest in time.
Sequence to sequence - video to text
S. Venugopalan, M. Rohrbach, J. Donahue, R. J. Mooney, T. Darrell, and K. Saenko · 2015
Closest in time.
Translating videos to natural language using deep recurrent neural networks
S. Venugopalan, H. Xu, J. Donahue, M. Rohrbach, R. J. Mooney, and K. Saenko · 2015
Closest in time.
A neural conversational model
O. Vinyals and Q. V. Le · 2015
Closest in time.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2015
Closest in time.
Semantically conditioned lstm-based natural language generation for spoken dialogue systems
T. Wen, M. Gasic, N. Mrksic, P. Su, D. Vandyke, and S. J. Young · 2015
Closest in time.
A multi-scale multiple instance video description network
H. Xu, S. Venugopalan, V. Ramanishka, M. Rohrbach, and K. Saenko · 2015
Closest in time.
Jointly modeling deep video and compositional text to bridge vision and language in a unified framework
R. Xu, C. Xiong, W. Chen, and J. J. Corso · 2015
Closest in time.
Describing videos by exploiting temporal structure
L. Yao, A. Torabi, K. Cho, N. Ballas, C. Pal, H. Larochelle, and A. Courville · 2015
Closest in time.
Learning to describe video with weak supervision by exploiting negative sentential information
H. Yu and J. M. Siskind · 2015
Closest in time.