Fetching the paper…
Reading the bibliography…
Real-world videos often have complex dynamics; and methods for generating open-domain video descriptions should be sensitive to temporal structure and allow both input (sequence of frames) and output (sequence of words) of variable length.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu · 2002
Earlier work this paper cites.
High accuracy optical flow estimation based on a theory for warping
T. Brox, A. Bruhn, N. Papenberg, and J. Weickert · 2004
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
C.-Y. Lin · 2004
Earlier work this paper cites.
Video2text: Learning to annotate video content
H. Aradhye, G. Toderici, and J. Yagnik · 2009
Earlier work this paper cites.
Multimodal fusion for video search reranking
S. Wei, Y. Zhao, Z. Zhu, and N. Liu · 2010
Earlier work this paper cites.
Collecting highly parallel data for paraphrase evaluation
D. L. Chen and W. B. Dolan · 2011
Earlier work this paper cites.
TRECVID 2012 – an overview of the goals, tasks, data, evaluation mechanisms and metrics
P. Over, G. Awad, M. Michel, J. Fiscus, G. Sanders, B. Shaw, A. F. Smeaton, and G. Quéenot · 2012
Earlier work this paper cites.
Youtube2text: Recognizing and describing arbitrary activities using semantic hierarchies and zero-shoot recognition
S. Guadarrama, N. Krishnamoorthy, G. Malkarnenkar, S. Venugopalan, R. Mooney, T. Darrell, and K. Saenko · 2013
Earlier work this paper cites.
A multi-modal clustering method for web videos
H. Huang, Y. Lu, F. Zhang, and S. Sun · 2013
Earlier work this paper cites.
Generating natural-language video descriptions using text-mined knowledge
N. Krishnamoorthy, G. Malkarnenkar, R. J. Mooney, K. Saenko, and S. Guadarrama · 2013
Earlier work this paper cites.
Translating video content to natural language descriptions
M. Rohrbach, W. Qiu, I. Titov, S. Thater, M. Pinkal, and B. Schiele · 2013
Earlier work this paper cites.
Action recognition with improved trajectories
H. Wang and C. Schmid · 2013
Earlier work this paper cites.
On the properties of neural machine translation: Encoder-decoder approaches
K. Cho, B. van Merriënboer, D. Bahdanau, and Y. Bengio · 2014
Earlier work this paper cites.
Meteor universal: Language specific translation evaluation for any target language
M. Denkowski and A. Lavie · 2014
Earlier work this paper cites.
Finding action tubes
G. Gkioxari and J. Malik · 2014
Cited alongside, same era.
Towards end-to-end speech recognition with recurrent neural networks
A. Graves and N. Jaitly · 2014
Cited alongside, same era.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
P. Hodosh, A. Young, M. Lai, and J. Hockenmaier · 2014
Cited alongside, same era.
Caffe: Convolutional architecture for fast feature embedding
Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2014
Cited alongside, same era.
Unifying visual-semantic embeddings with multimodal neural language models
Microsoft COCO captions: Data collection and evaluation server
X. Chen, H. Fang, T.-Y. Lin, R. Vedantam, S. Gupta, P. Dollar, and C. L. Zitnick · 2015
Closest in time.
Learning a recurrent visual representation for image caption generation
X. Chen and C. L. Zitnick · 2015
Closest in time.
Long-term recurrent convolutional networks for visual recognition and description
J. Donahue, L. A. Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell · 2015
Closest in time.
Deep visual-semantic alignments for generating image descriptions
A. Karpathy and L. Fei-Fei · 2015
Closest in time.
Beyond short snippets: Deep networks for video classification
J. Y. Ng, M. J. Hausknecht, S. Vijayanarasimhan, O. Vinyals, R. Monga, and G. Toderici · 2015
Closest in time.
The long-short story of movie description
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
R. Kiros, R. Salakhutdinov, and R. S. Zemel · 2014
Cited alongside, same era.
Treetalk: Composition and compression of trees for image descriptions
P. Kuznetsova, V. Ordonez, T. L. Berg, U. C. Hill, and Y. Choi · 2014
Cited alongside, same era.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Cited alongside, same era.
Deep captioning with multimodal recurrent neural networks (m-rnn)
J. Mao, W. Xu, Y. Yang, J. Wang, and A. L. Yuille · 2014
Cited alongside, same era.
ILSVRC, 2014
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei · 2014
Cited alongside, same era.
Two-stream convolutional networks for action recognition in videos
K. Simonyan and A. Zisserman · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Cited alongside, same era.
A. Rohrbach, M. Rohrbach, and B. Schiele · 2015
Closest in time.
A dataset for movie description
A. Rohrbach, M. Rohrbach, N. Tandon, and B. Schiele · 2015
Closest in time.
Unsupervised learning of video representations using LSTMs
N. Srivastava, E. Mansimov, and R. Salakhutdinov · 2015
Closest in time.
Going deeper with convolutions
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2015
Closest in time.
Using descriptive video services to create a large data source for video annotation research
A. Torabi, C. Pal, H. Larochelle, and A. Courville · 2015
Closest in time.
CIDEr: Consensus-based image description evaluation
R. Vedantam, C. L. Zitnick, and D. Parikh · 2015
Closest in time.
Translating videos to natural language using deep recurrent neural networks
S. Venugopalan, H. Xu, J. Donahue, M. Rohrbach, R. Mooney, and K. Saenko · 2015
Closest in time.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2015
Closest in time.
Describing videos by exploiting temporal structure
L. Yao, A. Torabi, K. Cho, N. Ballas, C. Pal, H. Larochelle, and A. Courville · 2015
Closest in time.