Fetching the paper…
Reading the bibliography…
In this work, we introduce a dataset of video annotated with high quality natural language phrases describing the visual content in a given segment of time.
Feature-rich part-of-speech tagging with a cyclic dependency network
K. Toutanova, D. Klein, C. D. Manning, and Y. Singer · 2003
Earlier work this paper cites.
Collecting highly parallel data for paraphrase evaluation
D. L. Chen and W. B. Dolan · 2011
Earlier work this paper cites.
Large scale image annotations on amazon mechanical turk
S. Maji · 2011
Earlier work this paper cites.
Im2text: Describing images using 1 million captioned photographs
V. Ordonez, G. Kulkarni, and T. L. Berg · 2011
Earlier work this paper cites.
Youtube2text: Recognizing and describing arbitrary activities using semantic hierarchies and zero-shot recognition
S. Guadarrama, N. Krishnamoorthy, G. Malkarnenkar, S. Venugopalan, R. Mooney, T. Darrell, and K. Saenko · 2013
Earlier work this paper cites.
Translating video content to natural language descriptions
M. Rohrbach, W. Qiu, I. Titov, S. Thater, M. Pinkal, and B. Schiele · 2013
Earlier work this paper cites.
Long-term recurrent convolutional networks for visual recognition and description
J. Donahue, L. A. Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell · 2014
Cited alongside, same era.
Large-scale video classification with convolutional neural networks
A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei · 2014
Cited alongside, same era.
Unifying visual-semantic embeddings with multimodal neural language models
R. Kiros, R. Salakhutdinov, and R. S. Zemel · 2014
Cited alongside, same era.
Microsoft COCO: common objects in context
T. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Cited alongside, same era.
Coherent multi-sentence video description with variable level of detail
A. Senina, M. Rohrbach, W. Qiu, A. Friedrich, S. Amin, M. Andriluka, M. Pinkal, and B. Schiele · 2014
Translating videos to natural language using deep recurrent neural networks
S. Venugopalan, H. Xu, J. Donahue, M. Rohrbach, R. J. Mooney, and K. Saenko · 2014
Later among the works it cites.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2014
Later among the works it cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
P. Young, A. Lai, M. Hodosh, and J. Hockenmaier · 2014
Later among the works it cites.
A dataset for movie description
N. T. B. S. A Rohrbach, M Rohrbach · 2015
Closest in time.
Video description generation incorporating spatio-temporal features and a soft-attention mechanism
L. Yao, A. Torabi, K. Cho, N. Ballas, C. Pal, H. Larochelle, and A. Courville · 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
https://creative.adobe.com/products/audition
Adobe audition, audio editing software
Cited in the paper.
http://www.amazon.com/
Amazon.com: Online shopping
Cited in the paper.
http://www.acb.org/
American council of blind
Cited in the paper.
http://www.crtc.gc.ca/eng/publications/reports/rp120229.htm
Canadian radio-television and telecommunications commission (crtc) list of tv channels with described video
Cited in the paper.
https://www.hmv.ca/
Hmvl home of entetainment
Cited in the paper.
http://main.wgbh.org/wgbh/pages/mag/services/description/dvs-faq.html
Media access group at wgbh
Cited in the paper.
http://transcribeme.com
Transcribeme professional transctiption
Cited in the paper.
Closest in time.