Fetching the paper…
Reading the bibliography…
We present a general approach to video understanding, inspired by semantic transfer techniques that have been successfully used for 2D image analysis.
Canonical ridge and econometrics of joint production
H. D. Vinod · 1976
Earlier work this paper cites.
A tutorial on hidden Markov models and selected applications in speech recognition
L. R. Rabiner · 1989
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
An integrated system for content-based video retrieval and browsing
H. J. Zhang, J. Wu, D. Zhong, and S. W. Smoliar · 1997
Earlier work this paper cites.
BLEU: a method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu · 2002
Earlier work this paper cites.
Automatic evaluation of machine translation quality using longest common subsequence and skip-bigram statistics
C.-Y. Lin and F. J. Och · 2004
Earlier work this paper cites.
On space-time interest points
I. Laptev · 2005
Earlier work this paper cites.
Video2text: Learning to annotate video content
H. Aradhye, G. Toderici, and J. Yagnik · 2009
Earlier work this paper cites.
Improving the fisher kernel for large-scale image classification
F. Perronnin, J. Sánchez, and T. Mensink · 2010
Earlier work this paper cites.
Convolutional learning of spatio-temporal features
G. W. Taylor, R. Fergus, Y. LeCun, and C. Bregler · 2010
Earlier work this paper cites.
Multimodal fusion for video search reranking
S. Wei, Y. Zhao, Z. Zhu, and N. Liu · 2010
Earlier work this paper cites.
Collecting highly parallel data for paraphrase evaluation
D. L. Chen and W. B. Dolan · 2011
Earlier work this paper cites.
Vsumm: A mechanism designed to produce static video summaries and a novel evaluation method
S. E. F. De Avila, A. P. B. Lopes, A. da Luz, and A. de Albuquerque Araújo · 2011
Earlier work this paper cites.
Hmdb: a large video database for human motion recognition
H. Kuehne, H. Jhuang, E. Garrote, T. Poggio, and T. Serre · 2011
Earlier work this paper cites.
SIFT flow: Dense correspondence across scenes and its applications
C. Liu, J. Yuen, and A. Torralba · 2011
Earlier work this paper cites.
A multi-modal clustering method for web videos
H. Huang, Y. Lu, F. Zhang, and S. Sun · 2012
Earlier work this paper cites.
Motion interchange patterns for action recognition in unconstrained videos
O. Kliper-Gross, Y. Gurovich, T. Hassner, and L. Wolf · 2012
Earlier work this paper cites.
The action similarity labeling challenge
O. Kliper-Gross, T. Hassner, and L. Wolf · 2012
Earlier work this paper cites.
Determinantal point processes for machine learning
A. Kulesza and B. Taskar · 2012
Earlier work this paper cites.
TRECVID 2012–an overview of the goals, tasks, data, evaluation mechanisms and metrics
P. Over, G. Awad, M. Michel, J. Fiscus, G. Sanders, B. Shaw, A. F. Smeaton, and G. Quénot · 2012
Earlier work this paper cites.
Ucf101: A dataset of 101 human actions classes from videos in the wild
K. Soomro, A. R. Zamir, and M. Shah · 2012
Earlier work this paper cites.
Youtube2text: Recognizing and describing arbitrary activities using semantic hierarchies and zero-shot recognition
S. Guadarrama, N. Krishnamoorthy, G. Malkarnenkar, S. Venugopalan, R. Mooney, T. Darrell, and K. Saenko · 2013
Earlier work this paper cites.
A critical review of action recognition benchmarks
T. Hassner · 2013
Earlier work this paper cites.
3d convolutional neural networks for human action recognition
S. Ji, W. Xu, M. Yang, and K. Yu · 2013
Earlier work this paper cites.
Large-scale video summarization using web-image priors
A. Khosla, R. Hamid, C.-J. Lin, and N. Sundaresan · 2013
Earlier work this paper cites.
Generating natural-language video descriptions using text-mined knowledge
N. Krishnamoorthy, G. Malkarnenkar, R. J. Mooney, K. Saenko, and S. Guadarrama · 2013
Cited alongside, same era.
Efficient estimation of word representations in vector space
T. Mikolov, K. Chen, G. Corrado, and J. Dean · 2013
Cited alongside, same era.
Deep fisher networks for large-scale image classification
K. Simonyan, A. Vedaldi, and A. Zisserman · 2013
Cited alongside, same era.
Action recognition with improved trajectories
H. Wang and C. Schmid · 2013
Cited alongside, same era.
Meteor universal: Language specific translation evaluation for any target language
M. Denkowski and A. Lavie · 2014
Cited alongside, same era.
Diverse sequential subset selection for supervised video summarization
A dataset for movie description
A. Rohrbach, M. Rohrbach, N. Tandon, and B. Schiele · 2015
Later among the works it cites.
Imagenet large scale visual recognition challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, et al · 2015
Later among the works it cites.
R. Shetty and J. Laaksonen · 2015
Later among the works it cites.
Tvsum: Summarizing web videos using titles
Y. Song, J. Vallmitjana, A. Stent, and A. Jaimes · 2015
Later among the works it cites.
Unsupervised learning of video representations using LSTMs
N. Srivastava, E. Mansimov, and R. Salakhutdinov · 2015
Later among the works it cites.
Using descriptive video services to create a large data source for video annotation research
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
B. Gong, W.-L. Chao, K. Grauman, and F. Sha · 2014
Cited alongside, same era.
Creating summaries from user videos
M. Gygli, H. Grabner, H. Riemenschneider, and L. Van Gool · 2014
Cited alongside, same era.
THUMOS challenge: Action recognition with a large number of classes
Y.-G. Jiang, J. Liu, A. Roshan Zamir, G. Toderici, I. Laptev, M. Shah, and R. Sukthankar · 2014
Cited alongside, same era.
Large-scale video classification with convolutional neural networks
A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei · 2014
Cited alongside, same era.
The lear submission at thumos 2014
D. Oneata, J. Verbeek, and C. Schmid · 2014
Cited alongside, same era.
Category-specific video summarization
D. Potapov, M. Douze, Z. Harchaoui, and C. Schmid · 2014
Cited alongside, same era.
Two-stream convolutional networks for action recognition in videos
K. Simonyan and A. Zisserman · 2014
Cited alongside, same era.
A. Torabi, C. Pal, H. Larochelle, and A. Courville · 2015
Later among the works it cites.
Learning spatiotemporal features with 3d convolutional networks
D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri · 2015
Later among the works it cites.
Cider: Consensus-based image description evaluation
R. Vedantam, C. Lawrence Zitnick, and D. Parikh · 2015
Later among the works it cites.
Sequence to sequence-video to text
S. Venugopalan, M. Rohrbach, J. Donahue, R. Mooney, T. Darrell, and K. Saenko · 2015
Later among the works it cites.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2015
Later among the works it cites.
Action recognition with trajectory-pooled deep-convolutional descriptors
L. Wang, Y. Qiao, and X. Tang · 2015
Later among the works it cites.
Describing videos by exploiting temporal structure
L. Yao, A. Torabi, K. Cho, N. Ballas, C. Pal, H. Larochelle, and A. Courville · 2015
Later among the works it cites.
Fast temporal activity proposals for efficient detection of human actions in untrimmed videos
F. Caba Heilbron, J. Carlos Niebles, and B. Ghanem · 2016
Closest in time.
Daps: Deep action proposals for action understanding
V. Escorcia, F. C. Heilbron, J. C. Niebles, and B. Ghanem · 2016
Closest in time.
Vlad3: Encoding dynamics of deep features for action recognition
Y. Li, W. Li, V. Mahadevan, and N. Vasconcelos · 2016
Closest in time.
Visually indicated sounds
A. Owens, P. Isola, J. McDermott, A. Torralba, E. H. Adelson, and W. T. Freeman · 2016
Closest in time.
Temporal action detection using a statistical language model
A. Richard and J. Gall · 2016
Closest in time.
The large scale movie description and understanding challenge (LSMDC 2016), howpublished = Available:
A. Rohrbach, A. Torabi, T. Maharaj, M. Rohrbach, C. Pal, A. Courville, and B. Schiele · 2016
Closest in time.
Movie description
A. Rohrbach, A. Torabi, M. Rohrbach, N. Tandon, P. Chris, L. Hugo, C. Aaron, and B. Schiele · 2016
Closest in time.
Temporal action localization in untrimmed videos via multi-stage cnns
Z. Shou, D. Wang, and S.-F. Chang · 2016
Closest in time.
End-to-end learning of action detection from frame glimpses in videos
S. Yeung, O. Russakovsky, G. Mori, and L. Fei-Fei · 2016
Closest in time.
Video captioning and retrieval models with semantic attention
Y. Yu, H. Ko, J. Choi, and G. Kim · 2016
Closest in time.
Temporal action localization with pyramid of score distribution features
J. Yuan, B. Ni, X. Yang, and A. A. Kassim · 2016
Closest in time.
Summary transfer: Exemplar-based subset selection for video summarizatio
K. Zhang, W.-L. Chao, F. Sha, and K. Grauman · 2016
Closest in time.
Video summarization with long short-term memory
K. Zhang, W.-L. Chao, F. Sha, and K. Grauman · 2016
Closest in time.