Fetching the paper…
Reading the bibliography…
We address the problem of fine-grained action localization from temporally untrimmed web videos.
An efficient gradient-based algorithm for on-line training of recurrent network trajectories
R. J. Williams and J. Peng · 1990
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Distinctive image features from scale-invariant keypoints
D. G. Lowe · 2004
Earlier work this paper cites.
Recognizing human actions: A local SVM approach
C. Schuldt, I. Laptev, and B. Caputo · 2004
Earlier work this paper cites.
Fisher kernels on visual vocabularies for image categorization
F. Perronnin and C. Dance · 2007
Earlier work this paper cites.
Offline handwriting recognition with multidimensional recurrent neural networks
A. Graves and J. Schmidhuber · 2008
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
HMDB: a large video database for human motion recognition
H. Kuehne, H. Jhuang, E. Garrote, T. Poggio, and T. Serre · 2011
Earlier work this paper cites.
Discriminative tag learning on youtube videos with latent sub-tags
W. Yang and G. Toderici · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
A database for fine grained activity detection of cooking activities
M. Rohrbach, S. Amin, M. Andriluka, and B. Schiele · 2012
Earlier work this paper cites.
Recognizing human-object interactions in still images by modeling the mutual context of objects and human poses
B. Yao and F. Li · 2012
Earlier work this paper cites.
Speech recognition with deep recurrent neural networks
A. Graves, A. Mohamed, and G. E. Hinton · 2013
Earlier work this paper cites.
Recommendations for video event recognition using concept vocabularies
A. Habibian, K. E. A. van de Sande, and C. G. M. Snoek · 2013
Cited alongside, same era.
Action and Event Recognition with Fisher Vectors on a Compact Feature Set
D. Oneata, J. Verbeek, and C. Schmid · 2013
Cited alongside, same era.
TRECVID 2013 – an overview of the goals, tasks, data, evaluation mechanisms and metrics
P. Over, G. Awad, M. Michel, J. Fiscus, G. Sanders, W. Kraaij, A. F. Smeaton, and G. Queenot · 2013
Cited alongside, same era.
Spatiotemporal deformable part models for action detection
Y. Tian, R. Sukthankar, and M. Shah · 2013
Cited alongside, same era.
Dense trajectories and motion boundary descriptors for action recognition
H. Wang, A. Kläser, C. Schmid, and C.-L. Liu · 2013
Cited alongside, same era.
Action Recognition with Improved Trajectories
H. Wang and C. Schmid · 2013
Category-specific video summarization
D. Potapov, M. Douze, Z. Harchaoui, and C. Schmid · 2014
Later among the works it cites.
ImageNet Large Scale Visual Recognition Challenge, 2014
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei · 2014
Later among the works it cites.
H. Sak, A. Senior, and F. Beaufays · 2014
Later among the works it cites.
Two-stream convolutional networks for action recognition in videos
K. Simonyan and A. Zisserman · 2014
Later among the works it cites.
ISOMER: Informative segment observations for multimedia event recounting
C. Sun, B. Burns, R. Nevatia, C. Snoek, B. Bolles, G. Myers, W. Wang, and E. Yeh · 2014
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Effective transfer tagging from image to video
Y. Yang, Y. Yang, and H. T. Shen · 2013
Cited alongside, same era.
Event-driven semantic concept discovery by exploiting weakly tagged internet images
J. Chen, Y. Cui, G. Ye, D. Liu, and S. Chang · 2014
Cited alongside, same era.
Learning everything about anything: Webly-supervised visual concept learning
S. K. Divvala, A. Farhadi, and C. Guestrin · 2014
Cited alongside, same era.
Action localization with tubelets from motion
M. Jain, J. van Gemert, H. Jégou, P. Bouthemy, and C. G. M. Snoek · 2014
Cited alongside, same era.
THUMOS challenge: Action recognition with a large number of classes
Y.-G. Jiang, J. Liu, A. Roshan Zamir, G. Toderici, I. Laptev, M. Shah, and R. Sukthankar · 2014
Cited alongside, same era.
Large-scale video classification with convolutional neural networks
A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei · 2014
Cited alongside, same era.
I. Sutskever, O. Vinyals, and Q. V. Le · 2014
Later among the works it cites.
Video action detection with relational dynamic-poselets
L. Wang, Y. Qiao, and X. Tang · 2014
Later among the works it cites.
Long-term recurrent convolutional networks for visual recognition and description
J. Donahue, L. A. Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell · 2015
Closest in time.
Unifying visual-semantic embeddings with multimodal neural language models
R. Kiros, R. Salakhutdinov, and R. S. Zemel · 2015
Closest in time.
Unsupervised learning of video representations using LSTMs
N. Srivastava, E. Mansimov, and R. Salakhutdinov · 2015
Closest in time.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2015
Closest in time.
Beyond short snippets: Deep networks for video classification
J. Yue-Hei Ng, M. Hausknecht, S. Vijayanarasimhan, O. Vinyals, R. Monga, and G. Toderici · 2015
Closest in time.