Fetching the paper…
Reading the bibliography…
Many recent advancements in Computer Vision are attributed to large datasets.
Hierarchical mixtures of experts and the em algorithm
M. I. Jordan · 1994
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Space-time interest points
I. Laptev and T. Lindeberg · 2003
Earlier work this paper cites.
Actions as space-time shapes
M. Blank, L. Gorelick, E. Shechtman, M. Irani, and R. Basri · 2005
Earlier work this paper cites.
One-shot learning of object categories
L. Fei-fei, R. Fergus, and P. Perona · 2006
Earlier work this paper cites.
Textonboost: Joint appearance, shape and context modeling for multi-class object
J. Shotton, J. Winn, C. Rother, and A. Criminisi · 2006
Earlier work this paper cites.
Caltech-256 object category dataset
G. Griffin, A. Holub, and P. Perona · 2007
Earlier work this paper cites.
Fisher kernels on visual vocabularies for image categorization
F. Perronnin and C. Dance · 2007
Earlier work this paper cites.
Learning realistic human actions from movies
I. Laptev, M. Marszalek, C. Schmid, and B. Rozenfeld · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L. jia Li, K. Li, and L. Fei-fei · 2009
Earlier work this paper cites.
The pascal visual object classes (voc) challenge, 2009
M. Everingham, L. V. Gool, C. K. I. Williams, J. Winn, and A. Zisserman · 2009
Earlier work this paper cites.
Recognizing indoor scenes
A. Quattoni and A. Torralba · 2009
Earlier work this paper cites.
Evaluation of local spatio-temporal features for action recognition
H. Wang, M. M. Ullah, A. Kläser, I. Laptev, and C. Schmid · 2009
Earlier work this paper cites.
Hmdb: a large video database for human motion recognition
H. Kuehne, H. Jhuang, E. Garrote, T. Poggio, and T. Serre · 2011
Cited alongside, same era.
Aggregating local image descriptors into compact codes
H. Jegou, F. Perronnin, M. Douze, J. Sanchez, P. Perez, and C. Schmid · 2012
Cited alongside, same era.
ImageNet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Cited alongside, same era.
Learning to label aerial images from noisy data
V. Mnih and G. Hinton · 2012
Cited alongside, same era.
UCF101: A dataset of 101 human actions classes from videos in the wild
K. Soomro, A. R. Zamir, and M. Shah · 2012
Cited alongside, same era.
https://www.youtube.com/watch?v=wf_77z1H-vQ
Google I/O 2013 - semantic video annotations in the Youtube Topics API: Theory and applications · 2013
Cited alongside, same era.
Mean-normalized stochastic gradient for large-scale deep learning
S. Wiesler, A. Richard, R. Schlüter, and H. Ney · 2014
Later among the works it cites.
Large-scale multi-label learning with missing labels
H.-F. Yu, P. Jain, P. Kar, and I. Dhillon · 2014
Later among the works it cites.
Fast R-CNN
R. Girshick · 2015
Later among the works it cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Later among the works it cites.
Activitynet: A large-scale video benchmark for human activity understanding
F. C. Heilbron, V. Escorcia, B. Ghanem, and J. C. Niebles · 2015
Later among the works it cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sun database: Exploring a large collection of scene categories, 2013
J. Xiao, K. A. Ehinger, J. Hays, A. Torralba, A. Oliva, and J. Xiao · 2013
Cited alongside, same era.
Visualizing and understanding convolutional networks
M. D. Zeiler and R. Fergus · 2013
Cited alongside, same era.
THUMOS challenge: Action recognition with a large number of classes
Y. Jiang, J. Liu, A. Roshan Zamir, G. Toderici, I. Laptev, M. Shah, and R. Sukthankar · 2014
Cited alongside, same era.
Large-scale video classification with convolutional neural networks
A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei · 2014
Cited alongside, same era.
Training deep neural networks on noisy labels with bootstrapping
S. Reed, H. Lee, D. Anguelov, C. Szegedy, D. Erhan, and A. Rabinovich · 2014
Cited alongside, same era.
C3D: generic features for video analysis
D. Tran, L. D. Bourdev, R. Fergus, L. Torresani, and M. Paluri · 2014
Cited alongside, same era.
Later among the works it cites.
Y.-G. Jiang, Z. Wu, J. Wang, X. Xue, and S.-F. Chang · 2015
Later among the works it cites.
Do less and achieve more: Training cnns for action recognition utilizing action images from the web
S. Ma, S. A. Bargal, J. Zhang, L. Sigal, and S. Sclaroff · 2015
Later among the works it cites.
Beyond short snippets: Deep networks for video classification
J. Y.-H. Ng, M. J. Hausknecht, S. Vijayanarasimhan, O. Vinyals, R. Monga, and G. Toderici · 2015
Later among the works it cites.
ImageNet Large Scale Visual Recognition Challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei · 2015
Later among the works it cites.
The new data and new challenges in multimedia research
B. Thomee, D. A. Shamma, G. Friedland, B. Elizalde, K. Ni, D. Poland, D. Borth, and L. Li · 2015
Later among the works it cites.
A discriminative cnn video representation for event detection
Z. Xu, Y. Yang, and A. G. Hauptmann · 2015
Later among the works it cites.