Fetching the paper…
Reading the bibliography…
Training robust deep video representations has proven to be much more challenging than learning deep image representations.
Mpeg: A video compression standard for multimedia applications
D. Le Gall · 1991
Earlier work this paper cites.
Rapid scene analysis on compressed video
B.-L. Yeo and B. Liu · 1995
Earlier work this paper cites.
Fast object detection and segmentation in MPEG compressed domain
O. Sukmarg and K. R. Rao · 2000
Earlier work this paper cites.
Video codec design: developing image and video compression systems
I. E. Richardson · 2002
Earlier work this paper cites.
Histograms of oriented gradients for human detection
N. Dalal and B. Triggs · 2005
Earlier work this paper cites.
Moving object detection in wavelet compressed video
B. U. Töreyin, A. E. Cetin, A. Aksay, and M. B. Akhan · 2005
Earlier work this paper cites.
A duality based approach for realtime tv-l 1 optical flow
C. Zach, T. Pock, and H. Bischof · 2007
Earlier work this paper cites.
Learning realistic human actions from movies
I. Laptev, M. Marszalek, C. Schmid, and B. Rozenfeld · 2008
Earlier work this paper cites.
Visualizing data using t-SNE
L. v. d. Maaten and G. Hinton · 2008
Earlier work this paper cites.
Detailed real-time urban 3d reconstruction from video
M. Pollefeys, D. Nistér, J.-M. Frahm, A. Akbarzadeh, P. Mordohai, B. Clipp, C. Engels, D. Gallup, S.-J. Kim, P. Merrell, et al · 2008
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Evaluation of local spatio-temporal features for action recognition
H. Wang, M. M. Ullah, A. Klaser, I. Laptev, and C. Schmid · 2009
Earlier work this paper cites.
HMDB: a large video database for human motion recognition
H. Kuehne, H. Jhuang, E. Garrote, T. Poggio, and T. Serre · 2011
Earlier work this paper cites.
UCF101: A dataset of 101 human actions classes from videos in the wild
K. Soomro, A. Roshan Zamir, and M. Shah · 2012
Earlier work this paper cites.
Dense trajectories and motion boundary descriptors for action recognition
H. Wang, A. Kläser, C. Schmid, and C.-L. Liu · 2013
Earlier work this paper cites.
Action recognition with improved trajectories
H. Wang and C. Schmid · 2013
Earlier work this paper cites.
Efficient feature extraction, encoding and classification for action recognition
V. Kantorov and I. Laptev · 2014
Earlier work this paper cites.
Large-scale video classification with convolutional neural networks
A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2014
Cited alongside, same era.
Action recognition with stacked fisher vectors
X. Peng, C. Zou, Y. Qiao, and Q. Peng · 2014
Cited alongside, same era.
Two-stream convolutional networks for action recognition in videos
K. Simonyan and A. Zisserman · 2014
Cited alongside, same era.
Long-term recurrent convolutional networks for visual recognition and description
J. Donahue, L. Anne Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Forecast and methodology, 2016-2021, white paper
C. V. networking Index · 2016
Later among the works it cites.
Hollywood in homes: Crowdsourcing data collection for activity understanding
G. A. Sigurdsson, G. Varol, X. Wang, A. Farhadi, I. Laptev, and A. Gupta · 2016
Later among the works it cites.
Temporal segment networks: Towards good practices for deep action recognition
L. Wang, Y. Xiong, Z. Wang, Y. Qiao, D. Lin, X. Tang, and L. Van Gool · 2016
Later among the works it cites.
Real-time action recognition with enhanced motion vector CNNs
B. Zhang, L. Wang, Z. Wang, Y. Qiao, and H. Wang · 2016
Later among the works it cites.
Quo vadis, action recognition? a new model and the kinetics dataset
J. Carreira and A. Zisserman · 2017
Closest in time.
Deep temporal linear encoding networks
A. Diba, V. Sharma, and L. Van Gool · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Ioffe and C. Szegedy · 2015
Cited alongside, same era.
ImageNet large scale visual recognition challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei · 2015
Cited alongside, same era.
Unsupervised learning of video representations using lstms
N. Srivastava, E. Mansimov, and R. Salakhudinov · 2015
Cited alongside, same era.
Learning spatiotemporal features with 3d convolutional networks
D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri · 2015
Cited alongside, same era.
Convolutional LSTM network: A machine learning approach for precipitation nowcasting
S. Xingjian, Z. Chen, H. Wang, D.-Y. Yeung, W.-K. Wong, and W.-c. Woo · 2015
Cited alongside, same era.
Beyond short snippets: Deep networks for video classification
J. Yue-Hei Ng, M. Hausknecht, S. Vijayanarasimhan, O. Vinyals, R. Monga, and G. Toderici · 2015
Cited alongside, same era.
YouTube-8M: A large-scale video classification benchmark
S. Abu-El-Haija, N. Kothari, J. Lee, P. Natsev, G. Toderici, B. Varadarajan, and S. Vijayanarasimhan · 2016
Cited alongside, same era.
Spatiotemporal multiplier networks for video action recognition
C. Feichtenhofer, A. Pinz, and R. P. Wildes · 2017
Closest in time.
Attentional pooling for action recognition
R. Girdhar and D. Ramanan · 2017
Closest in time.
Actionvlad: Learning spatio-temporal aggregation for action classification
R. Girdhar, D. Ramanan, A. Gupta, J. Sivic, and B. Russell · 2017
Closest in time.
TS-LSTM and temporal-inception: Exploiting spatiotemporal dynamics for activity recognition
C.-Y. Ma, M.-H. Chen, Z. Kira, and G. AlRegib · 2017
Closest in time.
Learning spatio-temporal representation with pseudo-3d residual networks
Z. Qiu, T. Yao, and T. Mei · 2017
Closest in time.
Learning long-term dependencies for action recognition with a biologically-inspired deep network
Y. Shi, Y. Tian, Y. Wang, W. Zeng, and T. Huang · 2017
Closest in time.
Asynchronous temporal fields for action recognition
G. A. Sigurdsson, S. Divvala, A. Farhadi, and A. Gupta · 2017
Closest in time.
Lattice long short-term memory for human action recognition
L. Sun, K. Jia, K. Chen, D.-Y. Yeung, B. E. Shi, and S. Savarese · 2017
Closest in time.
ConvNet architecture search for spatiotemporal feature learning
D. Tran, J. Ray, Z. Shou, S.-F. Chang, and M. Paluri · 2017
Closest in time.
Spatiotemporal pyramid network for video action recognition
Y. Wang, M. Long, J. Wang, and P. S. Yu · 2017
Closest in time.