Fetching the paper…
Reading the bibliography…
Capturing spatiotemporal dynamics is an essential topic in video recognition.
Histograms of oriented gradients for human detection
Navneet Dalal and Bill Triggs · 2005
Earlier work this paper cites.
Conditional random fields for object recognition
Ariadna Quattoni, Michael Collins, and Trevor Darrell · 2005
Earlier work this paper cites.
Conditional random fields for activity recognition
Douglas L Vail, Manuela M Veloso, and John D Lafferty · 2007
Earlier work this paper cites.
Learning realistic human actions from movies
Ivan Laptev, Marcin Marszałek, Cordelia Schmid, and Benjamin Rozenfeld · 2008
Earlier work this paper cites.
Actions in context
Marcin Marszałek, Ivan Laptev, and Cordelia Schmid · 2009
Earlier work this paper cites.
Hierarchical spatio-temporal context modeling for action recognition
Ju Sun, Xiao Wu, Shuicheng Yan, Loong-Fah Cheong, Tat-Seng Chua, and Jintao Li · 2009
Earlier work this paper cites.
Context based object categorization: A critical survey
Carolina Galleguillos and Serge Belongie · 2010
Earlier work this paper cites.
Learning a hierarchy of discriminative space-time neighborhood features for human action recognition
Adriana Kovashka and Kristen Grauman · 2010
Earlier work this paper cites.
Action recognition by dense trajectories
Heng Wang, Alexander Kläser, Cordelia Schmid, and Liu Cheng-Lin · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Actionness ranking with lattice conditional ordinal random fields
Wei Chen, Caiming Xiong, Ran Xu, and Jason J Corso · 2014
Earlier work this paper cites.
Action recognition with stacked fisher vectors
Xiaojiang Peng, Changqing Zou, Yu Qiao, and Qiang Peng · 2014
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
K. Simonyan and A. Zisserman · 2014
Earlier work this paper cites.
Spatio-temporal triangular-chain crf for activity recognition
Congqi Cao, Yifan Zhang, and Hanqing Lu · 2015
Earlier work this paper cites.
Action recognition using visual attention
Shikhar Sharma, Ryan Kiros, and Ruslan Salakhutdinov · 2015
Earlier work this paper cites.
Learning spatiotemporal features with 3d con- volutional networks
D. Tran, L. D. Bourdev, R. Fergus, L. Torresani, and M. Paluri · 2015
Cited alongside, same era.
David Ha, Andrew Dai, and Quoc V Le · 2016
Cited alongside, same era.
Hollywood in homes: Crowdsourcing data collection for activity understanding
G. A. Sigurdsson, G. Varol, X. Wang, A. Farhadi, I. Laptev, and A. Gupta · 2016
Cited alongside, same era.
Temporal segment networks: Towards good practices for deep action recognition
L. Wang, Y. Xiong, Z. Wang, Y. Qiao, D. Lin, X. Tang, and L. Van Gool · 2016
Cited alongside, same era.
Quo vadis, action recognition? a new model and the kinetics dataset
J. Carreira and A. Zisserman · 2017
Cited alongside, same era.
Deformable convolutional networks
Generating neural networks with neural networks
Lior Deutsch · 2018
Closest in time.
Fine-grained video classification and captioning
F. Mahdisoltani, G. Berger, W. Gharbieh, D. Fleet, and R. Memisevic · 2018
Closest in time.
A closer look at spatiotem- poral convolutions for action recognition
D. Tran, H. Wang, L. Torresani, J. Ray, Y. LeCun, and M. Paluri · 2018
Closest in time.
Dividing and aggregating network for multi-view action recognition
Dongang Wang, Wanli Ouyang, Wen Li, and Dong Xu · 2018
Closest in time.
Appearance-and-relation networks for video classification
Limin Wang, Wei Li, Wen Li, and Luc Van Gool · 2018
Closest in time.
Non-local neural networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jifeng Dai, Haozhi Qi, Yuwen Xiong, Yi Li, Guodong Zhang, Han Hu, and Yichen Wei · 2017
Cited alongside, same era.
Attentional pooling for action recognition
Rohit Girdhar and Deva Ramanan · 2017
Cited alongside, same era.
Self-normalizing neural networks
Günter Klambauer, Thomas Unterthiner, Andreas Mayr, and Sepp Hochreiter · 2017
Cited alongside, same era.
Learnable pooling with context gating for video classification
Antoine Miech, Ivan Laptev, and Josef Sivic · 2017
Cited alongside, same era.
Learning spatio-temporal representation with pseudo-3d residual networks
Z. Qiu, T. Yao, and T. Mei · 2017
Cited alongside, same era.
Learning spatio-temporal representation with pseudo-3d residual networks
Zhaofan Qiu, Ting Yao, and Tao Mei · 2017
Cited alongside, same era.
Asynchronous temporal fields for action recognition
G. A. Sigurdsson, S. Divvala, A. Farhadi, and A. Gupta · 2017
Cited alongside, same era.
Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaiming He · 2018
Closest in time.
Videos as space-time region graphs
X. Wang and A. Gupta · 2018
Closest in time.
Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification
S. Xie, C. Sun, J. Huang, Z. Tu, and K. Murphy · 2018
Closest in time.
Temporal relational reasoning in videos
Andonian A. Torralba A. Zhou, B · 2018
Closest in time.
Eco: Efficient convolutional network for online video understanding
M. Zolfaghari, K. Singh, and T. Brox · 2018
Closest in time.
Gcnet: Non-local networks meet squeeze-excitation networks and beyond
Yue Cao, Jiarui Xu, Stephen Lin, Fangyun Wei, and Han Hu · 2019
Closest in time.
Mars: Motion-augmented rgb stream for action recognition
Nieves Crasto, Philippe Weinzaepfel, Karteek Alahari, and Cordelia Schmid · 2019
Closest in time.
Video action transformer network
Rohit Girdhar, Joao Carreira, Carl Doersch, and Andrew Zisserman · 2019
Closest in time.
Learning spatio-temporal representation with local and global diffusion
Zhaofan Qiu, Ting Yao, Chong-Wah Ngo, Xinmei Tian, and Tao Mei · 2019
Closest in time.
Pay less attention with lightweight and dynamic convolutions
Felix Wu, Angela Fan, Alexei Baevski, Yann N Dauphin, and Michael Auli · 2019
Closest in time.