Fetching the paper…
Reading the bibliography…
Video action recognition is one of the representative tasks for video understanding.
Determining Optical Flow
Berthold K.P. Horn and Brian G. Rhunck · 1981
Earlier work this paper cites.
Separate Visual Pathways for Perception and Action
M. A. Goodale and A. D. Milner · 1992
Earlier work this paper cites.
Long Short-Term Memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
A Duality based Approach for Realtime TV-L1 Optical Flow
Christopher Zach, Thomas Pock, and Horst Bischof · 2007
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Aggregating Local Descriptors into a Compact Image Representation
Hervé Jégou, Matthijs Douze, Cordelia Schmid, and Patrick Pérez · 2010
Earlier work this paper cites.
Convolutional Learning of Spatio-temporal Features
Graham W. Taylor, Rob Fergus, Yann LeCun, and Christoph Bregler · 2010
Earlier work this paper cites.
Sequential Deep Learning for Human Action Recognition
Moez Baccouche, Franck Mamalet, Christian Wolf, Christophe Garcia, and Atilla Baskurt · 2011
Earlier work this paper cites.
HMDB: A Large Video Database for Human Motion Recognition
Hildegard Kuehne, Hueihan Jhuang, Estíbaliz Garrote, Tomaso Poggio, and Thomas Serre · 2011
Earlier work this paper cites.
Multimodal deep learning
Jiquan Ngiam, Aditya Khosla, Mingyu Kim, Juhan Nam, Honglak Lee, and Andrew Y Ng · 2011
Earlier work this paper cites.
Action Recognition by Dense Trajectories
Heng Wang, Alexander Kläser, Cordelia Schmid, and Liu Cheng-Lin · 2011
Earlier work this paper cites.
3D Convolutional Neural Networks for Human Action Recognition
Shuiwang Ji, Wei Xu, Ming Yang, and Kai Yu · 2012
Earlier work this paper cites.
ImageNet Classification with Deep Convolutional Neural Networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton · 2012
Earlier work this paper cites.
UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah · 2012
Earlier work this paper cites.
Image Classification with the Fisher Vector: Theory and Practice
Jorge Sanchez, Florent Perronnin, Thomas Mensink, and Jakob Verbeek · 2013
Earlier work this paper cites.
Intriguing Properties of Neural Networks
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus · 2013
Earlier work this paper cites.
Action Recognition with Improved Trajectories
Heng Wang and Cordelia Schmid · 2013
Earlier work this paper cites.
Large-Scale Video Classification with Convolutional Neural Networks
Andrej Karpathy, George Toderici, Sanketh Shetty, Thomas Leung, Rahul Sukthankar, and Li Fei-Fei · 2014
Earlier work this paper cites.
The language of actions: Recovering the syntax and semantics of goal-directed human activities
Hilde Kuehne, Ali Arslan, and Thomas Serre · 2014
Earlier work this paper cites.
Bag of Visual Words and Fusion Methods for Action Recognition: Comprehensive Study and Good Practice
Xiaojiang Peng, Limin Wang, Xingxing Wang, and Yu Qiao · 2014
Earlier work this paper cites.
Action Recognition with Stacked Fisher Vectors
Xiaojiang Peng, Changqing Zou, Yu Qiao, and Qiang Peng · 2014
Earlier work this paper cites.
Two-Stream Convolutional Networks for Action Recognition in Videos
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Human Action Recognition across Datasets by Foreground-Weighted Histogram Decomposition
Waqas Sultani and Imran Saleemi · 2014
Earlier work this paper cites.
Webly Supervised Learning of Convolutional Networks
Xinlei Chen and Abhinav Gupta · 2015
Earlier work this paper cites.
P-CNN: Pose-based CNN Features for Action Recognition
Guilhem Cheron, Ivan Laptev, and Cordelia Schmid · 2015
Earlier work this paper cites.
Long-term Recurrent Convolutional Networks for Visual Recognition and Description
Jeff Donahue, Lisa Anne Hendricks, Marcus Rohrbach, Subhashini Venugopalan, Sergio Guadarrama, Kate Saenko, and Trevor Darrell · 2015
Earlier work this paper cites.
ActivityNet: A Large-Scale Video Benchmark for Human Activity Understanding
Bernard Ghanem Fabian Caba Heilbron, Victor Escorcia and Juan Carlos Niebles · 2015
Earlier work this paper cites.
Modeling Video Evolution For Action Recognition
Basura Fernando, Efstratios Gavves, Jose Oramas M., Amir Ghodrati, and Tinne Tuytelaars · 2015
Earlier work this paper cites.
Exploring Semantic Inter-Class Relationships (SIR) for Zero-Shot Action Recognition
Chuang Gan, Ming Lin, Yi Yang, Yueting Zhuang, and Alexander G.Hauptmann · 2015
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Ian Goodfellow, Jonathon Shlens, and Christian Szegedy · 2015
Earlier work this paper cites.
Objects2action: Classifying and Localizing Actions without Any Video Example
Mihir Jain, Jan C van Gemert, Thomas Mensink, and Cees GM Snoek · 2015
Earlier work this paper cites.
Beyond Gaussian Pyramid: Multi-skip Feature Stacking for Action Recognition
Zhenzhong Lan, Ming Lin, Xuanchong Li, Alexander G. Hauptmann, and Bhiksha Raj · 2015
Earlier work this paper cites.
Zhenzhong Lan, Dezhong Yao, Ming Lin, Shoou-I Yu, and Alexander Hauptmann · 2015
Earlier work this paper cites.
Joint Action Recognition and Pose Estimation From Video
Bruce Xiaohan Nie, Caiming Xiong, and Song-Chun Zhu · 2015
Earlier work this paper cites.
Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Earlier work this paper cites.
Recognizing fine-grained and composite activities using hand-centric features and script data
Marcus Rohrbach, Anna Rohrbach, Michaela Regneri, Sikandar Amin, Mykhaylo Andriluka, Manfred Pinkal, and Bernt Schiele · 2015
Earlier work this paper cites.
Very Deep Convolutional Networks for Large-Scale Image Recognition
Karen Simonyan and Andrew Zisserman · 2015
Earlier work this paper cites.
Going Deeper with Convolutions
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich · 2015
Earlier work this paper cites.
Learning Spatiotemporal Features with 3D Convolutional Networks
Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri · 2015
Earlier work this paper cites.
Action Recognition With Trajectory-Pooled Deep-Convolutional Descriptors
Limin Wang, Yu Qiao, and Xiaoou Tang · 2015
Earlier work this paper cites.
CUHK and SIAT Submission for THUMOS15 Action Recognition Challenge
Limin Wang, Zhe Wang, Yuanjun Xiong, and Yu Qiao · 2015
Earlier work this paper cites.
Towards Good Practices for Very Deep Two-Stream ConvNets
Limin Wang, Yuanjun Xiong, Zhe Wang, and Yu Qiao · 2015
Earlier work this paper cites.
A Discriminative CNN Video Representation for Event Detection
Zhongwen Xu, Yi Yang, and Alexander G. Hauptmann · 2015
Earlier work this paper cites.
Describing Videos by Exploiting Temporal Structure
Li Yao, Atousa Torabi, Kyunghyun Cho, Nicolas Ballas, Christopher Pal, Hugo Larochelle, and Aaron Courville · 2015
Earlier work this paper cites.
Beyond Short Snippets: Deep Networks for Video Classification
Joe Yue-Hei Ng, Matthew Hausknecht, Sudheendra Vijayanarasimhan, Oriol Vinyals, Rajat Monga, and George Toderici · 2015
Earlier work this paper cites.
YouTube-8M: A Large-Scale Video Classification Benchmark
Sami Abu-El-Haija, Nisarg Kothari, Joonseok Lee, Paul Natsev, George Toderici, Balakrishnan Varadarajan, and Sudheendra Vijayanarasimhan · 2016
Earlier work this paper cites.
NetVLAD: CNN Architecture for Weakly Supervised Place Recognition
Relja Arandjelović, Petr Gronat, Akihiko Torii, Tomas Pajdla, and Josef Sivic · 2016
Earlier work this paper cites.
Dynamic Image Networks for Action Recognition
Hakan Bilen, Basura Fernando, Efstratios Gavves, Andrea Vedaldi, and Stephen Gould · 2016
Earlier work this paper cites.
Efficient Two-Stream Motion and Appearance 3D CNNs for Video Classification
Ali Diba, Ali Mohammad Pazandeh, and Luc Van Gool · 2016
Earlier work this paper cites.
Spatiotemporal Residual Networks for Video Action Recognition
Christoph Feichtenhofer, Axel Pinz, and Richard P. Wildes · 2016
Earlier work this paper cites.
Convolutional Two-Stream Network Fusion for Video Action Recognition
Christoph Feichtenhofer, Axel Pinz, and Andrew Zisserman · 2016
Earlier work this paper cites.
Discriminative Hierarchical Rank Pooling for Activity Recognition
Basura Fernando, Peter Anderson, Marcus Hutter, and Stephen Gould · 2016
Earlier work this paper cites.
Learning End-to-end Video Classification with Rank-Pooling
Basura Fernando and Stephen Gould · 2016
Earlier work this paper cites.
Webly-supervised video recognition by mutually voting for relevant web images and web video frames
Chuang Gan, Chen Sun, Lixin Duan, and Boqing Gong · 2016
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Going Deeper into Action Recognition: A Survey
Samitha Herath, Mehrtash Harandi, and Fatih Porikli · 2016
Earlier work this paper cites.
Action Recognition by Learning Deep Multi-Granular Spatio-Temporal Video Representation
Qing Li, Zhaofan Qiu, Ting Yao, Tao Mei, Yong Rui, and Jiebo Luo · 2016
Earlier work this paper cites.
VLAD3: Encoding Dynamics of Deep Features for Action Recognition
Yingwei Li, Weixin Li, Vijay Mahadevan, and Nuno Vasconcelos · 2016
Earlier work this paper cites.
Going deeper into first-person activity recognition
Minghuang Ma, Haoqi Fan, and Kris M Kitani · 2016
Earlier work this paper cites.
Shuffle and learn: unsupervised learning using temporal order verification
Ishan Misra, C Lawrence Zitnick, and Martial Hebert · 2016
Earlier work this paper cites.
Context encoders: Feature learning by inpainting
Deepak Pathak, Philipp Krähenbühl, Jeff Donahue, Trevor Darrell, and Alexei A. Efros · 2016
Earlier work this paper cites.
Hollywood in Homes: Crowdsourcing Data Collection for Activity Understanding
Gunnar A. Sigurdsson, Gül Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta · 2016
Earlier work this paper cites.
A Multi-Stream Bi-Directional Recurrent Neural Network for Fine-Grained Action Detection
Bharat Singh, Tim K. Marks, Michael Jones, Oncel Tuzel, and Ming Shao · 2016
Earlier work this paper cites.
Anticipating visual representations from unlabeled video
Carl Vondrick, Hamed Pirsiavash, and Antonio Torralba · 2016
Earlier work this paper cites.
Temporal Segment Networks: Towards Good Practices for Deep Action Recognition
Limin Wang, Yuanjun Xiong, Zhe Wang, Yu Qiao, Dahua Lin, Xiaoou Tang, and Luc Van Gool · 2016
Earlier work this paper cites.
Harnessing Object and Scene Semantics for Large-Scale Video Understanding
Zuxuan Wu, Yanwei Fu, Yu-Gang Jiang, and Leonid Sigal · 2016
Earlier work this paper cites.
Multi-Stream Multi-Class Fusion of Deep Networks for Video Classification
Zuxuan Wu, Yu-Gang Jiang, Xi Wang, Hao Ye, and Xiangyang Xue · 2016
Earlier work this paper cites.
Deep Learning for Video Classification and Captioning
Zuxuan Wu, Ting Yao, Yanwei Fu, and Yu-Gang Jiang · 2016
Earlier work this paper cites.
Dual Many-to-One-Encoder-based Transfer Learning for Cross-Dataset Human Action Recognition
Tiantian Xu, Fan Zhu, Edward K. Wong, and Yi Fang · 2016
Earlier work this paper cites.
Multi-Task Zero-Shot Action Recognition with Prioritised Data Augmentation
Xun Xu, Timothy Hospedales, and Shaogang Gong · 2016
Earlier work this paper cites.
End-to-end learning of action detection from frame glimpses in videos
Serena Yeung, Olga Russakovsky, Greg Mori, and Li Fei-Fei · 2016
Earlier work this paper cites.
Two-Stream SR-CNNs for Action Recognition in Videos
Wang Yifan, Jie Song, Limin Wang, Luc Van Gool, and Otmar Hilliges · 2016
Earlier work this paper cites.
Real-time Action Recognition with Enhanced Motion Vector CNNs
Bowen Zhang, Limin Wang, Zhe Wang, Yu Qiao, and Hanli Wang · 2016
Earlier work this paper cites.
Colorful image colorization
Richard Zhang, Phillip Isola, and Alexei A Efros · 2016
Earlier work this paper cites.
A Key Volume Mining Deep Framework for Action Recognition
Wangjiang Zhu, Jie Hu, Gang Sun, Xudong Cao, and Yu Qiao · 2016
Earlier work this paper cites.
Depth2Action: Exploring Embedded Depth for Large-Scale Action Recognition
Yi Zhu and Shawn Newsam · 2016
Earlier work this paper cites.
Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset
Joao Carreira and Andrew Zisserman · 2017
Earlier work this paper cites.
Generalized Rank Pooling for Activity Recognition
Anoop Cherian, Basura Fernando, Mehrtash Harandi, and Stephen Gould · 2017
Earlier work this paper cites.
Improved Regularization of Convolutional Neural Networks with Cutout
Terrance DeVries and Graham W Taylor · 2017
Earlier work this paper cites.
Temporal 3D ConvNets: New Architecture and Transfer Learning for Video Classification
Ali Diba, Mohsen Fayyaz, Vivek Sharma, Amir Hossein Karami, Mohammad Mahdi Arzani, Rahman Yousefzadeh, and Luc Van Gool · 2017
Earlier work this paper cites.
Deep Temporal Linear Encoding Networks
Ali Diba, Vivek Sharma, and Luc Van Gool · 2017
Earlier work this paper cites.
Spatiotemporal Multiplier Networks for Video Action Recognition
Christoph Feichtenhofer, Axel Pinz, and Richard P Wildes · 2017
Earlier work this paper cites.
Self-supervised video representation learning with odd-one-out networks
Basura Fernando, Hakan Bilen, Efstratios Gavves, and Stephen Gould · 2017
Earlier work this paper cites.
Two Stream LSTM: A Deep Fusion Framework for Human Action Recognition
Harshala Gammulle, Simon Denman, Sridha Sridharan, and Clinton Fookes · 2017
Earlier work this paper cites.
ActionVLAD: Learning Spatio-Temporal Aggregation for Action Classification
Rohit Girdhar, Deva Ramanan, Abhinav Gupta, Josef Sivic, and Bryan Russell · 2017
Earlier work this paper cites.
Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Earlier work this paper cites.
The” Something Something” Video Database for Learning and Evaluating Visual Common Sense
Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michalski, Joanna Materzynska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fruend, Peter Yianilos, Moritz Mueller-Freitag, et al · 2017
Earlier work this paper cites.
Densely Connected Convolutional Networks
Gao Huang, Zhuang Liu, Laurens van der Maaten, and Kilian Q. Weinberger · 2017
Earlier work this paper cites.
FlowNet 2.0: Evolution of Optical Flow Estimation with Deep Networks
E. Ilg, N. Mayer, T. Saikia, M. Keuper, A. Dosovitskiy, and T. Brox · 2017
Earlier work this paper cites.
AdaScan: Adaptive Scan Pooling in Deep Convolutional Neural Networks for Human Action Recognition in Videos
Amlan Kar, Nishant Rai, Karan Sikka, and Gaurav Sharma · 2017
Earlier work this paper cites.
The Kinetics Human Action Video Dataset
Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, Mustafa Suleyman, and Andrew Zisserman · 2017
Earlier work this paper cites.
Deep Local Video Feature for Action Recognition
Zhenzhong Lan, Yi Zhu, Alexander G. Hauptmann, and Shawn Newsam · 2017
Earlier work this paper cites.
Unsupervised Representation Learning by Sorting Sequence, 2017
Hsin-Ying Lee, Jia-Bin Huang, Maneesh Kumar Singh, and Ming-Hsuan Yang · 2017
Earlier work this paper cites.
Unsupervised learning of long-term motion dynamics for videos
Zelun Luo, Boya Peng, De-An Huang, Alexandre Alahi, and Li Fei-Fei · 2017
Earlier work this paper cites.
Spatial-Aware Object Embeddings for Zero-Shot Localization and Classification of Actions
Pascal Mettes and Cees G. M. Snoek · 2017
Cited alongside, same era.
Zero-Shot Action Recognition with Error-Correcting Output Codes
Jie Qin, Li Liu, Ling Shao, Fumin Shen, Bingbing Ni, Jiaxin Chen, and Yunhong Wang · 2017
Cited alongside, same era.
Learning Spatio-Temporal Representation with Pseudo-3D Residual Networks
Zhaofan Qiu, Ting Yao, and Tao Mei · 2017
Cited alongside, same era.
Weakly supervised action learning with rnn based fine-to-coarse modeling
Alexander Richard, Hilde Kuehne, and Juergen Gall · 2017
Cited alongside, same era.
Learning Long-Term Dependencies for Action Recognition With a Biologically-Inspired Deep Network
Yemin Shi, Yonghong Tian, Yaowei Wang, Wei Zeng, and Tiejun Huang · 2017
Cited alongside, same era.
Lattice Long Short-Term Memory for Human Action Recognition
SCSampler: Sampling Salient Clips From Video for Efficient Action Recognition
Bruno Korbar, Du Tran, and Lorenzo Torresani · 2019
Later among the works it cites.
Resource Efficient 3D Convolutional Neural Networks
Okan Köpüklü, Neslihan Kose, Ahmet Gunduz, and Gerhard Rigoll · 2019
Later among the works it cites.
Joint-task self-supervised learning for temporal correspondence
Xueting Li, Sifei Liu, Shalini De Mello, Xiaolong Wang, Jan Kautz, and Ming-Hsuan Yang · 2019
Later among the works it cites.
Fast AutoAugment
Sungbin Lim, Ildoo Kim, Taesup Kim, Chiheon Kim, and Sungwoong Kim · 2019
Later among the works it cites.
Training Kinetics in 15 Minutes: Large-scale Distributed Training on Videos
Ji Lin, Chuang Gan, and Song Han · 2019
Later among the works it cites.
TSM: Temporal Shift Module for Efficient Video Understanding
Ji Lin, Chuang Gan, and Song Han · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lin Sun, Kui Jia, Kevin Chen, Dit-Yan Yeung, Bertram E. Shi, and Silvio Savarese · 2017
Cited alongside, same era.
Action Recognition in Video Sequences using Deep Bi-Directional LSTM With CNN Features
A. Ullah, J. Ahmad, K. Muhammad, M. Sajjad, and S. W. Baik · 2017
Cited alongside, same era.
Attention is All You Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Untrimmednets for weakly supervised action recognition and detection
Limin Wang, Yuanjun Xiong, Dahua Lin, and Luc Van Gool · 2017
Cited alongside, same era.
Spatiotemporal Pyramid Network for Video Action Recognition
Yunbo Wang, Mingsheng Long, Jianmin Wang, and Philip S. Yu · 2017
Cited alongside, same era.
Aggregated Residual Transformations for Deep Neural Networks
Saining Xie, Ross Girshick, Piotr Dollar, Zhuowen Tu, and Kaiming He · 2017
Cited alongside, same era.
Transductive Zero-Shot Action Recognition by Word-Vector Embedding
Xun Xu, Timothy Hospedales, and Shaogang Gong · 2017
Cited alongside, same era.
Later among the works it cites.
Use What You Have: Video Retrieval using Representations from Collaborative Experts
Liu et al · 2019
Later among the works it cites.
DARTS: Differentiable Architecture Search
Hanxiao Liu, Karen Simonyan, and Yiming Yang · 2019
Later among the works it cites.
TS-LSTM and Temporal-Inception: Exploiting Spatiotemporal Dynamics for Activity Recognition
Chih-Yao Ma, Min-Hung Chen, Zsolt Kira, and Ghassan AlRegib · 2019
Later among the works it cites.
HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips
Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic · 2019
Later among the works it cites.
Moments in Time Dataset: One Million Videos for Event Understanding
Mathew Monfort, Alex Andonian, Bolei Zhou, Kandan Ramakrishnan, Sarah Adel Bargal, Tom Yan, Lisa Brown, Quanfu Fan, Dan Gutfruend, Carl Vondrick, et al · 2019
Later among the works it cites.
Multi-moments in time: Learning and interpreting models for multi-action video understanding
Mathew Monfort, Kandan Ramakrishnan, Alex Andonian, Barry A McNamara, Alex Lascelles, Bowen Pan, Quanfu Fan, Dan Gutfreund, Rogerio Feris, and Aude Oliva · 2019
Later among the works it cites.
Video Action Recognition Via Neural Architecture Searching
Wei Peng, Xiaopeng Hong, and Guoying Zhao · 2019
Later among the works it cites.
DDLSTM: Dual-Domain LSTM for Cross-Dataset Action Recognition
Toby Perrett and Dima Damen · 2019
Later among the works it cites.
AJ Piergiovanni, Anelia Angelova, and Michael S. Ryoo · 2019
Later among the works it cites.
Evolving Space-Time Neural Architectures for Videos
AJ Piergiovanni, Anelia Angelova, Alexander Toshev, and Michael S. Ryoo · 2019
Later among the works it cites.
Representation Flow for Action Recognition
AJ Piergiovanni and Michael S. Ryoo · 2019
Later among the works it cites.
Video Activity Recognition: State-of-the-Art
Itsaso Rodríguez-Moreno, José María Martínez-Otzeta, Basilio Sierra, Igor Rodriguez, and Ekaitz Jauregi · 2019
Later among the works it cites.
Multi-modal pyramid feature combination for human action recognition
C. Roig, M. Sarmiento, D. Varas, I. Masuda, J. C. Riveiro, and E. Bou-Balust · 2019
Later among the works it cites.
DMC-Net: Generating Discriminative Motion Cues for Fast Compressed Video Action Recognition
Zheng Shou, Xudong Lin, Yannis Kalantidis, Laura Sevilla-Lara, Marcus Rohrbach, Shih-Fu Chang, and Zhicheng Yan · 2019
Later among the works it cites.
Lsta: Long short-term attention for egocentric action recognition
Swathikiran Sudhakaran, Sergio Escalera, and Oswald Lanz · 2019
Later among the works it cites.
Learning video representations using contrastive bidirectional transformer, 2019
Chen Sun, Fabien Baradel, Kevin Murphy, and Cordelia Schmid · 2019
Later among the works it cites.
VideoBERT: A Joint Model for Video and Language Representation Learning
Chen Sun, Austin Myers, Carl Vondrick, Kevin Murphy, and Cordelia Schmid · 2019
Later among the works it cites.
Yonglong Tian, Dilip Krishnan, and Phillip Isola · 2019
Later among the works it cites.
Video Classification With Channel-Separated Convolutional Networks
Du Tran, Heng Wang, Lorenzo Torresani, and Matt Feiszli · 2019
Later among the works it cites.
Self-supervised spatio-temporal representation learning for videos by predicting motion and appearance statistics
Jiangliu Wang, Jianbo Jiao, Linchao Bao, Shengfeng He, Yunhui Liu, and Wei Liu · 2019
Later among the works it cites.
Hallucinating idt descriptors and i3d optical flow features for action recognition with cnns
Lei Wang, Piotr Koniusz, and Du Huynh · 2019
Later among the works it cites.
Learning Correspondence from the Cycle-Consistency of Time
Xiaolong Wang, Allan Jabri, and Alexei A. Efros · 2019
Later among the works it cites.
Heuristic black-box adversarial attacks on video recognition models
Zhipeng Wei, Jingjing Chen, Xingxing Wei, Linxi Jiang, Tat-Seng Chua, Fengfeng Zhou, and Yu-Gang Jiang · 2019
Later among the works it cites.
Mimetics: Towards understanding human actions out of context
Philippe Weinzaepfel and Grégory Rogez · 2019
Later among the works it cites.
Long-Term Feature Banks for Detailed Video Understanding
Chao-Yuan Wu, Christoph Feichtenhofer, Haoqi Fan, Kaiming He, Philipp Krahenbuhl, and Ross Girshick · 2019
Later among the works it cites.
AdaFrame: Adaptive Frame Selection for Fast Video Recognition
Zuxuan Wu, Caiming Xiong, Chih-Yao Ma, Richard Socher, and Larry S. Davis · 2019
Later among the works it cites.
Self-supervised spatiotemporal learning via video clip order prediction
Dejing Xu, Jun Xiao, Zhou Zhao, Jian Shao, Di Xie, and Yueting Zhuang · 2019
Later among the works it cites.
CutMix: Regularization Strategy to Train Strong Classifiers With Localizable Features
Sangdoo Yun, Dongyoon Han, Seong Joon Oh, Sanghyuk Chun, Junsuk Choe, and Youngjoon Yoo · 2019
Later among the works it cites.
HACS: Human Action Clips and Segments Dataset for Recognition and Temporal Localization
Hang Zhao, Zhicheng Yan, Lorenzo Torresani, and Antonio Torralba · 2019
Later among the works it cites.
Self-Supervised MultiModal Versatile Networks, 2020
Jean-Baptiste Alayrac, Adrià Recasens, Rosalia Schneider, Relja Arandjelović, Jason Ramapuram, Jeffrey De Fauw, Lucas Smaira, Sander Dieleman, and Andrew Zisserman · 2020
Closest in time.
Self-Supervised Learning by Cross-Modal Audio-Video Clustering
Humam Alwassel, Dhruv Mahajan, Bruno Korbar, Lorenzo Torresani, Bernard Ghanem, and Du Tran · 2020
Closest in time.
SpeedNet: Learning the Speediness in Videos
Sagie Benaim, Ariel Ephrat, Oran Lang, Inbar Mosseri, William T. Freeman, Michael Rubinstein, Michal Irani, and Tali Dekel · 2020
Closest in time.
Rethinking Zero-Shot Video Classification: End-to-End Training for Realistic Applications
Biagio Brattoli, Joseph Tighe, Fedor Zhdanov, Pietro Perona, and Krzysztof Chalupka · 2020
Closest in time.
RSPNet: Relative Speed Perception for Unsupervised Video Representation Learning, 2020
Peihao Chen, Deng Huang, Dongliang He, Xiang Long, Runhao Zeng, Shilei Wen, Mingkui Tan, and Chuang Gan · 2020
Closest in time.
A Simple Framework for Contrastive Learning of Visual Representations
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton · 2020
Closest in time.
Improved Baselines with Momentum Contrastive Learning
Xinlei Chen, Haoqi Fan, Ross Girshick, and Kaiming He · 2020
Closest in time.
The EPIC-KITCHENS Dataset: Collection, Challenges and Baselines
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler, Antonino Furnari, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, and Michael Wray · 2020
Closest in time.
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Antonino Furnari, Evangelos Kazakos, Jian Ma, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, et al · 2020
Closest in time.
Large Scale Holistic Video Understanding
Ali Diba, Mohsen Fayyaz, Vivek Sharma, Manohar Paluri, Jürgen Gall, Rainer Stiefelhagen, and Luc Van Gool · 2020
Closest in time.
Omni-sourced Webly-supervised Learning for Video Recognition
Hao Dong Duan, Yue Zhao, Yuanjun Xiong, Wentao Liu, and Dahu Lin · 2020
Closest in time.
X3D: Expanding Architectures for Efficient Video Recognition
Christoph Feichtenhofer · 2020
Closest in time.
Multi-modal Transformer for Video Retrieval
Gabeur et al · 2020
Closest in time.
Coot: Cooperative hierarchical transformer for video-text representation learning
Simon Ging, Mohammadreza Zolfaghari, Hamed Pirsiavash, and Thomas Brox · 2020
Closest in time.
Watching the World Go By: Representation Learning from Unlabeled Videos
Daniel Gordon, Kiana Ehsani, Dieter Fox, and Ali Farhadi · 2020
Closest in time.
Memory-augmented dense predictive coding for video representation learning
Tengda Han, Weidi Xie, and Andrew Zisserman · 2020
Closest in time.
Self-supervised co-training for video representation learning
Tengda Han, Weidi Xie, and Andrew Zisserman · 2020
Closest in time.
Space-time correspondence as a contrastive random walk
Allan Jabri, Andrew Owens, and Alexei A Efros · 2020
Closest in time.
Action genome: Actions as compositions of spatio-temporal scene graphs
Jingwei Ji, Ranjay Krishna, Li Fei-Fei, and Juan Carlos Niebles · 2020
Closest in time.
Shuffle and Attend: Video Domain Adaptation
Samuel Schulter Jinwoo Choi, Gaurav Sharma and Jia-Bin Huang · 2020
Closest in time.
Motionsqueeze: Neural motion feature learning for video understanding
Heeseung Kwon, Manjin Kim, Suha Kwak, and Minsu Cho · 2020
Closest in time.
The ava-kinetics localized human actions video dataset
Ang Li, Meghana Thotakuri, David A Ross, João Carreira, Alexander Vostrikov, and Andrew Zisserman · 2020
Closest in time.
Directional temporal modeling for action recognition
Xinyu Li, Bing Shuai, and Joseph Tighe · 2020
Closest in time.
TEA: Temporal Excitation and Aggregation for Action Recognition
Yan Li, Bin Ji, Xintian Shi, Jianguo Zhang, Bin Kang, and Limin Wang · 2020
Closest in time.
Forecasting human object interaction: Joint prediction of motor attention and egocentric activity
Miao Liu, Siyu Tang, Yin Li, and James Rehg · 2020
Closest in time.
TEINet: Towards an Efficient Architecture for Video Recognition
Zhaoyang Liu, Donghao Luo, Yabiao Wang, Limin Wang, Ying Tai, Chengjie Wang, Jilin Li, Feiyue Huang, and Tong Lu · 2020
Closest in time.
TAM: Temporal Adaptive Module for Video Recognition
Zhaoyang Liu, Limin Wang, Wayne Wu, Chen Qian, and Tong Lu · 2020
Closest in time.
End-to-End Learning of Visual Representations from Uncurated Instructional Videos
Antoine Miech, Jean-Baptiste Alayrac, Lucas Smaira, Ivan Laptev, Josef Sivic, and Andrew Zisserman · 2020
Closest in time.
Audio-visual instance discrimination with cross-modal agreement
Pedro Morgado, Nuno Vasconcelos, and Ishan Misra · 2020
Closest in time.
Multi-Modal Domain Adaptation for Fine-Grained Action Recognition
Jonathan Munro and Dima Damen · 2020
Closest in time.
Adversarial Cross-Domain Action Recognition with Co-Attention
Boxiao Pan, Zhangjie Cao, Ehsan Adeli, and Juan Carlos Niebles · 2020
Closest in time.
Multi-modal self-supervision from generalized data transformations
Mandela Patrick, Yuki M. Asano, Ruth Fong, João F. Henriques, Geoffrey Zweig, and Andrea Vedaldi · 2020
Closest in time.
Evolving Losses for Unsupervised Video Representation Learning
AJ Piergiovanni, Anelia Angelova, and Michael S Ryoo · 2020
Closest in time.
Avid dataset: Anonymized videos from diverse countries, 2020
AJ Piergiovanni and Michael S. Ryoo · 2020
Closest in time.
Learning multimodal representations for unseen activities, 2020
AJ Piergiovanni and Michael S. Ryoo · 2020
Closest in time.
Spatiotemporal contrastive video representation learning
Rui Qian, Tianjian Meng, Boqing Gong, Ming-Hsuan Yang, Huisheng Wang, Serge Belongie, and Yin Cui · 2020
Closest in time.
Designing Network Design Spaces
Ilija Radosavovic, Raj Prateek Kosaraju, Ross Girshick, Kaiming He, and Piotr Dollár · 2020
Closest in time.
Avlnet: Learning audio-visual language representations from instructional videos, 2020
Andrew Rouditchenko, Angie Boggust, David Harwath, Dhiraj Joshi, Samuel Thomas, Kartik Audhkhasi, Rogerio Feris, Brian Kingsbury, Michael Picheny, Antonio Torralba, and James Glass · 2020
Closest in time.
Assemblenet++: Assembling modality representations via attention connections
Michael S Ryoo, AJ Piergiovanni, Juhana Kangaspunta, and Anelia Angelova · 2020
Closest in time.
AssembleNet: Searching for Multi-Stream Neural Connectivity in Video Architectures
Michael S. Ryoo, AJ Piergiovanni, Mingxing Tan, and Anelia Angelova · 2020
Closest in time.
Temporal aggregate representations for long-range video understanding
Fadime Sener, Dipika Singhania, and Angela Yao · 2020
Closest in time.
FineGym: A Hierarchical Video Dataset for Fine-grained Action Understanding
Dian Shao, Yue Zhao, Bo Dai, and Dahua Lin · 2020
Closest in time.
Temporal interlacing network, 2020
Hao Shao, Shengju Qian, and Yu Liu · 2020
Closest in time.
D3D: Distilled 3D Networks for Video Action Recognition
Jonathan C. Stroud, David A. Ross, Chen Sun, Jia Deng, and Rahul Sukthankar · 2020
Closest in time.
Attentionnas: Spatiotemporal attention cell search for video classification, 2020
Xiaofang Wang, Xuehan Xiong, Maxim Neumann, AJ Piergiovanni, Michael S. Ryoo, Anelia Angelova, Kris M. Kitani, and Wei Hua · 2020
Closest in time.
Symbiotic attention for egocentric action recognition with object-centric alignment
Xiaohan Wang, Linchao Zhu, Yu Wu, and Yi Yang · 2020
Closest in time.
A Multigrid Method for Efficiently Training Video Models
Chao-Yuan Wu, Ross Girshick, Kaiming He, Christoph Feichtenhofer, and Philipp Krähenbühl · 2020
Closest in time.
Audiovisual SlowFast Networks for Video Recognition
Fanyi Xiao, Yong Jae Lee, Kristen Grauman, Jitendra Malik, and Christoph Feichtenhofer · 2020
Closest in time.
Sparse black-box video attack with reinforcement learning
Huanqian Yan, Xingxing Wei, and Bo Li · 2020
Closest in time.
Video Representation Learning with Visual Tempo Consistency
Ceyuan Yang, Yinghao Xu, Bo Dai, and Bolei Zhou · 2020
Closest in time.
Temporal Pyramid Network for Action Recognition
Ceyuan Yang, Yinghao Xu, Jianping Shi, Bo Dai, and Bolei Zhou · 2020
Closest in time.
Hierarchical contrastive motion learning for video action recognition, 2020
Xitong Yang, Xiaodong Yang, Sifei Liu, Deqing Sun, Larry Davis, and Jan Kautz · 2020
Closest in time.
Pan: Towards fast action recognition via learning persistence of appearance, 2020
Can Zhang, Yuexian Zou, Guang Chen, and Lei Gan · 2020
Closest in time.
ResNeSt: Split-Attention Networks
Hang Zhang, Chongruo Wu, Zhongyue Zhang, Yi Zhu, Zhi Zhang, Haibin Lin, Yue Sun, Tong He, Jonas Muller, R. Manmatha, Mu Li, and Alexander Smola · 2020
Closest in time.
Motion-Excited Sampler: Video Adversarial Attack with Sparked Prior
Hu Zhang, Linchao Zhu, Yi Zhu, and Yi Yang · 2020
Closest in time.
V4D:4D Convolutional Neural Networks for Video-level Representation Learning
Shiwen Zhang, Sheng Guo, Weilin Huang, Matthew R. Scott, and Limin Wang · 2020
Closest in time.
Hierarchically decoupled spatial-temporal contrast for self-supervised video representation learning, 2020
Zehua Zhang and David Crandall · 2020
Closest in time.
ActBERT: Learning Global-Local Video-Text Representations
Linchao Zhu and Yi Yang · 2020
Closest in time.
A3d: Adaptive 3d networks for video action recognition, 2020
Sijie Zhu, Taojiannan Yang, Matias Mendieta, and Chen Chen · 2020
Closest in time.