Fetching the paper…
Reading the bibliography…
We explore a new perspective on video understanding by casting the video recognition problem as an image recognition task.
Assemblenet: Searching for multi-stream neural connectivity in video architectures
Michael S. Ryoo, A. J. Piergiovanni, Mingxing Tan, and Anelia Angelova · 1905
Earlier work this paper cites.
AVD: adversarial video distillation
Mohammad Tavakolian, Mohammad Sabokrou, and Abdenour Hadid · 1907
Earlier work this paper cites.
The representation and recognition of action using temporal templates
James Davis and Aaron Bobick · 1997
Earlier work this paper cites.
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Antonino Furnari, Evangelos Kazakos, Jian Ma, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, and Michael Wray · 2006
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Learning Spatiotemporal Features With 3D Convolutional Networks
Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri · 2015
Earlier work this paper cites.
Dynamic image networks for action recognition
Hakan Bilen, Basura Fernando, Efstratios Gavves, Andrea Vedaldi, and Stephen Gould · 2016
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna · 2016
Earlier work this paper cites.
Temporal segment networks: Towards good practices for deep action recognition
Limin Wang, Yuanjun Xiong, Zhe Wang, Yu Qiao, Dahua Lin, Xiaoou Tang, and Luc Van Gool · 2016
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Joao Carreira, Andrew Zisserman, and xxx · 2017
Earlier work this paper cites.
The" something something" video database for learning and evaluating visual common sense
Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michalski, Joanna Materzynska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fruend, Peter Yianilos, Moritz Mueller-Freitag, et al · 2017
Earlier work this paper cites.
The kinetics human action video dataset
Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al · 2017
Earlier work this paper cites.
Learning efficient convolutional networks through network slimming
Zhuang Liu, Jianguo Li, Zhiqiang Shen, Gao Huang, Shoumeng Yan, and Changshui Zhang · 2017
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
Single image action recognition using semantic body part actions
Zhichen Zhao, Huimin Ma, and Shaodi You · 2017
Earlier work this paper cites.
Resound: Towards action recognition without representation bias
Yingwei Li, Yi Li, and Nuno Vasconcelos · 2018
Earlier work this paper cites.
Non-local neural networks
Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaiming He · 2018
Earlier work this paper cites.
Rethinking Spatiotemporal Feature Learning: Speed-Accuracy Trade-offs in Video Classification
Saining Xie, Chen Sun, Jonathan Huang, Zhuowen Tu, and Kevin Murphy · 2018
Earlier work this paper cites.
mixup: Beyond empirical risk minimization
Hongyi Zhang, Moustapha Cisse, Yann N. Dauphin, and David Lopez-Paz · 2018
Cited alongside, same era.
Temporal relational reasoning in videos
Bolei Zhou, Alex Andonian, Aude Oliva, and Antonio Torralba · 2018
Cited alongside, same era.
More Is Less: Learning Efficient Video Representations by Temporal Aggregation Modules
Quanfu Fan, Chun-Fu (Ricarhd) Chen, Hilde Kuehne, Marco Pistoia, and David Cox · 2019
Cited alongside, same era.
Stm: Spatiotemporal and motion encoding for action recognition
Boyuan Jiang, Mengmeng Wang, Weihao Gan, Wei Wu, and Junjie Yan · 2019
Cited alongside, same era.
Collaborative Spatiotemporal Feature Learning for Video Action Recognition
Chao Li, Qiaoyong Zhong, Di Xie, and Shiliang Pu · 2019
Cited alongside, same era.
Temporal Shift Module for Efficient Video Understanding
Ji Lin, Chuang Gan, and Song Han · 2019
Ar-net: Adaptive frame resolution for efficient action recognition
Yue Meng, Chung-Ching Lin, Rameswar Panda, Prasanna Sattigeri, Leonid Karlinsky, Aude Oliva, Kate Saenko, and Rogerio Feris · 2020
Later among the works it cites.
Assemblenet: Searching for multi-stream neural connectivity in video architectures
Michael S. Ryoo, AJ Piergiovanni, Mingxing Tan, and Anelia Angelova · 2020
Later among the works it cites.
Learn to cycle: Time-consistent feature discovery for action recognition
Alexandros Stergiou and Ronald Poppe · 2020
Later among the works it cites.
Training data-efficient image transformers & distillation through attention
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou · 2020
Later among the works it cites.
Temporal pyramid network for action recognition
Ceyuan Yang, Yinghao Xu, Jianping Shi, Bo Dai, and Bolei Zhou · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2019
Cited alongside, same era.
The jester dataset: A large-scale video dataset of human gestures
Joanna Materzynska, Guillaume Berger, Ingo Bax, and Roland Memisevic · 2019
Cited alongside, same era.
Moments in time dataset: one million videos for event understanding
Mathew Monfort, Alex Andonian, Bolei Zhou, Kandan Ramakrishnan, Sarah Adel Bargal, Yan Yan, Lisa Brown, Quanfu Fan, Dan Gutfreund, Carl Vondrick, et al · 2019
Cited alongside, same era.
Still image action recognition by predicting spatial-temporal pixel evolution
Marjaneh Safaei and Hassan Foroosh · 2019
Cited alongside, same era.
EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks
Mingxing Tan and Quoc Le · 2019
Cited alongside, same era.
Video classification with channel-separated convolutional networks
Du Tran, Heng Wang, Lorenzo Torresani, and Matt Feiszli · 2019
Cited alongside, same era.
Later among the works it cites.
PAN: Towards Fast Action Recognition via Learning Persistence of Appearance
Can Zhang, Yuexian Zou, Guang Chen, and Lei Gan · 2020
Later among the works it cites.
Vivit: A video vision transformer, 2021
Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lučić, and Cordelia Schmid · 2021
Closest in time.
Deep analysis of cnn-based spatio-temporal representations for action recognition, June 2021
Chun-Fu Chen, Rameswar Panda, Kandan Ramakrishnan, Rogerio Feris, John Cohn, Aude Oliva, and Quanfu Fan · 2021
Closest in time.
Twins: Revisiting Spatial Attention Design in Vision Transformers
Xiangxiang Chu, Zhi Tian, Yuqing Wang, Bo Zhang, Haibing Ren, Xiaolin Wei, Huaxia Xia, and Chunhua Shen · 2021
Closest in time.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Closest in time.
SSTVOS: sparse spatiotemporal transformers for video object segmentation
Brendan Duke, Abdalla Ahmed, Christian Wolf, Parham Aarabi, and Graham W. Taylor · 2021
Closest in time.
Multiscale vision transformers, 2021
Haoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li, Zhicheng Yan, Jitendra Malik, and Christoph Feichtenhofer · 2021
Closest in time.
VidTr: Video Transformer Without Convolutions
Xinyu Li, Yanyi Zhang, Chunhui Liu, Bing Shuai, Yi Zhu, Biagio Brattoli, Hao Chen, Ivan Marsic, and Joseph Tighe · 2021
Closest in time.
Swin Transformer: Hierarchical Vision Transformer using Shifted Windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo · 2021
Closest in time.
Adafuse: Adaptive temporal fusion network for efficient action recognition
Yue Meng, Rameswar Panda, Chung-Ching Lin, Prasanna Sattigeri, Leonid Karlinsky, Kate Saenko, Aude Oliva, and Rogerio Feris · 2021
Closest in time.
Video transformer network, 2021
Daniel Neimark, Omri Bar, Maya Zohar, and Dotan Asselmann · 2021
Closest in time.
Condensing a sequence to one informative frame for video recognition
Zhaofan Qiu, Ting Yao, Yan Shu, Chong-Wah Ngo, and Tao Mei · 2021
Closest in time.
Dynamic network quantization for efficient video inference
Ximeng Sun, Rameswar Panda, Chun-Fu Richard Chen, Aude Oliva, Rogerio Feris, and Kate Saenko · 2021
Closest in time.
Multi-Scale Vision Longformer: A New Vision Transformer for High-Resolution Image Encoding
Pengchuan Zhang, Xiyang Dai, Jianwei Yang, Bin Xiao, Lu Yuan, Lei Zhang, and Jianfeng Gao · 2021
Closest in time.