Fetching the paper…
Reading the bibliography…
Video understanding requires reasoning at multiple spatiotemporal resolutions -- from short fine-grained motions to events taking place over longer durations.
An iterative image registration technique with an application to stereo vision
Bruce D Lucas, Takeo Kanade, et al · 1981
Earlier work this paper cites.
Pyramid methods in image processing
Edward H Adelson, Charles H Anderson, James R Bergen, Peter J Burt, and Joan M Ogden · 1984
Earlier work this paper cites.
The laplacian pyramid as a compact image code
Peter J Burt and Edward H Adelson · 1987
Earlier work this paper cites.
Backpropagation applied to handwritten zip code recognition
Yann LeCun, Bernhard Boser, John S Denker, Donnie Henderson, Richard E Howard, Wayne Hubbard, and Lawrence D Jackel · 1989
Earlier work this paper cites.
Statistics of natural images and models
Jinggang Huang and David Mumford · 1999
Earlier work this paper cites.
Pyramidal implementation of the affine lucas kanade feature tracker description of the algorithm
Jean-Yves Bouguet et al · 2001
Earlier work this paper cites.
Statistics of natural image categories
Antonio Torralba and Aude Oliva · 2003
Earlier work this paper cites.
Distinctive image features from scale-invariant keypoints
David G Lowe · 2004
Earlier work this paper cites.
Histograms of oriented gradients for human detection
Navneet Dalal and Bill Triggs · 2005
Earlier work this paper cites.
On space-time interest points
Ivan Laptev · 2005
Earlier work this paper cites.
Beyond bags of features: Spatial pyramid matching for recognizing natural scene categories
Svetlana Lazebnik, Cordelia Schmid, and Jean Ponce · 2006
Earlier work this paper cites.
A spatio-temporal descriptor based on 3d-gradients
Alexander Klaser, Marcin Marszałek, and Cordelia Schmid · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Dense trajectories and motion boundary descriptors for action recognition
Heng Wang, Alexander Kläser, Cordelia Schmid, and Cheng-Lin Liu · 2013
Earlier work this paper cites.
Large-scale video classification with convolutional neural networks
Andrej Karpathy, George Toderici, Sanketh Shetty, Thomas Leung, Rahul Sukthankar, and Li Fei-Fei · 2014
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Beyond short snippets: Deep networks for video classification
Joe Yue-Hei Ng, Matthew Hausknecht, Sudheendra Vijayanarasimhan, Oriol Vinyals, Rajat Monga, and George Toderici · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2015
Earlier work this paper cites.
Human action recognition using factorized spatio-temporal convolutional networks
Lin Sun, Kui Jia, Dit-Yan Yeung, and Bertram E Shi · 2015
Earlier work this paper cites.
Going deeper with convolutions
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich · 2015
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks
Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
Spatiotemporal residual networks for video action recognition
Christoph Feichtenhofer, Axel Pinz, and Richard Wildes · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Gaussian error linear units (gelus)
Dan Hendrycks and Kevin Gimpel · 2016
Earlier work this paper cites.
Deep networks with stochastic depth
Gao Huang, Yu Sun, Zhuang Liu, Daniel Sedra, and Kilian Weinberger · 2016
Earlier work this paper cites.
Ssd: Single shot multibox detector
Wei Liu, Dragomir Anguelov, Dumitru Erhan, Christian Szegedy, Scott Reed, Cheng-Yang Fu, and Alexander C Berg · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna · 2016
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Joao Carreira and Andrew Zisserman · 2017
Cited alongside, same era.
Xception: Deep learning with depthwise separable convolutions
François Chollet · 2017
Cited alongside, same era.
The” something something” video database for learning and evaluating visual common sense
Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michalski, Joanna Materzynska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fruend, Peter Yianilos, Moritz Mueller-Freitag, et al · 2017
Cited alongside, same era.
The kinetics human action video dataset
Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al · 2017
Cited alongside, same era.
Feature pyramid networks for object detection
Tsung-Yi Lin, Piotr Dollár, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie · 2017
Cited alongside, same era.
X3d: Expanding architectures for efficient video recognition
Christoph Feichtenhofer · 2020
Later among the works it cites.
Tea: Temporal excitation and aggregation for action recognition
Yan Li, Bin Ji, Xintian Shi, Jianguo Zhang, Bin Kang, and Limin Wang · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 2020
Later among the works it cites.
Learning video representations from textual web supervision
Jonathan C. Stroud, Zhichao Lu, Chen Sun, Jia Deng, Rahul Sukthankar, Cordelia Schmid, and David A. Ross · 2020
Later among the works it cites.
Attentionnas: Spatiotemporal attention cell search for video classification
Xiaofang Wang, Xuehan Xiong, Maxim Neumann, AJ Piergiovanni, Michael S Ryoo, Anelia Angelova, Kris M Kitani, and Wei Hua · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Large-scale evolution of image classifiers
Esteban Real, Sherry Moore, Andrew Selle, Saurabh Saxena, Yutaka Leon Suematsu, Jie Tan, Quoc V Le, and Alexey Kurakin · 2017
Cited alongside, same era.
Revisiting Unreasonable Effectiveness of Data in Deep Learning Era
Chen Sun, Abhinav Shrivastava, Saurabh Singh, and Abhinav Gupta · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Pyramid scene parsing network
Hengshuang Zhao, Jianping Shi, Xiaojuan Qi, Xiaogang Wang, and Jiaya Jia · 2017
Cited alongside, same era.
Neural architecture search with reinforcement learning
Barret Zoph and Quoc V Le · 2017
Cited alongside, same era.
A closer look at spatiotemporal convolutions for action recognition
Du Tran, Heng Wang, Lorenzo Torresani, Jamie Ray, Yann LeCun, and Manohar Paluri · 2018
Cited alongside, same era.
Non-local neural networks
Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaiming He · 2018
Cited alongside, same era.
Chao-Yuan Wu, Ross Girshick, Kaiming He, Christoph Feichtenhofer, and Philipp Krahenbuhl · 2020
Later among the works it cites.
Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text
Hassan Akbari, Linagzhe Yuan, Rui Qian, Wei-Hong Chuang, Shih-Fu Chang, Yin Cui, and Boqing Gong · 2021
Later among the works it cites.
Vivit: A video vision transformer
Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lučić, and Cordelia Schmid · 2021
Later among the works it cites.
Unified graph structured models for video understanding
Anurag Arnab, Chen Sun, and Cordelia Schmid · 2021
Later among the works it cites.
Is space-time attention all you need for video understanding?
Gedas Bertasius, Heng Wang, and Lorenzo Torresani · 2021
Later among the works it cites.
Crossvit: Cross-attention multi-scale vision transformer for image classification
Chun-Fu Chen, Quanfu Fan, and Rameswar Panda · 2021
Later among the works it cites.
Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, , Antonino Furnari, Jian Ma, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, and Michael Wray · 2021
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Later among the works it cites.
Revisiting 3d resnets for video recognition
Xianzhi Du, Yeqing Li, Yin Cui, Rui Qian, Jing Li, and Irwan Bello · 2021
Later among the works it cites.
Multiscale vision transformers
Haoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li, Zhicheng Yan, Jitendra Malik, and Christoph Feichtenhofer · 2021
Later among the works it cites.
Perceiver io: A general architecture for structured inputs & outputs
Andrew Jaegle, Sebastian Borgeaud, Jean-Baptiste Alayrac, Carl Doersch, Catalin Ionescu, David Ding, Skanda Koppula, Daniel Zoran, Andrew Brock, Evan Shelhamer, Olivier Hénaff, Matthew M. Botvinick, Andrew Zisserman, Oriol Vinyals, and João Carreira · 2021
Later among the works it cites.
Movinets: Mobile video networks for efficient video recognition
Dan Kondratyuk, Liangzhe Yuan, Yandong Li, Li Zhang, Mingxing Tan, Matthew Brown, and Boqing Gong · 2021
Later among the works it cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo · 2021
Later among the works it cites.
Attention bottlenecks for multimodal fusion
Arsha Nagrani, Shan Yang, Anurag Arnab, Aren Jansen, Cordelia Schmid, and Chen Sun · 2021
Later among the works it cites.
Keeping your eye on the ball: Trajectory attention in video transformers
Mandela Patrick, Dylan Campbell, Yuki M Asano, Ishan Misra Florian Metze, Christoph Feichtenhofer, Andrea Vedaldi, Jo Henriques, et al · 2021
Later among the works it cites.
Tokenlearner: What can 8 learned tokens do for images and videos?
Michael S Ryoo, AJ Piergiovanni, Anurag Arnab, Mostafa Dehghani, and Anelia Angelova · 2021
Later among the works it cites.
How to train your vit? data, augmentation, and regularization in vision transformers
Andreas Steiner, Alexander Kolesnikov, Xiaohua Zhai, Ross Wightman, Jakob Uszkoreit, and Lucas Beyer · 2021
Later among the works it cites.
Training data-efficient image transformers & distillation through attention
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou · 2021
Later among the works it cites.
Pyramid vision transformer: A versatile backbone for dense prediction without convolutions
Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao · 2021
Later among the works it cites.
Florence: A new foundation model for computer vision
Lu Yuan, Dongdong Chen, Yi-Ling Chen, Noel Codella, Xiyang Dai, Jianfeng Gao, Houdong Hu, Xuedong Huang, Boxin Li, Chunyuan Li, et al · 2021
Later among the works it cites.
Xiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, and Lucas Beyer · 2021
Later among the works it cites.
Co-training transformer with videos and images improves action recognition
Bowen Zhang, Jiahui Yu, Christopher Fifty, Wei Han, Andrew M Dai, Ruoming Pang, and Fei Sha · 2021
Later among the works it cites.
Vidtr: Video transformer without convolutions
Yanyi Zhang, Xinyu Li, Chunhui Liu, Bing Shuai, Yi Zhu, Biagio Brattoli, Hao Chen, Ivan Marsic, and Joseph Tighe · 2021
Later among the works it cites.