Fetching the paper…
Reading the bibliography…
While today's video recognition systems parse snapshots or short clips accurately, they cannot connect the dots and reason across a longer range of time yet.
Recognizing action at a distance
Alexei A Efros, Alexander C Berg, Greg Mori, and Jitendra Malik · 2003
Earlier work this paper cites.
Histograms of oriented gradients for human detection
Navneet Dalal and Bill Triggs · 2005
Earlier work this paper cites.
Behavior recognition via sparse spatio-temporal features
Piotr Dollár, Vincent Rabaud, Garrison Cottrell, and Serge Belongie · 2005
Earlier work this paper cites.
A spatio-temporal descriptor based on 3d-gradients
Alexander Klaser, Marcin Marszałek, and Cordelia Schmid · 2008
Earlier work this paper cites.
Learning realistic human actions from movies
Ivan Laptev, Marcin Marszalek, Cordelia Schmid, and Benjamin Rozenfeld · 2008
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Evaluation of local spatio-temporal features for action recognition
Heng Wang, Muhammad Muneeb Ullah, Alexander Klaser, Ivan Laptev, and Cordelia Schmid · 2009
Earlier work this paper cites.
Convolutional learning of spatio-temporal features
Graham W Taylor, Rob Fergus, Yann LeCun, and Christoph Bregler · 2010
Earlier work this paper cites.
Dense trajectories and motion boundary descriptors for action recognition
Heng Wang, Alexander Kläser, Cordelia Schmid, and Cheng-Lin Liu · 2013
Earlier work this paper cites.
Action recognition with improved trajectories
Heng Wang and Cordelia Schmid · 2013
Earlier work this paper cites.
Action recognition with stacked fisher vectors
Xiaojiang Peng, Changqing Zou, Yu Qiao, and Qiang Peng · 2014
Earlier work this paper cites.
Long-term recurrent convolutional networks for visual recognition and description
Jeff Donahue, Lisa Anne Hendricks, Sergio Guadarrama, Marcus Rohrbach, Subhashini Venugopalan, Kate Saenko, and Trevor Darrell · 2015
Earlier work this paper cites.
Beyond short snippets: Deep networks for video classification
Joe Yue-Hei Ng, Matthew Hausknecht, Sudheendra Vijayanarasimhan, Oriol Vinyals, Rajat Monga, and George Toderici · 2015
Earlier work this paper cites.
Faster R-CNN: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Earlier work this paper cites.
Learning spatiotemporal features with 3D convolutional networks
Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri · 2015
Earlier work this paper cites.
Action recognition with trajectory-pooled deep-convolutional descriptors
Limin Wang, Yu Qiao, and Xiaoou Tang · 2015
Earlier work this paper cites.
Beyond short snippets: Deep networks for video classification
Joe Yue-Hei Ng, Matthew Hausknecht, Sudheendra Vijayanarasimhan, Oriol Vinyals, Rajat Monga, and George Toderici · 2015
Earlier work this paper cites.
Youtube-8m: A large-scale video classification benchmark
Sami Abu-El-Haija, Nisarg Kothari, Joonseok Lee, Paul Natsev, George Toderici, Balakrishnan Varadarajan, and Sudheendra Vijayanarasimhan · 2016
Earlier work this paper cites.
Temporal segment networks: Towards good practices for deep action recognition
Limin Wang, Yuanjun Xiong, Zhe Wang, Yu Qiao, Dahua Lin, Xiaoou Tang, and Luc Val Gool · 2016
Earlier work this paper cites.
ActionVLAD: Learning spatio-temporal aggregation for action classification
Rohit Girdhar, Deva Ramanan, Abhinav Gupta, Josef Sivic, and Bryan Russell · 2017
Earlier work this paper cites.
Accurate, large minibatch SGD: training ImageNet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Earlier work this paper cites.
The kinetics human action video dataset
Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al · 2017
Earlier work this paper cites.
Feature pyramid networks for object detection
Tsung-Yi Lin, Piotr Dollár, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
Learning spatio-temporal representation with pseudo-3d residual networks
Zhaofan Qiu, Ting Yao, and Tao Mei · 2017
Earlier work this paper cites.
Lattice long short-term memory for human action recognition
Lin Sun, Kui Jia, Kevin Chen, Dit-Yan Yeung, Bertram E Shi, and Silvio Savarese · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Aggregated residual transformations for deep neural networks
Saining Xie, Ross Girshick, Piotr Dollár, Zhuowen Tu, and Kaiming He · 2017
Earlier work this paper cites.
Rethinking spatiotemporal feature learning for video understanding
Saining Xie, Chen Sun, Jonathan Huang, Zhuowen Tu, and Kevin Murphy · 2017
Earlier work this paper cites.
Flow-guided feature aggregation for video object detection
Xizhou Zhu, Yujie Wang, Jifeng Dai, Lu Yuan, and Yichen Wei · 2017
Earlier work this paper cites.
A short note about Kinetics-600
Joao Carreira, Eric Noland, Andras Banki-Horvath, Chloe Hillier, and Andrew Zisserman · 2018
Earlier work this paper cites.
Massively parallel video networks
Joao Carreira, Viorica Patraucean, Laurent Mazare, Andrew Zisserman, and Simon Osindero · 2018
Cited alongside, same era.
Scaling egocentric vision: The epic-kitchens dataset
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler, Antonino Furnari, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, and Michael Wray · 2018
Cited alongside, same era.
Leveraging uncertainty to rethink loss functions and evaluation measures for egocentric action anticipation
Antonino Furnari, Sebastiano Battiato, and Giovanni Maria Farinella · 2018
Cited alongside, same era.
AVA: A video dataset of spatio-temporally localized atomic visual actions
Chunhui Gu, Chen Sun, David A. Ross, Carl Vondrick, Caroline Pantofaru, Yeqing Li, Sudheendra Vijayanarasimhan, George Toderici, Susanna Ricco, Rahul Sukthankar, Cordelia Schmid, and Jitendra Malik · 2018
Cited alongside, same era.
A memory network approach for story-based temporal summarization of 360 videos
Sangho Lee, Jinyoung Sung, Youngjae Yu, and Gunhee Kim · 2018
Cited alongside, same era.
Do transformers need deep long-range memory
Jack W Rae and Ali Razavi · 2020
Later among the works it cites.
Equalization loss for long-tailed object recognition
Jingru Tan, Changbao Wang, Buyu Li, Quanquan Li, Wanli Ouyang, Changqing Yin, and Junjie Yan · 2020
Later among the works it cites.
1st place solution of lvis challenge 2020: A good box is not a guarantee of a good mask
Jingru Tan, Gang Zhang, Hanming Deng, Changbao Wang, Lewei Lu, Quanquan Li, and Jifeng Dai · 2020
Later among the works it cites.
Asynchronous interaction aggregation for action detection
Jiajun Tang, Jin Xia, Xinzhi Mu, Bo Pang, and Cewu Lu · 2020
Later among the works it cites.
ViViT: A video vision transformer
Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lučić, and Cordelia Schmid · 2021
Later among the works it cites.
Is space-time attention all you need for video understanding?
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Recurrent tubelet proposal and recognition networks for action detection
Dong Li, Zhaofan Qiu, Qi Dai, Ting Yao, and Tao Mei · 2018
Cited alongside, same era.
VideoLSTM convolves, attends and flows for action recognition
Zhenyang Li, Kirill Gavrilyuk, Efstratios Gavves, Mihir Jain, and Cees GM Snoek · 2018
Cited alongside, same era.
Mobile video object detection with temporally-aware feature maps
Mason Liu and Menglong Zhu · 2018
Cited alongside, same era.
Self-attention with relative position representations
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani · 2018
Cited alongside, same era.
Non-local neural networks
Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaiming He · 2018
Cited alongside, same era.
Compressed video action recognition
Chao-Yuan Wu, Manzil Zaheer, Hexiang Hu, R Manmatha, Alexander J Smola, and Philipp Krähenbühl · 2018
Cited alongside, same era.
Temporal relational reasoning in videos
Bolei Zhou, Alex Andonian, Aude Oliva, and Antonio Torralba · 2018
Cited alongside, same era.
Gedas Bertasius, Heng Wang, and Lorenzo Torresani · 2021
Later among the works it cites.
Watch only once: An end-to-end video action detection framework
Shoufa Chen, Peize Sun, Enze Xie, Chongjian Ge, Jiannan Wu, Lan Ma, Jiajun Shen, and Ping Luo · 2021
Later among the works it cites.
Visformer: The vision-friendly transformer
Zhengsu Chen, Lingxi Xie, Jianwei Niu, Xuefeng Liu, Longhui Wei, and Qi Tian · 2021
Later among the works it cites.
The epic-kitchens dataset: Collection, challenges and baselines
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler, Antonino Furnari, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, and Michael Wray · 2021
Later among the works it cites.
Cswin transformer: A general vision transformer backbone with cross-shaped windows
Xiaoyi Dong, Jianmin Bao, Dongdong Chen, Weiming Zhang, Nenghai Yu, Lu Yuan, Dong Chen, and Baining Guo · 2021
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2021
Later among the works it cites.
Multiscale vision transformers
Haoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li, Zhicheng Yan, Jitendra Malik, and Christoph Feichtenhofer · 2021
Later among the works it cites.
Anticipative Video Transformer
Rohit Girdhar and Kristen Grauman · 2021
Later among the works it cites.
LeViT: A vision transformer in ConvNet’s clothing for faster inference
Benjamin Graham, Alaaeldin El-Nouby, Hugo Touvron, Pierre Stock, Armand Joulin, Herve Jegou, and Matthijs Douze · 2021
Later among the works it cites.
MoViNets: Mobile video networks for efficient video recognition
Dan Kondratyuk, Liangzhe Yuan, Yandong Li, Li Zhang, Mingxing Tan, Matthew Brown, and Boqing Gong · 2021
Later among the works it cites.
Video prediction recalling long-term motion context via memory alignment learning
Sangmin Lee, Hak Gu Kim, Dae Hwi Choi, Hyung-Il Kim, and Yong Man Ro · 2021
Later among the works it cites.
Ego-exo: Transferring visual representations from third-person to first-person videos
Yanghao Li, Tushar Nagarajan, Bo Xiong, and Kristen Grauman · 2021
Later among the works it cites.
Improved multiscale vision transformers for classification and detection
Yanghao Li, Chao-Yuan Wu, Haoqi Fan, Karttikeya Mangalam, Bo Xiong, Jitendra Malik, and Christoph Feichtenhofer · 2021
Later among the works it cites.
Daniel Neimark, Omri Bar, Maya Zohar, and Dotan Asselmann · 2021
Later among the works it cites.
Actor-context-actor relation network for spatio-temporal action localization
Junting Pan, Siyu Chen, Mike Zheng Shou, Yu Liu, Jing Shao, and Hongsheng Li · 2021
Later among the works it cites.
Keeping your eye on the ball: Trajectory attention in video transformers
Mandela Patrick, Dylan Campbell, Yuki M Asano, Ishan Misra Florian Metze, Christoph Feichtenhofer, Andrea Vedaldi, Jo Henriques, et al · 2021
Later among the works it cites.
Technical report: Temporal aggregate representations
Fadime Sener, Dibyadip Chatterjee, and Angela Yao · 2021
Later among the works it cites.
Not all memories are created equal: Learning to forget by expiring
Sainbayar Sukhbaatar, Da Ju, Spencer Poff, Stephen Roller, Arthur Szlam, Jason Weston, and Angela Fan · 2021
Later among the works it cites.
Training data-efficient image transformers and distillation through attention
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Herve Jegou · 2021
Later among the works it cites.
Going deeper with image transformers
Hugo Touvron, Matthieu Cord, Alexandre Sablayrolles, Gabriel Synnaeve, and Hervé Jégou · 2021
Later among the works it cites.
Pyramid vision transformer: A versatile backbone for dense prediction without convolutions
Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao · 2021
Later among the works it cites.
Interactive prototype learning for egocentric action recognition
Xiaohan Wang, Linchao Zhu, Heng Wang, and Yi Yang · 2021
Later among the works it cites.
Towards long-form video understanding
Chao-Yuan Wu and Philipp Krahenbuhl · 2021
Later among the works it cites.
Tokens-to-token ViT: Training vision transformers from scratch on imagenet
Li Yuan, Yunpeng Chen, Tao Wang, Weihao Yu, Yujun Shi, Francis EH Tay, Jiashi Feng, and Shuicheng Yan · 2021
Later among the works it cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo · 2022
Closest in time.