Fetching the paper…
Reading the bibliography…
In video transformers, the time dimension is often treated in the same way as the two spatial dimensions.
Maximum entropy sampling
Michael C Shewry and Henry P Wynn · 1987
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Strike a pose: Tracking people by finding stylized poses
Deva Ramanan, David A. Forsyth, and Andrew Zisserman · 2005
Earlier work this paper cites.
A 3-dimensional sift descriptor and its application to action recognition
P. Scovanner, S. Ali, and M. Shah · 2007
Earlier work this paper cites.
A spatio-temporal descriptor based on 3d-gradients
Alexander Klaser, Marcin Marszałek, and Cordelia Schmid · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton · 2012
Earlier work this paper cites.
Action recognition with improved trajectories
Heng Wang and Cordelia Schmid · 2013
Earlier work this paper cites.
High-speed tracking with kernelized correlation filters
João F Henriques, Rui Caseiro, Pedro Martins, and Jorge Batista · 2014
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei · 2015
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks
Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri · 2015
Earlier work this paper cites.
Learning to track for spatio-temporal action localization
Philippe Weinzaepfel, Zaid Harchaoui, and Cordelia Schmid · 2015
Earlier work this paper cites.
Layer normalization, 2016
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Flownet 2.0: Evolution of optical flow estimation with deep networks
Eddy Ilg, Nikolaus Mayer, Tonmoy Saikia, Margret Keuper, Alexey Dosovitskiy, and Thomas Brox · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jonathon Shlens, and Zbigniew Wojna · 2016
Earlier work this paper cites.
Temporal segment networks: Towards good practices for deep action recognition
Limin Wang, Yuanjun Xiong, Yu Qiao, Dahua Lin, Xiaoou Tang, and Luc Van Gool · 2016
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Joao Carreira and Andrew Zisserman · 2017
Earlier work this paper cites.
The “something something” video database for learning and evaluating visual common sense
Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michalski, Joanna Materzynska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fruend, Peter Yianilos, Moritz Mueller-Freitag, and et al · 2017
Earlier work this paper cites.
The kinetics human action video dataset, 2017
Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al · 2017
Earlier work this paper cites.
Learning spatio-temporal representation with pseudo-3d residual networks
Zhaofan Qiu, Ting Yao, and Tao Mei · 2017
Earlier work this paper cites.
Online real-time multiple spatiotemporal action localisation and prediction
Gurkirt Singh, Suman Saha, Michael Sapienza, Philip HS Torr, and Fabio Cuzzolin · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
a 2 a^{2} -nets: Double attention networks
Yunpeng Chen, Yannis Kalantidis, Jianshu Li, Shuicheng Yan, and Jiashi Feng · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Kenton Lee Jacob Devlin, Ming-Wei Chang and Kristina Toutanova · 2018
Earlier work this paper cites.
Fixing weight decay regularization in adam, 2018
Ilya Loshchilov and Frank Hutter · 2018
Earlier work this paper cites.
Mixed precision training
Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, David Garcia, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, and Hao Wu · 2018
Earlier work this paper cites.
Image transformer
Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Łukasz Kaiser, Noam Shazeer, Alexander Ku, and Dustin Tran · 2018
Earlier work this paper cites.
Pwc-net: Cnns for optical flow using pyramid, warping, and cost volume
Deqing Sun, Xiaodong Yang, Ming-Yu Liu, and Jan Kautz · 2018
Earlier work this paper cites.
A closer look at spatiotemporal convolutions for action recognition
Du Tran, Heng Wang, Lorenzo Torresani, Jamie Ray, Yann LeCun, and Manohar Paluri · 2018
Earlier work this paper cites.
Non-local neural networks
Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaiming He · 2018
Earlier work this paper cites.
Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification
Saining Xie, Chen Sun, Jonathan Huang, Zhuowen Tu, and Kevin Murphy · 2018
Earlier work this paper cites.
Trajectory convolution for action recognition
Yue Zhao, Yuanjun Xiong, and Dahua Lin · 2018
Earlier work this paper cites.
Temporal relational reasoning in videos
Bolei Zhou, Alex Andonian, Aude Oliva, and Antonio Torralba · 2018
Cited alongside, same era.
Generating long sequences with sparse transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever · 2019
Cited alongside, same era.
More Is Less: Learning Efficient Video Representations by Temporal Aggregation Modules
Quanfu Fan, Chun-Fu (Ricarhd) Chen, Hilde Kuehne, Marco Pistoia, and David Cox · 2019
Cited alongside, same era.
Slowfast networks for video recognition
Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He · 2019
Cited alongside, same era.
Stm: Spatiotemporal and motion encoding for action recognition
Boyuan Jiang, Mengmeng Wang, Weihao Gan, Wei Wu, and Junjie Yan · 2019
Cited alongside, same era.
Epic-fusion: Audio-visual temporal binding for egocentric action recognition
Evangelos Kazakos, Arsha Nagrani, Andrew Zisserman, and Dima Damen · 2019
Reformer: The efficient transformer
Nikita Kitaev, Lukasz Kaiser, and Anselm Levskaya · 2020
Later among the works it cites.
Motionsqueeze: Neural motion feature learning for video understanding
Heeseung Kwon, Manjin Kim, Suha Kwak, and Minsu Cho · 2020
Later among the works it cites.
Mast: A memory-augmented self-supervised tracker
Zihang Lai, Erika Lu, and Weidi Xie · 2020
Later among the works it cites.
Tea: Temporal excitation and aggregation for action recognition
Yan Li, Bin Ji, Xintian Shi, Jianguo Zhang, Bin Kang, and Limin Wang · 2020
Later among the works it cites.
Object-centric learning with slot attention
Francesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran, Georg Heigold, Jakob Uszkoreit, Alexey Dosovitskiy, and Thomas Kipf · 2020
Later among the works it cites.
Learning visual representations with caption annotations
Mert Bulent Sariyildiz, Julien Perez, and Diane Larlus · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Visualbert: A simple and performant baseline for vision and language, 2019
Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang · 2019
Cited alongside, same era.
Tsm: Temporal shift module for efficient video understanding
Ji Lin, Chuang Gan, and Song Han · 2019
Cited alongside, same era.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee · 2019
Cited alongside, same era.
Video object segmentation using space-time memory networks
Seoung Wug Oh, Joon-Young Lee, Ning Xu, and Seon Joo Kim · 2019
Cited alongside, same era.
Learning spatio-temporal representation with local and global diffusion
Zhaofan Qiu, Ting Yao, Chong-Wah Ngo, Xinmei Tian, and Tao Mei · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Cited alongside, same era.
Later among the works it cites.
Vl-bert: Pre-training of generic visual-linguistic representations
Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai · 2020
Later among the works it cites.
Vokenization: Improving language understanding with contextualized, visual-grounded supervision
Hao Tan and Mohit Bansal · 2020
Later among the works it cites.
RAFT: recurrent all-pairs field transforms for optical flow
Zachary Teed and Jia Deng · 2020
Later among the works it cites.
Linformer: Self-attention with linear complexity
Sinong Wang, Belinda Li, Madian Khabsa, Han Fang, and Hao Ma · 2020
Later among the works it cites.
Attentionnas: Spatiotemporal attention cell search for video classification
Xiaofang Wang, Xuehan Xiong, Maxim Neumann, AJ Piergiovanni, Michael S Ryoo, Anelia Angelova, Kris M Kitani, and Wei Hua · 2020
Later among the works it cites.
Visual transformers: Token-based image representation and processing for computer vision, 2020
Bichen Wu, Chenfeng Xu, Xiaoliang Dai, Alvin Wan, Peizhao Zhang, Masayoshi Tomizuka, Kurt Keutzer, and Peter Vajda · 2020
Later among the works it cites.
Actbert: Learning global-local video-text representations
Linchao Zhu and Yi Yang · 2020
Later among the works it cites.
Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text, 2021
Hassan Akbari, Linagzhe Yuan, Rui Qian, Wei-Hong Chuang, Shih-Fu Chang, Yin Cui, and Boqing Gong · 2021
Closest in time.
Vivit: A video vision transformer, 2021
Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lučić, and Cordelia Schmid · 2021
Closest in time.
Frozen in time: A joint video and image encoder for end-to-end retrieval, 2021
Max Bain, Arsha Nagrani, Gül Varol, and Andrew Zisserman · 2021
Closest in time.
Is space-time attention all you need for video understanding?
Gedas Bertasius, Heng Wang, and Lorenzo Torresani · 2021
Closest in time.
Emerging properties in self-supervised vision transformers
Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin · 2021
Closest in time.
Rethinking attention with performers
Krzysztof Marcin Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Quincy Davis, Afroz Mohiuddin, Lukasz Kaiser, David Benjamin Belanger, Lucy J Colwell, and Adrian Weller · 2021
Closest in time.
VirTex: Learning Visual Representations from Textual Annotations
Karan Desai and Justin Johnson · 2021
Closest in time.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Closest in time.
Multiscale vision transformers, 2021
Haoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li, Zhicheng Yan, Jitendra Malik, and Christoph Feichtenhofer · 2021
Closest in time.
Multilingual multimodal pre-training for zero-shot cross-lingual transfer of vision-language models
Po-Yao Huang, Mandela Patrick, Junjie Hu, Graham Neubig, Florian Metze, and Alexander Hauptmann · 2021
Closest in time.
Learning to estimate hidden motions with global motion aggregation
Shihao Jiang, Dylan Campbell, Yao Lu, Hongdong Li, and Richard Hartley · 2021
Closest in time.
Parameter efficient multimodal transformers for video representation learning
Sangho Lee, Youngjae Yu, Gunhee Kim, Thomas Breuel, Jan Kautz, and Yale Song · 2021
Closest in time.
Vidtr: Video transformer without convolutions, 2021
Xinyu Li, Yanyi Zhang, Chunhui Liu, Bing Shuai, Yi Zhu, Biagio Brattoli, Hao Chen, Ivan Marsic, and Joseph Tighe · 2021
Closest in time.
Video transformer network, 2021
Daniel Neimark, Omri Bar, Maya Zohar, and Dotan Asselmann · 2021
Closest in time.
Space-time crop & attend: Improving cross-modal video representation learning, 2021
Mandela Patrick, Yuki M. Asano, Bernie Huang, Ishan Misra, Florian Metze, Joao Henriques, and Andrea Vedaldi · 2021
Closest in time.
Support-set bottlenecks for video-text representation learning
Mandela Patrick, Po-Yao Huang, Yuki Asano, Florian Metze, Alexander G Hauptmann, Joao F. Henriques, and Andrea Vedaldi · 2021
Closest in time.
Long range arena : A benchmark for efficient transformers
Yi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen, Dara Bahri, Philip Pham, Jinfeng Rao, Liu Yang, Sebastian Ruder, and Donald Metzler · 2021
Closest in time.
Training data-efficient image transformers & distillation through attention
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou · 2021
Closest in time.
Nyströmformer: A nyström-based algorithm for approximating self-attention
Yunyang Xiong, Zhanpeng Zeng, Rudrasis Chakraborty, Mingxing Tan, Glenn Fung, Yin Li, and Vikas Singh · 2021
Closest in time.
Long short-term transformer for online action detection
Mingze Xu, Yuanjun Xiong, Hao Chen, Xinyu Li, Wei Xia, Zhuowen Tu, and Stefano Soatto · 2021
Closest in time.