Fetching the paper…
Reading the bibliography…
Self-attention learns pairwise interactions to model long-range dependencies, yielding great improvements for video action recognition.
Context-based vision: recognizing objects using information from both 2 d and 3 d imagery
Thomas M Strat and Martin A Fischler · 1991
Earlier work this paper cites.
Texture synthesis by non-parametric sampling
Alexei A Efros and Thomas K Leung · 1999
Earlier work this paper cites.
A non-local algorithm for image denoising
Antoni Buades, Bartomeu Coll, and J-M Morel · 2005
Earlier work this paper cites.
Learning spatial context: Using stuff to find things
Geremy Heitz and Daphne Koller · 2008
Earlier work this paper cites.
3d convolutional neural networks for human action recognition
Shuiwang Ji, Wei Xu, Ming Yang, and Kai Yu · 2012
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
Large-scale video classification with convolutional neural networks
Andrej Karpathy, George Toderici, Sanketh Shetty, Thomas Leung, Rahul Sukthankar, and Li Fei-Fei · 2014
Earlier work this paper cites.
The role of context for object detection and semantic segmentation in the wild
Roozbeh Mottaghi, Xianjie Chen, Xiaobai Liu, Nam-Gyu Cho, Seong-Whan Lee, Sanja Fidler, Raquel Urtasun, and Alan Yuille · 2014
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Attention-based models for speech recognition
Jan K Chorowski, Dzmitry Bahdanau, Dmitriy Serdyuk, Kyunghyun Cho, and Yoshua Bengio · 2015
Earlier work this paper cites.
Long-term recurrent convolutional networks for visual recognition and description
Jeff Donahue, Lisa Anne Hendricks, Sergio Guadarrama, Marcus Rohrbach, Subhashini Venugopalan, Trevor Darrell, and Kate Saenko · 2015
Earlier work this paper cites.
Exploring person context and local scene context for object detection
Saurabh Gupta, Bharath Hariharan, and Jitendra Malik · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks
Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri · 2015
Earlier work this paper cites.
Beyond short snippets: Deep networks for video classification
Joe Yue-Hei Ng, Matthew Hausknecht, Sudheendra Vijayanarasimhan, Oriol Vinyals, Rajat Monga, and George Toderici · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Temporal segment networks: Towards good practices for deep action recognition
Limin Wang, Yuanjun Xiong, Zhe Wang, Yu Qiao, Dahua Lin, Xiaoou Tang, and Luc Van Gool · 2016
Cited alongside, same era.
Quo vadis, action recognition? a new model and the kinetics dataset
Joao Carreira and Andrew Zisserman · 2017
Cited alongside, same era.
Actionvlad: Learning spatio-temporal aggregation for action classification
Rohit Girdhar, Deva Ramanan, Abhinav Gupta, Josef Sivic, and Bryan Russell · 2017
Cited alongside, same era.
Eco: Efficient convolutional network for online video understanding
Mohammadreza Zolfaghari, Kamaljeet Singh, and Thomas Brox · 2018
Later among the works it cites.
Gcnet: Non-local networks meet squeeze-excitation networks and beyond
Yue Cao, Jiarui Xu, Stephen Lin, Fangyun Wei, and Han Hu · 2019
Later among the works it cites.
Graph-based global reasoning networks
Yunpeng Chen, Marcus Rohrbach, Zhicheng Yan, Yan Shuicheng, Jiashi Feng, and Yannis Kalantidis · 2019
Later among the works it cites.
More is less: Learning efficient video representations by big-little network and depthwise temporal aggregation
Quanfu Fan, Chun-Fu Richard Chen, Hilde Kuehne, Marco Pistoia, and David Cox · 2019
Later among the works it cites.
Slowfast networks for video recognition
Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He · 2019
Later among the works it cites.
Dual attention network for scene segmentation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al · 2017
Cited alongside, same era.
A simple neural network module for relational reasoning
Adam Santoro, David Raposo, David G. T. Barrett, Mateusz Malinowski, Razvan Pascanu, Peter W. Battaglia, and Timothy P. Lillicrap · 2017
Cited alongside, same era.
An end-to-end spatio-temporal attention model for human action recognition from skeleton data
Sijie Song, Cuiling Lan, Junliang Xing, Wenjun Zeng, and Jiaying Liu · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Non-local neural networks
Xiaolong Wang, Ross B. Girshick, Abhinav Gupta, and Kaiming He · 2017
Cited alongside, same era.
A closer look at spatiotemporal convolutions for action recognition
Du Tran, Heng Wang, Lorenzo Torresani, Jamie Ray, Yann LeCun, and Manohar Paluri · 2018
Cited alongside, same era.
Videos as space-time region graphs
Xiaolong Wang and Abhinav Gupta · 2018
Cited alongside, same era.
Jun Fu, Jing Liu, Haijie Tian, Yong Li, Yongjun Bao, Zhiwei Fang, and Hanqing Lu · 2019
Later among the works it cites.
Ccnet: Criss-cross attention for semantic segmentation
Zilong Huang, Xinggang Wang, Lichao Huang, Chang Huang, Yunchao Wei, and Wenyu Liu · 2019
Later among the works it cites.
Timeception for complex action recognition
Noureldien Hussein, Efstratios Gavves, and Arnold WM Smeulders · 2019
Later among the works it cites.
Stm: Spatiotemporal and motion encoding for action recognition
Boyuan Jiang, MengMeng Wang, Weihao Gan, Wei Wu, and Junjie Yan · 2019
Later among the works it cites.
Tsm: Temporal shift module for efficient video understanding
Ji Lin, Chuang Gan, and Song Han · 2019
Later among the works it cites.
On the relationship between self-attention and convolutional layers
Jean-Baptiste Cordonnier, Andreas Loukas, and Martin Jaggi · 2020
Closest in time.
Motionsqueeze: Neural motion feature learning for video understanding
Heeseung Kwon, Manjin Kim, Suha Kwak, and Minsu Cho · 2020
Closest in time.
Tea: Temporal excitation and aggregation for action recognition
Yan Li, Bin Ji, Xintian Shi, Jianguo Zhang, Bin Kang, and Limin Wang · 2020
Closest in time.
Video modeling with correlation networks
Heng Wang, Du Tran, Lorenzo Torresani, and Matt Feiszli · 2020
Closest in time.