Fetching the paper…
Reading the bibliography…
It is a challenging task to learn discriminative representation from images and videos, due to large local redundancy and complex global dependency in these visual data.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Ucf101: A dataset of 101 human actions classes from videos in the wild
Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah · 2012
Earlier work this paper cites.
Action recognition with improved trajectories
Heng Wang and Cordelia Schmid · 2013
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Going deeper with convolutions
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott E. Reed, Dragomir Anguelov, D. Erhan, Vincent Vanhoucke, and Andrew Rabinovich · 2015
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks
Du Tran, Lubomir D. Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri · 2015
Earlier work this paper cites.
Jimmy Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, X. Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Temporal segment networks: Towards good practices for deep action recognition
L. Wang, Yuanjun Xiong, Zhe Wang, Y. Qiao, D. Lin, X. Tang, and L. Gool · 2016
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
João Carreira and Andrew Zisserman · 2017
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
João Carreira and Andrew Zisserman · 2017
Earlier work this paper cites.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross B. Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Earlier work this paper cites.
The “something something” video database for learning and evaluating visual common sense
Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michalski, Joanna Materzynska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fründ, Peter Yianilos, Moritz Mueller-Freitag, Florian Hoppe, Christian Thurau, Ingo Bax, and Roland Memisevic · 2017
Earlier work this paper cites.
Mask r-cnn
Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick · 2017
Earlier work this paper cites.
Mobilenets: Efficient convolutional neural networks for mobile vision applications
Andrew G. Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, M. Andreetto, and Hartwig Adam · 2017
Earlier work this paper cites.
Densely connected convolutional networks
Gao Huang, Zhuang Liu, and Kilian Q. Weinberger · 2017
Earlier work this paper cites.
Fixing weight decay regularization in adam
I. Loshchilov and F. Hutter · 2017
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
Learning spatio-temporal representation with pseudo-3d residual networks
Zhaofan Qiu, Ting Yao, and Tao Mei · 2017
Earlier work this paper cites.
Revisiting unreasonable effectiveness of data in deep learning era
Chen Sun, Abhinav Shrivastava, Saurabh Singh, and Abhinav Kumar Gupta · 2017
Earlier work this paper cites.
Ashish Vaswani, Noam M. Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Aggregated residual transformations for deep neural networks
Saining Xie, Ross B. Girshick, Piotr Dollár, Zhuowen Tu, and Kaiming He · 2017
Earlier work this paper cites.
A short note about kinetics-600
João Carreira, Eric Noland, Andras Banki-Horvath, Chloe Hillier, and Andrew Zisserman · 2018
Earlier work this paper cites.
Shufflenet v2: Practical guidelines for efficient cnn architecture design
Ningning Ma, Xiangyu Zhang, Haitao Zheng, and Jian Sun · 2018
Earlier work this paper cites.
Mobilenetv2: Inverted residuals and linear bottlenecks
Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen · 2018
Earlier work this paper cites.
A closer look at spatiotemporal convolutions for action recognition
Du Tran, Hong xiu Wang, Lorenzo Torresani, Jamie Ray, Yann LeCun, and Manohar Paluri · 2018
Earlier work this paper cites.
Non-local neural networks
X. Wang, Ross B. Girshick, Abhinav Gupta, and Kaiming He · 2018
Earlier work this paper cites.
Simple baselines for human pose estimation and tracking
Bin Xiao, Haiping Wu, and Yichen Wei · 2018
Earlier work this paper cites.
Unified perceptual parsing for scene understanding
Tete Xiao, Yingcheng Liu, Bolei Zhou, Yuning Jiang, and Jian Sun · 2018
Earlier work this paper cites.
mixup: Beyond empirical risk minimization
Hongyi Zhang, Moustapha Cissé, Yann Dauphin, and David Lopez-Paz · 2018
Earlier work this paper cites.
Shufflenet: An extremely efficient convolutional neural network for mobile devices
Xiangyu Zhang, Xinyu Zhou, Mengxiao Lin, and Jian Sun · 2018
Earlier work this paper cites.
Cascade r-cnn: High quality object detection and instance segmentation
Zhaowei Cai and Nuno Vasconcelos · 2019
Earlier work this paper cites.
MMDetection: Open mmlab detection toolbox and benchmark
Kai Chen, Jiaqi Wang, Jiangmiao Pang, Yuhang Cao, Yu Xiong, Xiaoxiao Li, Shuyang Sun, Wansen Feng, Ziwei Liu, Jiarui Xu, Zheng Zhang, Dazhi Cheng, Chenchen Zhu, Tianheng Cheng, Qijie Zhao, Buyu Li, Xin Lu, Rui Zhu, Yue Wu, Jifeng Dai, Jingdong Wang, Jianping Shi, Wanli Ouyang, Chen Change Loy, and Dahua Lin · 2019
Earlier work this paper cites.
Slowfast networks for video recognition
Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He · 2019
Earlier work this paper cites.
Searching for mobilenetv3
Andrew G. Howard, Mark Sandler, Grace Chu, Liang-Chieh Chen, Bo Chen, Mingxing Tan, Weijun Wang, Yukun Zhu, Ruoming Pang, Vijay Vasudevan, Quoc V. Le, and Hartwig Adam · 2019
Earlier work this paper cites.
Stm: Spatiotemporal and motion encoding for action recognition
Boyuan Jiang, Mengmeng Wang, Weihao Gan, Wei Wu, and Junjie Yan · 2019
Earlier work this paper cites.
Panoptic feature pyramid networks
Alexander Kirillov, Ross Girshick, Kaiming He, and Piotr Dollár · 2019
Earlier work this paper cites.
Tsm: Temporal shift module for efficient video understanding
Ji Lin, Chuang Gan, and Song Han · 2019
Earlier work this paper cites.
Grouped spatial-temporal aggregation for efficient action recognition
Chenxu Luo and Alan L. Yuille · 2019
Earlier work this paper cites.
Learning spatio-temporal representation with local and global diffusion
Zhaofan Qiu, Ting Yao, C. Ngo, Xinmei Tian, and Tao Mei · 2019
Cited alongside, same era.
Stand-alone self-attention in vision models
Prajit Ramachandran, Niki Parmar, Ashish Vaswani, Irwan Bello, Anselm Levskaya, and Jonathon Shlens · 2019
Cited alongside, same era.
Grad-cam: Visual explanations from deep networks via gradient-based localization
Ramprasaath R. Selvaraju, Abhishek Das, Ramakrishna Vedantam, Michael Cogswell, Devi Parikh, and Dhruv Batra · 2019
Cited alongside, same era.
Efficientnet: Rethinking model scaling for convolutional neural networks
Mingxing Tan and Quoc V. Le · 2019
Cited alongside, same era.
Video classification with channel-separated convolutional networks
Du Tran, Heng Wang, L. Torresani, and Matt Feiszli · 2019
Cited alongside, same era.
All tokens matter: Token labeling for training better vision transformers
Zihang Jiang, Qibin Hou, Li Yuan, Daquan Zhou, Yujun Shi, Xiaojie Jin, Anran Wang, and Jiashi Feng · 2021
Later among the works it cites.
Trseg: Transformer for semantic segmentation
Youngsaeng Jin, David K. Han, and Hanseok Ko · 2021
Later among the works it cites.
Movinets: Mobile video networks for efficient video recognition
D. Kondratyuk, Liangzhe Yuan, Yandong Li, Li Zhang, Mingxing Tan, Matthew A. Brown, and Boqing Gong · 2021
Later among the works it cites.
Pose recognition with cascade transformers
Ke Li, Shijie Wang, Xiang Zhang, Yifan Xu, Weijian Xu, and Zhuowen Tu · 2021
Later among the works it cites.
Ct-net: Channel tensorization network for video classification
Kunchang Li, Xianhang Li, Yali Wang, Jun Wang, and Y. Qiao · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cutmix: Regularization strategy to train strong classifiers with localizable features
Sangdoo Yun, Dongyoon Han, Seong Joon Oh, Sanghyuk Chun, Junsuk Choe, and Young Joon Yoo · 2019
Cited alongside, same era.
Semantic understanding of scenes through the ade20k dataset
Bolei Zhou, Hang Zhao, Xavier Puig, Tete Xiao, Sanja Fidler, Adela Barriuso, and Antonio Torralba · 2019
Cited alongside, same era.
End-to-end object detection with transformers
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko · 2020
Cited alongside, same era.
Openmmlab pose estimation toolbox and benchmark
MMPose Contributors · 2020
Cited alongside, same era.
MMSegmentation: Openmmlab semantic segmentation toolbox and benchmark
MMSegmentation Contributors · 2020
Cited alongside, same era.
On the relationship between self-attention and convolutional layers
Jean-Baptiste Cordonnier, Andreas Loukas, and Martin Jaggi · 2020
Cited alongside, same era.
Pyslowfast
Haoqi Fan, Yanghao Li, Bo Xiong, Wan-Yen Lo, and Christoph Feichtenhofer · 2020
Cited alongside, same era.
Xinyu Li, Yanyi Zhang, Chunhui Liu, Bing Shuai, Yi Zhu, Biagio Brattoli, Hao Chen, Ivan Marsic, and Joseph Tighe · 2021
Later among the works it cites.
Tokenpose: Learning keypoint tokens for human pose estimation
Yanjie Li, Shoukui Zhang, Zhicheng Wang, Sen Yang, Wankou Yang, Shutao Xia, and Erjin Zhou · 2021
Later among the works it cites.
Swinir: Image restoration using swin transformer
Jingyun Liang, Jie Cao, Guolei Sun, K. Zhang, Luc Van Gool, and Radu Timofte · 2021
Later among the works it cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, S. Lin, and B. Guo · 2021
Later among the works it cites.
Mobilevit: light-weight, general-purpose, and mobile-friendly vision transformer
Sachin Mehta and Mohammad Rastegari · 2021
Later among the works it cites.
Video transformer network
Daniel Neimark, Omri Bar, Maya Zohar, and Dotan Asselmann · 2021
Later among the works it cites.
Keeping your eye on the ball: Trajectory attention in video transformers
Mandela Patrick, Dylan Campbell, Yuki M. Asano, Ishan Misra Florian Metze, Christoph Feichtenhofer, A. Vedaldi, and João F. Henriques · 2021
Later among the works it cites.
Dynamicvit: Efficient vision transformers with dynamic token sparsification
Yongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu, Jie Zhou, and Cho-Jui Hsieh · 2021
Later among the works it cites.
An image is worth 16x16 words, what is a video worth?
Gilad Sharir, Asaf Noy, and Lihi Zelnik-Manor · 2021
Later among the works it cites.
Bottleneck transformers for visual recognition
A. Srinivas, Tsung-Yi Lin, Niki Parmar, Jonathon Shlens, P. Abbeel, and Ashish Vaswani · 2021
Later among the works it cites.
Efficientnetv2: Smaller models and faster training
Mingxing Tan and Quoc V. Le · 2021
Later among the works it cites.
Training data-efficient image transformers & distillation through attention
Hugo Touvron, M. Cord, M. Douze, Francisco Massa, Alexandre Sablayrolles, and Herv’e J’egou · 2021
Later among the works it cites.
Going deeper with image transformers
Hugo Touvron, M. Cord, Alexandre Sablayrolles, Gabriel Synnaeve, and Herv’e J’egou · 2021
Later among the works it cites.
Deep high-resolution representation learning for visual recognition
Jingdong Wang, Ke Sun, Tianheng Cheng, Borui Jiang, Chaorui Deng, Yang Zhao, D. Liu, Yadong Mu, Mingkui Tan, Xinggang Wang, Wenyu Liu, and Bin Xiao · 2021
Later among the works it cites.
Tdn: Temporal difference networks for efficient action recognition
Limin Wang, Zhan Tong, Bin Ji, and Gangshan Wu · 2021
Later among the works it cites.
Transformer meets tracker: Exploiting temporal context for robust visual tracking
Ning Wang, Wen gang Zhou, Jie Wang, and Houqaing Li · 2021
Later among the works it cites.
Pyramid vision transformer: A versatile backbone for dense prediction without convolutions
Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, P. Luo, and L. Shao · 2021
Later among the works it cites.
Pvtv2: Improved baselines with pyramid vision transformer
Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao · 2021
Later among the works it cites.
Cvt: Introducing convolutions to vision transformers
Haiping Wu, Bin Xiao, N. Codella, Mengchen Liu, Xiyang Dai, Lu Yuan, and Lei Zhang · 2021
Later among the works it cites.
Early convolutions help transformers see better
Tete Xiao, Mannat Singh, Eric Mintun, Trevor Darrell, Piotr Dollár, and Ross B. Girshick · 2021
Later among the works it cites.
Segformer: Simple and efficient design for semantic segmentation with transformers
Enze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar, Jose M Alvarez, and Ping Luo · 2021
Later among the works it cites.
Evo-vit: Slow-fast token evolution for dynamic vision transformer
Yifan Xu, Zhijie Zhang, Mengdan Zhang, Kekai Sheng, Ke Li, Weiming Dong, Liqing Zhang, Changsheng Xu, and Xing Sun · 2021
Later among the works it cites.
Focal self-attention for local-global interactions in vision transformers
Jianwei Yang, Chunyuan Li, Pengchuan Zhang, Xiyang Dai, Bin Xiao, Lu Yuan, and Jianfeng Gao · 2021
Later among the works it cites.
Incorporating convolution designs into visual transformers
Kun Yuan, Shaopeng Guo, Ziwei Liu, Aojun Zhou, Fengwei Yu, and Wei Wu · 2021
Later among the works it cites.
Tokens-to-token vit: Training vision transformers from scratch on imagenet
Li Yuan, Y. Chen, Tao Wang, Weihao Yu, Yujun Shi, Francis E. H. Tay, Jiashi Feng, and Shuicheng Yan · 2021
Later among the works it cites.
Volo: Vision outlooker for visual recognition
Li Yuan, Qibin Hou, Zihang Jiang, Jiashi Feng, and Shuicheng Yan · 2021
Later among the works it cites.
Hrformer: High-resolution transformer for dense prediction
Yuhui Yuan, Rao Fu, Lang Huang, Weihong Lin, Chao Zhang, Xilin Chen, and Jingdong Wang · 2021
Later among the works it cites.
Morphmlp: A self-attention free, mlp-like backbone for image and video
David Junhao Zhang, Kunchang Li, Yunpeng Chen, Yali Wang, Shashwat Chandra, Yu Qiao, Luoqi Liu, and Mike Zheng Shou · 2021
Later among the works it cites.
Multi-scale vision longformer: A new vision transformer for high-resolution image encoding
Pengchuan Zhang, Xiyang Dai, Jianwei Yang, Bin Xiao, Lu Yuan, Lei Zhang, and Jianfeng Gao · 2021
Later among the works it cites.
Convnets vs. transformers: Whose visual representations are more transferable?
Hong-Yu Zhou, Chixiang Lu, Sibei Yang, and Yizhou Yu · 2021
Later among the works it cites.
Deformable detr: Deformable transformers for end-to-end object detection
Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai · 2021
Later among the works it cites.
Self-slimmed vision transformer
Zhuofan Zong, Kunchang Li, Guanglu Song, Yali Wang, Y. Qiao, Biao Leng, and Yu Liu · 2021
Later among the works it cites.
Mobile-former: Bridging mobilenet and transformer
Yinpeng Chen, Xiyang Dai, Dongdong Chen, Mengchen Liu, Xiaoyi Dong, Lu Yuan, and Zicheng Liu · 2022
Closest in time.
Cswin transformer: A general vision transformer backbone with cross-shaped windows
Xiaoyi Dong, Jianmin Bao, Dongdong Chen, Weiming Zhang, Nenghai Yu, Lu Yuan, Dong Chen, and B. Guo · 2022
Closest in time.
Evit: Expediting vision transformers via token reorganizations
Youwei Liang, Chongjian Ge, Zhan Tong, Yibing Song, Jue Wang, and Pengtao Xie · 2022
Closest in time.
Video swin transformer
Ze Liu, Jia Ning, Yue Cao, Yixuan Wei, Zheng Zhang, S. Lin, and Han Hu · 2022
Closest in time.