Fetching the paper…
Reading the bibliography…
Since the introduction of Vision Transformers, the landscape of many computer vision tasks (e.g., semantic segmentation), which has been overwhelmingly dominated by CNNs, recently has significantly revolutionized.
Heo B, Kim J, Yun S, Park H, Kwak N, Choi JY (2019) A comprehensive overhaul of feature distillation. In: Proceedings of the IEEE/CVF international conference on computer vision, pp 1921–1930
1930
Earlier work this paper cites.
Deng J, Dong W, Socher R, Li LJ, Li K, Fei-Fei L (2009) Imagenet: A large-scale hierarchical image database. In: 2009 IEEE conference on computer vision and pattern recognition, pp 248–255
2009
Earlier work this paper cites.
Kingma DP (2014) Adam: A method for stochastic optimization. arXiv preprint arXiv:14126980
2014
Earlier work this paper cites.
Mottaghi R, Chen X, Liu X, Cho NG, Lee SW, Fidler S, Urtasun R, Yuille A (2014) The role of context for object detection and semantic segmentation in the wild. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 891–898
2014
Earlier work this paper cites.
Romero A, Ballas N, Kahou SE, Chassang A, Gatta C, Bengio Y (2014) Fitnets: Hints for thin deep nets. arXiv preprint arXiv:14126550
2014
Earlier work this paper cites.
Hinton G (2015) Distilling the knowledge in a neural network. arXiv preprint arXiv:150302531
2015
Earlier work this paper cites.
Ioffe S, Szegedy C (2015) Batch normalization: Accelerating deep network training by reducing internal covariate shift. In: International Conference on Machine Learning
2015
Earlier work this paper cites.
Long J, Shelhamer E, Darrell T (2015) Fully convolutional networks for semantic segmentation. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 3431–3440
2015
Earlier work this paper cites.
Luong MT (2015) Effective approaches to attention-based neural machine translation. arXiv preprint arXiv:150804025
2015
Earlier work this paper cites.
Cordts M, Omran M, Ramos S, Rehfeld T, Enzweiler M, Benenson R, Franke U, Roth S, Schiele B (2016) The cityscapes dataset for semantic urban scene understanding. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 3213–3223
2016
Earlier work this paper cites.
He K, Zhang X, Ren S, Sun J (2016) Deep residual learning for image recognition. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 770–778
2016
Earlier work this paper cites.
Chen LC (2017) Rethinking atrous convolution for semantic image segmentation. arXiv preprint arXiv:170605587
2017
Earlier work this paper cites.
Lin T (2017) Focal loss for dense object detection. arXiv preprint arXiv:170802002
2017
Earlier work this paper cites.
Vaswani A (2017) Attention is all you need. Advances in Neural Information Processing Systems
2017
Earlier work this paper cites.
Yim J, Joo D, Bae J, Kim J (2017) A gift from knowledge distillation: Fast optimization, network minimization and transfer learning. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 4133–4141
2017
Earlier work this paper cites.
Zhou B, Zhao H, Puig X, Fidler S, Barriuso A, Torralba A (2017) Scene parsing through ade20k dataset. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 633–641
2017
Earlier work this paper cites.
Caesar H, Uijlings J, Ferrari V (2018) Coco-stuff: Thing and stuff classes in context. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 1209–1218
2018
Earlier work this paper cites.
Kim J, Park S, Kwak N (2018) Paraphrasing complex network: Network compression via factor transfer. Advances in neural information processing systems 31
2018
Earlier work this paper cites.
Liu PJ, Saleh M, Pot E, Goodrich B, Sepassi R, Kaiser L, Shazeer N (2018) Generating wikipedia by summarizing long sequences. In: International Conference on Learning Representations
2018
Earlier work this paper cites.
Ma N, Zhang X, Zheng HT, Sun J (2018) Shufflenet v2: Practical guidelines for efficient cnn architecture design. In: Proceedings of the European conference on computer vision (ECCV), pp 116–131
2018
Earlier work this paper cites.
Sandler M, Howard A, Zhu M, Zhmoginov A, Chen LC (2018) Mobilenetv2: Inverted residuals and linear bottlenecks. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 4510–4520
2018
Earlier work this paper cites.
Woo S, Park J, Lee JY, Kweon IS (2018) Cbam: Convolutional block attention module. In: Proceedings of the European conference on computer vision (ECCV), pp 3–19
2018
Earlier work this paper cites.
Yu C, Wang J, Peng C, Gao C, Yu G, Sang N (2018) Bisenet: Bilateral segmentation network for real-time semantic segmentation. In: Proceedings of the European conference on computer vision (ECCV), pp 325–341
2018
Earlier work this paper cites.
Zhang Y, Xiang T, Hospedales TM, Lu H (2018) Deep mutual learning. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 4320–4328
2018
Cited alongside, same era.
Zhao H, Qi X, Shen X, Shi J, Jia J (2018) Icnet for real-time semantic segmentation on high-resolution images. In: Proceedings of the European conference on computer vision (ECCV), pp 405–420
2018
Cited alongside, same era.
Cao Y, Xu J, Lin S, Wei F, Hu H (2019) Gcnet: Non-local networks meet squeeze-excitation networks and beyond. In: Proceedings of the IEEE/CVF international conference on computer vision workshops, pp 0–0
2019
Cited alongside, same era.
Cho JH, Hariharan B (2019) On the efficacy of knowledge distillation. In: Proceedings of the IEEE/CVF international conference on computer vision, pp 4794–4802
2019
Cited alongside, same era.
Choromanski K, Likhosherstov V, Dohan D, Song X, Gane A, Sarlos T, Hawkins P, Davis J, Mohiuddin A, Kaiser L, et al. (2021) Rethinking attention with performers. In: International Conference on Learning Representations
2021
Later among the works it cites.
Dosovitskiy A, Beyer L, Kolesnikov A, Weissenborn D, Zhai X, Unterthiner T, Dehghani M, Minderer M, Heigold G, Gelly S, et al. (2021) An image is worth 16x16 words: Transformers for image recognition at scale. In: International Conference on Learning Representations
2021
Later among the works it cites.
Hong Y, Pan H, Sun W, Jia Y (2021) Deep dual-resolution networks for real-time and accurate semantic segmentation of road scenes. arXiv preprint arXiv:210106085
2021
Later among the works it cites.
Hou Q, Zhou D, Feng J (2021) Coordinate attention for efficient mobile network design. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp 13713–13722
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
He T, Shen C, Tian Z, Gong D, Sun C, Yan Y (2019) Knowledge adaptation for efficient semantic segmentation. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp 578–587
2019
Cited alongside, same era.
Ho J, Kalchbrenner N, Weissenborn D, Salimans T (2019) Axial attention in multidimensional transformers. arXiv preprint arXiv:191212180
2019
Cited alongside, same era.
Howard A, Sandler M, Chu G, Chen LC, Chen B, Tan M, Wang W, Zhu Y, Pang R, Vasudevan V, et al. (2019) Searching for mobilenetv3. In: Proceedings of the IEEE/CVF international conference on computer vision, pp 1314–1324
2019
Cited alongside, same era.
Kirillov A, Girshick R, He K, Dollár P (2019) Panoptic feature pyramid networks. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp 6399–6408
2019
Cited alongside, same era.
Li H, Xiong P, Fan H, Sun J (2019) Dfanet: Deep feature aggregation for real-time semantic segmentation. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp 9522–9531
2019
Cited alongside, same era.
Liu Y, Chen K, Liu C, Qin Z, Luo Z, Wang J (2019) Structured knowledge distillation for semantic segmentation. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp 2604–2613
2019
Cited alongside, same era.
Park W, Kim D, Lu Y, Cho M (2019) Relational knowledge distillation. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp 3967–3976
2019
Cited alongside, same era.
Liu Z, Lin Y, Cao Y, Hu H, Wei Y, Zhang Z, Lin S, Guo B (2021) Swin transformer: Hierarchical vision transformer using shifted windows. In: Proceedings of the IEEE/CVF international conference on computer vision, pp 10012–10022
2021
Later among the works it cites.
Qi L, Kuen J, Gu J, Lin Z, Wang Y, Chen Y, Li Y, Jia J (2021) Multi-scale aligned distillation for low-resolution detection. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 14443–14453
2021
Later among the works it cites.
Shen Z, Zhang M, Zhao H, Yi S, Li H (2021) Efficient attention: Attention with linear complexities. In: Proceedings of the IEEE/CVF winter conference on applications of computer vision, pp 3531–3539
2021
Later among the works it cites.
Wang W, Xie E, Li X, Fan DP, Song K, Liang D, Lu T, Luo P, Shao L (2021) Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. In: Proceedings of the IEEE/CVF international conference on computer vision, pp 568–578
2021
Later among the works it cites.
Xie E, Wang W, Yu Z, Anandkumar A, Alvarez JM, Luo P (2021) Segformer: Simple and efficient design for semantic segmentation with transformers. Advances in neural information processing systems 34:12077–12090
2021
Later among the works it cites.
Xu W, Xu Y, Chang T, Tu Z (2021) Co-scale conv-attentional image transformers. In: Proceedings of the IEEE/CVF international conference on computer vision, pp 9981–9990
2021
Later among the works it cites.
Yan H, Li Z, Li W, Wang C, Wu M, Zhang C (2021) Contnet: Why not use convolution and transformer at the same time? arXiv preprint arXiv:210413497
2021
Later among the works it cites.
Yu C, Gao C, Wang J, Yu G, Shen C, Sang N (2021) Bisenet v2: Bilateral network with guided aggregation for real-time semantic segmentation. International journal of computer vision 129:3051–3068
2021
Later among the works it cites.
Yuan Y, Fu R, Huang L, Lin W, Zhang C, Chen X, Wang J (2021) Hrformer: High-resolution transformer for dense prediction. arXiv preprint arXiv:211009408
2021
Later among the works it cites.
Zheng S, Lu J, Zhao H, Zhu X, Luo Z, Wang Y, Fu Y, Feng J, Xiang T, Torr PH, et al. (2021) Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp 6881–6890
2021
Later among the works it cites.
Hu B, Zhou S, Xiong Z, Wu F (2022) Cross-resolution distillation for efficient 3d medical image registration. IEEE Transactions on Circuits and Systems for Video Technology 32(10):7269–7283
2022
Later among the works it cites.
Li Y, Yuan G, Wen Y, Hu J, Evangelidis G, Tulyakov S, Wang Y, Ren J (2022) Efficientformer: Vision transformers at mobilenet speed. Advances in Neural Information Processing Systems 35:12934–12949
2022
Later among the works it cites.
Liu R, Yang K, Roitberg A, Zhang J, Peng K, Liu H, Wang Y, Stiefelhagen R (2022) Transkd: Transformer knowledge distillation for efficient semantic segmentation. arXiv preprint arXiv:220213393
2022
Later among the works it cites.
Zhao B, Cui Q, Song R, Qiu Y, Liang J (2022) Decoupled knowledge distillation. In: Proceedings of the IEEE/CVF Conference on computer vision and pattern recognition, pp 11953–11962
2022
Later among the works it cites.
Tang S, Sun T, Peng J, Chen G, Hao Y, Lin M, Xiao Z, You J, Liu Y (2023) Pp-mobileseg: Explore the fast and accurate semantic segmentation model on mobile devices. arXiv preprint arXiv:230405152
2023
Closest in time.
Vasu PKA, Gabriel J, Zhu J, Tuzel O, Ranjan A (2023) Mobileone: An improved one millisecond mobile backbone. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp 7907–7917
2023
Closest in time.
Wan Q, Huang Z, Lu J, Yu G, Zhang L (2023) Seaformer: Squeeze-enhanced axial transformer for mobile semantic segmentation. In: International Conference on Learning Representations
2023
Closest in time.
Zhang L, Chen M, Arnab A, Xue X, Torr PH (2023) Dynamic graph message passing networks. IEEE Transactions on Pattern Analysis & Machine Intelligence 45(05):5712–5730
2023
Closest in time.
Wang J, Chen Y, Zheng Z, Li X, Cheng MM, Hou Q (2024) Crosskd: Cross-head knowledge distillation for object detection. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 16520–16530
2024
Closest in time.