Fetching the paper…
Reading the bibliography…
Computer vision researchers are embracing two promising paradigms: Vision Transformers (ViTs) and Multi-task Learning (MTL), which both show great performance but are computation-intensive, given the quadratic complexity of self-attention in ViT and the need to activate an entire large MTL model for one task.
J. Ma, Z. Zhao, X. Yi, J. Chen, L. Hong, and E. H. Chi, “Modeling task relationships in multi-task learning with multi-gate mixture-of-experts,” in Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining , 2018, pp. 1930–1939
1939
Earlier work this paper cites.
M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Benenson, U. Franke, S. Roth, and B. Schiele, “The cityscapes dataset for semantic urban scene understanding,” in Proc. of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
I. Misra, A. Shrivastava, A. Gupta, and M. Hebert, “Cross-stitch networks for multi-task learning,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 3994–4003
2016
Earlier work this paper cites.
2018
Earlier work this paper cites.
D. Xu, W. Ouyang, X. Wang, and N. Sebe, “Pad-net: Multi-tasks guided prediction-and-distillation network for simultaneous depth estimation and scene parsing,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 675–684
2018
Earlier work this paper cites.
A. R. Zamir, A. Sax, W. Shen, L. J. Guibas, J. Malik, and S. Savarese, “Taskonomy: Disentangling task transfer learning,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 3712–3722
2018
Earlier work this paper cites.
Z. Zhang, Z. Cui, C. Xu, Z. Jie, X. Li, and J. Yang, “Joint task-recursive learning for semantic segmentation and depth estimation,” in Proceedings of the European Conference on Computer Vision (ECCV) , 2018, pp. 235–251
2018
Earlier work this paper cites.
Y. Gao, J. Ma, M. Zhao, W. Liu, and A. L. Yuille, “Nddr-cnn: Layerwise feature fusing in multi-task cnns by neural discriminative dimensionality reduction,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 3205–3214
2019
Earlier work this paper cites.
S. Liu, E. Johns, and A. J. Davison, “End-to-end multi-task learning with attention,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2019, pp. 1871–1880
2019
Earlier work this paper cites.
S. Ruder, J. Bingel, I. Augenstein, and A. Søgaard, “Latent multi-task architecture learning,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 33, no. 01, 2019, pp. 4822–4829
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko, “End-to-end object detection with transformers,” in European conference on computer vision . Springer, 2020, pp. 213–229
2020
Earlier work this paper cites.
2020
Cited alongside, same era.
B. Li, S. Pandey, H. Fang, Y. Lyv, J. Li, J. Chen, M. Xie, L. Wan, H. Liu, and C. Ding, “FTRANS: Energy-efficient acceleration of transformers using FPGA,” in Proceedings of the ACM/IEEE International Symposium on Low Power Electronics and Design , ser. ISLPED ’20. New York, NY, USA: Association for Computing Machinery, Aug. 2020, pp. 175–180
2020
Cited alongside, same era.
S. Lu, M. Wang, S. Liang, J. Lin, and Z. Wang, “Hardware accelerator for multi-head attention and position-wise feed-forward in the transformer,” in 2020 IEEE 33rd International System-on-Chip Conference (SOCC) . IEEE, 2020, pp. 84–89
2020
Cited alongside, same era.
S. Vandenhende, S. Georgoulis, and L. V. Gool, “Mti-net: Multi-scale task interaction networks for multi-task learning,” in European Conference on Computer Vision . Springer, 2020, pp. 527–543
P. Qi, Y. Song, H. Peng, S. Huang, Q. Zhuge, and E. H.-M. Sha, “Accommodating transformer onto FPGA: Coupling the balanced model compression and FPGA-implementation optimization,” in Proceedings of the 2021 on Great Lakes Symposium on VLSI . New York, NY, USA: Association for Computing Machinery, Jun. 2021, pp. 163–168
2021
Later among the works it cites.
I. O. Tolstikhin, N. Houlsby, A. Kolesnikov, L. Beyer, X. Zhai, T. Unterthiner, J. Yung, A. Steiner, D. Keysers, J. Uszkoreit et al. , “Mlp-mixer: An all-mlp architecture for vision,” Advances in Neural Information Processing Systems , vol. 34, pp. 24 261–24 272, 2021
2021
Later among the works it cites.
H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jegou, “Training data-efficient image transformers & distillation through attention,” in International Conference on Machine Learning , vol. 139, July 2021, pp. 10 347–10 357
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
A. Arnab, M. Dehghani, G. Heigold, C. Sun, M. Lučić, and C. Schmid, “Vivit: A video vision transformer,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 6836–6846
2021
Cited alongside, same era.
H. Chen, Y. Wang, T. Guo, C. Xu, Y. Deng, Z. Liu, S. Ma, C. Xu, C. Xu, and W. Gao, “Pre-trained image processing transformer,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 12 299–12 310
2021
Cited alongside, same era.
K. Han, A. Xiao, E. Wu, J. Guo, C. Xu, and Y. Wang, “Transformer in transformer,” Advances in Neural Information Processing Systems , vol. 34, pp. 15 908–15 919, 2021
2021
Cited alongside, same era.
J. He, J. Qiu, A. Zeng, Z. Yang, J. Zhai, and J. Tang, “FastMoE: A fast mixture-of-expert training system,” Mar. 2021
2021
Cited alongside, same era.
Z. Kong, P. Dong, X. Ma, X. Meng, W. Niu, M. Sun, B. Ren, M. Qin, H. Tang, and Y. Wang, “SPViT: Enabling faster vision transformers via soft token pruning,” Dec. 2021
2021
Cited alongside, same era.
D.-G. Lee, “Fast drivable areas estimation with multi-task learning for real-time autonomous driving assistant,” Applied Sciences , vol. 11, no. 22, p. 10713, 2021
2021
Cited alongside, same era.
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 10 012–10 022
2021
Cited alongside, same era.
H. Peng, S. Huang, T. Geng, A. Li, W. Jiang, H. Liu, S. Wang, and C. Ding, “Accelerating transformer-based deep learning models on FPGAs using column balanced block pruning,” in 2021 22nd International Symposium on Quality Electronic Design (ISQED) , Apr. 2021, pp. 142–148
2021
Cited alongside, same era.
2021
Later among the works it cites.
S. Vandenhende, S. Georgoulis, W. Van Gansbeke, M. Proesmans, D. Dai, and L. Van Gool, “Multi-task learning for dense prediction tasks: A survey,” IEEE transactions on pattern analysis and machine intelligence , 2021
2021
Later among the works it cites.
L. Yuan, Y. Chen, T. Wang, W. Yu, Y. Shi, Z.-H. Jiang, F. E. Tay, J. Feng, and S. Yan, “Tokens-to-token vit: Training vision transformers from scratch on imagenet,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 558–567
2021
Later among the works it cites.
X. Zhang, Y. Wu, P. Zhou, X. Tang, and J. Hu, “Algorithm-hardware co-design of attention mechanism on fpga devices,” ACM Transactions on Embedded Computing Systems (TECS) , vol. 20, no. 5s, pp. 1–24, 2021
2021
Later among the works it cites.
Q. Chen, C. Sun, Z. Lu, and C. Gao, “Enabling energy-efficient inference for self-attention mechanisms in neural networks,” in 2022 IEEE 4th International Conference on Artificial Intelligence Circuits and Systems (AICAS) . IEEE, 2022, pp. 25–28
2022
Later among the works it cites.
2022
Later among the works it cites.
H. Liang, Z. Fan, R. Sarkar, Z. Jiang, T. Chen, Y. Cheng, C. Hao, and Z. Wang, “M³vit: Mixture-of-experts vision transformer for efficient multi-task learning with model-accelerator co-design,” in 36th Annual Conference on Neural Information Processing System (NeurIPS 2022) , December 2022
2022
Later among the works it cites.
S. Park, G. Kim, Y. Oh, J. B. Seo, S. M. Lee, J. H. Kim, S. Moon, J.-K. Lim, and J. C. Ye, “Multi-task vision transformer using low-level chest x-ray feature corpus for covid-19 diagnosis and severity quantification,” Medical Image Analysis , vol. 75, p. 102299, 2022
2022
Later among the works it cites.
M. Sun, H. Ma, G. Kang, Y. Jiang, T. Chen, X. Ma, Z. Wang, and Y. Wang, “VAQF: Fully automatic software-hardware co-design framework for low-bit vision transformer,” Feb. 2022
2022
Later among the works it cites.