Fetching the paper…
Reading the bibliography…
Vision Transformer (ViT) architectures traditionally employ a grid-based approach to tokenization independent of the semantic content of an image.
Perona, P., Malik, J.: Scale-space and edge detection using anisotropic diffusion. IEEE Trans. Pattern Anal. Mach. Intell. 12
1990
Earlier work this paper cites.
Xiaohan, Y., Yla-Jaaski, J., Huttunen, O., Vehkomaki, T., Sipila, O., Katila, T.: Image segmentation combining region growing and edge detection. In: IEEE Inter. Conf. Pattern Recog. (ICPR), pp. 481–484 (1992), doi: 10.1109/ICPR.1992.202029
1992
Earlier work this paper cites.
Fellbaum, C.: WordNet: An Electronic Lexical Database. Bradford Books (1998), URL https://mitpress.mit.edu/9780262561167/
1998
Earlier work this paper cites.
Shi, J., Malik, J.: Normalized cuts and image segmentation. IEEE Trans. Pattern Anal. Mach. Intell. 22
2000
Earlier work this paper cites.
Leung, T., Malik, J.: Representing and recognizing the visual appearance of materials using three-dimensional textons. Inter. J. Comput. Vis. 43
2001
Earlier work this paper cites.
Dalal, N., Triggs, B.: Histograms of oriented gradients for human detection. In: IEEE/CVF Inter. Conf. Comput. Vis. Pattern Recog. (CVPR), vol. 1, pp. 886–893 vol. 1 (2005), doi: 10.1109/CVPR.2005.177
2005
Earlier work this paper cites.
2006
Earlier work this paper cites.
Scharr, H.: Optimal filters for extended optical flow. In: Jähne, B., Mester, R., Barth, E., Scharr, H. (eds.) Int. Wksp. Compl. Mot. (IWCM), pp. 14–29, Springer Berlin Heidelberg, Berlin, Heidelberg (2007), ISBN 978-3-540-69866-1
2007
Earlier work this paper cites.
Moore, A.P., Prince, S.J.D., Warrell, J., Mohammed, U., Jones, G.: Superpixel lattices. In: IEEE/CVF Inter. Conf. Comput. Vis. Pattern Recog. (CVPR) (2008), doi: 10.1109/CVPR.2008.4587471
2008
Earlier work this paper cites.
Vedaldi, A., Soatto, S.: Quick shift and kernel methods for mode seeking. In: Forsyth, D., Torr, P., Zisserman, A. (eds.) European Conf. Comput. Vis. (ECCV), pp. 705–718, Springer Berlin Heidelberg, Berlin, Heidelberg (2008), ISBN 978-3-540-88693-8
2008
Earlier work this paper cites.
Deng, J., Dong, W., Socher, R., Li, L., Li, K., Fei-Fei, L.: Imagenet: A large-scale hierarchical image database. In: IEEE/CVF Inter. Conf. Comput. Vis. Pattern Recog. (CVPR), pp. 248–255, IEEE Computer Society (2009), URL https://doi.org/10.1109/CVPR.2009.5206848
2009
Earlier work this paper cites.
Krizhevsky, A., Hinton, G., et al.: Learning multiple layers of features from tiny images (2009)
2009
Earlier work this paper cites.
Ladický, L., Russell, C., Kohli, P., Torr, P.H.: Associative hierarchical crfs for object class image segmentation. In: IEEE Inter. Conf. Comput. Vis. (ICCV), pp. 739–746 (2009), doi: 10.1109/ICCV.2009.5459248
2009
Earlier work this paper cites.
Achanta, R., Shaji, A., Smith, K., Lucchi, A., Fua, P., Süsstrunk, S.: Slic superpixels compared to state-of-the-art superpixel methods. IEEE Trans. Pattern Anal. Mach. Intell. 34
2012
Earlier work this paper cites.
Yan, Q., Xu, L., Shi, J., Jia, J.: Hierarchical saliency detection. In: IEEE/CVF Inter. Conf. Comput. Vis. Pattern Recog. (CVPR), pp. 1155–1162 (2013), doi: 10.1109/CVPR.2013.153
2013
Earlier work this paper cites.
Yang, C., Zhang, L., Lu, H., Ruan, X., Yang, M.H.: Saliency detection via graph-based manifold ranking. In: IEEE/CVF Inter. Conf. Comput. Vis. Pattern Recog. (CVPR), pp. 3166–3173, IEEE (2013)
2013
Earlier work this paper cites.
Yan, J., Yu, Y., Zhu, X., Lei, Z., Li, S.Z.: Object detection by labeling superpixels. In: IEEE/CVF Inter. Conf. Comput. Vis. Pattern Recog. (CVPR), pp. 5107–5116 (2015), doi: 10.1109/CVPR.2015.7299146
2015
Earlier work this paper cites.
Ribeiro, M.T., Singh, S., Guestrin, C.: “why should I trust you?” explaining the predictions of any classifier. In: ACM Conf. Knowl. Discov. Data Min. (ACM SIGKDD), pp. 1135–1144 (2016), URL https://doi.org/10.1145/2939672.2939778
2016
Earlier work this paper cites.
Sennrich, R., Haddow, B., Birch, A.: Neural machine translation of rare words with subword units. In: Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics, ACL 2016, August 7-12, 2016, Berlin, Germany, Volume 1: Long Papers, The Association for Computer Linguistics (2016), URL https://doi.org/10.18653/v1/p16-1162
2016
Cited alongside, same era.
Johnson, M., Schuster, M., Le, Q.V., Krikun, M., Wu, Y., Chen, Z., Thorat, N., Viégas, F.B., Wattenberg, M., Corrado, G., Hughes, M., Dean, J.: Google’s multilingual neural machine translation system: Enabling zero-shot translation. Trans. Assoc. Comput. Linguistics 5
2017
Cited alongside, same era.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, L., Polosukhin, I.: Attention is all you need. In: Guyon, I., von Luxburg, U., Bengio, S., Wallach, H.M., Fergus, R., Vishwanathan, S.V.N., Garnett, R. (eds.) Adv. Neural Inf. Process. Sys. (NeurIPS), pp. 5998–6008 (2017), URL https://proceedings.neurips.cc/paper/2017/hash/3f5ee243547dee91fbd053c1c4a845aa-Abstract.html
2017
Cited alongside, same era.
Touvron, H., Cord, M., Sablayrolles, A., Synnaeve, G., Jégou, H.: Going deeper with image transformers. In: IEEE Inter. Conf. Comput. Vis. (ICCV), pp. 32–42, IEEE (2021c), URL https://doi.org/10.1109/ICCV48922.2021.00010
2021
Later among the works it cites.
Wang, W., Xie, E., Li, X., Fan, D., Song, K., Liang, D., Lu, T., Luo, P., Shao, L.: Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. In: IEEE Inter. Conf. Comput. Vis. (ICCV), pp. 548–558, IEEE (2021), doi: 10.1109/ICCV48922.2021.00061
2021
Later among the works it cites.
Xie, E., Wang, W., Yu, Z., Anandkumar, A., Alvarez, J.M., Luo, P.: SegFormer: Simple and efficient design for semantic segmentation with transformers. In: Ranzato, M., Beygelzimer, A., Dauphin, Y.N., Liang, P., Vaughan, J.W. (eds.) Adv. Neural Inf. Process. Sys. (NeurIPS), pp. 12077–12090 (2021), URL https://proceedings.neurips.cc/paper/2021/hash/64f1f27bf1b4ec22924fd0acb550c235-Abstract.html
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wang, L., Lu, H., Wang, Y., Feng, M., Wang, D., Yin, B., Ruan, X.: Learning to detect salient objects with image-level supervision. In: IEEE/CVF Inter. Conf. Comput. Vis. Pattern Recog. (CVPR) (2017)
2017
Cited alongside, same era.
Adebayo, J., Gilmer, J., Muelly, M., Goodfellow, I.J., Hardt, M., Kim, B.: Sanity checks for saliency maps. In: Bengio, S., Wallach, H.M., Larochelle, H., Grauman, K., Cesa-Bianchi, N., Garnett, R. (eds.) Adv. Neural Inf. Process. Sys. (NeurIPS), pp. 9525–9536 (2018), URL https://proceedings.neurips.cc/paper/2018/hash/294a8ed24b1ad22ec2e7efea049b8737-Abstract.html
2018
Cited alongside, same era.
Stutz, D., Hermans, A., Leibe, B.: Superpixels: An evaluation of the state-of-the-art. Comput. Vis. Image Underst. 166
2018
Cited alongside, same era.
Wei, X., Yang, Q., Gong, Y., Ahuja, N., Yang, M.: Superpixel hierarchy. IEEE Trans. Image Process. 27
2018
Cited alongside, same era.
Devlin, J., Chang, M., Lee, K., Toutanova, K.: BERT: pre-training of deep bidirectional transformers for language understanding. In: Burstein, J., Doran, C., Solorio, T. (eds.) Conf. North Amer. Ch. Assoc. Comput. Ling. (NAACL), pp. 4171–4186, Association for Computational Linguistics (2019), URL https://doi.org/10.18653/v1/n19-1423
2019
Cited alongside, same era.
Yun, S., Han, D., Chun, S., Oh, S.J., Yoo, Y., Choe, J.: Cutmix: Regularization strategy to train strong classifiers with localizable features. In: IEEE Inter. Conf. Comput. Vis. (ICCV), pp. 6022–6031, IEEE (2019), URL https://doi.org/10.1109/ICCV.2019.00612
2019
Cited alongside, same era.
Abnar, S., Zuidema, W.H.: Quantifying attention flow in transformers. In: Jurafsky, D., Chai, J., Schluter, N., Tetreault, J.R. (eds.) Conf. Assoc. Comput. Ling. (ACL), pp. 4190–4197, Association for Computational Linguistics (2020), URL https://doi.org/10.18653/v1/2020.acl-main.385
2020
Cited alongside, same era.
Brown, T.B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D.M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., Amodei, D.: Language models are few-shot learners. In: Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., Lin, H. (eds.) Adv. Neural Inf. Process. Sys. (NeurIPS) (2020), URL https://proceedings.neurips.cc/paper/2020/hash/1457c0d6bfcb4967418bfb8ac142f64a-Abstract.html
2020
Cited alongside, same era.
Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., Zagoruyko, S.: End-to-end object detection with transformers. In: Vedaldi, A., Bischof, H., Brox, T., Frahm, J. (eds.) European Conf. Comput. Vis. (ECCV), Lecture Notes in Computer Science, vol. 12346, pp. 213–229, Springer (2020), doi: 10.1007/978-3-030-58452-8\_13
2020
Cited alongside, same era.
Yuan, L., Chen, Y., Wang, T., Yu, W., Shi, Y., Jiang, Z., Tay, F.E.H., Feng, J., Yan, S.: Tokens-to-Token ViT: Training vision transformers from scratch on imagenet. In: IEEE Inter. Conf. Comput. Vis. (ICCV), pp. 538–547, IEEE (2021), doi: 10.1109/ICCV48922.2021.00060
2021
Later among the works it cites.
Chan, C.S., Kong, H., Guanqing, L.: A comparative study of faithfulness metrics for model interpretability methods. In: Conf. Assoc. Comput. Ling. (ACL), pp. 5029–5038, Association for Computational Linguistics, Dublin, Ireland (2022), URL https://aclanthology.org/2022.acl-long.345
2022
Later among the works it cites.
Griffin, G., Holub, A., Perona, P.: Caltech 256 (Apr 2022), doi: 10.22002/D1.20087
2022
Later among the works it cites.
Hamilton, M., Zhang, Z., Hariharan, B., Snavely, N., Freeman, W.T.: Unsupervised semantic segmentation by distilling feature correspondences. In: Inter. Conf. Learn. Represent. (ICLR) (2022), URL https://openreview.net/forum?id=SaKO6z6Hl0c
2022
Later among the works it cites.
Huang, H., Zhou, X., Cao, J., He, R., Tan, T.: Vision transformer with super token sampling (2022)
2022
Later among the works it cites.
Steiner, A., Kolesnikov, A., Zhai, X., Wightman, R., Uszkoreit, J., Beyer, L.: How to train your vit? data, augmentation, and regularization in vision transformers. Trans. Mach. Learn. Res. (2022), URL https://openreview.net/forum?id=4nPswr1KcP
2022
Later among the works it cites.
Touvron, H., Cord, M., Jégou, H.: Deit III: revenge of the vit. In: Avidan, S., Brostow, G.J., Cissé, M., Farinella, G.M., Hassner, T. (eds.) European Conf. Comput. Vis. (ECCV), Lecture Notes in Computer Science, vol. 13684, pp. 516–533, Springer (2022), doi: 10.1007/978-3-031-20053-3\_30
2022
Later among the works it cites.
Yan, T., Huang, X., Zhao, Q.: Hierarchical superpixel segmentation by parallel crtrees labeling. IEEE Trans. Image Process. 31
2022
Later among the works it cites.
Bolya, D., Fu, C.Y., Dai, X., Zhang, P., Feichtenhofer, C., Hoffman, J.: Token merging: Your vit but faster. In: Inter. Conf. Learn. Represent. (ICLR) (2023), URL https://openreview.net/forum?id=JroZRaRw7Eu
2023
Later among the works it cites.
Havtorn, J.D., Royer, A., Blankevoort, T., Bejnordi, B.E.: Msvit: Dynamic mixed-scale tokenization for vision transformers. In: IEEE Inter. Conf. Comput. Vis. (ICCV), pp. 838–848 (October 2023), URL https://doi.org/10.1109/ICCVW60793.2023.00091
2023
Later among the works it cites.
Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A.C., Lo, W., Dollár, P., Girshick, R.B.: Segment anything (2023), URL https://doi.org/10.1109/ICCV51070.2023.00371
2023
Later among the works it cites.
Ma, X., Zhou, Y., Wang, H., Qin, C., Sun, B., Liu, C., Fu, Y.: Image as set of points. In: Inter. Conf. Learn. Represent. (ICLR) (2023), URL https://openreview.net/forum?id=awnvqZja69
2023
Later among the works it cites.
Ronen, T., Levy, O., Golbert, A.: Vision transformers with mixed-resolution tokenization. In: IEEE/CVF Inter. Conf. Comput. Vis. Pattern Recog. (CVPR), pp. 4612–4621 (2023)
2023
Later among the works it cites.
Oquab, M., Darcet, T., Moutakanni, T., Vo, H.V., Szafraniec, M., Khalidov, V., Fernandez, P., Haziza, D., Massa, F., El-Nouby, A., Assran, M., Ballas, N., Galuba, W., Howes, R., Huang, P., Li, S., Misra, I., Rabbat, M., Sharma, V., Synnaeve, G., Xu, H., Jégou, H., Mairal, J., Labatut, P., Joulin, A., Bojanowski, P.: Dinov2: Learning robust visual features without supervision. Trans. Mach. Learn. Res. (2024), URL https://openreview.net/forum?id=a68SUt6zFt
2024
Closest in time.