Fetching the paper…
Reading the bibliography…
Semantic segmentation tasks naturally require high-resolution information for pixel-wise segmentation and global context information for class prediction.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” Adv. Neural Inform. Process. Syst. , vol. 25, pp. 1097–1105, 2012
2012
Earlier work this paper cites.
2014
Earlier work this paper cites.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in Eur. Conf. Comput. Vis. Springer, 2014, pp. 740–755
2014
Earlier work this paper cites.
S. Kazemzadeh, V. Ordonez, M. Matten, and T. Berg, “Referitgame: Referring to objects in photographs of natural scenes,” in Proceedings of the 2014 conference on empirical methods in natural language processing , 2014, pp. 787–798
2014
Earlier work this paper cites.
J. Long, E. Shelhamer, and T. Darrell, “Fully convolutional networks for semantic segmentation,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2015, pp. 3431–3440
2015
Earlier work this paper cites.
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich, “Going deeper with convolutions,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2015, pp. 1–9
2015
Earlier work this paper cites.
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein et al. , “ImageNet large scale visual recognition challenge,” Int. J. Comput. Vis. , vol. 115, no. 3, pp. 211–252, 2015
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2016, pp. 770–778
2016
Earlier work this paper cites.
F. Yu and V. Koltun, “Multi-scale context aggregation by dilated convolutions,” in Int. Conf. Learn. Represent. , 2016
2016
Earlier work this paper cites.
M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Benenson, U. Franke, S. Roth, and B. Schiele, “The cityscapes dataset for semantic urban scene understanding,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2016, pp. 3213–3223
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
A. Shrivastava, A. Gupta, and R. Girshick, “Training region-based object detectors with online hard example mining,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2016, pp. 761–769
2016
Earlier work this paper cites.
L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille, “Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 40, no. 4, pp. 834–848, 2017
2017
Earlier work this paper cites.
H. Zhao, J. Shi, X. Qi, X. Wang, and J. Jia, “Pyramid scene parsing network,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2017, pp. 2881–2890
2017
Earlier work this paper cites.
S. Xie, R. Girshick, P. Dollár, Z. Tu, and K. He, “Aggregated residual transformations for deep neural networks,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2017, pp. 1492–1500
2017
Earlier work this paper cites.
B. Zhou, H. Zhao, X. Puig, S. Fidler, A. Barriuso, and A. Torralba, “Scene parsing through ade20k dataset,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2017, pp. 633–641
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
G. Huang, Z. Liu, L. Van Der Maaten, and K. Q. Weinberger, “Densely connected convolutional networks,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2017, pp. 4700–4708
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Adv. Neural Inform. Process. Syst. , 2017, pp. 5998–6008
2017
Earlier work this paper cites.
M. Yang, K. Yu, C. Zhang, Z. Li, and K. Yang, “DenseASPP for semantic segmentation in street scenes,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2018, pp. 3684–3692
2018
Earlier work this paper cites.
H. Zhao, Y. Zhang, S. Liu, J. Shi, C. C. Loy, D. Lin, and J. Jia, “PSANet: Point-wise spatial attention network for scene parsing,” in Eur. Conf. Comput. Vis. , 2018, pp. 267–283
2018
Earlier work this paper cites.
H. Zhang, K. Dana, J. Shi, Z. Zhang, X. Wang, A. Tyagi, and A. Agrawal, “Context encoding for semantic segmentation,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2018, pp. 7151–7160
2018
Earlier work this paper cites.
C. Yu, J. Wang, C. Peng, C. Gao, G. Yu, and N. Sang, “Learning a discriminative feature network for semantic segmentation,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2018, pp. 1857–1866
2018
Earlier work this paper cites.
H. Caesar, J. Uijlings, and V. Ferrari, “COCO-Stuff: Thing and stuff classes in context,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2018, pp. 1209–1218
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
J. Hu, L. Shen, and G. Sun, “Squeeze-and-excitation networks,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2018, pp. 7132–7141
2018
Earlier work this paper cites.
I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” in Int. Conf. Learn. Represent. , 2018
2018
Earlier work this paper cites.
T. Xiao, Y. Liu, B. Zhou, Y. Jiang, and J. Sun, “Unified perceptual parsing for scene understanding,” in Eur. Conf. Comput. Vis. , 2018, pp. 418–434
2018
Earlier work this paper cites.
J. Fu, J. Liu, H. Tian, Y. Li, Y. Bao, Z. Fang, and H. Lu, “Dual attention network for scene segmentation,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2019, pp. 3146–3154
2019
Earlier work this paper cites.
Z. Zhu, M. Xu, S. Bai, T. Huang, and X. Bai, “Asymmetric non-local neural networks for semantic segmentation,” in Int. Conf. Comput. Vis. , 2019, pp. 593–602
2019
Earlier work this paper cites.
T. Takikawa, D. Acuna, V. Jampani, and S. Fidler, “Gated-scnn: Gated shape cnns for semantic segmentation,” in Int. Conf. Comput. Vis. , 2019, pp. 5229–5238
2019
Earlier work this paper cites.
H. Ding, X. Jiang, A. Q. Liu, N. M. Thalmann, and G. Wang, “Boundary-aware feature propagation for scene segmentation,” in Int. Conf. Comput. Vis. , 2019, pp. 6819–6829
2019
Earlier work this paper cites.
S.-H. Gao, M.-M. Cheng, K. Zhao, X.-Y. Zhang, M.-H. Yang, and P. Torr, “Res2Net: A new multi-scale backbone architecture,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 43, no. 2, pp. 652–662, 2019
2019
Cited alongside, same era.
X. Li, W. Wang, X. Hu, and J. Yang, “Selective kernel networks,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2019, pp. 510–519
2019
Cited alongside, same era.
2020
Cited alongside, same era.
Y. Yuan, X. Chen, and J. Wang, “Object-contextual representations for semantic segmentation,” in Eur. Conf. Comput. Vis. Springer, 2020, pp. 173–190
2020
Cited alongside, same era.
Q. Yu, H. Wang, S. Qiao, M. Collins, Y. Zhu, H. Adam, A. Yuille, and L.-C. Chen, “K-means mask transformer,” in Eur. Conf. Comput. Vis. Springer, 2022, pp. 288–307
2022
Later among the works it cites.
H. Zhang, C. Wu, Z. Zhang, Y. Zhu, H. Lin, Z. Zhang, Y. Sun, T. He, J. Mueller, R. Manmatha et al. , “ResNeSt: Split-attention networks,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2022, pp. 2736–2746
2022
Later among the works it cites.
Z. Liu, H. Mao, C.-Y. Wu, C. Feichtenhofer, T. Darrell, and S. Xie, “A convnet for the 2020s,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2022, pp. 11 976–11 986
2022
Later among the works it cites.
X. Ding, X. Zhang, J. Han, and G. Ding, “Scaling up your kernels to 31x31: Revisiting large kernel design in CNNs,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2022, pp. 11 963–11 975
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
Y. Yuan, J. Xie, X. Chen, and J. Wang, “Segfix: Model-agnostic boundary refinement for segmentation,” in Eur. Conf. Comput. Vis. Springer, 2020, pp. 489–506
2020
Cited alongside, same era.
M. Zhen, J. Wang, L. Zhou, S. Li, T. Shen, J. Shang, T. Fang, and L. Quan, “Joint semantic segmentation and boundary detection using iterative pyramid contexts,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2020, pp. 13 666–13 675
2020
Cited alongside, same era.
J. Wang, K. Sun, T. Cheng, B. Jiang, C. Deng, Y. Zhao, D. Liu, Y. Mu, M. Tan, X. Wang et al. , “Deep high-resolution representation learning for visual recognition,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 43, no. 10, pp. 3349–3364, 2020
2020
Cited alongside, same era.
I. Radosavovic, R. P. Kosaraju, R. Girshick, K. He, and P. Dollár, “Designing network design spaces,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2020, pp. 10 428–10 436
2020
Cited alongside, same era.
E. Xie, W. Wang, Z. Yu, A. Anandkumar, J. M. Alvarez, and P. Luo, “Segformer: Simple and efficient design for semantic segmentation with transformers,” Adv. Neural Inform. Process. Syst. , vol. 34, pp. 12 077–12 090, 2021
2021
Cited alongside, same era.
Y. Yuan, R. Fu, L. Huang, W. Lin, C. Zhang, X. Chen, and J. Wang, “Hrformer: High-resolution vision transformer for dense predict,” Adv. Neural Inform. Process. Syst. , vol. 34, pp. 7281–7293, 2021
2021
Cited alongside, same era.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly et al. , “An image is worth 16x16 words: Transformers for image recognition at scale,” in Int. Conf. Learn. Represent. , 2021
2021
Cited alongside, same era.
2022
Later among the works it cites.
D. Zhou, Z. Yu, E. Xie, C. Xiao, A. Anandkumar, J. Feng, and J. M. Alvarez, “Understanding the robustness in vision transformers,” in Int. Conf. Mach. Learn. PMLR, 2022, pp. 27 378–27 394
2022
Later among the works it cites.
W. Yu, M. Luo, P. Zhou, C. Si, Y. Zhou, X. Wang, J. Feng, and S. Yan, “Metaformer is actually what you need for vision,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2022, pp. 10 819–10 829
2022
Later among the works it cites.
Z. Xia, X. Pan, S. Song, L. E. Li, and G. Huang, “Vision transformer with deformable attention,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2022, pp. 4794–4803
2022
Later among the works it cites.
M.-H. Guo, T.-X. Xu, J.-J. Liu, Z.-N. Liu, P.-T. Jiang, T.-J. Mu, S.-H. Zhang, R. R. Martin, M.-M. Cheng, and S.-M. Hu, “Attention mechanisms in computer vision: A survey,” Computational visual media , vol. 8, no. 3, pp. 331–368, 2022
2022
Later among the works it cites.
B. Cheng, I. Misra, A. G. Schwing, A. Kirillov, and R. Girdhar, “Masked-attention mask transformer for universal image segmentation,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2022, pp. 1290–1299
2022
Later among the works it cites.
M. Ding, B. Xiao, N. Codella, P. Luo, J. Wang, and L. Yuan, “Davit: Dual attention vision transformers,” in Eur. Conf. Comput. Vis. Springer, 2022, pp. 74–92
2022
Later among the works it cites.
Y. Li, C.-Y. Wu, H. Fan, K. Mangalam, B. Xiong, J. Malik, and C. Feichtenhofer, “MViTv2: Improved multiscale vision transformers for classification and detection,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2022, pp. 4804–4814
2022
Later among the works it cites.
J. Yang, C. Li, and J. Gao, “Focal modulation networks,” arXiv preprint arXiv:2203.11926 , 2022
2022
Later among the works it cites.
J. Mao, J. Huang, A. Toshev, O. Camburu, A. L. Yuille, and K. Murphy, “Generation and comprehension of unambiguous object descriptions,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2016, pp. 11–20
2022
Later among the works it cites.
Z. Huang, X. Wang, Y. Wei, L. Huang, H. Shi, W. Liu, and T. S. Huang, “CCNet: Criss-cross attention for semantic segmentation,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 45, no. 06, pp. 6896–6908, 2023
2023
Closest in time.
Y.-H. Wu, Y. Liu, X. Zhan, and M.-M. Cheng, “P2T: Pyramid pooling transformer for scene understanding,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 45, no. 11, pp. 12 760–12 771, 2023
2023
Closest in time.
F. Li, H. Zhang, H. Xu, S. Liu, L. Zhang, L. M. Ni, and H.-Y. Shum, “Mask DINO: Towards a unified transformer-based framework for object detection and segmentation,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2023, pp. 3041–3050
2023
Closest in time.
J.-h. Shim, H. Yu, K. Kong, and S.-J. Kang, “Feedformer: Revisiting transformer decoder for efficient semantic segmentation,” in AAAI Conf. Artif. Intell. , vol. 37, no. 2, 2023, pp. 2263–2271
2023
Closest in time.
Y. Fang, W. Wang, B. Xie, Q. Sun, L. Wu, X. Wang, T. Huang, X. Wang, and Y. Cao, “EVA: Exploring the limits of masked visual representation learning at scale,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2023, pp. 19 358–19 369
2023
Closest in time.
J. Jain, J. Li, M. T. Chiu, A. Hassani, N. Orlov, and H. Shi, “Oneformer: One transformer to rule universal image segmentation,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2023, pp. 2989–2998
2023
Closest in time.
2023
Closest in time.
X. Chu, Z. Tian, B. Zhang, X. Wang, X. Wei, H. Xia, and C. Shen, “Conditional positional encodings for vision transformers,” in Int. Conf. Learn. Represent. , 2023
2023
Closest in time.
W. Wang, J. Dai, Z. Chen, Z. Huang, Z. Li, X. Zhu, X. Hu, T. Lu, L. Lu, H. Li et al. , “Internimage: Exploring large-scale vision foundation models with deformable convolutions,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2023, pp. 14 408–14 419
2023
Closest in time.
J. Jain, A. Singh, N. Orlov, Z. Huang, J. Li, S. Walton, and H. Shi, “Semask: Semantically masked transformers for semantic segmentation,” in International Conference on Computer Vision Workshops , 2023, pp. 752–761
2023
Closest in time.
H. Liu, C. Li, Q. Wu, and Y. J. Lee, “Visual instruction tuning,” in Adv. Neural Inform. Process. Syst. , 2023, pp. 34 892–34 916
2023
Closest in time.
Y. Wang, Y. Li, J. H. Elder, R. Wu, and H. Lu, “Class-conditional domain adaptation for semantic segmentation,” Computational Visual Media , vol. 10, no. 5, pp. 1013–1030, 2024
2024
Closest in time.
D. Liang, Y. Sun, Y. Du, S. Chen, and S.-J. Huang, “Relative difficulty distillation for semantic segmentation,” Science China Information Sciences , vol. 67, no. 9, p. 192105, 2024
2024
Closest in time.
Y. Liu, Y.-H. Wu, G. Sun, L. Zhang, A. Chhatkuli, and L. Van Gool, “Vision transformers with hierarchical attention,” Machine Intelligence Research , vol. 21, no. 4, pp. 670–683, 2024
2024
Closest in time.
L. Zhu, B. Liao, Q. Zhang, X. Wang, W. Liu, and X. Wang, “Vision mamba: Efficient visual representation learning with bidirectional state space model,” in Int. Conf. Mach. Learn. , 2024, pp. 62 429–62 442
2024
Closest in time.
A. Hatamizadeh, G. Heinrich, H. Yin, A. Tao, J. M. Alvarez, J. Kautz, and P. Molchanov, “FasterViT: Fast vision transformers with hierarchical attention,” in Int. Conf. Learn. Represent. , 2024
2024
Closest in time.
X. Lai, Z. Tian, Y. Chen, Y. Li, Y. Yuan, S. Liu, and J. Jia, “LISA: Reasoning segmentation via large language model,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2024, pp. 9579–9589
2024
Closest in time.
J. Li, Y. Huang, M. Wu, B. Zhang, X. Ji, and C. Zhang, “CLIP-SP: Vision-language model with adaptive prompting for scene parsing,” Computational Visual Media , vol. 10, no. 4, pp. 741–752, 2024
2024
Closest in time.