Fetching the paper…
Reading the bibliography…
This paper investigates the capability of plain Vision Transformers (ViTs) for semantic segmentation using the encoder-decoder framework and introduces \textbf{SegViTv2}.
G. Lin, A. Milan, C. Shen, and I. Reid, “RefineNet: Multi-path refinement networks for high-resolution semantic segmentation,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn. , 2017, pp. 1925–1934
1934
Earlier work this paper cites.
R. M. French, “Catastrophic forgetting in connectionist networks,” Trends in cognitive sciences , vol. 3, no. 4, pp. 128–135, 1999
1999
Earlier work this paper cites.
R. Mottaghi, X. Chen, X. Liu, N.-G. Cho, S.-W. Lee, S. Fidler, R. Urtasun, and A. Yuille, “The role of context for object detection and semantic segmentation in the wild,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn. , 2014, pp. 891–898
2014
Earlier work this paper cites.
J. Long, E. Shelhamer, and T. Darrell, “Fully convolutional networks for semantic segmentation,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn. , 2015, pp. 3431–3440
2015
Earlier work this paper cites.
O. Ronneberger, P. Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” in Medical Image Computing and Computer-Assisted Intervention . Springer, 2015, pp. 234–241
2015
Earlier work this paper cites.
Z. Chen and B. Liu, Lifelong Machine Learning . Synthesis Lectures on Artificial Intelligence and Machine Learning, 2016
2016
Earlier work this paper cites.
F. Milletari, N. Navab, and S.-A. Ahmadi, “V-net: Fully convolutional neural networks for volumetric medical image segmentation,” in 3DV . IEEE, 2016, pp. 565–571
2016
Earlier work this paper cites.
J. Kirkpatrick, R. Pascanu, N. Rabinowitz, J. Veness, G. Desjardins, A. A. Rusu, K. Milan, J. Quan, T. Ramalho, A. Grabska-Barwinska et al. , “Overcoming catastrophic forgetting in neural networks,” Proceedings of the national academy of sciences , vol. 114, no. 13, pp. 3521–3526, 2017
2017
Earlier work this paper cites.
H. Zhao, J. Shi, X. Qi, X. Wang, and J. Jia, “Pyramid scene parsing network,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn. , 2017
2017
Earlier work this paper cites.
T.-Y. Lin, P. Dollár, R. Girshick, K. He, B. Hariharan, and S. Belongie, “Feature pyramid networks for object detection,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn. , 2017, pp. 2117–2125
2017
Earlier work this paper cites.
T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár, “Focal loss for dense object detection,” in Proc. IEEE Int. Conf. Comp. Vis. , 2017, pp. 2980–2988
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Proc. Adv. Neural Inf. Process. Syst. , vol. 30, 2017
2017
Earlier work this paper cites.
B. Zhou, H. Zhao, X. Puig, S. Fidler, A. Barriuso, and A. Torralba, “Scene parsing through ade20k dataset,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn. , 2017, pp. 633–641
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
L.-C. Chen, Y. Zhu, G. Papandreou, F. Schroff, and H. Adam, “Encoder-decoder with atrous separable convolution for semantic image segmentation,” in Proc. Eur. Conf. Comp. Vis. , 2018, pp. 801–818
2018
Earlier work this paper cites.
T. Xiao, Y. Liu, B. Zhou, Y. Jiang, and J. Sun, “Unified perceptual parsing for scene understanding,” in Proc. Eur. Conf. Comp. Vis. , 2018, pp. 418–434
2018
Earlier work this paper cites.
Z. Li and D. Hoiem, “Learning without forgetting,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 40, pp. 2935–2947, 2018
2018
Earlier work this paper cites.
H. Caesar, J. Uijlings, and V. Ferrari, “Coco-stuff: Thing and stuff classes in context,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn. , 2018, pp. 1209–1218
2018
Earlier work this paper cites.
Z. Zhou, M. M. R. Siddiquee, N. Tajbakhsh, and J. Liang, “Unet++: A nested U-net architecture for medical image segmentation,” in Proc. Deep Learning in Medical Image Analysis Workshop , 2018, pp. 3–11
2018
Earlier work this paper cites.
H. Ding, X. Jiang, B. Shuai, A. Q. Liu, and G. Wang, “Context contrasted feature and gated multi-scale aggregation for scene segmentation,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn. , 2018, pp. 2393–2402
2018
Earlier work this paper cites.
H. Zhang, K. Dana, J. Shi, Z. Zhang, X. Wang, A. Tyagi, and A. Agrawal, “Context encoding for semantic segmentation,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn. , 2018, pp. 7151–7160
2018
Earlier work this paper cites.
2019
Earlier work this paper cites.
J. Fu, J. Liu, H. Tian, Y. Li, Y. Bao, Z. Fang, and H. Lu, “Dual attention network for scene segmentation,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn. , 2019, pp. 3146–3154
2019
Earlier work this paper cites.
X. Li, Z. Zhong, J. Wu, Y. Yang, Z. Lin, and H. Liu, “Expectation-maximization attention networks for semantic segmentation,” in Proc. IEEE Int. Conf. Comp. Vis. , 2019, pp. 9167–9176
2019
Earlier work this paper cites.
K. Sun, Y. Zhao, B. Jiang, T. Cheng, B. Xiao, D. Liu, Y. Mu, X. Wang, W. Liu, and J. Wang, “High-resolution representations for labeling pixels and regions,” 2019
2019
Earlier work this paper cites.
U. Michieli and P. Zanuttigh, “Incremental learning techniques for semantic segmentation,” in Proc. IEEE Int. Conf. Comp. Vis. Workshops , 2019, pp. 3205–3212
2019
Earlier work this paper cites.
J. Wang, K. Sun, T. Cheng, B. Jiang, C. Deng, Y. Zhao, D. Liu, Y. Mu, M. Tan, X. Wang et al. , “Deep high-resolution representation learning for visual recognition,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 43, no. 10, pp. 3349–3364, 2020
2020
Earlier work this paper cites.
Y. Yuan, X. Chen, and J. Wang, “Object-contextual representations for semantic segmentation,” in Proc. Eur. Conf. Comp. Vis. Springer, 2020, pp. 173–190
2020
Earlier work this paper cites.
F. Cermelli, M. Mancini, S. R. Bulò, E. Ricci, and B. Caputo, “Modeling the background for incremental learning in semantic segmentation,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn. , 2020, pp. 9230–9239
2020
Earlier work this paper cites.
A. Douillard, M. Cord, C. Ollion, T. Robert, and E. Valle, “Podnet: Pooled outputs distillation for small-tasks incremental learning,” in Proc. Eur. Conf. Comp. Vis. Springer, 2020, pp. 86–102
2020
Earlier work this paper cites.
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko, “End-to-end object detection with transformers,” in Proc. Eur. Conf. Comp. Vis. Springer, 2020, pp. 213–229
2020
Earlier work this paper cites.
MMSegmentation, “MMSegmentation: OpenMMLab semantic segmentation toolbox and benchmark,” https://github.com/open-mmlab/mmsegmentation , 2020
2020
Cited alongside, same era.
X. Li, Y. Yang, Q. Zhao, T. Shen, Z. Lin, and H. Liu, “Spatial pyramid based graph reasoning for semantic segmentation,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn. , 2020, pp. 8950–8959
2020
Cited alongside, same era.
T. Wu, Y. Lu, Y. Zhu, C. Zhang, M. Wu, Z. Ma, and G. Guo, “Ginet: Graph interaction network for scene parsing,” in Proc. Eur. Conf. Comp. Vis. Springer, 2020, pp. 34–51
2020
Cited alongside, same era.
W. Chen, X. Zhu, R. Sun, J. He, R. Li, X. Shen, and B. Yu, “Tensor low-rank reconstruction for semantic segmentation,” in Proc. Eur. Conf. Comp. Vis. Springer, 2020, pp. 52–69
2020
Cited alongside, same era.
A. Douillard, Y. Chen, A. Dapogny, and M. Cord, “Plop: Learning without forgetting for continual semantic segmentation,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn. , 2021
2021
Later among the works it cites.
2022
Later among the works it cites.
H. Bao, L. Dong, S. Piao, and F. Wei, “BEiT: BERT pre-training of image transformers,” in International Conference on Learning Representations , 2022. [Online]. Available: https://openreview.net/forum?id=p-BhZSz59o4
2022
Later among the works it cites.
Z. Peng, L. Dong, H. Bao, Q. Ye, and F. Wei, “BEiT v2: Masked image modeling with vector-quantized visual tokenizers,” 2022
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
J. Liu, J. He, J. Zhang, J. Ren, and H. Li, “EfficientFCN: Holistically-guided decoding for semantic segmentation,” in Proc. Eur. Conf. Comp. Vis. , 2020
2020
Cited alongside, same era.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly et al. , “An image is worth 16x16 words: Transformers for image recognition at scale,” Proc. Int. Conf. Learn. Repren. , 2021
2021
Cited alongside, same era.
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in Proc. IEEE Int. Conf. Comp. Vis. , 2021, pp. 10 012–10 022
2021
Cited alongside, same era.
W. Wang, E. Xie, X. Li, D.-P. Fan, K. Song, D. Liang, T. Lu, P. Luo, and L. Shao, “Pyramid vision transformer: A versatile backbone for dense prediction without convolutions,” in Proc. IEEE Int. Conf. Comp. Vis. , 2021, pp. 568–578
2021
Cited alongside, same era.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in International conference on machine learning . PMLR, 2021, pp. 8748–8763
2021
Cited alongside, same era.
R. Ranftl, A. Bochkovskiy, and V. Koltun, “Vision transformers for dense prediction,” in Proc. IEEE Int. Conf. Comp. Vis. , 2021, pp. 12 179–12 188
2021
Cited alongside, same era.
S. Zheng, J. Lu, H. Zhao, X. Zhu, Z. Luo, Y. Wang, Y. Fu, J. Feng, T. Xiang, P. H. Torr et al. , “Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn. , 2021, pp. 6881–6890
2021
Cited alongside, same era.
2022
Later among the works it cites.
H. Lu, N. Fei, Y. Huo, Y. Gao, Z. Lu, and J.-R. Wen, “Cots: Collaborative two-stream vision-language pre-training model for cross-modal retrieval,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn. , 2022, pp. 15 692–15 701
2022
Later among the works it cites.
2022
Later among the works it cites.
Z. Wang, Z. Zhang, C.-Y. Lee, H. Zhang, R. Sun, X. Ren, G. Su, V. Perot, J. Dy, and T. Pfister, “Learning to prompt for continual learning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 139–149
2022
Later among the works it cites.
Z. Wang, Z. Zhang, S. Ebrahimi, R. Sun, H. Zhang, C.-Y. Lee, X. Ren, G. Su, V. Perot, J. Dy et al. , “Dualprompt: Complementary prompting for rehearsal-free continual learning,” in Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part XXVI . Springer, 2022, pp. 631–648
2022
Later among the works it cites.
M. H. Phan, S. L. Phung, L. Tran-Thanh, A. Bouzerdoum et al. , “Class similarity weighted knowledge distillation for continual semantic segmentation,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn. , 2022, pp. 16 866–16 875
2022
Later among the works it cites.
O. Ostapenko, T. Lesort, P. Rodríguez, M. R. Arefin, A. Douillard, I. Rish, and L. Charlin, “Continual learning with foundation models: An empirical study of latent replay,” in Conference on Lifelong Learning Agents . PMLR, 2022, pp. 60–91
2022
Later among the works it cites.
V. V. Ramasesh, A. Lewkowycz, and E. Dyer, “Effect of scale on catastrophic forgetting in neural networks,” in Proc. Int. Conf. Learn. Repren. , 2022
2022
Later among the works it cites.
T. Wu, M. Caccia, Z. Li, Y.-F. Li, G. Qi, and G. Haffari, “Pretrained language model in continual learning: A comparative study,” in Proc. Int. Conf. Learn. Repren. , 2022
2022
Later among the works it cites.
C.-B. Zhang, J.-W. Xiao, X. Liu, Y.-C. Chen, and M.-M. Cheng, “Representation compensation networks for continual semantic segmentation,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn. , 2022, pp. 7053–7064
2022
Later among the works it cites.
X. Dong, J. Bao, D. Chen, W. Zhang, N. Yu, L. Yuan, D. Chen, and B. Guo, “Cswin transformer: A general vision transformer backbone with cross-shaped windows,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn. , 2022, pp. 12 124–12 134
2022
Later among the works it cites.
2022
Later among the works it cites.
B. Cheng, I. Misra, A. G. Schwing, A. Kirillov, and R. Girdhar, “Masked-attention mask transformer for universal image segmentation,” 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
K. He, X. Chen, S. Xie, Y. Li, P. Dollár, and R. Girshick, “Masked autoencoders are scalable vision learners,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 16 000–16 009
2022
Later among the works it cites.
2022
Later among the works it cites.
H. Touvron, M. Cord, and H. Jégou, “Deit iii: Revenge of the vit,” in Proc. Eur. Conf. Comp. Vis. Springer, 2022, pp. 516–533
2022
Later among the works it cites.
2022
Later among the works it cites.
Y.-H. Wu, Y. Liu, X. Zhan, and M.-M. Cheng, “P2t: Pyramid pooling transformer for scene understanding,” IEEE Trans. Pattern Anal. Mach. Intell. , 2022
2022
Later among the works it cites.
J. Zhou, C. Wei, H. Wang, W. Shen, C. Xie, A. Yuille, and T. Kong, “ibot: Image bert pre-training with online tokenizer,” Proc. Int. Conf. Learn. Repren. , 2022
2022
Later among the works it cites.
M. Kang, J. Park, and B. Han, “Class-incremental learning by knowledge distillation with adaptive feature consolidation,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn. , 2022, pp. 16 071–16 080
2022
Later among the works it cites.
A. Douillard, A. Ramé, G. Couairon, and M. Cord, “Dytox: Transformers for continual learning with dynamic token expansion,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn. , 2022, pp. 9285–9295
2022
Later among the works it cites.
Z. Wang, L. Liu, Y. Duan, Y. Kong, and D. Tao, “Continual learning with lifelong vision transformer,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn. , 2022, pp. 171–181
2022
Later among the works it cites.
Z. Wang, L. Liu, Y. Kong, J. Guo, and D. Tao, “Online continual learning with contrastive vision transformer,” in Proc. Eur. Conf. Comp. Vis. Springer, 2022, pp. 631–650
2022
Later among the works it cites.
Z. Kong, P. Dong, X. Ma, X. Meng, W. Niu, M. Sun, X. Shen, G. Yuan, B. Ren, H. Tang et al. , “Spvit: Enabling faster vision transformers via latency-aware soft token pruning,” in Proc. Eur. Conf. Comp. Vis. Springer, 2022, pp. 620–640
2022
Later among the works it cites.
B. Zhang, Z. Tian, Q. Tang, X. Chu, X. Wei, C. Shen, and Y. Liu, “Segvit: Semantic segmentation with plain vision transformers,” in Proc. Adv. Neural Inf. Process. Syst. , 2022
2022
Later among the works it cites.
F. Lin, Z. Liang, J. He, M. Zheng, S. Tian, and K. Chen, “Structtoken: Rethinking semantic segmentation with structural prior,” 2022
2022
Later among the works it cites.