Fetching the paper…
Reading the bibliography…
The perception capability of robotic systems relies on the richness of the dataset.
S. Antol et al. , “VQA: Visual question answering,” in Proc. ICCV , 2015, pp. 2425–2433
2015
Earlier work this paper cites.
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan, “Show and tell: A neural image caption generator,” in Proc. CVPR , 2015, pp. 3156–3164
2015
Earlier work this paper cites.
O. Ronneberger, P. Fischer, and T. Brox, “U-Net: Convolutional networks for biomedical image segmentation,” in Proc. MICCAI , 2015, pp. 234–241
2015
Earlier work this paper cites.
V. Badrinarayanan, A. Kendall, and R. Cipolla, “SegNet: A deep convolutional encoder-decoder architecture for image segmentation,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 39, no. 12, pp. 2481–2495, 2017
2017
Earlier work this paper cites.
Q. Ha, K. Watanabe, T. Karasawa, Y. Ushiku, and T. Harada, “MFNet: Towards real-time semantic segmentation for autonomous vehicles with multi-spectral scenes,” in Proc. IROS , 2017, pp. 5108–5115
2017
Earlier work this paper cites.
T.-Y. Lin, P. Dollár, R. Girshick, K. He, B. Hariharan, and S. Belongie, “Feature pyramid networks for object detection,” in Proc. CVPR , 2017, pp. 936–944
2017
Earlier work this paper cites.
H. Zhao, J. Shi, X. Qi, X. Wang, and J. Jia, “Pyramid scene parsing network,” in Proc. CVPR , 2017, pp. 6230–6239
2017
Earlier work this paper cites.
L. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille, “DeepLab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected CRFs,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 40, no. 4, pp. 834–848, 2018
2018
Earlier work this paper cites.
H. Tan and M. Bansal, “LXMERT: Learning cross-modality encoder representations from transformers,” in Proc. EMNLP-IJCNLP , 2019, pp. 5099–5110
2019
Earlier work this paper cites.
S. S. Shivakumar, N. Rodrigues, A. Zhou, I. D. Miller, V. Kumar, and C. J. Taylor, “PST900: RGB-thermal calibration, dataset and segmentation network,” in Proc. ICRA , 2020, pp. 9441–9447
2020
Earlier work this paper cites.
U. Michieli, E. Borsato, L. Rossi, and P. Zanuttigh, “GMNet: Graph matching network for large scale part semantic segmentation in the wild,” in Proc. ECCV , vol. 12353, 2020, pp. 397–414
2020
Earlier work this paper cites.
Z. Zhao, S. Xu, C. Zhang, J. Liu, P. Li, and J. Zhang, “DIDFuse: Deep image decomposition for infrared and visible image fusion,” in Proc. IJCAI , 2020, pp. 970–976
2020
Earlier work this paper cites.
E. Xie, W. Wang, Z. Yu, A. Anandkumar, J. M. Alvarez, and P. Luo, “SegFormer: Simple and efficient design for semantic segmentation with transformers,” in Proc. NeurIPS , vol. 34, 2021, pp. 12 077–12 090
2021
Earlier work this paper cites.
F. Deng et al. , “FEANet: Feature-enhanced attention network for RGB-thermal real-time semantic segmentation,” in Proc. IROS , 2021, pp. 4467–4473
2021
Earlier work this paper cites.
A. Radford et al. , “Learning transferable visual models from natural language supervision,” in Proc. ICML , 2021, pp. 8748–8763
2021
Earlier work this paper cites.
Y. Wang, X. Chen, L. Cao, W. Huang, F. Sun, and Y. Wang, “Multimodal token fusion for vision transformers,” in Proc. CVPR , 2022, pp. 12 176–12 185
2022
Earlier work this paper cites.
X. Lan, X. Gu, and X. Gu, “MMNet: Multi-modal multi-stage network for RGB-T image semantic segmentation,” Applied Intelligence , vol. 52, no. 5, pp. 5817–5829, 2022
2022
Cited alongside, same era.
Z. Liu, H. Mao, C.-Y. Wu, C. Feichtenhofer, T. Darrell, and S. Xie, “A ConvNet for the 2020s,” in Proc. CVPR , 2022, pp. 11 966–11 976
2022
Cited alongside, same era.
B. Li, K. Q. Weinberger, S. Belongie, V. Koltun, and R. Ranftl, “Language-driven semantic segmentation,” in Proc. ICLR , 2022
2022
Cited alongside, same era.
Z. Huang, J. Liu, X. Fan, R. Liu, W. Zhong, and Z. Luo, “ReCoNet: Recurrent correction network for fast and efficient multi-modality image fusion,” in Proc. ECCV , vol. 13678, 2022, pp. 539–555
2022
Cited alongside, same era.
H. Xu, J. Ma, J. Jiang, X. Guo, and H. Ling, “U2Fusion: A unified unsupervised image fusion network,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 44, no. 1, pp. 502–518, 2022
2024
Later among the works it cites.
D. Jia et al. , “GeminiFusion: Efficient pixel-wise multimodal fusion for vision transformer,” in Proc. ICML , 2024
2024
Later among the works it cites.
D. Peng and W. Kameyama, “Simple and efficient vision backbone adapter for image semantic segmentation,” in Proc. ACML , 2024, pp. 1071–1086
2024
Later among the works it cites.
B. Zhu et al. , “LanguageBind: Extending video-language pretraining to N-modality by language-based semantic alignment,” in Proc. ICLR , 2024
2024
Later among the works it cites.
H. Yuan, X. Li, C. Zhou, Y. Li, K. Chen, and C. C. Loy, “Open-vocabulary SAM: Segment and recognize twenty-thousand classes interactively,” in Proc. ECCV , vol. 15101, 2024, pp. 419–437
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2022
Cited alongside, same era.
A. Kirillov et al. , “Segment anything,” in Proc. ICCV , 2023, pp. 3992–4003
2023
Cited alongside, same era.
2023
Cited alongside, same era.
T. Chen et al. , “SAM-adapter: Adapting segment anything in underperformed scenes,” in Proc. ICCVW , 2023, pp. 3359–3367
2023
Cited alongside, same era.
J. Zhang et al. , “Delivering arbitrary-modal semantic segmentation,” in Proc. CVPR , 2023, pp. 1136–1147
2023
Cited alongside, same era.
J. Zhang, H. Liu, K. Yang, X. Hu, R. Liu, and R. Stiefelhagen, “CMX: Cross-modal fusion for RGB-X semantic segmentation with transformers,” IEEE Transactions on Intelligent Transportation Systems , vol. 24, no. 12, pp. 14 679–14 694, 2023
2023
Cited alongside, same era.
J. Liu et al. , “Multi-interactive feature learning and a full-time multi-modality benchmark for image fusion and segmentation,” in Proc. ICCV , 2023, pp. 8081–8090
2023
Cited alongside, same era.
C. Ryali et al. , “Hiera: A hierarchical vision transformer without the bells-and-whistles,” in Proc. ICML , vol. 202, 2023, pp. 29 441–29 454
2023
Cited alongside, same era.
2024
Later among the works it cites.
M. K. Reza, A. Prater-Bennette, and M. S. Asif, “MMSFormer: Multimodal transformer for material and semantic segmentation,” IEEE Open Journal of Signal Processing , 2024
2024
Later among the works it cites.
S. Dong, W. Zhou, C. Xu, and W. Yan, “EGFNet: Edge-aware guidance fusion network for RGB-thermal urban scene parsing,” IEEE Transactions on Intelligent Transportation Systems , vol. 25, no. 1, pp. 657–669, 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
H. Yuan, X. Li, C. Zhou, Y. Li, K. Chen, and C. C. Loy, “Open-vocabulary SAM: Segment and recognize twenty-thousand classes interactively,” in Proc. ECCV , vol. 15101, 2024, pp. 419–437
2024
Later among the works it cites.
B. Xie, J. Cao, J. Xie, F. S. Khan, and Y. Pang, “SED: A simple encoder-decoder for open-vocabulary semantic segmentation,” in Proc. CVPR , 2024, pp. 3426–3436
2024
Later among the works it cites.
T. Shao, Z. Tian, H. Zhao, and J. Su, “Explore the potential of CLIP for training-free open vocabulary semantic segmentation,” in Proc. ECCV , vol. 15144, 2024, pp. 139–156
2024
Later among the works it cites.
W. Zhou, S. Dong, M. Fang, and L. Yu, “CACFNet: Cross-modal attention cascaded fusion network for RGB-T urban scene parsing,” IEEE Transactions on Intelligent Vehicles , vol. 9, no. 1, pp. 1919–1929, 2024
2024
Later among the works it cites.
S. Dong, Y. Feng, Q. Yang, Y. Huang, D. Liu, and H. Fan, “Efficient multimodal semantic segmentation via dual-prompt learning,” in Proc. IROS , 2024, pp. 14 196–14 203
2024
Later among the works it cites.
U. Shin, K. Lee, I. S. Kweon, and J. Oh, “Complementary random masking for RGB-thermal semantic segmentation,” in Proc. ICRA , 2024, pp. 11 110–11 117
2024
Later among the works it cites.