Fetching the paper…
Reading the bibliography…
Multi-modal semantic segmentation significantly enhances AI agents' perception and scene understanding, especially under adverse conditions like low-light or overexposed environments.
Indoor segmentation and support inference from rgbd images
Nathan Silberman, Derek Hoiem, Pushmeet Kohli, and Rob Fergus · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al · 2015
Earlier work this paper cites.
Sun rgb-d: A rgb-d scene understanding benchmark suite
Shuran Song, Samuel P Lichtenberg, and Jianxiong Xiao · 2015
Earlier work this paper cites.
Fusenet: Incorporating depth into semantic segmentation via fusion-based cnn architecture
Caner Hazirbas, Lingni Ma, Csaba Domokos, and Daniel Cremers · 2016
Earlier work this paper cites.
Mfnet: Towards real-time semantic segmentation for autonomous vehicles with multi-spectral scenes
Qishen Ha, Kohei Watanabe, Takumi Karasawa, Yoshitaka Ushiku, and Tatsuya Harada · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
Full-resolution residual networks for semantic segmentation in street scenes
Tobias Pohlen, Alexander Hermans, Markus Mathias, and Bastian Leibe · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
A review of semantic segmentation using deep neural networks
Yanming Guo, Yu Liu, Theodoros Georgiou, and Michael S Lew · 2018
Earlier work this paper cites.
Unified perceptual parsing for scene understanding
Tete Xiao, Yingcheng Liu, Bolei Zhou, Yuning Jiang, and Jian Sun · 2018
Earlier work this paper cites.
Bisenet: Bilateral segmentation network for real-time semantic segmentation
Changqian Yu, Jingbo Wang, Chao Peng, Changxin Gao, Gang Yu, and Nong Sang · 2018
Earlier work this paper cites.
Learning a discriminative feature network for semantic segmentation
Changqian Yu, Jingbo Wang, Chao Peng, Changxin Gao, Gang Yu, and Nong Sang · 2018
Earlier work this paper cites.
Acnet: Attention based network to exploit complementary features for rgbd semantic segmentation
Xinxin Hu, Kailun Yang, Lei Fei, and Kaiwei Wang · 2019
Earlier work this paper cites.
Rtfnet: Rgb-thermal fusion network for semantic segmentation of urban scenes
Yuxiang Sun, Weixun Zuo, and Ming Liu · 2019
Earlier work this paper cites.
Bi-directional cross-modality feature propagation with separation-and-aggregation gate for rgb-d semantic segmentation
Xiaokang Chen, Kwan-Yee Lin, Jingbo Wang, Wayne Wu, Chen Qian, Hongsheng Li, and Gang Zeng · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Earlier work this paper cites.
Deep multi-modal object detection and semantic segmentation for autonomous driving: Datasets, methods, and challenges
Di Feng, Christian Haase-Schütz, Lars Rosenbaum, Heinz Hertlein, Claudius Glaeser, Fabian Timm, Werner Wiesbeck, and Klaus Dietmayer · 2020
Earlier work this paper cites.
Pst900: Rgb-thermal calibration, dataset and segmentation network
Shreyas S Shivakumar, Neil Rodrigues, Alex Zhou, Ian D Miller, Vijay Kumar, and Camillo J Taylor · 2020
Earlier work this paper cites.
Fuseseg: Semantic segmentation of urban scenes based on rgb and thermal data fusion
Yuxiang Sun, Weixun Zuo, Peng Yun, Hengli Wang, and Ming Liu · 2020
Earlier work this paper cites.
Deep multimodal fusion by channel exchanging
Yikai Wang, Wenbing Huang, Fuchun Sun, Tingyang Xu, Yu Rong, and Junzhou Huang · 2020
Earlier work this paper cites.
U2fusion: A unified unsupervised image fusion network
Han Xu, Jiayi Ma, Junjun Jiang, Xiaojie Guo, and Haibin Ling · 2020
Earlier work this paper cites.
Shapeconv: Shape-aware convolutional layer for indoor rgb-d semantic segmentation
Jinming Cao, Hanchao Leng, Dani Lischinski, Daniel Cohen-Or, Changhe Tu, and Yangyan Li · 2021
Earlier work this paper cites.
Crossvit: Cross-attention multi-scale vision transformer for image classification
Chun-Fu Richard Chen, Quanfu Fan, and Rameswar Panda · 2021
Earlier work this paper cites.
Spatial information guided convolution for real-time rgbd semantic segmentation
Lin-Zhuo Chen, Zheng Lin, Ziqin Wang, Yong-Liang Yang, and Ming-Ming Cheng · 2021
Earlier work this paper cites.
Feanet: Feature-enhanced attention network for rgb-thermal real-time semantic segmentation
Fuqin Deng, Hua Feng, Mingjian Liang, Hongmin Wang, Yong Yang, Yuan Gao, Junfeng Chen, Junjie Hu, Xiyue Guo, and Tin Lun Lam · 2021
Cited alongside, same era.
Efficiently modeling long sequences with structured state spaces
Albert Gu, Karan Goel, and Christopher Ré · 2021
Cited alongside, same era.
Combining recurrent, convolutional, and continuous-time models with linear state space layers
Albert Gu, Isys Johnson, Karan Goel, Khaled Saab, Tri Dao, Atri Rudra, and Christopher Ré · 2021
Cited alongside, same era.
Swin transformer: Hierarchical vision transformer using shifted windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo · 2021
Cited alongside, same era.
Efficient rgb-d semantic segmentation for indoor scene analysis
Daniel Seichter, Mona Köhler, Benjamin Lewandowski, Tim Wengefeld, and Horst-Michael Gross · 2021
Cited alongside, same era.
Explicit attention-enhanced fusion for rgb-thermal perception tasks
Mingjian Liang, Junjie Hu, Chenyu Bao, Hua Feng, Fuqin Deng, and Tin Lun Lam · 2023
Later among the works it cites.
Multi-interactive feature learning and a full-time multi-modality benchmark for image fusion and segmentation
Jinyuan Liu, Zhu Liu, Guanyao Wu, Long Ma, Risheng Liu, Wei Zhong, Zhongxuan Luo, and Xin Fan · 2023
Later among the works it cites.
Paif: Perception-aware infrared-visible image fusion for attack-tolerant semantic segmentation
Zhu Liu, Jinyuan Liu, Benzhuang Zhang, Long Ma, Xin Fan, and Risheng Liu · 2023
Later among the works it cites.
Transy-net: Learning fully transformer networks for change detection of remote sensing images
Tianyu Yan, Zifu Wan, Pingping Zhang, Gong Cheng, and Huchuan Lu · 2023
Later among the works it cites.
Pixel difference convolutional network for rgb-d semantic segmentation
Jun Yang, Lizhi Bai, Yaoru Sun, Chunqi Tian, Maoyu Mao, and Guorun Wang · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Segformer: Simple and efficient design for semantic segmentation with transformers
Enze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar, Jose M Alvarez, and Ping Luo · 2021
Cited alongside, same era.
Attention fusion network for multi-spectral semantic segmentation
Jiangtao Xu, Kaige Lu, and Han Wang · 2021
Cited alongside, same era.
Abmdrnet: Adaptive-weighted bi-directional modality difference reduction network for rgb-t semantic segmentation
Qiang Zhang, Shenlu Zhao, Yongjiang Luo, Dingwen Zhang, Nianchang Huang, and Jungong Han · 2021
Cited alongside, same era.
Deep multimodal fusion for semantic image segmentation: A survey
Yifei Zhang, Désiré Sidibé, Olivier Morel, and Fabrice Mériaudeau · 2021
Cited alongside, same era.
Gmnet: graded-feature multilabel-learning network for rgb-thermal urban scene semantic segmentation
Wujie Zhou, Jinfu Liu, Jingsheng Lei, Lu Yu, and Jenq-Neng Hwang · 2021
Cited alongside, same era.
Multimae: Multi-modal multi-task masked autoencoders
Roman Bachmann, David Mizrahi, Andrei Atanov, and Amir Zamir · 2022
Cited alongside, same era.
Omnivore: A single model for many visual modalities
Rohit Girdhar, Mannat Singh, Nikhila Ravi, Laurens van der Maaten, Armand Joulin, and Ishan Misra · 2022
Cited alongside, same era.
Swin3d: A pretrained transformer backbone for 3d indoor scene understanding, 2023
Yu-Qi Yang, Yu-Xiao Guo, Jian-Yu Xiong, Yang Liu, Hao Pan, Peng-Shuai Wang, Xin Tong, and Baining Guo · 2023
Later among the works it cites.
Taskprompter: Spatial-channel multi-task prompting for dense scene understanding
Hanrong Ye and Dan Xu · 2023
Later among the works it cites.
Robust hierarchical scene graph generation
Ce Zhang, Simon Stepputtis, Joseph Campbell, Katia Sycara, and Yaqi Xie · 2023
Later among the works it cites.
Delivering arbitrary-modal semantic segmentation
Jiaming Zhang, Ruiping Liu, Hao Shi, Kailun Yang, Simon Reiß, Kunyu Peng, Haodong Fu, Kaiwei Wang, and Rainer Stiefelhagen · 2023
Later among the works it cites.
Cacfnet: Cross-modal attention cascaded fusion network for rgb-t urban scene parsing
Wujie Zhou, Shaohua Dong, Meixin Fang, and Lu Yu · 2023
Later among the works it cites.
2023 low-power computer vision challenge (lpcvc) summary
Leo Chen, Benjamin Boardley, Ping Hu, Yiru Wang, Yifan Pu, Xin Jin, Yongqiang Yao, Ruihao Gong, Bo Li, Gao Huang, et al · 2024
Closest in time.
Understanding dark scenes by contrasting multi-modal observations
Xiaoyu Dong and Naoto Yokoya · 2024
Closest in time.
Mambair: A simple baseline for image restoration with state-space model, 2024
Hang Guo, Jinmin Li, Tao Dai, Zhihao Ouyang, Xudong Ren, and Shu-Tao Xia · 2024
Closest in time.
Pointmamba: A simple state space model for point cloud analysis, 2024
Dingkang Liang, Xin Zhou, Xinyu Wang, Xingkui Zhu, Wei Xu, Zhikang Zou, Xiaoqing Ye, and Xiang Bai · 2024
Closest in time.
Vmamba: Visual state space model
Yue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu, Lingxi Xie, Yaowei Wang, Qixiang Ye, and Yunfan Liu · 2024
Closest in time.
Context-aware interaction network for rgb-t semantic segmentation
Ying Lv, Zhi Liu, and Gongyang Li · 2024
Closest in time.
U-mamba: Enhancing long-range dependency for biomedical image segmentation
Jun Ma, Feifei Li, and Bo Wang · 2024
Closest in time.
Vm-unet: Vision mamba unet for medical image segmentation
Jiacheng Ruan and Suncheng Xiang · 2024
Closest in time.
Weak-mamba-unet: Visual mamba makes cnn and vit work better for scribble-based medical image segmentation, 2024
Ziyang Wang and Chao Ma · 2024
Closest in time.
Segmamba: Long-range sequential modeling mamba for 3d medical image segmentation
Zhaohu Xing, Tian Ye, Yijun Yang, Guang Liu, and Lei Zhu · 2024
Closest in time.
Vivim: a video vision mamba for medical video object segmentation, 2024
Yijun Yang, Zhaohu Xing, and Lei Zhu · 2024
Closest in time.
Hiker-sgg: Hierarchical knowledge enhanced robust scene graph generation
Ce Zhang, Simon Stepputtis, Joseph Campbell, Katia Sycara, and Yaqi Xie · 2024
Closest in time.
Vision mamba: Efficient visual representation learning with bidirectional state space model
Lianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang, Wenyu Liu, and Xinggang Wang · 2024
Closest in time.
Spectral-aware global fusion for rgb-thermal semantic segmentation
Ce Zhang, Zifu Wan, Simon Stepputtis, Katia Sycara, and Yaqi Xie · 2025
Closest in time.