Fetching the paper…
Reading the bibliography…
We present DFormer, a novel RGB-D pretraining framework to learn transferable representations for RGB-D segmentation tasks.
Bilateral attention network for rgb-d salient object detection
Zhao Zhang, Zheng Lin, Jun Xu, Wen-Da Jin, Shao-Ping Lu, and Deng-Ping Fan · 1961
Earlier work this paper cites.
Leveraging stereopsis for saliency analysis
Yuzhen Niu, Yujie Geng, Xueqing Li, and Feng Liu · 2012
Earlier work this paper cites.
Saliency filters: Contrast based filtering for salient region detection
Federico Perazzi, Philipp Krähenbühl, Yael Pritch, and Alexander Hornung · 2012
Earlier work this paper cites.
Indoor segmentation and support inference from RGBD images
Nathan Silberman, Derek Hoiem, Pushmeet Kohli, and Rob Fergus · 2012
Earlier work this paper cites.
Depth enhanced saliency detection method
Yupeng Cheng, Huazhu Fu, Xingxing Wei, Jiangjian Xiao, and Xiaochun Cao · 2014
Earlier work this paper cites.
Depth saliency based on anisotropic center-surround difference
Ran Ju, Ling Ge, Wenjing Geng, Tongwei Ren, and Gangshan Wu · 2014
Earlier work this paper cites.
How to evaluate foreground maps?
Ran Margolin, Lihi Zelnik-Manor, and Ayellet Tal · 2014
Earlier work this paper cites.
Rgbd salient object detection: A benchmark and algorithms
Houwen Peng, Bing Li, Weihua Xiong, Weiming Hu, and Rongrong Ji · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
ImageNet large scale visual recognition challenge
Olga Russakovsky et al · 2015
Earlier work this paper cites.
SUN RGB-D: A RGB-D scene understanding benchmark suite
Shuran Song, Samuel P. Lichtenberg, and Jianxiong Xiao · 2015
Earlier work this paper cites.
Learning aligned cross-modal representations from weakly aligned data
Lluis Castrejon, Yusuf Aytar, Carl Vondrick, Hamed Pirsiavash, and Antonio Torralba · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Structure-measure: A new way to evaluate foreground maps
Deng-Ping Fan, Ming-Ming Cheng, Yun Liu, Tao Li, and Ali Borji · 2017
Earlier work this paper cites.
MFNet: Towards real-time semantic segmentation for autonomous vehicles with multi-spectral scenes
Qishen Ha, Kohei Watanabe, Takumi Karasawa, Yoshitaka Ushiku, and Tatsuya Harada · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Enhanced-alignment measure for binary foreground map evaluation
Deng-Ping Fan, Cheng Gong, Yang Cao, Bo Ren, Ming-Ming Cheng, and Ali Borji · 2018
Earlier work this paper cites.
Depth-aware cnn for rgb-d segmentation
Weiyue Wang and Ulrich Neumann · 2018
Earlier work this paper cites.
ACNet: Attention based network to exploit complementary features for RGBD semantic segmentation
Xinxin Hu, Kailun Yang, Lei Fei, and Kaiwei Wang · 2019
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2019
Earlier work this paper cites.
Pseudo-lidar from visual depth estimation: Bridging the gap in 3d object detection for autonomous driving
Yan Wang, Wei-Lun Chao, Divyansh Garg, Bharath Hariharan, Mark Campbell, and Kilian Q Weinberger · 2019
Earlier work this paper cites.
Rethinking rgb-d salient object detection: Models, data sets, and large-scale benchmarks
Deng-Ping Fan, Zheng Lin, Zhao Zhang, Menglong Zhu, and Ming-Ming Cheng · 2020
Earlier work this paper cites.
Learning densities in feature space for reliable segmentation of indoor scenes
Nicolas Marchal, Charlotte Moraldo, Hermann Blum, Roland Siegwart, Cesar Cadena, and Abel Gawel · 2020
Cited alongside, same era.
Exploring language prior for mode-sensitive visual attention modeling
Xiaoshuai Sun, Xuying Zhang, Liujuan Cao, Yongjian Wu, Feiyue Huang, and Rongrong Ji · 2020
Cited alongside, same era.
Deep multimodal fusion by channel exchanging
Yikai Wang, Wenbing Huang, Fuchun Sun, Tingyang Xu, Yu Rong, and Junzhou Huang · 2020
Cited alongside, same era.
Depth-adapted cnn for rgb-d cameras
Zongwei Wu, Guillaume Allibert, Christophe Stolz, and Cédric Demonceaux · 2020
Cited alongside, same era.
Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text
Hassan Akbari, Liangzhe Yuan, Rui Qian, Wei-Hong Chuang, Shih-Fu Chang, Yin Cui, and Boqing Gong · 2021
Cited alongside, same era.
Adabins: Depth estimation using adaptive bins
Rgb-d salient object detection with ubiquitous target awareness
Yifan Zhao, Jiawei Zhao, Jia Li, and Xiaowu Chen · 2021
Later among the works it cites.
Perception-aware multi-sensor fusion for 3d lidar semantic segmentation
Zhuangwei Zhuang, Rong Li, Kui Jia, Qicheng Wang, Yuanqing Li, and Mingkui Tan · 2021
Later among the works it cites.
MultiMAE: Multi-modal multi-task masked autoencoders
Roman Bachmann, David Mizrahi, Andrei Atanov, and Amir Zamir · 2022
Later among the works it cites.
HRFuser: A multi-resolution sensor fusion architecture for 2D object detection
Tim Broedermann, Christos Sakaridis, Dengxin Dai, and Luc Van Gool · 2022
Later among the works it cites.
3-d convolutional neural networks for rgb-d salient object detection and beyond
Qian Chen, Zhenxi Zhang, Yanye Lu, Keren Fu, and Qijun Zhao · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Shariq Farooq Bhat, Ibraheem Alhashim, and Peter Wonka · 2021
Cited alongside, same era.
ShapeConv: Shape-aware convolutional layer for indoor RGB-D semantic segmentation
Jinming Cao, Hanchao Leng, Dani Lischinski, Danny Cohen-Or, Changhe Tu, and Yangyan Li · 2021
Cited alongside, same era.
FEANet: Feature-enhanced attention network for RGB-thermal real-time semantic segmentation
Fuqin Deng et al · 2021
Cited alongside, same era.
Omnidata: A scalable pipeline for making multi-task mid-level vision datasets from 3d scans
Ainaz Eftekhar, Alexander Sax, Jitendra Malik, and Amir Zamir · 2021
Cited alongside, same era.
Is attention better than matrix decomposition?
Zhengyang Geng, Meng-Hao Guo, Hongxu Chen, Xia Li, Ke Wei, and Zhouchen Lin · 2021
Cited alongside, same era.
Cross modal focal loss for rgbd face anti-spoofing
Anjith George and Sebastien Marcel · 2021
Cited alongside, same era.
Calibrated rgb-d salient object detection
Wei Ji, Jingjing Li, Shuang Yu, Miao Zhang, Yongri Piao, Shunyu Yao, Qi Bi, Kai Ma, Yefeng Zheng, Huchuan Lu, et al · 2021
Cited alongside, same era.
Doodlenet: Double deeplab enhanced feature fusion for thermal-color semantic segmentation
Oriel Frigo, Lucien Martin-Gaffe, and Catherine Wacongne · 2022
Later among the works it cites.
Omnivore: A single model for many visual modalities
Rohit Girdhar, Mannat Singh, Nikhila Ravi, Laurens van der Maaten, Armand Joulin, and Ishan Misra · 2022
Later among the works it cites.
Conv2former: A simple transformer-style convnet for visual recognition
Qibin Hou, Cheng-Ze Lu, Ming-Ming Cheng, and Jiashi Feng · 2022
Later among the works it cites.
Multi-modal sensor fusion for auto driving perception: A survey
Keli Huang, Botian Shi, Xiang Li, Xin Li, Siyuan Huang, and Yikang Li · 2022
Later among the works it cites.
Spsn: Superpixel prototype sampling network for rgb-d salient object detection
Minhyeok Lee, Chaewon Park, Suhwan Cho, and Sangyoun Lee · 2022
Later among the works it cites.
Rgb-t semantic segmentation with location, activation, and sharpening
Gongyang Li, Yike Wang, Zhi Liu, Xinpeng Zhang, and Dan Zeng · 2022
Later among the works it cites.
A convnet for the 2020s
Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie · 2022
Later among the works it cites.
Efficient multi-task rgb-d scene analysis for indoor environments
Daniel Seichter, Söhnke Benedikt Fischedick, Mona Köhler, and Horst-Michael Groß · 2022
Later among the works it cites.
Multimodal token fusion for vision transformers
Yikai Wang, Xinghao Chen, Lele Cao, Wenbing Huang, Fuchun Sun, and Yunhe Wang · 2022
Later among the works it cites.
Difnet: Boosting visual information flow for image captioning
Mingrui Wu, Xuying Zhang, Xiaoshuai Sun, Yiyi Zhou, Chao Chen, Jiaxin Gu, Xing Sun, and Rongrong Ji · 2022
Later among the works it cites.
Bi-directional progressive guidance network for rgb-d salient object detection
Yang Yang, Qi Qin, Yongjiang Luo, Yi Liu, Qiang Zhang, and Jungong Han · 2022
Later among the works it cites.
Camoformer: Masked separable attention for camouflaged object detection
Bowen Yin, Xuying Zhang, Qibin Hou, Bo-Yuan Sun, Deng-Ping Fan, and Luc Van Gool · 2022
Later among the works it cites.
C 2 dfnet: Criss-cross dynamic filter network for rgb-d salient object detection
Miao Zhang, Shunyu Yao, Beiqi Hu, Yongri Piao, and Wei Ji · 2022
Later among the works it cites.
Yolo-ms: Rethinking multi-scale representation learning for real-time object detection
Yuming Chen, Xinbin Yuan, Ruiqi Wu, Jiabao Wang, Qibin Hou, and Ming-Ming Cheng · 2023
Closest in time.
Co-slam: Joint coordinate and sparse parametric encodings for neural real-time slam
Hengyi Wang, Jingwen Wang, and Lourdes Agapito · 2023
Closest in time.
Hidanet: Rgb-d salient object detection via hierarchical depth awareness
Zongwei Wu, Guillaume Allibert, Fabrice Meriaudeau, Chao Ma, and Cédric Demonceaux · 2023
Closest in time.