An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Later among the works it cites.
Guiding monocular depth estimation using depth-attention volume
Lam Huynh, Phong Nguyen-Ha, Jiri Matas, Esa Rahtu, and Janne Heikkilä · 2020
Later among the works it cites.
Self-supervised monocular depth estimation: Solving the dynamic object problem by semantic guidance
Marvin Klingner, Jan-Aike Termöhlen, Jonas Mikolajczyk, and Tim Fingscheidt · 2020
Later among the works it cites.
Bifuse: Monocular 360 depth estimation via bi-projection fusion
Fu-En Wang, Yu-Hsuan Yeh, Min Sun, Wei-Chen Chiu, and Yi-Hsuan Tsai · 2020
Later among the works it cites.
Mfusenet: Robust depth estimation with learned multiscopic fusion
Weihao Yuan, Rui Fan, Michael Yu Wang, and Qifeng Chen · 2020
Later among the works it cites.
Bidirectional attention network for monocular depth estimation
Shubhra Aich, Jean Marie Uwabeza Vianney, Md Amirul Islam, Mannat Kaur, and Bingbing Liu · 2021
Later among the works it cites.
Adabins: Depth estimation using adaptive bins
Shariq Farooq Bhat, Ibraheem Alhashim, and Peter Wonka · 2021
Later among the works it cites.
Dro: Deep recurrent optimizer for structure-from-motion
Original
Xiaodong Gu, Weihao Yuan, Zuozhuo Dai, Chengzhou Tang, Siyu Zhu, and Ping Tan · 2021
Later among the works it cites.
Sparse auxiliary networks for unified monocular depth prediction and completion
Vitor Guizilini, Rares Ambrus, Wolfram Burgard, and Adrien Gaidon · 2021
Later among the works it cites.
Unifuse: Unidirectional fusion for 360 panorama depth estimation
Hualie Jiang, Zhe Sheng, Siyu Zhu, Zilong Dong, and Rui Huang · 2021
Later among the works it cites.
Patch-wise attention network for monocular depth estimation
Sihaeng Lee, Janghyeon Lee, Byungju Kim, Eojindl Yi, and Junmo Kim · 2021
Later among the works it cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Original
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo · 2021
Later among the works it cites.
Slicenet: deep dense depth estimation from a single indoor panorama using a slice-based representation
Giovanni Pintore, Marco Agus, Eva Almansa, Jens Schneider, and Enrico Gobbetti · 2021
Later among the works it cites.
Vip-deeplab: Learning visual perception with depth-aware video panoptic segmentation
Siyuan Qiao, Yukun Zhu, Hartwig Adam, Alan Yuille, and Liang-Chieh Chen · 2021
Later among the works it cites.
Vision transformers for dense prediction
René Ranftl, Alexey Bochkovskiy, and Vladlen Koltun · 2021
Later among the works it cites.
Hohonet: 360 indoor holistic understanding with latent horizontal features
Cheng Sun, Min Sun, and Hwann-Tzong Chen · 2021
Later among the works it cites.
Stereo matching by self-supervision of multiscopic vision
Weihao Yuan, Yazhan Zhang, Bingkun Wu, Siyu Zhu, Ping Tan, Michael Yu Wang, and Qifeng Chen · 2021
Later among the works it cites.
Adabins: Depth estimation using adaptive bins
Shariq Farooq Bhat, Ibraheem Alhashim, and Peter Wonka · 2021
Later among the works it cites.