Fetching the paper…
Reading the bibliography…
Accurate monocular metric depth estimation (MMDE) is crucial to solving downstream tasks in 3D perception and modeling.
A. Bemporad and M. Morari, “Robust model predictive control: A survey,” in Robustness in identification and control . Springer, 2007, pp. 207–226
2007
Earlier work this paper cites.
J. Nickolls, I. Buck, M. Garland, and K. Skadron, “Scalable parallel programming with cuda: Is cuda the parallel programming model that application developers have been waiting for?” Queue , vol. 6, no. 2, pp. 40–53, 2008
2008
Earlier work this paper cites.
P. K. Nathan Silberman, Derek Hoiem and R. Fergus, “Indoor segmentation and support inference from rgbd images,” in The European Conference on Computer Vision (ECCV) , 2012
2012
Earlier work this paper cites.
A. Geiger, P. Lenz, and R. Urtasun, “Are we ready for autonomous driving? The KITTI vision benchmark suite,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2012
2012
Earlier work this paper cites.
R. Achanta, A. Shaji, K. Smith, A. Lucchi, P. Fua, and S. Süsstrunk, “Slic superpixels compared to state-of-the-art superpixel methods,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 34, no. 11, pp. 2274–2282, 2012
2012
Earlier work this paper cites.
J. Sturm, N. Engelhard, F. Endres, W. Burgard, and D. Cremers, “A benchmark for the evaluation of rgb-d slam systems,” in Proc. of the International Conference on Intelligent Robot Systems (IROS) , 2012
2012
Earlier work this paper cites.
D. J. Butler, J. Wulff, G. B. Stanley, and M. J. Black, “A naturalistic open source movie for optical flow evaluation,” in The European Conference on Computer Vision (ECCV) , ser. Part IV, LNCS 7577. Springer, 2012, pp. 611–625
2012
Earlier work this paper cites.
D. Eigen, C. Puhrsch, and R. Fergus, “Depth map prediction from a single image using a multi-scale deep network,” in Advances in Neural Information Processing Systems (NeurIPS) , vol. 3. Neural information processing systems foundation, 6 2014, pp. 2366–2374
2014
Earlier work this paper cites.
W. Hu, X. Gao, X. Li, S. Zhao, X. Cun, Y. Zhang, L. Quan, and Y. Shan, “Depthcrafter: Generating consistent long depth sequences for open-world videos,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , June 2025, pp. 2005–2015
2015
Earlier work this paper cites.
F. Liu, C. Shen, G. Lin, and I. Reid, “Learning depth from single monocular images using deep convolutional neural fields,” IEEE Transactions on Pattern Analysis and Machine Intelligence (T-PAMI) , vol. 38, pp. 2024–2039, 2 2015
2015
Earlier work this paper cites.
S. Song, S. P. Lichtenberg, and J. Xiao, “Sun rgb-d: A rgb-d scene understanding benchmark suite,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , vol. 07-12-June-2015, pp. 567–576, 10 2015
2015
Earlier work this paper cites.
A. Mesbah, “Stochastic model predictive control: An overview and perspectives for future research,” IEEE Control Systems Magazine , vol. 36, no. 6, pp. 30–44, 2016
2016
Earlier work this paper cites.
I. Laina, C. Rupprecht, V. Belagiannis, F. Tombari, and N. Navab, “Deeper depth prediction with fully convolutional residual networks,” Proceedings of the International Conference on 3D Vision (3DV) , pp. 239–248, 6 2016
2016
Earlier work this paper cites.
A. Chang, A. Dai, T. Funkhouser, M. Halber, M. Niessner, M. Savva, S. Song, A. Zeng, and Y. Zhang, “Matterport3d: Learning from rgb-d data in indoor environments,” in Proceedings of the International Conference on 3D Vision (3DV) , 2017
2017
Earlier work this paper cites.
A. Dai, A. X. Chang, M. Savva, M. Halber, T. Funkhouser, and M. Nießner, “Scannet: Richly-annotated 3d reconstructions of indoor scenes,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2017
2017
Earlier work this paper cites.
T. Schöps, J. L. Schönberger, S. Galliani, T. Sattler, K. Schindler, M. Pollefeys, and A. Geiger, “A multi-view stereo benchmark with high-resolution images and multi-camera videos,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2017
2017
Earlier work this paper cites.
I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” 7th International Conference on Learning Representations, ICLR 2019 , 11 2017
2017
Earlier work this paper cites.
H. Fu, M. Gong, C. Wang, K. Batmanghelich, and D. Tao, “Deep ordinal regression network for monocular depth estimation,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pp. 2002–2011, 6 2018
2018
Earlier work this paper cites.
A. R. Zamir, A. Sax, W. B. Shen, L. Guibas, J. Malik, and S. Savarese, “Taskonomy: Disentangling task transfer learning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . IEEE, 2018
2018
Earlier work this paper cites.
B. Zhou, P. Krähenbühl, and V. Koltun, “Does computer vision matter for action?” Science Robotics , vol. 4, 5 2019
2019
Earlier work this paper cites.
Y. Wang, W.-L. Chao, D. Garg, B. Hariharan, M. Campbell, and K. Q. Weinberger, “Pseudo-lidar from visual depth estimation: Bridging the gap in 3d object detection for autonomous driving,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 8445–8453
2019
Earlier work this paper cites.
J. M. Facil, B. Ummenhofer, H. Zhou, L. Montesano, T. Brox, and J. Civera, “Cam-convs: Camera-aware multi-scale convolutions for single-view depth,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2019, pp. 11 826–11 835
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
G. Yang, X. Song, C. Huang, Z. Deng, J. Shi, and B. Zhou, “Drivingstereo: A large-scale dataset for stereo matching in autonomous driving scenarios,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2019
2019
Earlier work this paper cites.
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Kopf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala, “Pytorch: An imperative style, high-performance deep learning library,” in Advances in Neural Information Processing Systems (NeurIPS) . Curran Associates, Inc., 2019, pp. 8024–8035
2019
Earlier work this paper cites.
S. F. Bhat, I. Alhashim, and P. Wonka, “Adabins: Depth estimation using adaptive bins,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pp. 4008–4017, 11 2020
2020
Earlier work this paper cites.
Y. Wang, X. Chen, Y. You, L. E. Li, B. Hariharan, M. Campbell, K. Q. Weinberger, and W. L. Chao, “Train in germany, test in the usa: Making 3d object detectors generalize,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pp. 11 710–11 720, 5 2020
2020
Earlier work this paper cites.
R. Ranftl, K. Lasinger, D. Hafner, K. Schindler, and V. Koltun, “Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer,” IEEE Transactions on Pattern Analysis and Machine Intelligence (T-PAMI) , vol. 44, no. 3, pp. 1623–1637, 2020
2020
Earlier work this paper cites.
M. L. Antequera, P. Gargallo, M. Hofinger, S. R. Bulò, Y. Kuang, and P. Kontschieder, “Mapillary planet-scale depth dataset,” in The European Conference on Computer Vision (ECCV) . Springer International Publishing, 2020, pp. 589–604
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
Y. Yao, Z. Luo, S. Li, J. Zhang, Y. Ren, L. Zhou, T. Fang, and L. Quan, “Blendedmvs: A large-scale dataset for generalized multi-view stereo networks,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2020, pp. 1790–1799
2020
Cited alongside, same era.
W. Wang, D. Zhu, X. Wang, Y. Hu, Y. Qiu, C. Wang, Y. Hu, A. Kapoor, and S. Scherer, “Tartanair: A dataset to push the limits of visual slam,” in 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, 2020, pp. 4909–4916
2020
Cited alongside, same era.
P. Sun, H. Kretzschmar, X. Dotiwalla, A. Chouard, V. Patnaik, P. Tsui, J. Guo, Y. Zhou, Y. Chai, B. Caine et al. , “Scalability in perception for autonomous driving: Waymo open dataset,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2020, pp. 2446–2454
2020
Cited alongside, same era.
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo, P. Dollár, and R. Girshick, “Segment anything,” 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
S. Gasperini, V. Olsson, M. Poggi, F. Tosi, S. Salti, L. Di Stefano, K. AAstr ”om, J. Gonfaus, L. Van Gool, R. Timofte, A. N ”aslund, and L. Bitti, “Robust monocular depth estimation under challenging conditions,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , October 2023, pp. 7897–7908
2023
Later among the works it cites.
A. Costanzino, F. Tosi, M. Poggi, S. Salti, S. Mattoccia, and L. Di Stefano, “Learning depth estimation for transparent and mirror surfaces,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , October 2023, pp. 17 770–17 780
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
V. Guizilini, R. Ambrus, S. Pillai, A. Raventos, and A. Gaidon, “3d packing for self-supervised monocular depth estimation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2020
2020
Cited alongside, same era.
H. Caesar, V. Bankiti, A. H. Lang, S. Vora, V. E. Liong, Q. Xu, A. Krishnan, Y. Pan, G. Baldan, and O. Beijbom, “nuscenes: A multimodal dataset for autonomous driving,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2020
2020
Cited alongside, same era.
M. Poggi, F. Aleotti, F. Tosi, and S. Mattoccia, “On the uncertainty of self-supervised monocular depth estimation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2020, pp. 3227–3237
2020
Cited alongside, same era.
D. Park, R. Ambrus, V. Guizilini, J. Li, and A. Gaidon, “Is pseudo-lidar needed for monocular 3d object detection?” in IEEE/CVF International Conference on Computer Vision (ICCV) , 2021
2021
Cited alongside, same era.
R. Ranftl, A. Bochkovskiy, and V. Koltun, “Vision transformers for dense prediction,” Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , pp. 12 159–12 168, 3 2021
2021
Cited alongside, same era.
A. Eftekhar, A. Sax, J. Malik, and A. Zamir, “Omnidata: A scalable pipeline for making multi-task mid-level vision datasets from 3d scans,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , 2021, pp. 10 786–10 796
2021
Cited alongside, same era.
W. Yin, J. Zhang, O. Wang, S. Niklaus, L. Mai, S. Chen, and C. Shen, “Learning to recover 3d scene shape from a single image,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2021, pp. 204–213
2021
Cited alongside, same era.
G. Yang, H. Tang, M. Ding, N. Sebe, and E. Ricci, “Transformer-based attention networks for continuous pixel-wise prediction,” Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , pp. 16 249–16 259, 3 2021
2021
Cited alongside, same era.
2023
Later among the works it cites.
M. J. Black, P. Patel, J. Tesch, and J. Yang, “BEDLAM: A synthetic dataset of bodies exhibiting detailed lifelike animated motion,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2023, pp. 8726–8737
2023
Later among the works it cites.
N. Karaev, I. Rocco, B. Graham, N. Neverova, A. Vedaldi, and C. Rupprecht, “Dynamicstereo: Consistent dynamic depth from stereo videos,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2023
2023
Later among the works it cites.
Y. Li, L. Jiang, L. Xu, Y. Xiangli, Z. Wang, D. Lin, and B. Dai, “Matrixcity: A large-scale city dataset for city-scale neural rendering and beyond,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , 2023, pp. 3205–3215
2023
Later among the works it cites.
Y. Zheng, A. W. Harley, B. Shen, G. Wetzstein, and L. J. Guibas, “Pointodyssey: A large-scale synthetic dataset for long-term point tracking,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , 2023, pp. 19 855–19 865
2023
Later among the works it cites.
C. Yeshwanth, Y.-C. Liu, M. Nießner, and A. Dai, “Scannet++: A high-fidelity dataset of 3d indoor scenes,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , 2023
2023
Later among the works it cites.
Y. Randall and T. Treibitz, “Flsea: Underwater visual-inertial and stereo-vision forward-looking datasets,” 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2024
Later among the works it cites.
L. Piccinelli, Y.-H. Yang, C. Sakaridis, M. Segu, S. Li, L. Van Gool, and F. Yu, “Unidepth: Universal monocular metric depth estimation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2024, pp. 10 106–10 116
2024
Later among the works it cites.
B. Ke, A. Obukhov, S. Huang, N. Metzger, R. C. Daudt, and K. Schindler, “Repurposing diffusion-based image generators for monocular depth estimation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2024, pp. 9492–9502
2024
Later among the works it cites.
L. Yang, B. Kang, Z. Huang, X. Xu, J. Feng, and H. Zhao, “Depth anything: Unleashing the power of large-scale unlabeled data,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2024, pp. 10 371–10 381
2024
Later among the works it cites.
2024
Later among the works it cites.
X. Fu, W. Yin, M. Hu, K. Wang, Y. Ma, P. Tan, S. Shen, D. Lin, and X. Long, “Geowizard: Unleashing the diffusion priors for 3d geometry estimation from a single image,” in European Conference on Computer Vision (ECCV) , 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
F. Tosi, P. Zama Ramirez, and M. Poggi, “Diffusion models for monocular depth estimation: Overcoming challenging conditions,” in European Conference on Computer Vision (ECCV) , 2024
2024
Later among the works it cites.
P. Z. Ramirez, A. Costanzino, F. Tosi, M. Poggi, S. Salti, S. Mattoccia, and L. Di Stefano, “Booster: A benchmark for depth from images of specular and transparent surfaces,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 46, no. 1, pp. 85–102, 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
L. Ling, Y. Sheng, Z. Tu, W. Zhao, C. Xin, K. Wan, L. Yu, Q. Guo, Z. Yu, Y. Lu et al. , “DL3DV-10k: A large-scale scene dataset for deep learning-based 3d vision,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2024, pp. 22 160–22 169
2024
Later among the works it cites.
H. Xia, Y. Fu, S. Liu, and X. Wang, “Rgbd objects in the wild: Scaling real-world 3d object learning from rgb-d videos,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2024, pp. 22 378–22 389
2024
Later among the works it cites.
A. Veicht, P. Sarlin, P. Lindenberger, and M. Pollefeys, “Geocalib: Learning single-image calibration with geometric optimization,” in European Conference on Computer Vision (ECCV) , ser. Lecture Notes in Computer Science, vol. 15098. Springer, 2024, pp. 1–20
2024
Later among the works it cites.
J. He, H. Li, W. Yin, Y. Liang, L. Li, K. Zhou, H. Zhang, B. Liu, and Y.-C. Chen, “Lotus: Diffusion-based visual foundation model for high-quality dense prediction,” in International Conference on Learning Representations (ICLR) , 2025
2025
Closest in time.
S. Chen, H. Guo, S. Zhu, F. Zhang, Z. Huang, J. Feng, and B. Kang, “Video depth anything: Consistent depth estimation for super-long videos,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , June 2025, pp. 22 831–22 840
2025
Closest in time.
2025
Closest in time.
L. Piccinelli, C. Sakaridis, M. Segu, Y.-H. Yang, S. Li, W. Abbeloos, and L. Van Gool, “UniK3D: Universal camera monocular 3d estimation,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2025
2025
Closest in time.
Z. Li and N. Snavely, “Megadepth: Learning single-view depth prediction from internet photos,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2018, pp. 2041–2050
2050
Closest in time.
A. D. Bonzanini, A. Mesbah, and S. Di Cairano, “Perception-aware chance-constrained model predictive control for uncertain environments,” in 2021 American Control Conference (ACC) . IEEE, 2021, pp. 2082–2087
2087
Closest in time.