Fetching the paper…
Reading the bibliography…
We introduce Metric3D v2, a geometric foundation model for zero-shot metric depth and surface normal estimation from a single image, which is crucial for metric 3D recovery.
P. Besl and N. McKay, “Method for registration of 3-d shapes,” in
1992
Earlier work this paper cites.
A. Saxena, M. Sun, and A. Y. Ng, “Make3d: Learning 3d scene structure from a single still image,”
2008
Earlier work this paper cites.
N. Silberman, D. Hoiem, P. Kohli, and R. Fergus, “Indoor segmentation and support inference from rgbd images,” in
2012
Earlier work this paper cites.
A. Geiger, P. Lenz, and R. Urtasun, “Are we ready for autonomous driving? the kitti vision benchmark suite,” in
2012
Earlier work this paper cites.
N. Silberman, D. Hoiem, P. Kohli, and R. Fergus, “Indoor segmentation and support inference from rgbd images,” in
2012
Earlier work this paper cites.
A. Geiger, P. Lenz, C. Stiller, and R. Urtasun, “Vision meets robotics: The kitti dataset,”
2013
Earlier work this paper cites.
A. Geiger, P. Lenz, C. Stiller, and R. Urtasun, “Vision meets robotics: The kitti dataset,”
2013
Earlier work this paper cites.
J. T. Barron and J. Malik, “Shape, illumination, and reflectance from shading,”
2014
Earlier work this paper cites.
D. Eigen, C. Puhrsch, and R. Fergus, “Depth map prediction from a single image using a multi-scale deep network,” in
2014
Earlier work this paper cites.
L. Ladickỳ, B. Zeisl, and M. Pollefeys, “Discriminatively trained dense surface normal estimation,” in
2014
Earlier work this paper cites.
D. F. Fouhey, A. Gupta, and M. Hebert, “Unfolding an indoor origami world,” in
2014
Earlier work this paper cites.
X. Wang, D. Fouhey, and A. Gupta, “Designing deep networks for surface normal estimation,” in
2015
Earlier work this paper cites.
D. Eigen and R. Fergus, “Predicting depth, surface normals and semantic labels with a common multi-scale convolutional architecture,” in
2015
Earlier work this paper cites.
B. Li, C. Shen, Y. Dai, A. Van Den Hengel, and M. He, “Depth and surface normal estimation from monocular images using regression on deep features and hierarchical crfs,” in
2015
Earlier work this paper cites.
S. Song, M. Chandraker, and C. C. Guest, “High accuracy monocular sfm and scale correction for autonomous driving,”
2015
Earlier work this paper cites.
W. Chen, Z. Fu, D. Yang, and J. Deng, “Single-image depth perception in the wild,” in
2016
Earlier work this paper cites.
I. Laina, C. Rupprecht, V. Belagiannis, F. Tombari, and N. Navab, “Deeper depth prediction with fully convolutional residual networks,” in
2016
Earlier work this paper cites.
A. Bansal, B. Russell, and A. Gupta, “Marr revisited: 2d-3d alignment via surface normal prediction,” in
2016
Earlier work this paper cites.
P. Wang, X. Shen, B. Russell, S. Cohen, B. Price, and A. L. Yuille, “Surge: Surface regularized geometry estimation from a single image,”
2016
Earlier work this paper cites.
J. L. Schönberger, E. Zheng, M. Pollefeys, and J.-M. Frahm, “Pixelwise view selection for unstructured multi-view stereo,” in
2016
Earlier work this paper cites.
M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Benenson, U. Franke, S. Roth, and B. Schiele, “The cityscapes dataset for semantic urban scene understanding,” in
2016
Earlier work this paper cites.
M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Benenson, U. Franke, S. Roth, and B. Schiele, “The cityscapes dataset for semantic urban scene understanding,” in
2016
Earlier work this paper cites.
R. Mur-Artal and J. D. Tardós, “ORB-SLAM2: an open-source SLAM system for monocular, stereo and RGB-D cameras,”
2017
Earlier work this paper cites.
R. Mur-Artal and J. D. Tardós, “Orb-slam2: An open-source slam system for monocular, stereo, and rgb-d cameras,”
2017
Earlier work this paper cites.
W. Chen, D. Xiang, and J. Deng, “Surface normals in the wild,” in
2017
Earlier work this paper cites.
J. Li, R. Klein, and A. Yao, “A two-streamed network for estimating fine-scaled depth maps from single rgb images,” in
2017
Earlier work this paper cites.
S. Xie, R. Girshick, P. Dollár, Z. Tu, and K. He, “Aggregated residual transformations for deep neural networks,” in
2017
Earlier work this paper cites.
A. Dai, A. X. Chang, M. Savva, M. Halber, T. Funkhouser, and M. Nießner, “Scannet: Richly-annotated 3d reconstructions of indoor scenes,” in
2017
Earlier work this paper cites.
A. Knapitsch, J. Park, Q.-Y. Zhou, and V. Koltun, “Tanks and temples: Benchmarking large-scale scene reconstruction,”
2017
Earlier work this paper cites.
T. Schops, J. L. Schonberger, S. Galliani, T. Sattler, K. Schindler, M. Pollefeys, and A. Geiger, “A multi-view stereo benchmark with high-resolution images and multi-camera videos,” in
2017
Earlier work this paper cites.
A. Dai, A. X. Chang, M. Savva, M. Halber, T. Funkhouser, and M. Nießner, “Scannet: Richly-annotated 3d reconstructions of indoor scenes,” in
2017
Earlier work this paper cites.
T. Schops, J. L. Schonberger, S. Galliani, T. Sattler, K. Schindler, M. Pollefeys, and A. Geiger, “A multi-view stereo benchmark with high-resolution images and multi-camera videos,” in
2017
Earlier work this paper cites.
B. Yang, S. Rosa, A. Markham, N. Trigoni, and H. Wen, “Dense 3d object reconstruction from a single depth view,”
2018
Earlier work this paper cites.
J. Behley and C. Stachniss, “Efficient surfel-based slam using 3d laser range data in urban environments.,” in
2018
Earlier work this paper cites.
K. Xian, C. Shen, Z. Cao, H. Lu, Y. Xiao, R. Li, and Z. Luo, “Monocular relative depth perception with web stereo data supervision,” in
2018
Earlier work this paper cites.
X. Qi, R. Liao, Z. Liu, R. Urtasun, and J. Jia, “Geonet: Geometric neural network for joint depth and surface normal estimation,” in
2018
Earlier work this paper cites.
N. Wang, Y. Zhang, Z. Li, Y. Fu, W. Liu, and Y.-G. Jiang, “Pixel2mesh: Generating 3d mesh models from single RGB images,” in
2018
Earlier work this paper cites.
J. Wu, C. Zhang, X. Zhang, Z. Zhang, W. Freeman, and J. Tenenbaum, “Learning shape priors for single-view 3d completion and reconstruction,” in
2018
Earlier work this paper cites.
D. Sun, X. Yang, M.-Y. Liu, and J. Kautz, “Pwc-net: Cnns for optical flow using pyramid, warping, and cost volume,” in
2018
Earlier work this paper cites.
X. Cheng, P. Wang, and R. Yang, “Depth estimation via affinity learned with convolutional spatial propagation network,” in
2018
Earlier work this paper cites.
Y. Hold-Geoffroy, K. Sunkavalli, J. Eisenmann, M. Fisher, E. Gambaretto, S. Hadap, and J.-F. Lalonde, “A perceptual measure for deep single image camera calibration,” in
2018
Earlier work this paper cites.
X. Guo, H. Li, S. Yi, J. Ren, and X. Wang, “Learning monocular depth by distilling cross-domain stereo networks,” in
2018
Earlier work this paper cites.
T. Koch, L. Liebel, F. Fraundorfer, and M. Korner, “Evaluation of cnn-based single-image depth estimation methods,” in
2018
Earlier work this paper cites.
A. Zamir, A. Sax, , W. Shen, L. Guibas, J. Malik, and S. Savarese, “Taskonomy: Disentangling task transfer learning,” in
2018
Earlier work this paper cites.
Z. Yin and J. Shi, “Geonet: Unsupervised learning of dense depth, optical flow and camera pose,” in
2018
Earlier work this paper cites.
A. Zamir, A. Sax, , W. Shen, L. Guibas, J. Malik, and S. Savarese, “Taskonomy: Disentangling task transfer learning,” in
2018
Earlier work this paper cites.
T. Koch, L. Liebel, F. Fraundorfer, and M. Korner, “Evaluation of cnn-based single-image depth estimation methods,” in
2018
Earlier work this paper cites.
L. Mescheder, M. Oechsle, M. Niemeyer, S. Nowozin, and A. Geiger, “Occupancy networks: Learning 3d reconstruction in function space,” in
2019
Earlier work this paper cites.
T. Schops, T. Sattler, and M. Pollefeys, “Bad slam: Bundle adjusted direct rgb-d slam,” in
2019
Earlier work this paper cites.
W. Yin, Y. Liu, C. Shen, and Y. Yan, “Enforcing geometric constraints of virtual normal for depth prediction,” in
2019
Earlier work this paper cites.
J. Facil, B. Ummenhofer, H. Zhou, L. Montesano, T. Brox, and J. Civera, “CAM-Convs: camera-aware multi-scale convolutions for single-view depth,” in
2019
Earlier work this paper cites.
J. Huang, Y. Zhou, T. Funkhouser, and L. J. Guibas, “Framenet: Learning local canonical frames of 3d surfaces from a single rgb image,” in
2019
Earlier work this paper cites.
S. Im, H.-G. Jeon, S. Lin, and I.-S. Kweon, “Dpsnet: End-to-end deep plane sweep stereo,” in
2019
Earlier work this paper cites.
S. Saito, Z. Huang, R. Natsume, S. Morishima, A. Kanazawa, and H. Li, “Pifu: Pixel-aligned implicit function for high-resolution clothed human digitization,” in
2019
Earlier work this paper cites.
K. Wang, F. Gao, and S. Shen, “Real-time scalable dense surfel mapping,” in
2019
Earlier work this paper cites.
S. Liao, E. Gavves, and C. G. Snoek, “Spherical regression: Learning viewpoints, surface normals and 3d rotations on n-spheres,” in
2019
Earlier work this paper cites.
D. Singh and B. Singh, “Investigating the impact of data normalization on classification performance,”
2019
Earlier work this paper cites.
Z. Zhang, Z. Cui, C. Xu, Y. Yan, N. Sebe, and J. Yang, “Pattern-affinitive propagation across depth, surface normal and semantic segmentation,” in
2019
Earlier work this paper cites.
J. H. Lee, M.-K. Han, D. W. Ko, and I. H. Suh, “From big to small: Multi-scale local planar guidance for monocular depth estimation,”
2019
Cited alongside, same era.
I. Vasiljevic, N. Kolkin, S. Zhang, R. Luo, H. Wang, F. Z. Dai, A. F. Daniele, M. Mostajabi, S. Basart, M. R. Walter,
2019
Cited alongside, same era.
R. Kesten, M. Usman, J. Houston, T. Pandya, K. Nadhamuni, A. Ferreira, M. Yuan, B. Low, A. Jain, P. Ondruska, S. Omari, S. Shah, A. Kulkarni, A. Kazakova, C. Tao, L. Platinsky, W. Jiang, and V. Shet, “Level 5 perception dataset 2020.”
2019
Cited alongside, same era.
G. Yang, X. Song, C. Huang, Z. Deng, J. Shi, and B. Zhou, “Drivingstereo: A large-scale dataset for stereo matching in autonomous driving scenarios,” in
2019
Cited alongside, same era.
Z. Bauer, F. Gomez-Donoso, E. Cruz, S. Orts-Escolano, and M. Cazorla, “Uasol, a large-scale high-resolution outdoor stereo dataset,”
S. F. Bhat, I. Alhashim, and P. Wonka, “Adabins: Depth estimation using adaptive bins,” in
2021
Later among the works it cites.
R. Ranftl, A. Bochkovskiy, and V. Koltun, “Vision transformers for dense prediction,” in
2021
Later among the works it cites.
L. Lipson, Z. Teed, and J. Deng, “Raft-stereo: Multilevel recurrent field transforms for stereo matching,” in
2021
Later among the works it cites.
J. Cho, D. Min, Y. Kim, and K. Sohn, “DIML/CVL RGB-D dataset: 2m RGB-D images of natural indoor and outdoor scenes,”
2021
Later among the works it cites.
B. Wilson, W. Qi, T. Agarwal, J. Lambert, J. Singh, S. Khandelwal, B. Pan, R. Kumar, A. Hartnett, J. K. Pontes, D. Ramanan, P. Carr, and J. Hays, “Argoverse 2: Next generation datasets for self-driving perception and forecasting,” in
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
2019
Cited alongside, same era.
R. Kesten, M. Usman, J. Houston, T. Pandya, K. Nadhamuni, A. Ferreira, M. Yuan, B. Low, A. Jain, P. Ondruska, S. Omari, S. Shah, A. Kulkarni, A. Kazakova, C. Tao, L. Platinsky, W. Jiang, and V. Shet, “Level 5 perception dataset 2020.”
2019
Cited alongside, same era.
G. Yang, X. Song, C. Huang, Z. Deng, J. Shi, and B. Zhou, “Drivingstereo: A large-scale dataset for stereo matching in autonomous driving scenarios,” in
2019
Cited alongside, same era.
Z. Bauer, F. Gomez-Donoso, E. Cruz, S. Orts-Escolano, and M. Cazorla, “Uasol, a large-scale high-resolution outdoor stereo dataset,”
2019
Cited alongside, same era.
2019
Cited alongside, same era.
I. Vasiljevic, N. Kolkin, S. Zhang, R. Luo, H. Wang, F. Z. Dai, A. F. Daniele, M. Mostajabi, S. Basart, M. R. Walter,
2019
Cited alongside, same era.
R. Fan, H. Wang, P. Cai, and M. Liu, “Sne-roadseg: Incorporating surface normal information into semantic segmentation for accurate freespace detection,” in
2020
Cited alongside, same era.
M. Gehrig, W. Aarents, D. Gehrig, and D. Scaramuzza, “Dsec: A stereo event camera dataset for driving scenarios,”
2021
Later among the works it cites.
P. Xiao, Z. Shao, S. Hao, Z. Zhang, X. Chai, J. Jiao, Z. Li, J. Wu, K. Sun, K. Jiang, Y. Wang, and D. Yang, “Pandaset: Advanced sensor suite dataset for autonomous driving,” in
2021
Later among the works it cites.
2021
Later among the works it cites.
M. Roberts, J. Ramapuram, A. Ranjan, A. Kumar, M. A. Bautista, N. Paczan, R. Webb, and J. M. Susskind, “Hypersim: A photorealistic synthetic dataset for holistic indoor scene understanding,” in
2021
Later among the works it cites.
W. Yin, J. Zhang, O. Wang, S. Niklaus, L. Mai, S. Chen, and C. Shen, “Learning to recover 3d scene shape from a single image,” in
2021
Later among the works it cites.
G. Bae, I. Budvytis, and R. Cipolla, “Estimating and exploiting the aleatoric uncertainty in surface normal estimation,” in
2021
Later among the works it cites.
A. Eftekhar, A. Sax, J. Malik, and A. Zamir, “Omnidata: A scalable pipeline for making multi-task mid-level vision datasets from 3d scans,” in
2021
Later among the works it cites.
Z. Teed and J. Deng, “Droid-slam: Deep visual slam for monocular, stereo, and rgb-d cameras,” vol. 34, pp. 16558–16569, 2021
2021
Later among the works it cites.
K. Deng, A. Liu, J.-Y. Zhu, and D. Ramanan, “Depth-supervised nerf: Fewer views and faster training for free,” in
2022
Later among the works it cites.
B. Roessle, J. T. Barron, B. Mildenhall, P. P. Srinivasan, and M. Nießner, “Dense depth priors for neural radiance fields from sparse input views,” in
2022
Later among the works it cites.
Z. Yu, S. Peng, M. Niemeyer, T. Sattler, and A. Geiger, “Monosdf: Exploring monocular geometric cues for neural implicit surface reconstruction,”
2022
Later among the works it cites.
Y. Li, Z. Ge, G. Yu, J. Yang, Z. Wang, Y. Shi, J. Sun, and Z. Li, “Bevdepth: Acquisition of reliable depth for multi-view 3d object detection,”
2022
Later among the works it cites.
W. Yuan, X. Gu, Z. Dai, S. Zhu, and P. Tan, “New CRFs: Neural window fully-connected CRFs for monocular depth estimation,” in
2022
Later among the works it cites.
W. Yin, J. Zhang, O. Wang, S. Niklaus, S. Chen, Y. Liu, and C. Shen, “Towards accurate reconstruction of 3d scene shape from a single monocular image,”
2022
Later among the works it cites.
C. Zhang, W. Yin, Z. Wang, G. Yu, B. Fu, and C. Shen, “Hierarchical normalization for robust monocular depth estimation,”
2022
Later among the works it cites.
S. Peng, S. Zhang, Z. Xu, C. Geng, B. Jiang, H. Bao, and X. Zhou, “Animatable neural implicit surfaces for creating avatars from videos,”
2022
Later among the works it cites.
2022
Later among the works it cites.
L. Sun, W. Yin, E. Xie, Z. Li, C. Sun, and C. Shen, “Improving monocular visual odometry using learned depth,”
2022
Later among the works it cites.
J. Wang, P. Wang, X. Long, C. Theobalt, T. Komura, L. Liu, and W. Wang, “Neuris: Neural reconstruction of indoor scenes using normal priors,” in
2022
Later among the works it cites.
W. Yin, Y. Liu, C. Shen, A. v. d. Hengel, and B. Sun, “The devil is in the labels: Semantic segmentation from sentences,”
2022
Later among the works it cites.
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” in
2022
Later among the works it cites.
O. F. Kar, T. Yeo, A. Atanov, and A. Zamir, “3d common corruptions and data augmentation,” in
2022
Later among the works it cites.
Z. Liu, H. Mao, C.-Y. Wu, C. Feichtenhofer, T. Darrell, and S. Xie, “A convnet for the 2020s,” in
2022
Later among the works it cites.
X. Zhai, A. Kolesnikov, N. Houlsby, and L. Beyer, “Scaling vision transformers,” in
2022
Later among the works it cites.
M. Sayed, J. Gibson, J. Watson, V. Prisacariu, M. Firman, and C. Godard, “Simplerecon: 3d reconstruction without 3d convolutions,” in
2022
Later among the works it cites.
Z. Liu, H. Mao, C.-Y. Wu, C. Feichtenhofer, T. Darrell, and S. Xie, “A convnet for the 2020s,” in
2022
Later among the works it cites.
W. Yin, J. Zhang, O. Wang, S. Niklaus, S. Chen, Y. Liu, and C. Shen, “Towards accurate reconstruction of 3d scene shape from a single monocular image,”
2022
Later among the works it cites.
W. Yuan, X. Gu, Z. Dai, S. Zhu, and P. Tan, “New CRFs: Neural window fully-connected CRFs for monocular depth estimation,” in
2022
Later among the works it cites.
J. Ju, C. W. Tseng, O. Bailo, G. Dikov, and M. Ghafoorian, “Dg-recon: Depth-guided neural 3d scene reconstruction,” in
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
H. Zhu, H. Yang, X. Wu, D. Huang, S. Zhang, X. He, T. He, H. Zhao, C. Shen, Y. Qiao,
2023
Later among the works it cites.
2023
Later among the works it cites.
H. Xu, J. Zhang, J. Cai, H. Rezatofighi, F. Yu, D. Tao, and A. Geiger, “Unifying flow, stereo and depth estimation,”
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
C. Zhang, W. Yin, G. Yu, Z. Wang, T. Chen, B. Fu, J. T. Zhou, and C. Shen, “Robust geometry-preserving depth estimation using differentiable rendering,” in
2023
Later among the works it cites.
V. Guizilini, I. Vasiljevic, D. Chen, R. Ambruș, and A. Gaidon, “Towards zero-shot scale-aware monocular depth estimation,” in
2023
Later among the works it cites.
2023
Later among the works it cites.
G. Xu, X. Wang, X. Ding, and X. Yang, “Iterative geometry encoding volume for stereo matching,” in
2023
Later among the works it cites.
2023
Later among the works it cites.
T. Darcet, M. Oquab, J. Mairal, and P. Bojanowski, “Vision transformers need registers,”
2023
Later among the works it cites.
2023
Later among the works it cites.
T. Darcet, M. Oquab, J. Mairal, and P. Bojanowski, “Vision transformers need registers,”
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
X. Yang, L. Yuan, K. Wilber, A. Sharma, X. Gu, S. Qiao, S. Debats, H. Wang, H. Adam, M. Sirotenko,
2024
Closest in time.
L. Piccinelli, Y.-H. Yang, C. Sakaridis, M. Segu, S. Li, L. Van Gool, and F. Yu, “Unidepth: Universal monocular metric depth estimation,” in
2024
Closest in time.
2024
Closest in time.
G. Bae and A. J. Davison, “Rethinking inductive biases for surface normal estimation,” in
2024
Closest in time.