Fetching the paper…
Reading the bibliography…
The success of deep learning in computer vision over the past decade has hinged on large labeled datasets and strong pretrained models.
B. K. Horn, “Shape from shading: A method for obtaining the shape of a smooth opaque object from one view,” 1970
1970
Earlier work this paper cites.
H. Barrow, J. Tenenbaum, A. Hanson, and E. Riseman, “Recovering intrinsic scene characteristics,” Comput. vis. syst , vol. 2, no. 3-26, p. 2, 1978
1978
Earlier work this paper cites.
Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli, “Image quality assessment: from error visibility to structural similarity,” IEEE TIP , vol. 13, no. 4, pp. 600–612, 2004
2004
Earlier work this paper cites.
W. Wagner, A. Ullrich, V. Ducic, T. Melzer, and N. Studnicka, “Gaussian decomposition and calibration of a novel small-footprint full-waveform digitising airborne laser scanner,” ISPRS journal of Photogrammetry and Remote Sensing , vol. 60, no. 2, pp. 100–112, 2006
2006
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in CVPR , 2009
2009
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” NeurIPS , 2012
2012
Earlier work this paper cites.
P. K. Nathan Silberman, Derek Hoiem and R. Fergus, “Indoor segmentation and support inference from RGBD images,” in ECCV , 2012
2012
Earlier work this paper cites.
A. Geiger, P. Lenz, and R. Urtasun, “Are we ready for autonomous driving? the KITTI vision benchmark suite,” in CVPR , 2012
2012
Earlier work this paper cites.
D. J. Butler, J. Wulff, G. B. Stanley, and M. J. Black, “A naturalistic open source movie for optical flow evaluation,” in ECCV , 2012
2012
Earlier work this paper cites.
A. Geiger, P. Lenz, C. Stiller, and R. Urtasun, “Vision meets robotics: The KITTI dataset,” International Journal of Robotics Research , 2013
2013
Earlier work this paper cites.
D. F. Fouhey, A. Gupta, and M. Hebert, “Data-driven 3D primitives for single image understanding,” in ICCV , 2013
2013
Earlier work this paper cites.
J. Yosinski, J. Clune, Y. Bengio, and H. Lipson, “How transferable are features in deep neural networks?” NeurIPS , 2014
2014
Earlier work this paper cites.
D. Eigen, C. Puhrsch, and R. Fergus, “Depth map prediction from a single image using a multi-scale deep network,” in NeurIPS , 2014
2014
Earlier work this paper cites.
L. Ladický, B. Zeisl, and M. Pollefeys, “Discriminatively trained dense surface normal estimation,” in ECCV , 2014
2014
Earlier work this paper cites.
D. Scharstein, H. Hirschmüller, Y. Kitajima, G. Krathwohl, N. Nešić, X. Wang, and P. Westling, “High-resolution stereo datasets with subpixel-accurate ground truth,” in GCPR . Springer, 2014
2014
Earlier work this paper cites.
K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” in ICLR , 2015
2015
Earlier work this paper cites.
J. Long, E. Shelhamer, and T. Darrell, “Fully convolutional networks for semantic segmentation,” in CVPR , 2015
2015
Earlier work this paper cites.
X. Wang, D. Fouhey, and A. Gupta, “Designing deep networks for surface normal estimation,” in CVPR , 2015
2015
Earlier work this paper cites.
D. Eigen and R. Fergus, “Predicting depth, surface normals and semantic labels with a common multi-scale convolutional architecture,” in ICCV , 2015
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in CVPR , 2016
2016
Earlier work this paper cites.
A. Bansal, B. Russell, and A. Gupta, “Marr Revisited: 2D-3D model alignment via surface normal prediction,” in CVPR , 2016
2016
Earlier work this paper cites.
A. Dai, A. X. Chang, M. Savva, M. Halber, T. Funkhouser, and M. Nießner, “ScanNet: Richly-annotated 3d reconstructions of indoor scenes,” in CVPR , 2017
2017
Earlier work this paper cites.
T. Schops, J. L. Schonberger, S. Galliani, T. Sattler, K. Schindler, M. Pollefeys, and A. Geiger, “A multi-view stereo benchmark with high-resolution images and multi-camera videos,” in CVPR , 2017
2017
Earlier work this paper cites.
H. Fu, M. Gong, C. Wang, K. Batmanghelich, and D. Tao, “Deep ordinal regression network for monocular depth estimation,” in CVPR , 2018
2018
Earlier work this paper cites.
Z. Li and N. Snavely, “MegaDepth: Learning single-view depth prediction from internet photos,” in CVPR , 2018
2018
Earlier work this paper cites.
R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang, “The unreasonable effectiveness of deep features as a perceptual metric,” in CVPR , 2018
2018
Earlier work this paper cites.
T. Koch, L. Liebel, F. Fraundorfer, and M. Korner, “Evaluation of cnn-based single-image depth estimation methods,” in ECCV workshop , 2018
2018
Earlier work this paper cites.
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga et al. , “Pytorch: An imperative style, high-performance deep learning library,” NeurIPS , 2019
2019
Earlier work this paper cites.
Y. Song and S. Ermon, “Generative modeling by estimating gradients of the data distribution,” NeurIPS , vol. 32, 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
J. Huang, Y. Zhou, T. Funkhouser, and L. Guibas, “Framenet: Learning local canonical frames of 3d surfaces from a single rgb image,” in ICCV , 2019
2019
Earlier work this paper cites.
T. Karras, S. Laine, and T. Aila, “A style-based generator architecture for generative adversarial networks,” in CVPR , 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
T. Koch, L. Liebel, F. Fraundorfer, and M. Körner, “Evaluation of cnn-based single-image depth estimation methods,” in ECCV workshop , 2019
2019
Earlier work this paper cites.
I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” in ICLR , 2019
2019
Earlier work this paper cites.
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” NeurIPS , 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
R. Ranftl, K. Lasinger, D. Hafner, K. Schindler, and V. Koltun, “Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer,” IEEE PAMI , 2020
2020
Earlier work this paper cites.
Y. Cabon, N. Murray, and M. Humenberger, “Virtual KITTI 2,” arXiv preprint arXiv:2001.10773 , 2020
2020
Earlier work this paper cites.
Z. Li, M. Shafiei, R. Ramamoorthi, K. Sunkavalli, and M. Chandraker, “Inverse rendering for complex indoor scenes: Shape, spatially-varying lighting and svbrdf from a single image,” in CVPR , 2020
2020
Earlier work this paper cites.
W. Chen, S. Qian, D. Fan, N. Kojima, M. Hamilton, and J. Deng, “OASIS: A large-scale dataset for single image 3d in the wild,” in CVPR , 2020
2020
Earlier work this paper cites.
A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. Voss, A. Radford, M. Chen, and I. Sutskever, “Zero-shot text-to-image generation,” OpenAI, 2021
2021
Earlier work this paper cites.
J. Song, C. Meng, and S. Ermon, “Denoising diffusion implicit models,” in ICLR , 2021
2021
Earlier work this paper cites.
Y. Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole, “Score-based generative modeling through stochastic differential equations,” in ICLR , 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
S. F. Bhat, I. Alhashim, and P. Wonka, “AdaBins: Depth estimation using adaptive bins,” in CVPR , 2021
2021
Cited alongside, same era.
G. Yang, H. Tang, M. Ding, N. Sebe, and E. Ricci, “Transformer-based attention networks for continuous pixel-wise prediction,” in ICCV , 2021
2021
Cited alongside, same era.
S. Aich, J. M. U. Vianney, M. A. Islam, M. Kaur, and B. Liu, “Bidirectional attention network for monocular depth estimation,” in ICRA , 2021
Z. Li, Z. Chen, X. Liu, and J. Jiang, “Depthformer: Exploiting long-range correlation and local information for accurate monocular depth estimation,” Machine Intelligence Research , pp. 1–18, 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
W. Yin, C. Zhang, H. Chen, Z. Cai, G. Yu, K. Wang, X. Chen, and C. Shen, “Metric3D: Towards zero-shot metric 3d prediction from a single image,” in ICCV , 2023
2023
Later among the works it cites.
V. Guizilini, I. Vasiljevic, D. Chen, R. Ambruș, and A. Gaidon, “Towards zero-shot scale-aware monocular depth estimation,” in ICCV , 2023
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
R. Ranftl, A. Bochkovskiy, and V. Koltun, “Vision transformers for dense prediction,” in ICCV , 2021
2021
Cited alongside, same era.
M. Roberts, J. Ramapuram, A. Ranjan, A. Kumar, M. A. Bautista, N. Paczan, R. Webb, and J. M. Susskind, “Hypersim: A photorealistic synthetic dataset for holistic indoor scene understanding,” in ICCV , 2021
2021
Cited alongside, same era.
G. Bae, I. Budvytis, and R. Cipolla, “Estimating and exploiting the aleatoric uncertainty in surface normal estimation,” in ICCV , 2021
2021
Cited alongside, same era.
Z. Wang, J. Philion, S. Fidler, and J. Kautz, “Learning indoor inverse rendering with 3d spatially-varying lighting,” in ICCV , 2021
2021
Cited alongside, same era.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in ICML , 2021, pp. 8748–8763
2021
Cited alongside, same era.
S. M. H. Miangoleh, S. Dille, L. Mai, S. Paris, and Y. Aksoy, “Boosting monocular depth estimation models to high-resolution via content-adaptive multi-resolution merging,” in CVPR , 2021
2021
Cited alongside, same era.
A. Eftekhar, A. Sax, J. Malik, and A. Zamir, “Omnidata: A scalable pipeline for making multi-task mid-level vision datasets from 3d scans,” in ICCV , 2021
2021
Cited alongside, same era.
2023
Later among the works it cites.
Y. Lipman, R. T. Q. Chen, H. Ben-Hamu, M. Nickel, and M. Le, “Flow matching for generative modeling,” in ICLR , 2023
2023
Later among the works it cites.
——, “Intrinsic image decomposition via ordinal shading,” ACM Transactions on Graphics , vol. 43, no. 1, pp. 1–24, 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
O. Bar-Tal, L. Yariv, Y. Lipman, and T. Dekel, “Multidiffusion: fusing diffusion paths for controlled image generation,” in ICML , 2023
2023
Later among the works it cites.
S. Huang, Z. Gojcic, Z. Wang, F. Williams, Y. Kasten, S. Fidler, K. Schindler, and O. Litany, “Neural lidar fields for novel view synthesis,” in ICCV , 2023
2023
Later among the works it cites.
B. Ke, A. Obukhov, S. Huang, N. Metzger, R. C. Daudt, and K. Schindler, “Repurposing diffusion-based image generators for monocular depth estimation,” in CVPR , 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
O. B. Bohan, “Tiny autoencoder for stable diffusion,” https://hf.co/madebyollin/taesd , 2024, last accessed 15.06.2024
2024
Later among the works it cites.
Z. Li, X. Wang, X. Liu, and J. Jiang, “Binsformer: Revisiting adaptive bins for monocular depth estimation,” IEEE TIP , vol. 33, pp. 3964–3976, 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
L. Yang, B. Kang, Z. Huang, X. Xu, J. Feng, and H. Zhao, “Depth anything: Unleashing the power of large-scale unlabeled data,” in CVPR , 2024
2024
Later among the works it cites.
M. Oquab, T. Darcet, T. Moutakanni, H. Vo, M. Szafraniec, V. Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby et al. , “DINOv2: Learning robust visual features without supervision,” Transactions on Machine Learning Research , 2024
2024
Later among the works it cites.
Y. Duan, X. Guo, and Z. Zhu, “DiffusionDepth: Diffusion denoising approach for monocular depth estimation,” in ECCV , 2024
2024
Later among the works it cites.
X. Fu, W. Yin, M. Hu, K. Wang, Y. Ma, P. Tan, S. Shen, D. Lin, and X. Long, “Geowizard: Unleashing the diffusion priors for 3d geometry estimation from a single image,” in ECCV , 2024
2024
Later among the works it cites.
X. Zhang, B. Ke, H. Riemenschneider, N. Metzger, A. Obukhov, M. Gross, K. Schindler, and C. Schroers, “Betterdepth: Plug-and-play diffusion refiner for zero-shot monocular depth estimation,” NeurIPS , 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
F. Tosi, P. Z. Ramirez, and M. Poggi, “Diffusion models for monocular depth estimation: Overcoming challenging conditions,” in ECCV , 2024
2024
Later among the works it cites.
Y. Jia, L. Hoyer, S. Huang, T. Wang, L. V. Gool, K. Schindler, and A. Obukhov, “Dginstyle: Domain-generalizable semantic segmentation with image diffusion models and stylized semantic control,” in ECCV , 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
G. Bae and A. J. Davison, “Rethinking inductive biases for surface normal estimation,” in CVPR , 2024
2024
Later among the works it cites.
C. Ye, L. Qiu, X. Gu, Q. Zuo, Y. Wu, Z. Dong, L. Bo, Y. Xiu, and X. Han, “Stablenormal: Reducing diffusion variance for stable and sharp normal,” ACM Transactions on Graphics (TOG) , 2024
2024
Later among the works it cites.
J. Luo, D. Ceylan, J. S. Yoon, N. Zhao, J. Philip, A. Frühstück, W. Li, C. Richardt, and T. Wang, “Intrinsicdiffusion: Joint intrinsic layers from latent diffusion models,” in ACM SIGGRAPH , 2024, pp. 1–11
2024
Later among the works it cites.
C. Careaga and Y. Aksoy, “Colorful diffuse intrinsic image decomposition in the wild,” ACM Trans. Graph. , vol. 43, no. 6, 2024
2024
Later among the works it cites.
A. Bhattad, D. McKee, D. Hoiem, and D. Forsyth, “Stylegan knows normal, depth, albedo, and more,” NeurIPS , 2024
2024
Later among the works it cites.
H.-Y. Lee, H.-Y. Tseng, and M.-H. Yang, “Exploiting diffusion prior for generalizable dense prediction,” in CVPR , 2024
2024
Later among the works it cites.
P. Kocsis, V. Sitzmann, and M. Nießner, “Intrinsic image diffusion for indoor single-view material estimation,” in CVPR , 2024
2024
Later among the works it cites.
Z. Zeng, V. Deschaintre, I. Georgiev, Y. Hold-Geoffroy, Y. Hu, F. Luan, L.-Q. Yan, and M. Hašan, “Rgb-x: Image decomposition and synthesis using material-and lighting-aware diffusion models,” in SIGGRAPH conference , 2024
2024
Later among the works it cites.
Z. Li, S. F. Bhat, and P. Wonka, “Patchfusion: An end-to-end tile-based framework for high-resolution monocular metric depth estimation,” in CVPR , 2024
2024
Later among the works it cites.
——, “Patchrefiner: Leveraging synthetic data for real-domain high-resolution monocular metric depth estimation,” in ECCV , 2024
2024
Later among the works it cites.
A. Bochkovskii, A. Delaunoy, H. Germain, M. Santos, Y. Zhou, S. R. Richter, and V. Koltun, “Depth pro: Sharp monocular metric depth in less than a second,” arXiv , 2024
2024
Later among the works it cites.
Y. Song and P. Dhariwal, “Improved techniques for training consistency models,” in ICLR , 2024
2024
Later among the works it cites.
P. Z. Ramirez, F. Tosi, L. Di Stefano, R. Timofte, A. Costanzino, M. Poggi, S. Salti, S. Mattoccia, Y. Zhang, C. Wu et al. , “Ntire 2024 challenge on hr depth from images of specular and transparent surfaces,” in CVPR , 2024
2024
Later among the works it cites.