Fetching the paper…
Reading the bibliography…
We present a foundation model for zero-shot metric monocular depth estimation.
Numerische Isotropieoptimierung von FIR-Filtern mittels Querglättung
Hanno Scharr, Stefan Körkel, and Bernd Jähne · 1997
Earlier work this paper cites.
Variation and extrema of human interpupillary distance
Neil A. Dodgson · 2004
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Make3D: Learning 3D scene structure from a single still image
Ashutosh Saxena, Min Sun, and Andrew Y. Ng · 2009
Earlier work this paper cites.
Learning photographic global tonal adjustment with a database of input / output image pairs
Vladimir Bychkovsky, Sylvain Paris, Eric Chan, and Frédo Durand · 2011
Earlier work this paper cites.
A naturalistic open source movie for optical flow evaluation
Daniel J. Butler, Jonas Wulff, Garrett B. Stanley, and Michael J. Black · 2012
Earlier work this paper cites.
Indoor segmentation and support inference from RGBD images
Nathan Silberman, Derek Hoiem, Pushmeet Kohli, and Rob Fergus · 2012
Earlier work this paper cites.
Vision meets robotics: The KITTI dataset
Andreas Geiger, Philip Lenz, Christoph Stiller, and Raquel Urtasun · 2013
Earlier work this paper cites.
Depth map prediction from a single image using a multi-scale deep network
David Eigen, Christian Puhrsch, and Rob Fergus · 2014
Earlier work this paper cites.
Pulling things out of perspective
Lubor Ladicky, Jianbo Shi, and Marc Pollefeys · 2014
Earlier work this paper cites.
High-resolution stereo datasets with subpixel-accurate ground truth
Daniel Scharstein, Heiko Hirschmüller, York Kitajima, Greg Krathwohl, Nera Nesic, Xi Wang, and Porter Westling · 2014
Earlier work this paper cites.
RAISE: A raw images dataset for digital image forensics
Duc-Tien Dang-Nguyen, Cecilia Pasquini, Valentina Conotter, and Giulia Boato · 2015
Earlier work this paper cites.
Predicting depth, surface normals and semantic labels with a common multi-scale convolutional architecture
David Eigen and Rob Fergus · 2015
Earlier work this paper cites.
SUN RGB-D: A RGB-D scene understanding benchmark suite
Shuran Song, Samuel P. Lichtenberg, and Jianxiong Xiao · 2015
Earlier work this paper cites.
Single-image depth perception in the wild
Weifeng Chen, Zhao Fu, Dawei Yang, and Jia Deng · 2016
Earlier work this paper cites.
Virtual worlds as proxy for multi-object tracking analysis
Adrien Gaidon, Qiao Wang, Yohann Cabon, and Eleonora Vig · 2016
Earlier work this paper cites.
Structure selective depth superresolution for RGB-D cameras
Youngjung Kim, Bumsub Ham, Changjae Oh, and Kwanghoon Sohn · 2016
Earlier work this paper cites.
YFCC100M: The new data in multimedia research
Bart Thomee, David A Shamma, Gerald Friedland, Benjamin Elizalde, Karl Ni, Douglas Poland, Damian Borth, and Li-Jia Li · 2016
Earlier work this paper cites.
ScanNet: Richly-annotated 3D reconstructions of indoor scenes
Angela Dai, Angel X. Chang, Manolis Savva, Maciej Halber, Thomas A. Funkhouser, and Matthias Nießner · 2017
Earlier work this paper cites.
Casual 3D photography
Peter Hedman, Suhib Alsisan, Richard Szeliski, and Johannes Kopf · 2017
Earlier work this paper cites.
A multi-view stereo benchmark with high-resolution images and multi-camera videos
Thomas Schöps, Johannes L. Schönberger, S. Galliani, Torsten Sattler, Konrad Schindler, Marc Pollefeys, and Andreas Geiger · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Deep ordinal regression network for monocular depth estimation
Huan Fu, Mingming Gong, Chaohui Wang, Kayhan Batmanghelich, and Dacheng Tao · 2018
Earlier work this paper cites.
DeepMVS: Learning multi-view stereopsis
Po-Han Huang, Kevin Matzen, Johannes Kopf, Narendra Ahuja, and Jia-Bin Huang · 2018
Earlier work this paper cites.
Evaluation of CNN-based single-image depth estimation methods
Tobias Koch, Lukas Liebel, Friedrich Fraundorfer, and Marco Körner · 2018
Earlier work this paper cites.
MegaDepth: Learning single-view depth prediction from internet photos
Zhengqi Li and Noah Snavely · 2018
Earlier work this paper cites.
Synscapes: A photorealistic synthetic dataset for street scene parsing
Magnus Wrenninge and Jonas Unger · 2018
Earlier work this paper cites.
Monocular relative depth perception with web stereo data supervision
Ke Xian, Chunhua Shen, Zhiguo Cao, Hao Lu, Yang Xiao, Ruibo Li, and Zhenbo Luo · 2018
Earlier work this paper cites.
UASOL, a large-scale high-resolution outdoor stereo dataset
Zuria Bauer, Francisco Gomez-Donoso, Edmanuel Cruz, Sergio Orts-Escolano, and Miguel Cazorla · 2019
Earlier work this paper cites.
CAM-Convs: Camera-aware multi-scale convolutions for single-view depth
José M. Fácil, Benjamin Ummenhofer, Huizhong Zhou, Luis Montesano, Thomas Brox, and Javier Civera · 2019
Earlier work this paper cites.
3D Ken Burns effect from a single image
Simon Niklaus, Long Mai, Jimei Yang, and Feng Liu · 2019
Earlier work this paper cites.
SharpNet: Fast and accurate recovery of occluding contours in monocular depth estimation
Michaël Ramamonjisoa and Vincent Lepetit · 2019
Earlier work this paper cites.
IRS: A large synthetic indoor robotics stereo dataset for disparity and surface normal estimation
Qiang Wang, Shizhen Zheng, Qingsong Yan, Fei Deng, Kaiyong Zhao, and Xiaowen Chu · 2019
Earlier work this paper cites.
Pytorch image models
Ross Wightman · 2019
Earlier work this paper cites.
Zoom to learn, learn to zoom
Xuaner Zhang, Qifeng Chen, Ren Ng, and Vladlen Koltun · 2019
Earlier work this paper cites.
Defocus deblurring using dual-pixel data
Abdullah Abuolaim and Michael S Brown · 2020
Earlier work this paper cites.
Height and uprightness invariance for 3D prediction from a single view
Manel Baradad and Antonio Torralba · 2020
Earlier work this paper cites.
nuScenes: A multimodal dataset for autonomous driving
Holger Caesar, Varun Bankiti, Alex H. Lang, Sourabh Vora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom · 2020
Cited alongside, same era.
Perceptual quality assessment of smartphone photography
Yuming Fang, Hanwei Zhu, Yan Zeng, Kede Ma, and Zhou Wang · 2020
Cited alongside, same era.
3D packing for self-supervised monocular depth estimation
Vitor Guizilini, Rares Ambrus, Sudeep Pillai, Allan Raventos, and Adrien Gaidon · 2020
Cited alongside, same era.
The ApolloScape open dataset for autonomous driving and its application
Xinyu Huang, Peng Wang, Xinjing Cheng, Dingfu Zhou, Qichuan Geng, and Ruigang Yang · 2020
Cited alongside, same era.
Predicting sharp and accurate occlusion boundaries in monocular depth estimation using displacement fields
Michaël Ramamonjisoa, Yuming Du, and Vincent Lepetit · 2020
Cited alongside, same era.
MPViT: Multi-path vision transformer for dense prediction
Youngwan Lee, Jonghee Kim, Jeffrey Willette, and Sung Ju Hwang · 2022
Later among the works it cites.
Exploiting pseudo labels in a self-supervised learning framework for improved monocular depth estimation
Andra Petrovai and Sergiu Nedevschi · 2022
Later among the works it cites.
Highly accurate dichotomous image segmentation
Xuebin Qin, Hang Dai, Xiaobin Hu, Deng-Ping Fan, Ling Shao, and Luc Van Gool · 2022
Later among the works it cites.
Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer
René Ranftl, Katrin Lasinger, David Hafner, Konrad Schindler, and Vladlen Koltun · 2022
Later among the works it cites.
Deconstructing self-supervised monocular reconstruction: The design decisions that matter
Jaime Spencer, Chris Russell, Simon Hadfield, and Richard Bowden · 2022
Later among the works it cites.
DeiT III: Revenge of the ViT
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Meng-Li Shih, Shih-Yang Su, Johannes Kopf, and Jia-Bin Huang · 2020
Cited alongside, same era.
TartanAir: A dataset to push the limits of visual SLAM
Wenshan Wang, Delong Zhu, Xiangwei Wang, Yaoyu Hu, Yuheng Qiu, Chen Wang, Yafei Hu, Ashish Kapoor, and Sebastian A. Scherer · 2020
Cited alongside, same era.
Structure-guided ranking loss for single image depth prediction
Ke Xian, Jianming Zhang, Oliver Wang, Long Mai, Zhe Lin, and Zhiguo Cao · 2020
Cited alongside, same era.
BlendedMVS: A large-scale dataset for generalized multi-view stereo networks
Yao Yao, Zixin Luo, Shiwei Li, Jingyang Zhang, Yufan Ren, Lei Zhou, Tian Fang, and Long Quan · 2020
Cited alongside, same era.
AdaBins: Depth estimation using adaptive bins
Shariq Farooq Bhat, Ibraheem Alhashim, and Peter Wonka · 2021
Cited alongside, same era.
CrossViT: Cross-attention multi-scale vision transformer for image classification
Chun-Fu (Richard) Chen, Quanfu Fan, and Rameswar Panda · 2021
Cited alongside, same era.
Twins: Revisiting the design of spatial attention in vision transformers
Xiangxiang Chu, Zhi Tian, Yuqing Wang, Bo Zhang, Haibing Ren, Xiaolin Wei, Huaxia Xia, and Chunhua Shen · 2021
Cited alongside, same era.
Hugo Touvron, Matthieu Cord, and Hervé Jégou · 2022
Later among the works it cites.
ZoeDepth: Zero-shot transfer by combining relative and metric depth
Shariq Farooq Bhat, Reiner Birkl, Diana Wofk, Peter Wonka, and Matthias Müller · 2023
Later among the works it cites.
MiDaS v3.1 - A model zoo for robust monocular relative depth estimation
Reiner Birkl, Diana Wofk, and Matthias Müller · 2023
Later among the works it cites.
BEDLAM: A synthetic dataset of bodies exhibiting detailed lifelike animated motion
Michael J. Black, Priyanka Patel, Joachim Tesch, and Jinlong Yang · 2023
Later among the works it cites.
EfficientViT: Lightweight multi-scale attention for high-resolution dense prediction
Han Cai, Junyan Li, Muyan Hu, Chuang Gan, and Song Han · 2023
Later among the works it cites.
Vision transformer adapter for dense predictions
Zhe Chen, Yuchen Duan, Wenhai Wang, Junjun He, Tong Lu, Jifeng Dai, and Yu Qiao · 2023
Later among the works it cites.
All for one, and one for all: UrbanSyn dataset, the third musketeer of synthetic driving scenes
Jose Luis Gómez, Manuel Silva, Antonio Seoane, Agnès Borràs, Mario Noriega, Germán Ros, José Antonio Iglesias Guitián, and Antonio M. López · 2023
Later among the works it cites.
Towards zero-shot scale-aware monocular depth estimation
Vitor Guizilini, Igor Vasiljevic, Dian Chen, Rares Ambrus, and Adrien Gaidon · 2023
Later among the works it cites.
DynamicStereo: Consistent dynamic depth from stereo videos
Nikita Karaev, Ignacio Rocco, Benjamin Graham, Natalia Neverova, Andrea Vedaldi, and Christian Rupprecht · 2023
Later among the works it cites.
Tiled multiplane images for practical 3D photography
Numair Khan, Lei Xiao, and Douglas Lanman · 2023
Later among the works it cites.
Segment anything
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C. Berg, Wan-Yen Lo, Piotr Dollar, and Ross Girshick · 2023
Later among the works it cites.
Deep image matting: A comprehensive survey
Jizhizi Li, Jing Zhang, and Dacheng Tao · 2023
Later among the works it cites.
EfficientViT: Memory efficient vision transformer with cascaded group attention
Xinyu Liu, Houwen Peng, Ningxin Zheng, Yuqing Yang, Han Hu, and Yixuan Yuan · 2023
Later among the works it cites.
Spring: A high-resolution high-detail dataset and benchmark for scene flow, optical flow and stereo
Lukas Mehl, Jenny Schmalfuss, Azin Jahedi, Yaroslava Nalivayko, and Andrés Bruhn · 2023
Later among the works it cites.
Guided depth super-resolution by deep anisotropic diffusion
Nando Metzger, Rodrigo Caye Daudt, and Konrad Schindler · 2023
Later among the works it cites.
Zero-shot metric depth with a field-of-view conditioned diffusion model
Saurabh Saxena, Junhwa Hur, Charles Herrmann, Deqing Sun, and David J. Fleet · 2023
Later among the works it cites.
Kick back & relax: Learning to reconstruct the world by watching SlowTV
Jaime Spencer, Simon Hadfield, Chris Russell, and Richard Bowden · 2023
Later among the works it cites.
EVA-CLIP: Improved training techniques for CLIP at scale
Quan Sun, Yuxin Fang, Ledell Wu, Xinlong Wang, and Yue Cao · 2023
Later among the works it cites.
ConvNeXt V2: Co-Designing and Scaling ConvNets With Masked Autoencoders
Sanghyun Woo, Shoubhik Debnath, Ronghang Hu, Xinlei Chen, Zhuang Liu, In So Kweon, and Saining Xie · 2023
Later among the works it cites.
Metric3D: Towards zero-shot metric 3D prediction from a single image
Wei Yin, Chi Zhang, Hao Chen, Zhipeng Cai, Gang Yu, Kaixuan Wang, Xiaozhi Chen, and Chunhua Shen · 2023
Later among the works it cites.
Sigmoid loss for language image pre-training
Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer · 2023
Later among the works it cites.
Guided depth map super-resolution: A survey
Zhiwei Zhong, Xianming Liu, Junjun Jiang, Debin Zhao, and Xiangyang Ji · 2023
Later among the works it cites.
Metric3D v2: A versatile monocular geometric foundation model for zero-shot metric depth and surface normal estimation
Mu Hu, Wei Yin, Chi Zhang, Zhipeng Cai, Xiaoxiao Long, Hao Chen, Kaixuan Wang, Gang Yu, Chunhua Shen, and Shaojie Shen · 2024
Closest in time.
Repurposing diffusion-based image generators for monocular depth estimation
Bingxin Ke, Anton Obukhov, Shengyu Huang, Nando Metzger, Rodrigo Caye Daudt, and Konrad Schindler · 2024
Closest in time.
DINOv2: Learning robust visual features without supervision
Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, Mahmoud Assran, Nicolas Ballas, Wojciech Galuba, Russell Howes, Po-Yao Huang, Shang-Wen Li, Ishan Misra, Michael Rabbat, Vasu Sharma, Gabriel Synnaeve, Hu Xu, Hervé Jegou, Julien Mairal, Patrick Labatut, Armand Joulin, and Piotr Bojanowski · 2024
Closest in time.
UniDepth: Universal monocular metric depth estimation
Luigi Piccinelli, Yung-Hsu Yang, Christos Sakaridis, Mattia Segù, Siyuan Li, Luc Van Gool, and Fisher Yu · 2024
Closest in time.
Booster: A benchmark for depth from images of specular and transparent surfaces
Pierluigi Zama Ramirez, Alex Costanzino, Fabio Tosi, Matteo Poggi, Samuele Salti, Stefano Mattoccia, and Luigi Di Stefano · 2024
Closest in time.
Kick back & relax++: Scaling beyond ground-truth depth with SlowTV & CribsTV
Jaime Spencer, Chris Russell, Simon Hadfield, and Richard Bowden · 2024
Closest in time.
ViT-CoMer: Vision transformer with convolutional multi-scale feature interaction for dense predictions
Chunlong Xia, Xinliang Wang, Feng Lv, Xin Hao, and Yifeng Shi · 2024
Closest in time.
Metaformer baselines for vision
Weihao Yu, Chenyang Si, Pan Zhou, Mi Luo, Yichen Zhou, Jiashi Feng, Shuicheng Yan, and Xinchao Wang · 2024
Closest in time.
DepthFM: Fast monocular depth estimation with flow matching
Ming Gui, Johannes S. Fischer, Ulrich Prestel, Pingchuan Ma, Dmytro Kotovenko, Olga Grebenkova, Stefan Andreas Baumann, Vincent Tao Hu, and Björn Ommer · 2025
Closest in time.