Fetching the paper…
Reading the bibliography…
We present Depth Anything at Any Condition (DepthAnything-AC), a foundation monocular depth estimation (MDE) model capable of handling diverse environmental conditions.
Computer rendering of stochastic models
Alain Fournier, Don Fussell, and Loren Carpenter · 1982
Earlier work this paper cites.
Depth estimation from image structure
Antonio Torralba and Aude Oliva · 2002
Earlier work this paper cites.
Learning depth from single monocular images
Ashutosh Saxena, Sung Chung, and Andrew Ng · 2005
Earlier work this paper cites.
Realtime depth estimation and obstacle detection from monocular video
Andreas Wedel, Uwe Franke, Jens Klappstein, Thomas Brox, and Daniel Cremers · 2006
Earlier work this paper cites.
The pascal visual object classes (voc) challenge
Mark Everingham, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman · 2010
Earlier work this paper cites.
Single image depth estimation from predicted semantic labels
Beyang Liu, Stephen Gould, and Daphne Koller · 2010
Earlier work this paper cites.
A naturalistic open source movie for optical flow evaluation
Daniel J Butler, Jonas Wulff, Garrett B Stanley, and Michael J Black · 2012
Earlier work this paper cites.
Indoor segmentation and support inference from rgbd images
Pushmeet Kohli Nathan Silberman, Derek Hoiem and Rob Fergus · 2012
Earlier work this paper cites.
Vision meets robotics: The kitti dataset
Andreas Geiger, Philip Lenz, Christoph Stiller, and Raquel Urtasun · 2013
Earlier work this paper cites.
Depth map prediction from a single image using a multi-scale deep network
David Eigen, Christian Puhrsch, and Rob Fergus · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Virtual worlds as proxy for multi-object tracking analysis
Adrien Gaidon, Qiao Wang, Yohann Cabon, and Eleonora Vig · 2016
Earlier work this paper cites.
Structure selective depth superresolution for rgb-d cameras
Youngjung Kim, Bumsub Ham, Changjae Oh, and Kwanghoon Sohn · 2016
Earlier work this paper cites.
Unsupervised monocular depth estimation with left-right consistency
Clément Godard, Oisin Mac Aodha, and Gabriel J Brostow · 2017
Earlier work this paper cites.
Deep stereo confidence prediction for depth estimation
Sunok Kim, Dongbo Min, Bumsub Ham, Seungryong Kim, and Kwanghoon Sohn · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
1 year, 1000 km: The oxford robotcar dataset
Will Maddern, Geoffrey Pascoe, Chris Linegar, and Paul Newman · 2017
Earlier work this paper cites.
A multi-view stereo benchmark with high-resolution images and multi-camera videos
Thomas Schöps, Johannes L. Schönberger, Silvano Galliani, Torsten Sattler, Konrad Schindler, Marc Pollefeys, and Andreas Geiger · 2017
Earlier work this paper cites.
Scene parsing through ade20k dataset
Bolei Zhou, Hang Zhao, Xavier Puig, Sanja Fidler, Adela Barriuso, and Antonio Torralba · 2017
Earlier work this paper cites.
Unpaired image-to-image translation using cycle-consistent adversarial networks
Jun-Yan Zhu, Taesung Park, Phillip Isola, and Alexei A Efros · 2017
Earlier work this paper cites.
Deep ordinal regression network for monocular depth estimation
Huan Fu, Mingming Gong, Chaohui Wang, Kayhan Batmanghelich, and Dacheng Tao · 2018
Earlier work this paper cites.
Deep monocular depth estimation via integration of global and local predictions
Youngjung Kim, Hyungjoo Jung, Dongbo Min, and Kwanghoon Sohn · 2018
Earlier work this paper cites.
Megadepth: Learning single-view depth prediction from internet photos
Zhengqi Li and Noah Snavely · 2018
Earlier work this paper cites.
Monocular depth estimation: A survey
Amlaan Bhoi · 2019
Earlier work this paper cites.
Unsupervised scale-consistent depth and ego-motion learning from monocular video
Jiawang Bian, Zhichao Li, Naiyan Wang, Huangying Zhan, Chunhua Shen, Ming-Ming Cheng, and Ian Reid · 2019
Earlier work this paper cites.
Digging into self-supervised monocular depth estimation
Clément Godard, Oisin Mac Aodha, Michael Firman, and Gabriel J Brostow · 2019
Earlier work this paper cites.
DIODE: A Dense Indoor and Outdoor DEpth Dataset
Igor Vasiljevic, Nick Kolkin, Shanyi Zhang, Ruotian Luo, Haochen Wang, Falcon Z. Dai, Andrea F. Daniele, Mohammadreza Mostajabi, Steven Basart, Matthew R. Walter, and Gregory Shakhnarovich · 2019
Earlier work this paper cites.
Drivingstereo: A large-scale dataset for stereo matching in autonomous driving scenarios
Guorun Yang, Xiao Song, Chaoqin Huang, Zhidong Deng, Jianping Shi, and Bolei Zhou · 2019
Earlier work this paper cites.
Pattern-affinitive propagation across depth, surface normal and semantic segmentation
Zhenyu Zhang, Zhen Cui, Chunyan Xu, Yan Yan, Nicu Sebe, and Jian Yang · 2019
Earlier work this paper cites.
Semantic understanding of scenes through the ade20k dataset
Bolei Zhou, Hang Zhao, Xavier Puig, Tete Xiao, Sanja Fidler, Adela Barriuso, and Antonio Torralba · 2019
Earlier work this paper cites.
Virtual kitti 2, 2020
Yohann Cabon, Naila Murray, and Martin Humenberger · 2020
Earlier work this paper cites.
nuscenes: A multimodal dataset for autonomous driving
Holger Caesar, Varun Bankiti, Alex H Lang, Sourabh Vora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom · 2020
Earlier work this paper cites.
Generative adversarial networks
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2020
Earlier work this paper cites.
Defeat-net: General monocular depth via simultaneous unsupervised representation learning
Jaime Spencer, Richard Bowden, and Simon Hadfield · 2020
Earlier work this paper cites.
Unsupervised monocular depth estimation for night-time images using adversarial domain feature adaptation
Madhu Vankadari, Sourav Garg, Anima Majumder, Swagat Kumar, and Ardhendu Behera · 2020
Earlier work this paper cites.
Structure-guided ranking loss for single image depth prediction
Ke Xian, Jianming Zhang, Oliver Wang, Long Mai, Zhe Lin, and Zhiguo Cao · 2020
Earlier work this paper cites.
Monocular depth estimation based on deep learning: An overview
Chaoqiang Zhao, Qiyu Sun, Chongzhen Zhang, Yang Tang, and Feng Qian · 2020
Cited alongside, same era.
Adabins: Depth estimation using adaptive bins
Shariq Farooq Bhat, Ibraheem Alhashim, and Peter Wonka · 2021
Cited alongside, same era.
Auto-rectify network for unsupervised indoor depth estimation
Jia-Wang Bian, Huangying Zhan, Naiyan Wang, Tat-Jin Chin, Chunhua Shen, and Ian Reid · 2021
Cited alongside, same era.
Deep monocular depth estimation leveraging a large-scale outdoor stereo dataset
Jaehoon Cho, Dongbo Min, Youngjung Kim, and Kwanghoon Sohn · 2021
Cited alongside, same era.
Diml/cvl rgb-d dataset: 2m rgb-d images of natural indoor and outdoor scenes
Jaehoon Cho, Dongbo Min, Youngjung Kim, and Kwanghoon Sohn · 2021
Cited alongside, same era.
Depthfm: Fast monocular depth estimation with flow matching, 2024
Ming Gui, Johannes Schusterbauer, Ulrich Prestel, Pingchuan Ma, Dmytro Kotovenko, Olga Grebenkova, Stefan Andreas Baumann, Vincent Tao Hu, and Björn Ommer · 2024
Later among the works it cites.
Lotus: Diffusion-based visual foundation model for high-quality dense prediction
Jing He, Haodong Li, Wei Yin, Yixun Liang, Leheng Li, Kaiqiang Zhou, Hongbo Liu, Bingbing Liu, and Ying-Cong Chen · 2024
Later among the works it cites.
Metric3d v2: A versatile monocular geometric foundation model for zero-shot metric depth and surface normal estimation
Mu Hu, Wei Yin, Chi Zhang, Zhipeng Cai, Xiaoxiao Long, Hao Chen, Kaixuan Wang, Gang Yu, Chunhua Shen, and Shaojie Shen · 2024
Later among the works it cites.
Drivingworld: Constructingworld model for autonomous driving via video gpt
Xiaotao Hu, Wei Yin, Mingkai Jia, Junyuan Deng, Xiaoyang Guo, Qian Zhang, Xiaoxiao Long, and Ping Tan · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Cited alongside, same era.
R4dyn: Exploring radar for self-supervised monocular depth estimation of dynamic scenes
Stefano Gasperini, Patrick Koch, Vinzenz Dallabetta, Nassir Navab, Benjamin Busam, and Federico Tombari · 2021
Cited alongside, same era.
Self-supervised monocular depth estimation for all day images using domain separation
Lina Liu, Xibin Song, Mengmeng Wang, Yong Liu, and Liangjun Zhang · 2021
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models, 2021
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2021
Cited alongside, same era.
Regularizing nighttime weirdness: Efficient self-supervised monocular depth estimation in the dark
Kun Wang, Zhenyu Zhang, Zhiqiang Yan, Xiang Li, Baobei Xu, Jun Li, and Jian Yang · 2021
Cited alongside, same era.
Depth-conditioned dynamic message propagation for monocular 3d object detection
Li Wang, Liang Du, Xiaoqing Ye, Yanwei Fu, Guodong Guo, Xiangyang Xue, Jianfeng Feng, and Li Zhang · 2021
Cited alongside, same era.
Learning to recover 3d scene shape from a single image
Wei Yin, Jianming Zhang, Oliver Wang, Simon Niklaus, Long Mai, Simon Chen, and Chunhua Shen · 2021
Cited alongside, same era.
Bingxin Ke, Anton Obukhov, Shengyu Huang, Nando Metzger, Rodrigo Caye Daudt, and Konrad Schindler · 2024
Later among the works it cites.
Visual robotic manipulation with depth-aware pretraining
Jinming Li, Wanying Wang, Yaxin Peng, Chaomin Shen, Yichen Zhu, and Zhiyuan Xu · 2024
Later among the works it cites.
Manipllm: Embodied multimodal large language model for object-centric robotic manipulation
Xiaoqi Li, Mingxu Zhang, Yiran Geng, Haoran Geng, Yuxing Long, Yan Shen, Renrui Zhang, Jiaming Liu, and Hao Dong · 2024
Later among the works it cites.
Patchfusion: An end-to-end tile-based framework for high-resolution monocular metric depth estimation
Zhenyu Li, Shariq Farooq Bhat, and Peter Wonka · 2024
Later among the works it cites.
Towards raw object detection in diverse conditions
Zhong-Yu Li, Xin Jin, Boyuan Sun, Chun-Le Guo, and Ming-Ming Cheng · 2024
Later among the works it cites.
Prompting depth anything for 4k resolution accurate metric depth estimation
Haotong Lin, Sida Peng, Jingxiao Chen, Songyou Peng, Jiaming Sun, Minghuan Liu, Hujun Bao, Jiashi Feng, Xiaowei Zhou, and Bingyi Kang · 2024
Later among the works it cites.
Depthlab: From partial to complete
Zhiheng Liu, Ka Leong Cheng, Qiuyu Wang, Shuzhe Wang, Hao Ouyang, Bin Tan, Kai Zhu, Yujun Shen, Qifeng Chen, and Ping Luo · 2024
Later among the works it cites.
Stealing stable diffusion prior for robust monocular depth estimation
Yifan Mao, Jian Liu, and Xianming Liu · 2024
Later among the works it cites.
Depth helps: Improving pre-trained rgb-based policy with depth information injection, 2024
Xincheng Pang, Wenke Xia, Zhigang Wang, Bin Zhao, Di Hu, Dong Wang, and Xuelong Li · 2024
Later among the works it cites.
Sharpdepth: Sharpening metric depth predictions using diffusion distillation
Duc-Hai Pham, Tung Do, Phong Nguyen, Binh-Son Hua, Khoi Nguyen, and Rang Nguyen · 2024
Later among the works it cites.
Deep learning-based depth estimation methods from monocular image and videos: A comprehensive survey
Uchitha Rajapaksha, Ferdous Sohel, Hamid Laga, Dean Diepeveen, and Mohammed Bennamoun · 2024
Later among the works it cites.
Corrmatch: Label propagation via correlation matching for semi-supervised semantic segmentation
Boyuan Sun, Yuqi Yang, Le Zhang, Ming-Ming Cheng, and Qibin Hou · 2024
Later among the works it cites.
Diffusion models for monocular depth estimation: Overcoming challenging conditions
Fabio Tosi, Pierluigi Zama Ramirez, and Matteo Poggi · 2024
Later among the works it cites.
Weatherdepth: Curriculum contrastive learning for self-supervised depth estimation under adverse weather conditions
Jiyuan Wang, Chunyu Lin, Lang Nie, Shujun Huang, Yao Zhao, Xing Pan, and Rui Ai · 2024
Later among the works it cites.
Digging into contrastive learning for robust depth estimation with diffusion models
Jiyuan Wang, Chunyu Lin, Lang Nie, Kang Liao, Shuwei Shao, and Yao Zhao · 2024
Later among the works it cites.
Depth anything: Unleashing the power of large-scale unlabeled data
Lihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu, Jiashi Feng, and Hengshuang Zhao · 2024
Later among the works it cites.
Depth anything v2
Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao, Xiaogang Xu, Jiashi Feng, and Hengshuang Zhao · 2024
Later among the works it cites.
Dformer: Rethinking rgbd representation learning for semantic segmentation
Bowen Yin, Xuying Zhang, Zhong-Yu Li, Li Liu, Ming-Ming Cheng, and Qibin Hou · 2024
Later among the works it cites.
Depth-centric dehazing and depth-estimation from real-world hazy driving video
Junkai Fan, Kun Wang, Zhiqiang Yan, Xiang Chen, Shangbing Gao, Jun Li, and Jian Yang · 2025
Closest in time.
Multi-view reconstruction via sfm-guided monocular depth estimation
Haoyu Guo, He Zhu, Sida Peng, Haotong Lin, Yunzhi Yan, Tao Xie, Wenguan Wang, Xiaowei Zhou, and Hujun Bao · 2025
Closest in time.
Distill any depth: Distillation creates a stronger monocular depth estimator
Xiankang He, Dongyan Guo, Hongji Li, Ruibo Li, Ying Cui, and Chi Zhang · 2025
Closest in time.
Realmdreamer: Text-driven 3d scene generation with inpainting and depth diffusion
Jaidev Shriram, Alex Trevithick, Lingjie Liu, and Ravi Ramamoorthi · 2025
Closest in time.
Depthmaster: Taming diffusion models for monocular depth estimation
Ziyang Song, Zerong Wang, Bo Li, Hao Zhang, Ruijie Zhu, Li Liu, Peng-Tao Jiang, and Tianzhu Zhang · 2025
Closest in time.
Llava-scissor: Token compression with semantic connected components for video llms
Boyuan Sun, Jiaxing Zhao, Xihan Wei, and Qibin Hou · 2025
Closest in time.
Vggt: Visual geometry grounded transformer
Jianyuan Wang, Minghao Chen, Nikita Karaev, Andrea Vedaldi, Christian Rupprecht, and David Novotny · 2025
Closest in time.
Tacodepth: Towards efficient radar-camera depth estimation with one-stage fusion
Yiran Wang, Jiaqi Li, Chaoyi Hong, Ruibo Li, Liusheng Sun, Xiao Song, Zhe Wang, Zhiguo Cao, and Guosheng Lin · 2025
Closest in time.
Depthsplat: Connecting gaussian splatting and depth
Haofei Xu, Songyou Peng, Fangjinhua Wang, Hermann Blum, Daniel Barath, Andreas Geiger, and Marc Pollefeys · 2025
Closest in time.
Weilong Yan, Ming Li, Haipeng Li, Shuwei Shao, and Robby T Tan · 2025
Closest in time.
Dformerv2: Geometry self-attention for rgbd semantic segmentation
Bo-Wen Yin, Jiao-Long Cao, Ming-Ming Cheng, and Qibin Hou · 2025
Closest in time.
Ms-nerf: Multi-space neural radiance fields
Ze-Xin Yin, Peng-Yi Jiao, Jiaxiong Qiu, Ming-Ming Cheng, and Bo Ren · 2025
Closest in time.
Llava-octopus: Unlocking instruction-driven adaptive projector fusion for video understanding
Jiaxing Zhao, Boyuan Sun, Xiang Chen, Xihan Wei, and Qibin Hou · 2025
Closest in time.