Fetching the paper…
Reading the bibliography…
We release MiDaS v3.1 for monocular depth estimation, offering a variety of new models based on different encoder backbones.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Indoor segmentation and support inference from rgbd images
Nathan Silberman, Derek Hoiem, Pushmeet Kohli, and Rob Fergus · 2012
Earlier work this paper cites.
A naturalistic open source movie for optical flow evaluation
Daniel J Butler, Jonas Wulff, Garrett B Stanley, and Michael J Black · 2012
Earlier work this paper cites.
A benchmark for the evaluation of rgb-d slam systems
Jürgen Sturm, Nikolas Engelhard, Felix Endres, Wolfram Burgard, and Daniel Cremers · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Object scene flow for autonomous vehicles
Moritz Menze and Andreas Geiger · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Single-image depth perception in the wild
Weifeng Chen, Zhao Fu, Dawei Yang, and Jia Deng · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
A multi-view stereo benchmark with high-resolution images and multi-camera videos
Thomas Schops, Johannes L Schonberger, Silvano Galliani, Torsten Sattler, Konrad Schindler, Marc Pollefeys, and Andreas Geiger · 2017
Earlier work this paper cites.
Monocular relative depth perception with web stereo data supervision
Ke Xian, Chunhua Shen, Zhiguo Cao, Hao Lu, Yang Xiao, Ruibo Li, and Zhenbo Luo · 2018
Earlier work this paper cites.
Exploring the limits of weakly supervised pretraining
Dhruv Kumar Mahajan, Ross B. Girshick, Vignesh Ramanathan, Kaiming He, Manohar Paluri, Yixuan Li, Ashwin Bharambe, and Laurens van der Maaten · 2018
Earlier work this paper cites.
Multi-task learning as multi-objective optimization
Ozan Sener and Vladlen Koltun · 2018
Earlier work this paper cites.
Deep monocular depth estimation via integration of global and local predictions
Youngjung Kim, Hyungjoo Jung, Dongbo Min, and Kwanghoon Sohn · 2018
Earlier work this paper cites.
Megadepth: Learning single-view depth prediction from internet photos
Zhengqi Li and Noah Snavely · 2018
Earlier work this paper cites.
Monocular depth estimation using relative depth maps
Jae-Han Lee and Chang-Su Kim · 2019
Earlier work this paper cites.
Efficientnet: Rethinking model scaling for convolutional neural networks
Mingxing Tan and Quoc V. Le · 2019
Earlier work this paper cites.
Pytorch image models
Ross Wightman · 2019
Earlier work this paper cites.
Efficientnet: Rethinking model scaling for convolutional neural networks
Mingxing Tan and Quoc Le · 2019
Earlier work this paper cites.
Searching for mobilenetv3
Andrew Howard, Mark Sandler, Grace Chu, Liang-Chieh Chen, Bo Chen, Mingxing Tan, Weijun Wang, Yukun Zhu, Ruoming Pang, Vijay Vasudevan, et al · 2019
Earlier work this paper cites.
Web stereo video supervision for depth prediction from dynamic scenes
Chaoyang Wang, Simon Lucey, Federico Perazzi, and Oliver Wang · 2019
Cited alongside, same era.
From big to small: Multi-scale local planar guidance for monocular depth estimation
Jin Han Lee, Myung-Kyu Han, Dong Wook Ko, and Il Hong Suh · 2019
Cited alongside, same era.
Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer
René Ranftl, Katrin Lasinger, David Hafner, Konrad Schindler, and Vladlen Koltun · 2020
Cited alongside, same era.
Self-training with noisy student improves imagenet classification
Qizhe Xie, Minh-Thang Luong, Eduard Hovy, and Quoc V Le · 2020
Cited alongside, same era.
Tartanair: A dataset to push the limits of visual slam
Wenshan Wang, Delong Zhu, Xiangwei Wang, Yaoyu Hu, Yuheng Qiu, Chen Wang, Yafei Hu, Ashish Kapoor, and Sebastian Scherer · 2020
Cited alongside, same era.
A convnet for the 2020s
Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie · 2022
Later among the works it cites.
Metaformer is actually what you need for vision
Weihao Yu, Mi Luo, Pan Zhou, Chenyang Si, Yichen Zhou, Xinchao Wang, Jiashi Feng, and Shuicheng Yan · 2022
Later among the works it cites.
Localbins: Improving depth estimation by learning local distributions
Shariq Farooq Bhat, Ibraheem Alhashim, and Peter Wonka · 2022
Later among the works it cites.
Depth map decomposition for monocular depth estimation
Jinyoung Jun, Jae-Han Lee, Chul Lee, and Chang-Su Kim · 2022
Later among the works it cites.
Binsformer: Revisiting adaptive bins for monocular depth estimation
Zhenyu Li, Xuyang Wang, Xianming Liu, and Junjun Jiang · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Structure-guided ranking loss for single image depth prediction
Ke Xian, Jianming Zhang, Oliver Wang, Long Mai, Zhe Lin, and Zhiguo Cao · 2020
Cited alongside, same era.
The apolloscape open dataset for autonomous driving and its application
Xinyu Huang, Peng Wang, Xinjing Cheng, Dingfu Zhou, Qichuan Geng, and Ruigang Yang · 2020
Cited alongside, same era.
Blendedmvs: A large-scale dataset for generalized multi-view stereo networks
Yao Yao, Zixin Luo, Shiwei Li, Jingyang Zhang, Yufan Ren, Lei Zhou, Tian Fang, and Long Quan · 2020
Cited alongside, same era.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models, 2021
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2021
Cited alongside, same era.
Nerf: Representing scenes as neural radiance fields for view synthesis
Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng · 2021
Cited alongside, same era.
M4depth: Monocular depth estimation for autonomous vehicles in unseen environments
Michaël Fonder, Damien Ernst, and Marc Van Droogenbroeck · 2021
Cited alongside, same era.
Weihao Yuan, Xiaodong Gu, Zuozhuo Dai, Siyu Zhu, and Ping Tan · 2022
Later among the works it cites.
BEiT v2: Masked image modeling with vector-quantized visual tokenizers
Zhiliang Peng, Li Dong, Hangbo Bao, Qixiang Ye, and Furu Wei · 2022
Later among the works it cites.
Swin transformer v2: Scaling up capacity and resolution
Ze Liu, Han Hu, Yutong Lin, Zhuliang Yao, Zhenda Xie, Yixuan Wei, Jia Ning, Yue Cao, Zheng Zhang, Li Dong, Furu Wei, and Baining Guo · 2022
Later among the works it cites.
Jiashi Li, Xin Xia, Wei Li, Huixia Li, Xing Wang, Xuefeng Xiao, Rui Wang, Min Zheng, and Xin Pan · 2022
Later among the works it cites.
Deit iii: Revenge of the vit
Hugo Touvron, Matthieu Cord, and Hervé Jégou · 2022
Later among the works it cites.
Separable self-attention for mobile vision transformers
Sachin Mehta and Mohammad Rastegari · 2022
Later among the works it cites.
Adding conditional control to text-to-image diffusion models
Lvmin Zhang and Maneesh Agrawala · 2023
Closest in time.
Text2room: Extracting textured 3d meshes from 2d text-to-image models
Lukas Höllein, Ang Cao, Andrew Owens, Justin Johnson, and Matthias Nießner · 2023
Closest in time.
Consistentnerf: Enhancing neural radiance fields with 3d consistency for sparse view synthesis
Shoukang Hu, Kaichen Zhou, Kaiyu Li, Longhui Yu, Lanqing Hong, Tianyang Hu, Zhenguo Li, Gim Hee Lee, and Ziwei Liu · 2023
Closest in time.
Improving neural radiance fields with depth-aware optimization for novel view synthesis
Shu Chen, Junyao Li, Yang Zhang, and Beiji Zou · 2023
Closest in time.
Summary of chatgpt/gpt-4 research and perspective towards the future of large language models, 2023
Yiheng Liu, Tianle Han, Siyuan Ma, Jiayue Zhang, Yuanyuan Yang, Jiaming Tian, Hao He, Antong Li, Mengshen He, Zhengliang Liu, Zihao Wu, Dajiang Zhu, Xiang Li, Ning Qiang, Dingang Shen, Tianming Liu, and Bao Ge · 2023
Closest in time.
Image as a foreign language: BEiT pretraining for vision and vision-language tasks
Wenhui Wang, Hangbo Bao, Li Dong, Johan Bjorck, Zhiliang Peng, Qiang Liu, Kriti Aggarwal, Owais Khan Mohammed, Saksham Singhal, Subhojit Som, and Furu Wei · 2023
Closest in time.
Zoedepth: Zero-shot transfer by combining relative and metric depth, 2023
Shariq Farooq Bhat, Reiner Birkl, Diana Wofk, Peter Wonka, and Matthias Müller · 2023
Closest in time.
Monocular Visual-Inertial Depth Estimation
Wofk, Diana and Ranftl, René and Müller, Matthias and Koltun, Vladlen · 2023
Closest in time.
Ldm3d: Latent diffusion model for 3d
Gabriela Ben Melech Stan, Diana Wofk, Scottie Fox, Alex Redden, Will Saxton, Jean Yu, Estelle Aflalo, Shao-Yen Tseng, Fabio Nonato, Matthias Muller, and Vasudev Lal · 2023
Closest in time.