Fetching the paper…
Reading the bibliography…
Three-dimensional (3D) reconstruction from a single image is an ill-posed problem with inherent ambiguities, i.e.
Unsupervised depth completion from visual inertial odometry
Alex Wong, Xiaohan Fei, Stephanie Tsuei, and Stefano Soatto · 1906
Earlier work this paper cites.
From big to small: Multi-scale local planar guidance for monocular depth estimation
Jin Han Lee, Myung-Kyu Han, Dong Wook Ko, and Il Hong Suh · 1907
Earlier work this paper cites.
From big to small: Multi-scale local planar guidance for monocular depth estimation
Jin Han Lee, Myung-Kyu Han, Dong Wook Ko, and Il Hong Suh · 1907
Earlier work this paper cites.
Surface versus edge-based determinants of visual recognition
Irving Biederman and Ginny Ju · 1988
Earlier work this paper cites.
The importance of shape in early lexical learning
Barbara Landau, Linda B Smith, and Susan S Jones · 1988
Earlier work this paper cites.
Object shape, object function, and object name
Barbara Landau, Linda Smith, and Susan Jones · 1998
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Deep ordinal regression network for monocular depth estimation
Huan Fu, Mingming Gong, Chaohui Wang, Kayhan Batmanghelich, and Dacheng Tao · 2011
Earlier work this paper cites.
Deep ordinal regression network for monocular depth estimation
Huan Fu, Mingming Gong, Chaohui Wang, Kayhan Batmanghelich, and Dacheng Tao · 2011
Earlier work this paper cites.
Are we ready for autonomous driving? the kitti vision benchmark suite
Andreas Geiger, Philip Lenz, and Raquel Urtasun · 2012
Earlier work this paper cites.
Indoor segmentation and support inference from rgbd images
Nathan Silberman, Derek Hoiem, Pushmeet Kohli, and Rob Fergus · 2012
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2013
Earlier work this paper cites.
Depth map prediction from a single image using a multi-scale deep network
David Eigen, Christian Puhrsch, and Rob Fergus · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Sun rgb-d: A rgb-d scene understanding benchmark suite
Shuran Song, Samuel P Lichtenberg, and Jianxiong Xiao · 2015
Earlier work this paper cites.
Unsupervised cnn for single view depth estimation: Geometry to the rescue
Ravi Garg, Vijay Kumar Bg, Gustavo Carneiro, and Ian Reid · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Unsupervised monocular depth estimation with left-right consistency
Clément Godard, Oisin Mac Aodha, and Gabriel J Brostow · 2017
Earlier work this paper cites.
Unsupervised learning of depth and ego-motion from video
Tinghui Zhou, Matthew Brown, Noah Snavely, and David G Lowe · 2017
Earlier work this paper cites.
Codeslam—learning a compact, optimisable representation for dense visual slam
Michael Bloesch, Jan Czarnowski, Ronald Clark, Stefan Leutenegger, and Andrew J Davison · 2018
Earlier work this paper cites.
Visual-inertial object detection and mapping
Xiaohan Fei and Stefano Soatto · 2018
Earlier work this paper cites.
Geonet: Geometric neural network for joint depth and surface normal estimation
Xiaojuan Qi, Renjie Liao, Zhengzhe Liu, Raquel Urtasun, and Jiaya Jia · 2018
Earlier work this paper cites.
Geo-supervised visual depth prediction
Xiaohan Fei, Alex Wong, and Stefano Soatto · 2019
Earlier work this paper cites.
Bilateral cyclic constraint and adaptive regularization for unsupervised monocular depth prediction
Alex Wong and Stefano Soatto · 2019
Earlier work this paper cites.
Dense depth posterior (ddp) from single image and sparse range
Yanchao Yang, Alex Wong, and Stefano Soatto · 2019
Earlier work this paper cites.
Enforcing geometric constraints of virtual normal for depth prediction
Wei Yin, Yifan Liu, Chunhua Shen, and Youliang Yan · 2019
Earlier work this paper cites.
Moving indoor: Unsupervised video depth learning in challenging environments
Junsheng Zhou, Yuwang Wang, Kaihuai Qin, and Wenjun Zeng · 2019
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Earlier work this paper cites.
Generating and exploiting probabilistic monocular depth estimates
Zhihao Xia, Patrick Sullivan, and Ayan Chakrabarti · 2020
Cited alongside, same era.
P 2 net: Patch-match and plane-regularization for unsupervised indoor depth estimation
Zehao Yu, Lei Jin, and Shenghua Gao · 2020
Cited alongside, same era.
Towards better generalization: Joint depth-pose learning without posenet
Wang Zhao, Shaohui Liu, Yezhi Shu, and Yong-Jin Liu · 2020
Cited alongside, same era.
Adabins: Depth estimation using adaptive bins
Shariq Farooq Bhat, Ibraheem Alhashim, and Peter Wonka · 2021
Cited alongside, same era.
Unsupervised scale-consistent depth learning from video
Jia-Wang Bian, Huangying Zhan, Naiyan Wang, Zhichao Li, Le Zhang, Chunhua Shen, Ming-Ming Cheng, and Ian Reid · 2021
Cited alongside, same era.
Emerging properties in self-supervised vision transformers
Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin · 2021
Monovit: Self-supervised monocular depth estimation with a vision transformer
Chaoqiang Zhao, Youmin Zhang, Matteo Poggi, Fabio Tosi, Xianda Guo, Zheng Zhu, Guan Huang, Yang Tang, and Stefano Mattoccia · 2022
Later among the works it cites.
Detecting twenty-thousand classes using image-level supervision
Xingyi Zhou, Rohit Girdhar, Armand Joulin, Phillip Krähenbühl, and Ishan Misra · 2022
Later among the works it cites.
Learning to prompt clip for monocular depth estimation: Exploring the limits of human language
Dylan Auty and Krystian Mikolajczyk · 2023
Later among the works it cites.
iquery: Instruments as queries for audio-visual sound separation
Jiaben Chen, Renrui Zhang, Dongze Lian, Jiaqi Yang, Ziyao Zeng, and Jianbo Shi · 2023
Later among the works it cites.
Imagebind: One embedding space to bind them all
Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, and Ishan Misra · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Transformer-based monocular depth estimation with attention supervision
Wenjie Chang, Yueyi Zhang, and Zhiwei Xiong · 2021
Cited alongside, same era.
Clip-adapter: Better vision-language models with feature adapters
Peng Gao, Shijie Geng, Renrui Zhang, Teli Ma, Rongyao Fang, Yongfeng Zhang, Hongsheng Li, and Yu Qiao · 2021
Cited alongside, same era.
Monoindoor: Towards good practice of self-supervised monocular depth estimation for indoor environments
Pan Ji, Runze Li, Bir Bhanu, and Yi Xu · 2021
Cited alongside, same era.
Structdepth: Leveraging the structural regularities for self-supervised indoor depth estimation
Boying Li, Yuan Huang, Zeyu Liu, Danping Zou, and Wenxian Yu · 2021
Cited alongside, same era.
Swin transformer: Hierarchical vision transformer using shifted windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo · 2021
Cited alongside, same era.
Adaptive surface normal constraint for depth estimation
Xiaoxiao Long, Cheng Lin, Lingjie Liu, Wei Li, Christian Theobalt, Ruigang Yang, and Wenping Wang · 2021
Cited alongside, same era.
Ziyu Guo, Renrui Zhang, Xiangyang Zhu, Yiwen Tang, Xianzheng Ma, Jiaming Han, Kexin Chen, Peng Gao, Xianzhi Li, Hongsheng Li, et al · 2023
Later among the works it cites.
Learning to adapt clip for few-shot monocular depth estimation
Xueting Hu, Ce Zhang, Yi Zhang, Bowen Hai, Ke Yu, and Zhihai He · 2023
Later among the works it cites.
Text-image alignment for diffusion-based perception
Neehar Kondapaneni, Markus Marks, Manuel Knott, Rogério Guimarães, and Pietro Perona · 2023
Later among the works it cites.
Va-depthnet: A variational approach to single image depth prediction
Ce Liu, Suryansh Kumar, Shuhang Gu, Radu Timofte, and Luc Van Gool · 2023
Later among the works it cites.
Dinov2: Learning robust visual features without supervision
Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, et al · 2023
Later among the works it cites.
Monocular depth estimation using diffusion models
Saurabh Saxena, Abhishek Kar, Mohammad Norouzi, and David J Fleet · 2023
Later among the works it cites.
Depth estimation from camera image and mmwave radar point cloud
Akash Deep Singh, Yunhao Ba, Ankur Sarker, Howard Zhang, Achuta Kadambi, Stefano Soatto, Mani Srivastava, and Alex Wong · 2023
Later among the works it cites.
Enhancing diffusion models with 3d perspective geometry constraints
Rishi Upadhyay, Howard Zhang, Yunhao Ba, Ethan Yang, Blake Gella, Sicheng Jiang, Alex Wong, and Achuta Kadambi · 2023
Later among the works it cites.
Boosting detection in crowd analysis via underutilized output features
Shaokai Wu and Fengyu Yang · 2023
Later among the works it cites.
Augundo: Scaling up augmentations for unsupervised depth completion
Yangchao Wu, Tian Yu Liu, Hyoungseob Park, Stefano Soatto, Dong Lao, and Alex Wong · 2023
Later among the works it cites.
Generating visual scenes from touch
Fengyu Yang, Jiacheng Zhang, and Andrew Owens · 2023
Later among the works it cites.
Implicit anatomical rendering for medical image segmentation with stochastic experts
Chenyu You, Weicheng Dai, Yifei Min, Lawrence Staib, and James S Duncan · 2023
Later among the works it cites.
Monocular depth estimation network based on swin transformer
Shangbin Yu, Renyan Zhang, Shuaiye Ma, and Xinfang Jiang · 2023
Later among the works it cites.
Llama-adapter: Efficient fine-tuning of language models with zero-init attention
Renrui Zhang, Jiaming Han, Aojun Zhou, Xiangfei Hu, Shilin Yan, Pan Lu, Hongsheng Li, Peng Gao, and Yu Qiao · 2023
Later among the works it cites.
Unleashing text-to-image diffusion models for visual perception
Wenliang Zhao, Yongming Rao, Zuyan Liu, Benlin Liu, Jie Zhou, and Jiwen Lu · 2023
Later among the works it cites.
Tactile-augmented radiance fields
Yiming Dou, Fengyu Yang, Yi Liu, Antonio Loquercio, and Andrew Owens · 2024
Closest in time.
Efficient vision-language pre-training by cluster masking
Zixuan Pan, Zihao Wei, and Andrew Owens · 2024
Closest in time.
Test-time adaptation for depth completion
Hyoungseob Park, Anjali Gupta, and Alex Wong · 2024
Closest in time.
The surprising effectiveness of diffusion models for optical flow and monocular depth estimation
Saurabh Saxena, Charles Herrmann, Junhwa Hur, Abhishek Kar, Mohammad Norouzi, Deqing Sun, and David J Fleet · 2024
Closest in time.
Binding touch to everything: Learning unified multimodal tactile representations
Fengyu Yang, Chao Feng, Ziyang Chen, Hyoungseob Park, Daniel Wang, Yiming Dou, Ziyao Zeng, Xien Chen, Rit Gangopadhyay, Andrew Owens, and Alex Wong · 2024
Closest in time.
Rethinking semi-supervised medical image segmentation: A variance-reduction perspective
Chenyu You, Weicheng Dai, Yifei Min, Fenglin Liu, David Clifton, S Kevin Zhou, Lawrence Staib, and James Duncan · 2024
Closest in time.
Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems?
Renrui Zhang, Dongzhi Jiang, Yichi Zhang, Haokun Lin, Ziyu Guo, Pengshuo Qiu, Aojun Zhou, Pan Lu, Kai-Wei Chang, Peng Gao, et al · 2024
Closest in time.
Iterated learning improves compositionality in large vision-language models
Chenhao Zheng, Jieyu Zhang, Aniruddha Kembhavi, and Ranjay Krishna · 2024
Closest in time.