Fetching the paper…
Reading the bibliography…
Self-supervised monocular depth estimation has seen significant progress in recent years, especially in outdoor environments.
A. Saxena, M. Sun, and A. Y. Ng, “Make3d: Learning 3d scene structure from a single still image,” TPAMI , vol. 31, no. 5, pp. 824–840, 2008
2008
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in CVPR , 2009
2009
Earlier work this paper cites.
A. Geiger, P. Lenz, and R. Urtasun, “Are we ready for autonomous driving? the kitti vision benchmark suite,” in CVPR , 2012
2012
Earlier work this paper cites.
N. Silberman, D. Hoiem, P. Kohli, and R. Fergus, “Indoor segmentation and support inference from rgb-d images,” in ECCV , 2012
2012
Earlier work this paper cites.
J. Shotton, B. Glocker, C. Zach, S. Izadi, A. Criminisi, and A. Fitzgibbon, “Scene coordinate regression forests for camera relocalization in rgb-d images,” in CVPR , 2013
2013
Earlier work this paper cites.
2014
Earlier work this paper cites.
K. Karsch, C. Liu, and S. B. Kang, “Depthtransfer: Depth extraction from video using non-parametric sampling,” TPAMI , 2014
2014
Earlier work this paper cites.
M. Liu, M. Salzmann, and X. He, “Discrete-continuous depth estimation from a single image,” in CVPR , 2014
2014
Earlier work this paper cites.
L. Ladicky, J. Shi, and M. Pollefeys, “Pulling things out of perspective,” in CVPR , 2014
2014
Earlier work this paper cites.
D. Eigen and R. Fergus, “Predicting depth, surface normals and semantic labels with a common multi-scale convolutional architecture,” in ICCV , 2015
2015
Earlier work this paper cites.
S. Ren, K. He, R. Girshick, and J. Sun, “Faster r-cnn: Towards real-time object detection with region proposal networks,” in NeurIPS , 2015
2015
Earlier work this paper cites.
M. Jaderberg, K. Simonyan, A. Zisserman, and K. Kavukcuoglu, “Spatial transformer networks,” in NeurIPS , 2015
2015
Earlier work this paper cites.
D. P. Kingma and J. L. Ba, “Adam: A method for stochastic gradient descent,” in ICLR , 2015
2015
Earlier work this paper cites.
B. Li, C. Shen, Y. Dai, A. van den Hengel, and M. He, “Depth and surface normal estimation from monocular images using regression on deep features and hierarchical crfs,” in CVPR , 2015
2015
Earlier work this paper cites.
F. Liu, C. Shen, and G. Lin, “Deep convolutional neural fields for depth estimation from a single image,” in CVPR , 2015
2015
Earlier work this paper cites.
P. Wang, X. Shen, Z. Lin, S. Cohen, B. Price, and A. L. Yuille, “Towards unified depth and semantic prediction from a single image,” in CVPR , 2015
2015
Earlier work this paper cites.
R. Garg, V. K. Bg, G. Carneiro, and I. Reid, “Unsupervised cnn for single view depth estimation: Geometry to the rescue,” in ECCV , 2016
2016
Earlier work this paper cites.
J. L. Schonberger and J.-M. Frahm, “Structure-from-motion revisited,” in CVPR , 2016
2016
Earlier work this paper cites.
M. Burri, J. Nikolic, P. Gohl, T. Schneider, J. Rehder, S. Omari, M. W. Achtelik, and R. Siegwart, “The euroc micro aerial vehicle datasets,” IJRR , vol. 35, no. 10, pp. 1157–1163, 2016
2016
Earlier work this paper cites.
J. L. Schönberger, E. Zheng, M. Pollefeys, and J.-M. Frahm, “Pixelwise view selection for unstructured multi-view stereo,” in ECCV , 2016
2016
Earlier work this paper cites.
I. Laina, C. Rupprecht, V. Belagiannis, F. Tombari, and N. Navab, “Deeper depth prediction with fully convolutional residual networks,” in 3DV , 2016
2016
Earlier work this paper cites.
A. Roy and S. Todorovic, “Monocular depth estimation using neural regression forest,” in CVPR , 2016
2016
Earlier work this paper cites.
A. Chakrabarti, J. Shao, and G. Shakhnarovich, “Depth from a single image by harmonizing overcomplete local network predictions,” in NeurIPS , 2016
2016
Earlier work this paper cites.
H. Alismail, B. Browning, and S. Lucey, “Photometric bundle adjustment for vision-based SLAM,” CoRR , 2016
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in CVPR , 2016
2016
Earlier work this paper cites.
T. Zhou, M. Brown, N. Snavely, and D. G. Lowe, “Unsupervised learning of depth and ego-motion from video,” in CVPR , 2017
2017
Earlier work this paper cites.
R. Mur-Artal and J. D. Tardós, “Orb-slam2: An open-source slam system for monocular, stereo, and rgb-d cameras,” TOR , vol. 33, no. 5, pp. 1255–1262, 2017
2017
Earlier work this paper cites.
A. Dai, A. X. Chang, M. Savva, M. Halber, T. Funkhouser, and M. Nießner, “Scannet: Richly-annotated 3d reconstructions of indoor scenes,” in CVPR , 2017
2017
Cited alongside, same era.
J. Li, R. Klein, and A. Yao, “A two-streamed network for estimating fine-scaled depth maps from single rgb images,” in ICCV , 2017
2017
Cited alongside, same era.
B. Ummenhofer, H. Zhou, J. Uhrig, N. Mayer, E. Ilg, A. Dosovitskiy, and T. Brox, “Demon: Depth and motion network for learning monocular stereo,” in CVPR , 2017, pp. 5038–5047
2017
Cited alongside, same era.
C. Godard, O. Mac Aodha, and G. J. Brostow, “Unsupervised monocular depth estimation with left-right consistency,” in CVPR , 2017
2017
Cited alongside, same era.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin, “Attention is all you need,” in NeurIPS , 2017
I. Bello, B. Zoph, A. Vaswani, J. Shlens, and Q. V. Le, “Attention augmented convolutional networks,” in ICCV , 2019
2019
Later among the works it cites.
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Kopf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala, “Pytorch: An imperative style, high-performance deep learning library,” in NeurIPS , 2019
2019
Later among the works it cites.
W. Zhao, S. Liu, Y. Shu, and Y.-J. Liu, “Towards better generalization: Joint depth-pose learning without posenet,” in CVPR , 2020
2020
Later among the works it cites.
Y. Zou, P. Ji, , Q.-H. Tran, J.-B. Huang, and M. Chandraker, “Learning monocular visual odometry via self-supervised long-term modeling,” in ECCV , 2020
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
C. R. Qi, H. Su, K. Mo, and L. J. Guibas, “Pointnet: Deep learning on point sets for 3d classification and segmentation,” CVPR , 2017
2017
Cited alongside, same era.
C. R. Qi, L. Yi, H. Su, and L. J. Guibas, “Pointnet++: Deep hierarchical feature learning on point sets in a metric space,” in NeurIPS , 2017
2017
Cited alongside, same era.
B. Ummenhofer, H. Zhou, J. Uhrig, N. Mayer, E. Ilg, A. Dosovitskiy, and T. Brox, “Demon: Depth and motion network for learning monocular stereo,” in CVPR , 2017
2017
Cited alongside, same era.
H. Fu, M. Gong, C. Wang, K. Batmanghelich, and D. Tao, “Deep ordinal regression network for monocular depth estimation,” in CVPR , 2018
2018
Cited alongside, same era.
X. Guo, H. Li, S. Yi, J. Ren, and X. Wang, “Learning monocular depth by distilling cross-domain stereo networks,” in ECCV , 2018
2018
Cited alongside, same era.
Y. Cao, Z. Wu, and C. Shen, “Estimating depth from monocular images as classification using deep fully convolutional residual networks,” T-CSVT , vol. 28, no. 11, pp. 3174–3182, 2018
2018
Cited alongside, same era.
Z. Teed and J. Deng, “Deepv2d: Video to depth with differentiable structure from motion,” in ICLR , 2018
2018
Cited alongside, same era.
Y. Cao, T. Zhao, K. Xian, C. Shen, Z. Cao, and S. Xu, “Monocular depth estimation with augmented ordinal depth relationships,” T-CSVT , vol. 30, no. 8, pp. 2674–2682, 2020
2020
Later among the works it cites.
L. Tiwari, P. Ji, Q.-H. Tran, B. Zhuang, S. Anand, and M. Chandraker, “Pseudo rgb-d for self-improving monocular slam and depth prediction,” in ECCV , 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng, “Nerf: Representing scenes as neural radiance fields for view synthesis,” in ECCV , 2020
2020
Later among the works it cites.
Z. Fang, X. Chen, Y. Chen, and L. V. Gool, “Towards good practice for cnn-based monocular depth estimation,” in WACV , 2020
2020
Later among the works it cites.
Z. Yu, L. Jin, and S. Gao, “P 2
2020
Later among the works it cites.
H. Zhan, C. S. Weerasekera, J.-W. Bian, and I. Reid, “Visual odometry revisited: What should be learnt?” in ICRA , 2020
2020
Later among the works it cites.
J. Bian, H. Zhan, N. Wang, T.-J. Chin, C. Shen, and I. Reid, “Auto-rectify network for unsupervised indoor depth estimation,” TPAMI , 2021
2021
Later among the works it cites.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” in ICLR , 2021
2021
Later among the works it cites.
P. Ji, R. Li, B. Bhanu, and Y. Xu, “Monoindoor: Towards good practice of self-supervised monocular depth estimation for indoor environments,” in ICCV , 2021
2021
Later among the works it cites.
S. F. Bhat, I. Alhashim, and P. Wonka, “Adabins: Depth estimation using adaptive bins,” in CVPR , 2021
2021
Later among the works it cites.
M. Song, S. Lim, and W. Kim, “Monocular depth estimation using laplacian pyramid-based depth residuals,” T-CSVT , vol. 31, no. 11, pp. 4381–4393, 2021
2021
Later among the works it cites.
Z. Teed and J. Deng, “DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D Cameras,” in NeurIPS , 2021
2021
Later among the works it cites.
R. Ranftl, A. Bochkovskiy, and V. Koltun, “Vision transformers for dense prediction,” in ICCV , 2021
2021
Later among the works it cites.
J. Watson, O. M. Aodha, V. Prisacariu, G. Brostow, and M. Firman, “The Temporal Opportunist: Self-Supervised Multi-Frame Monocular Depth,” in CVPR , 2021
2021
Later among the works it cites.
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in ICCV , 2021
2021
Later among the works it cites.
K. Park, U. Sinha, J. T. Barron, S. Bouaziz, D. B. Goldman, S. M. Seitz, and R. Martin-Brualla, “Nerfies: Deformable neural radiance fields,” in ICCV , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
J. Sun, Y. Xie, L. Chen, X. Zhou, and H. Bao, “NeuralRecon: Real-time coherent 3D reconstruction from monocular video,” CVPR , 2021
2021
Later among the works it cites.
R. Ranftl, K. Lasinger, D. Hafner, K. Schindler, and V. Koltun, “Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer,” PAMI , vol. 44, no. 3, pp. 1623–1637, 2022
2022
Closest in time.
Z. Liu, H. Hu, Y. Lin, Z. Yao, Z. Xie, Y. Wei, J. Ning, Y. Cao, Z. Zhang, L. Dong, F. Wei, and B. Guo, “Swin transformer v2: Scaling up capacity and resolution,” in CVPR , 2022
2022
Closest in time.
M. Tancik, V. Casser, X. Yan, S. Pradhan, B. Mildenhall, P. Srinivasan, J. T. Barron, and H. Kretzschmar, “Block-NeRF: Scalable large scene neural view synthesis,” in CVPR , 2022
2022
Closest in time.