Fetching the paper…
Reading the bibliography…
Self-supervised monocular depth estimation is a salient task for 3D scene understanding.
1907
Earlier work this paper cites.
Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli, “Image quality assessment: from error visibility to structural similarity,” IEEE transactions on image processing , vol. 13, no. 4, pp. 600–612, 2004
2004
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in 2009 IEEE conference on computer vision and pattern recognition . Ieee, 2009, pp. 248–255
2009
Earlier work this paper cites.
A. Geiger, P. Lenz, C. Stiller, and R. Urtasun, “Vision meets robotics: The kitti dataset,” The International Journal of Robotics Research , vol. 32, no. 11, pp. 1231–1237, 2013
2013
Earlier work this paper cites.
D. Eigen, C. Puhrsch, and R. Fergus, “Depth map prediction from a single image using a multi-scale deep network,” in Advances in neural information processing systems , 2014, pp. 2366–2374
2014
Earlier work this paper cites.
J. L. Schonberger and J.-M. Frahm, “Structure-from-motion revisited,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 4104–4113
2016
Earlier work this paper cites.
A. Kurakin, I. Goodfellow, S. Bengio, et al. , “Adversarial examples in the physical world,” 2016
2016
Earlier work this paper cites.
T. Zhou, M. Brown, N. Snavely, and D. G. Lowe, “Unsupervised learning of depth and ego-motion from video,” 2017
2017
Earlier work this paper cites.
C. Godard, O. Mac Aodha, and G. J. Brostow, “Unsupervised monocular depth estimation with left-right consistency,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2017, pp. 270–279
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
W. Yin, Y. Liu, C. Shen, and Y. Yan, “Enforcing geometric constraints of virtual normal for depth prediction,” in Proceedings of the IEEE International Conference on Computer Vision , 2019, pp. 5684–5693
2019
Earlier work this paper cites.
C. Godard, O. Mac Aodha, M. Firman, and G. J. Brostow, “Digging into self-supervised monocular depth estimation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2019, pp. 3828–3838
2019
Earlier work this paper cites.
A. Gordon, H. Li, R. Jonschkowski, and A. Angelova, “Depth from videos in the wild: Unsupervised monocular depth learning from unknown cameras,” 2019
2019
Cited alongside, same era.
A. Ranjan, V. Jampani, L. Balles, K. Kim, D. Sun, J. Wulff, and M. J. Black, “Competitive collaboration: Joint unsupervised learning of depth, camera motion, optical flow and motion segmentation,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2019, pp. 12 240–12 249
2019
Cited alongside, same era.
J. Bian, Z. Li, N. Wang, H. Zhan, C. Shen, M.-M. Cheng, and I. Reid, “Unsupervised scale-consistent depth and ego-motion learning from monocular video,” in Advances in Neural Information Processing Systems , 2019, pp. 35–45
2019
Cited alongside, same era.
J. D. M.-W. C. Kenton and L. K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” in Proceedings of NAACL-HLT , 2019, pp. 4171–4186
2019
Cited alongside, same era.
H. Bao, L. Dong, S. Piao, and F. Wei, “Beit: Bert pre-training of image transformers,” in International Conference on Learning Representations , 2021
2021
Later among the works it cites.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” ICLR , 2021
2021
Later among the works it cites.
A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. Voss, A. Radford, M. Chen, and I. Sutskever, “Zero-shot text-to-image generation,” in International Conference on Machine Learning . PMLR, 2021, pp. 8821–8831
2021
Later among the works it cites.
H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jégou, “Training data-efficient image transformers & distillation through attention,” in International Conference on Machine Learning . PMLR, 2021, pp. 10 347–10 357
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Tsipras, S. Santurkar, L. Engstrom, A. Turner, and A. Madry, “Robustness may be at odds with accuracy,” in International Conference on Machine Learning , 2019
2019
Cited alongside, same era.
V. Guizilini, R. Ambrus, S. Pillai, A. Raventos, and A. Gaidon, “3d packing for self-supervised monocular depth estimation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2020, pp. 2485–2494
2020
Cited alongside, same era.
M. Klingner, J.-A. Termöhlen, J. Mikolajczyk, and T. Fingscheidt, “Self-Supervised Monocular Depth Estimation: Solving the Dynamic Object Problem by Semantic Guidance,” in ECCV , 2020
2020
Cited alongside, same era.
M. Chen, A. Radford, R. Child, J. Wu, H. Jun, D. Luan, and I. Sutskever, “Generative pretraining from pixels,” in International conference on machine learning . PMLR, 2020, pp. 1691–1703
2020
Cited alongside, same era.
A. Wong, S. Cicek, and S. Soatto, “Targeted adversarial perturbations for monocular depth prediction,” in Advances in neural information processing systems , 2020
2020
Cited alongside, same era.
H. Chawla, A. Varma, E. Arani, and B. Zonooz, “Multimodal scale consistency and awareness for monocular self-supervised depth estimation,” in 2021 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2021
2021
Cited alongside, same era.
S. Lee, C. Lee, H. Kim, and H. J. Kim, “Realtime object-aware monocular depth estimation in onboard systems,” International Journal of Control, Automation and Systems , vol. 19, no. 9, pp. 3179–3189, 2021
2021
Cited alongside, same era.
H. Chawla, A. Varma, E. Arani, and B. Zonooz, “Adversarial attacks on monocular pose estimation,” in 2022 IEEE/RSJ International Conference on Intelligent Robotics and Systems (IROS) . IEEE (in press), 2022
2022
Closest in time.
A. Varma., H. Chawla., B. Zonooz., and E. Arani., “Transformers in self-supervised monocular depth estimation with unknown camera intrinsics,” in Proceedings of the 17th International Joint Conference on Computer Vision, Imaging and Computer Graphics Theory and Applications - Volume 4: VISAPP, , INSTICC. SciTePress, 2022, pp. 758–769
2022
Closest in time.
Z. Xie, Z. Zhang, Y. Cao, Y. Lin, J. Bao, Z. Yao, Q. Dai, and H. Hu, “Simmim: A simple framework for masked image modeling,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 9653–9663
2022
Closest in time.
K. He, X. Chen, S. Xie, Y. Li, P. Dollár, and R. Girshick, “Masked autoencoders are scalable vision learners,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 16 000–16 009
2022
Closest in time.
C. Wei, H. Fan, S. Xie, C.-Y. Wu, A. Yuille, and C. Feichtenhofer, “Masked feature prediction for self-supervised visual pre-training,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 14 668–14 678
2022
Closest in time.
V. Guizilini, R. Ambruș, D. Chen, S. Zakharov, and A. Gaidon, “Multi-frame self-supervised depth with transformers,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 160–170
2022
Closest in time.
J. Yang, L. An, A. Dixit, J. Koo, and S. I. Park, “Depth estimation with simplified transformer,” in CVPR Workshops , 2022
2022
Closest in time.