Fetching the paper…
Reading the bibliography…
Besides image classification, Contrastive Language-Image Pre-training (CLIP) has accomplished extraordinary success for a wide range of vision tasks, including object-level and 3D space understanding.
Make3D: Depth Perception from a Single Still Image.. In Aaai , Vol. 3. 1571–1576
Ashutosh Saxena, Min Sun, and Andrew Y Ng. 2008 · 2008
Earlier work this paper cites.
Deep ordinal regression network for monocular depth estimation. In Proceedings of the IEEE conference on computer vision and pattern recognition . 2002–2011
Huan Fu, Mingming Gong, Chaohui Wang, Kayhan Batmanghelich, and Dacheng Tao. 2018 · 2011
Earlier work this paper cites.
Indoor segmentation and support inference from rgbd images. In European conference on computer vision . Springer, 746–760
Nathan Silberman, Derek Hoiem, Pushmeet Kohli, and Rob Fergus. 2012 · 2012
Earlier work this paper cites.
Vision meets robotics: The kitti dataset
Andreas Geiger, Philip Lenz, Christoph Stiller, and Raquel Urtasun. 2013 · 2013
Earlier work this paper cites.
Depth map prediction from a single image using a multi-scale deep network
David Eigen, Christian Puhrsch, and Rob Fergus. 2014 · 2014
Earlier work this paper cites.
Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition . 770–778
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Unsupervised monocular depth estimation with left-right consistency. In Proceedings of the IEEE conference on computer vision and pattern recognition . 270–279
Clément Godard, Oisin Mac Aodha, and Gabriel J Brostow. 2017 · 2017
Earlier work this paper cites.
Unsupervised learning of depth and ego-motion from monocular video using 3d geometric constraints. In Proceedings of the IEEE conference on computer vision and pattern recognition . 5667–5675
Reza Mahjourian, Martin Wicke, and Anelia Angelova. 2018 · 2018
Earlier work this paper cites.
Inferring point clouds from single monocular images by depth intermediation
Wei Zeng, Sezer Karaoglu, and Theo Gevers. 2018 · 2018
Earlier work this paper cites.
From big to small: Multi-scale local planar guidance for monocular depth estimation
Jin Han Lee, Myung-Kyu Han, Dong Wook Ko, and Il Hong Suh. 2019 · 2019
Earlier work this paper cites.
End-to-end object detection with transformers. In European conference on computer vision . Springer, 213–229
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. 2020 · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Cited alongside, same era.
Smoke: Single-stage monocular 3d object detection via keypoint estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops . 996–997
Zechen Liu, Zizhang Wu, and Roland Tóth. 2020 · 2020
Cited alongside, same era.
Unsupervised depth estimation from monocular videos with hybrid geometric-refined loss and contextual attention
Mingliang Zhang, Xinchen Ye, Xin Fan, and Wei Zhong. 2020 · 2020
Cited alongside, same era.
Transformer-based Monocular Depth Estimation with Attention Supervision
Wenjie Chang, Yueyi Zhang, and Zhiwei Xiong. 2021 · 2021
Cited alongside, same era.
Clip-adapter: Better vision-language models with feature adapters
Monocular depth estimation using laplacian pyramid-based depth residuals
Minsoo Song, Seokjae Lim, and Wonjun Kim. 2021 · 2021
Later among the works it cites.
Fcos3d: Fully convolutional one-stage monocular 3d object detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 913–922
Tai Wang, Xinge Zhu, Jiangmiao Pang, and Dahua Lin. 2021 · 2021
Later among the works it cites.
Tip-Adapter: Training-free CLIP-Adapter for Better Vision-Language Modeling
Renrui Zhang, Rongyao Fang, Peng Gao, Wei Zhang, Kunchang Li, Jifeng Dai, Yu Qiao, and Hongsheng Li. 2021a · 2021
Later among the works it cites.
VT-CLIP: Enhancing Vision-Language Models with Visual-guided Texts
Renrui Zhang, Longtian Qiu, Wei Zhang, and Ziyao Zeng. 2021b · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Peng Gao, Shijie Geng, Renrui Zhang, Teli Ma, Rongyao Fang, Yongfeng Zhang, Hongsheng Li, and Yu Qiao. 2021 · 2021
Cited alongside, same era.
Scaling up visual and vision-language representation learning with noisy text supervision. In International Conference on Machine Learning . PMLR, 4904–4916
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. 2021 · 2021
Cited alongside, same era.
PLNet: Plane and Line Priors for Unsupervised Indoor Depth Estimation. In 2021 International Conference on 3D Vision (3DV) . IEEE, 741–750
Hualie Jiang, Laiyan Ding, Junjie Hu, and Rui Huang. 2021 · 2021
Cited alongside, same era.
Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 10012–10022
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. 2021 · 2021
Cited alongside, same era.
Dual-stream network for visual recognition
Mingyuan Mao, Renrui Zhang, Honghui Zheng, Teli Ma, Yan Peng, Errui Ding, Baochang Zhang, Shumin Han, et al · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision. In International Conference on Machine Learning . PMLR, 8748–8763
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Cited alongside, same era.
DenseCLIP: Language-Guided Dense Prediction with Context-Aware Prompting
Yongming Rao, Wenliang Zhao, Guangyi Chen, Yansong Tang, Zheng Zhu, Guan Huang, Jie Zhou, and Jiwen Lu. 2021 · 2021
Cited alongside, same era.
Pointclip: Point cloud understanding by clip. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 8552–8562
Renrui Zhang, Ziyu Guo, Wei Zhang, Kunchang Li, Xupeng Miao, Bin Cui, Yu Qiao, Peng Gao, and Hongsheng Li. 2022a
Cited in the paper.
Chong Zhou, Chen Change Loy, and Bo Dai. 2021a · 2021
Later among the works it cites.
Learning to prompt for vision-language models
Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu. 2021b · 2021
Later among the works it cites.
Zhenyu Li, Zehui Chen, Xianming Liu, and Junjun Jiang. 2022 · 2022
Closest in time.
End-to-end Learning for Joint Depth and Image Reconstruction from Diffracted Rotation
Mazen Mel, Muhammad Siddiqui, and Pietro Zanuttigh. 2022 · 2022
Closest in time.
MonoDETR: Depth-aware Transformer for Monocular 3D Object Detection
Renrui Zhang, Han Qiu, Tai Wang, Xuanzhuo Xu, Ziyu Guo, Yu Qiao, Peng Gao, and Hongsheng Li. 2022b · 2022
Closest in time.
Detecting Twenty-thousand Classes using Image-level Supervision
Xingyi Zhou, Rohit Girdhar, Armand Joulin, Phillip Krähenbühl, and Ishan Misra. 2022 · 2022
Closest in time.