Fetching the paper…
Reading the bibliography…
Pre-trained Vision-Language Models (VLMs), such as CLIP, have shown enhanced performance across a range of tasks that involve the integration of visual and linguistic modalities.
Make3d: Learning 3d scene structure from a single still image
Ashutosh Saxena, Min Sun, and Andrew Y Ng · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Indoor segmentation and support inference from rgbd images
Nathan Silberman, Derek Hoiem, Pushmeet Kohli, and Rob Fergus · 2012
Earlier work this paper cites.
Vision meets robotics: The kitti dataset
Andreas Geiger, Philip Lenz, Christoph Stiller, and Raquel Urtasun · 2013
Earlier work this paper cites.
Predicting deep zero-shot convolutional neural networks using textual descriptions
Jimmy Lei Ba, Kevin Swersky, Sanja Fidler, et al · 2015
Earlier work this paper cites.
Fast depth estimation from single image using structured forest
Shuai Fang, Ren Jin, and Yang Cao · 2016
Earlier work this paper cites.
Monocular depth estimation using neural regression forest
Anirban Roy and Sinisa Todorovic · 2016
Earlier work this paper cites.
Self-supervised learning of visual features through embedding images into text topic spaces
Lluis Gomez, Yash Patel, Marçal Rusinol, Dimosthenis Karatzas, and CV Jawahar · 2017
Earlier work this paper cites.
Deep ordinal regression network for monocular depth estimation
Huan Fu, Mingming Gong, Chaohui Wang, Kayhan Batmanghelich, and Dacheng Tao · 2018
Earlier work this paper cites.
Unsupervised learning of depth and ego-motion from monocular video using 3d geometric constraints
Reza Mahjourian, Martin Wicke, and Anelia Angelova · 2018
Earlier work this paper cites.
Depth prediction without the sensors: Leveraging structure for unsupervised learning from monocular videos
Vincent Casser, Soeren Pirk, Reza Mahjourian, and Anelia Angelova · 2019
Earlier work this paper cites.
Style augmentation: data augmentation via style randomization
Philip TG Jackson, Amir Atapour Abarghouei, Stephen Bonner, Toby P Breckon, and Boguslaw Obara · 2019
Earlier work this paper cites.
Few-shot learning for monocular depth estimation based on local object relationship
Shuai Li, Jiaying Shi, Wenfeng Song, Aimin Hao, and Hong Qin · 2019
Earlier work this paper cites.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Learning monocular depth estimation infusing traditional stereo knowledge
Fabio Tosi, Filippo Aleotti, Matteo Poggi, and Stefano Mattoccia · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Learning visual representations with caption annotations
Mert Bulent Sariyildiz, Julien Perez, and Diane Larlus · 2020
Earlier work this paper cites.
Unsupervised depth estimation from monocular videos with hybrid geometric-refined loss and contextual attention
Mingliang Zhang, Xinchen Ye, Xin Fan, and Wei Zhong · 2020
Cited alongside, same era.
Transformer-based monocular depth estimation with attention supervision
Wenjie Chang, Yueyi Zhang, and Zhiwei Xiong · 2021
Cited alongside, same era.
Virtex: Learning visual representations from textual annotations
Karan Desai and Justin Johnson · 2021
Cited alongside, same era.
Scaling up visual and vision-language representation learning with noisy text supervision
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig · 2021
Cited alongside, same era.
Less is more: Clipbert for video-and-language learning via sparse sampling
Jie Lei, Linjie Li, Luowei Zhou, Zhe Gan, Tamara L Berg, Mohit Bansal, and Jingjing Liu · 2021
Cited alongside, same era.
Grounded language-image pre-training
Liunian Harold Li, Pengchuan Zhang, Haotian Zhang, Jianwei Yang, Chunyuan Li, Yiwu Zhong, Lijuan Wang, Lu Yuan, Lei Zhang, Jenq-Neng Hwang, et al · 2022
Later among the works it cites.
Prompt distribution learning
Yuning Lu, Jianzhuang Liu, Yonggang Zhang, Yajing Liu, and Xinmei Tian · 2022
Later among the works it cites.
End-to-end learning for joint depth and image reconstruction from diffracted rotation
Mazen Mel, Muhammad Siddiqui, and Pietro Zanuttigh · 2022
Later among the works it cites.
Coca: Contrastive captioners are image-text foundation models
Jiahui Yu, Zirui Wang, Vijay Vasudevan, Legg Yeung, Mojtaba Seyedhosseini, and Yonghui Wu · 2022
Later among the works it cites.
Can language understand depth?
Renrui Zhang, Ziyao Zeng, Ziyu Guo, and Yafeng Li · 2022
Later among the works it cites.
Conditional prompt learning for vision-language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Brian Lester, Rami Al-Rfou, and Noah Constant · 2021
Cited alongside, same era.
Prefix-tuning: Optimizing continuous prompts for generation
Xiang Lisa Li and Percy Liang · 2021
Cited alongside, same era.
Deep learning for monocular depth estimation: A review
Yue Ming, Xuyang Meng, Chunxiao Fan, and Hui Yu · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Cited alongside, same era.
Exploiting cloze questions for few shot text classification and natural language inference
Timo Schick and Hinrich Schütze · 2021
Cited alongside, same era.
It’s not just size that matters: Small language models are also few-shot learners
Timo Schick and Hinrich Schütze · 2021
Cited alongside, same era.
Cpt: Colorful prompt tuning for pre-trained vision-language models
Yuan Yao, Ao Zhang, Zhengyan Zhang, Zhiyuan Liu, Tat-Seng Chua, and Maosong Sun · 2021
Cited alongside, same era.
Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu · 2022
Later among the works it cites.
Learning to prompt for vision-language models
Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu · 2022
Later among the works it cites.
Unsupervised vision-and-language pre-training via retrieval-based multi-granular alignment
Mingyang Zhou, Licheng Yu, Amanpreet Singh, Mengjiao Wang, Zhou Yu, and Ning Zhang · 2022
Later among the works it cites.
Prompt learning with optimal transport for vision-language models
Guangyi Chen, Weiran Yao, Xiangchen Song, Xinyue Li, Yongming Rao, and Kun Zhang · 2023
Closest in time.
Clip-adapter: Better vision-language models with feature adapters
Peng Gao, Shijie Geng, Renrui Zhang, Teli Ma, Rongyao Fang, Yongfeng Zhang, Hongsheng Li, and Yu Qiao · 2023
Closest in time.
Depthformer: Exploiting long-range correlation and local information for accurate monocular depth estimation
Zhenyu Li, Zehui Chen, Xianming Liu, and Junjun Jiang · 2023
Closest in time.
Cross-inferential networks for source-free unsupervised domain adaptation
Yushun Tang, Qinghai Guo, and Zhihai He · 2023
Closest in time.
Neuro-modulated hebbian learning for fully test-time adaptation
Yushun Tang, Ce Zhang, Heng Xu, Shuoshuo Chen, Jie Cheng, Luziwei Leng, Qinghai Guo, and Zhihai He · 2023
Closest in time.
Learning to decompose visual features with latent textual prompts
Feng Wang, Manling Li, Xudong Lin, Hairong Lv, Alexander G Schwing, and Heng Ji · 2023
Closest in time.
Efficient image captioning for edge devices
Ning Wang, Jiangrong Xie, Hang Luo, Qinglin Cheng, Jihao Wu, Mingbo Jia, and Linlin Li · 2023
Closest in time.
Unsupervised prototype adapter for vision-language models
Yi Zhang, Ce Zhang, Xueting Hu, and Zhihai He · 2023
Closest in time.
Bdc-adapter: Brownian distance covariance for better vision-language reasoning
Yi Zhang, Ce Zhang, Zihan Liao, Yushun Tang, and Zhihai He · 2023
Closest in time.
Cross-modal concept learning and inference for vision-language models
Yi Zhang, Ce Zhang, Yushun Tang, and Zhihai He · 2023
Closest in time.