Fetching the paper…
Reading the bibliography…
The Contrastive Language-Image Pre-training (CLIP) framework has become a widely used approach for multimodal representation learning, particularly in image-text retrieval and clustering.
Microsoft coco: Common objects in context. In ECCV
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. 2014 · 2014
Earlier work this paper cites.
Notes on kullback-leibler divergence and likelihood
Jonathon Shlens. 2014 · 2014
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models. In ICCV
Bryan A Plummer, Liwei Wang, Chris M Cervantes, Juan C Caicedo, Julia Hockenmaier, and Svetlana Lazebnik. 2015 · 2015
Earlier work this paper cites.
Representation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals. 2018 · 2018
Earlier work this paper cites.
Conceptbert: Concept-aware representation for visual question answering. In EMNLP
François Gardères, Maryam Ziaeefard, Baptiste Abeloos, and Freddy Lecue. 2020 · 2020
Earlier work this paper cites.
Hard negative mixing for contrastive learning
Yannis Kalantidis, Mert Bulent Sariyildiz, Noe Pion, Philippe Weinzaepfel, and Diane Larlus. 2020 · 2020
Earlier work this paper cites.
Dense Passage Retrieval for Open-Domain Question Answering.. In EMNLP
Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick SH Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020 · 2020
Earlier work this paper cites.
Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters. In SIGKDD
Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He. 2020 · 2020
Earlier work this paper cites.
Contrastive Learning with Hard Negative Samples. In ICLR
Joshua David Robinson, Ching-Yao Chuang, Suvrit Sra, and Stefanie Jegelka. 2020 · 2020
Earlier work this paper cites.
Probabilistic embeddings for cross-modal retrieval. In CVPR
Sanghyuk Chun, Seong Joon Oh, Rafael Sampaio De Rezende, Yannis Kalantidis, and Diane Larlus. 2021 · 2021
Earlier work this paper cites.
Scaling deep contrastive learning batch size under memory limited setup
Luyu Gao, Yunyi Zhang, Jiawei Han, and Jamie Callan. 2021b · 2021
Earlier work this paper cites.
Simcse: Simple contrastive learning of sentence embeddings
Tianyu Gao, Xingcheng Yao, and Danqi Chen. 2021a · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision. In ICML
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Earlier work this paper cites.
Reproducible scaling laws for contrastive language-image learning
Mehdi Cherti, Romain Beaumont, Ross Wightman, Mitchell Wortsman, Gabriel Ilharco, Cade Gordon, Christoph Schuhmann, Ludwig Schmidt, and Jenia Jitsev. 2022 · 2022
Earlier work this paper cites.
MTEB: Massive Text Embedding Benchmark
Niklas Muennighoff, Nouamane Tazi, Loïc Magne, and Nils Reimers. 2022 · 2022
Earlier work this paper cites.
When and why vision-language models behave like bags-of-words, and what to do about it?
Mert Yuksekgonul, Federico Bianchi, Pratyusha Kalluri, Dan Jurafsky, and James Zou. 2022 · 2022
Earlier work this paper cites.
Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al · 2023
Earlier work this paper cites.
Zero-shot composed image retrieval with textual inversion. In ICCV . 15338–15347
Alberto Baldrati, Lorenzo Agnolucci, Marco Bertini, and Alberto Del Bimbo. 2023 · 2023
Earlier work this paper cites.
Universal Information Extraction with Meta-Pretrained Self-Retrieval. In ACL
Xin Cong, Bowen Yu, Mengcheng Fang, Tingwen Liu, Haiyang Yu, Zhongkai Hu, Fei Huang, Yongbin Li, and Bin Wang. 2023 · 2023
Earlier work this paper cites.
Qlora: Efficient finetuning of quantized llms
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. 2023 · 2023
Earlier work this paper cites.
Sugarcrepe: Fixing hackable benchmarks for vision-language compositionality
Cheng-Yu Hsieh, Jieyu Zhang, Zixian Ma, Aniruddha Kembhavi, and Ranjay Krishna. 2023 · 2023
Earlier work this paper cites.
Active retrieval augmented generation. In EMNLP . 7969–7992
Zhengbao Jiang, Frank F Xu, Luyu Gao, Zhiqing Sun, Qian Liu, Jane Dwivedi-Yu, Yiming Yang, Jamie Callan, and Graham Neubig. 2023 · 2023
Earlier work this paper cites.
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In ICML
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023 · 2023
Cited alongside, same era.
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023 · 2023
Cited alongside, same era.
Kosmos-2: Grounding multimodal large language models to the world
Zhiliang Peng, Wenhui Wang, Li Dong, Yaru Hao, Shaohan Huang, Shuming Ma, and Furu Wei. 2023 · 2023
Cited alongside, same era.
Eva-clip: Improved training techniques for clip at scale
Quan Sun, Yuxin Fang, Ledell Wu, Xinlong Wang, and Yue Cao. 2023a · 2023
Cited alongside, same era.
EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters
Quan Sun, Jinsheng Wang, Qiying Yu, Yufeng Cui, Fan Zhang, Xiaosong Zhang, and Xinlong Wang. 2023b · 2023
E5-v: Universal embeddings with multimodal large language models
Ting Jiang, Minghui Song, Zihan Zhang, Haizhen Huang, Weiwei Deng, Feng Sun, Qi Zhang, Deqing Wang, and Fuzhen Zhuang. 2024b · 2024
Later among the works it cites.
Nv-embed: Improved techniques for training llms as generalist embedding models
Chankyu Lee, Rajarshi Roy, Mengyao Xu, Jonathan Raiman, Mohammad Shoeybi, Bryan Catanzaro, and Wei Ping. 2024 · 2024
Later among the works it cites.
Llava-onevision: Easy visual task transfer
Bo Li, Yuanhan Zhang, Dong Guo, Renrui Zhang, Feng Li, Hao Zhang, Kaichen Zhang, Peiyuan Zhang, Yanwei Li, Ziwei Liu, et al · 2024
Later among the works it cites.
Vila: On pre-training for visual language models. In CVPR
Ji Lin, Hongxu Yin, Wei Ping, Pavlo Molchanov, Mohammad Shoeybi, and Song Han. 2024 · 2024
Later among the works it cites.
LLaVA-NeXT: Improved reasoning, OCR, and world knowledge
Haotian Liu, Chunyuan Li, Yuheng Li, Bo Li, Yuanhan Zhang, Sheng Shen, and Yong Jae Lee. 2024c · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Cited alongside, same era.
Image captioners are scalable vision learners too
Michael Tschannen, Manoj Kumar, Andreas Steiner, Xiaohua Zhai, Neil Houlsby, and Lucas Beyer. 2023 · 2023
Cited alongside, same era.
CogVLM: Visual Expert for Pretrained Language Models
Weihan Wang, Qingsong Lv, Wenmeng Yu, Wenyi Hong, Ji Qi, Yan Wang, Junhui Ji, Zhuoyi Yang, Lei Zhao, Xixuan Song, Jiazheng Xu, Bin Xu, Juanzi Li, Yuxiao Dong, Ming Ding, and Jie Tang. 2023 · 2023
Cited alongside, same era.
Alip: Adaptive language-image pre-training with synthetic caption. In ICCV . 2922–2931
Kaicheng Yang, Jiankang Deng, Xiang An, Jiawei Li, Ziyong Feng, Jia Guo, Jing Yang, and Tongliang Liu. 2023 · 2023
Cited alongside, same era.
Sigmoid loss for language image pre-training. In ICCV
Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer. 2023 · 2023
Cited alongside, same era.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. 2023 · 2023
Cited alongside, same era.
Later among the works it cites.
LamRA: Large Multimodal Model as Your Advanced Retrieval Assistant
Yikun Liu, Pingan Chen, Jiayin Cai, Xiaolong Jiang, Yao Hu, Jiangchao Yao, Yanfeng Wang, and Weidi Xie. 2024a · 2024
Later among the works it cites.
Fine-tuning llama for multi-stage text retrieval. In SIGIR
Xueguang Ma, Liang Wang, Nan Yang, Furu Wei, and Jimmy Lin. 2024 · 2024
Later among the works it cites.
Ullme: A unified framework for large language model embeddings with generation-augmented learning. In EMNLP
Hieu Man, Nghia Ngo, Franck Dernoncourt, and Thien Nguyen. 2024 · 2024
Later among the works it cites.
Discriminative Fine-tuning of LVLMs
Yassine Ouali, Adrian Bulat, Alexandros Xenos, Anestis Zaganidis, Ioannis Maniadis Metaxas, Brais Martinez, and Georgios Tzimiropoulos. 2024 · 2024
Later among the works it cites.
Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution
Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, et al · 2024
Later among the works it cites.
Uniir: Training and benchmarking universal multimodal information retrievers. In ECCV . Springer, 387–404
Cong Wei, Yang Chen, Haonan Chen, Hexiang Hu, Ge Zhang, Jie Fu, Alan Ritter, and Wenhu Chen. 2024 · 2024
Later among the works it cites.
Show-o: One single transformer to unify multimodal understanding and generation
Jinheng Xie, Weijia Mao, Zechen Bai, David Junhao Zhang, Weihao Wang, Kevin Qinghong Lin, Yuchao Gu, Zhijie Chen, Zhenheng Yang, and Mike Zheng Shou. 2024a · 2024
Later among the works it cites.
Croc: Pretraining Large Multimodal Models with Cross-Modal Comprehension
Yin Xie, Kaicheng Yang, Ninghua Yang, Weimo Deng, Xiangzi Dai, Tiancheng Gu, Yumeng Wang, Xiang An, Yongle Zhao, Ziyong Feng, et al · 2024
Later among the works it cites.
An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, et al · 2024
Later among the works it cites.
Magiclens: Self-supervised image retrieval with open-ended instructions
Kai Zhang, Yi Luan, Hexiang Hu, Kenton Lee, Siyuan Qiao, Wenhu Chen, Yu Su, and Ming-Wei Chang. 2024b · 2024
Later among the works it cites.
CTRL-O: Language-Controllable Object-Centric Visual Representation Learning. In CVPR
Aniket Didolkar, Andrii Zadaianchuk, Rabiul Awal, Maximilian Seitzer, Efstratios Gavves, and Aishwarya Agrawal. 2025 · 2025
Closest in time.
Decoupled Global-Local Alignment for Improving Compositional Understanding
Xiaoxing Hu, Kaicheng Yang, Jun Wang, Haoran Xu, Ziyong Feng, and Yupei Wang. 2025 · 2025
Closest in time.
Vlm2vec: Training vision-language models for massive multimodal embedding tasks
Ziyan Jiang, Rui Meng, Xinyi Yang, Semih Yavuz, Yingbo Zhou, and Wenhu Chen. 2025 · 2025
Closest in time.
Generative Modeling of Class Probability for Multi-Modal Representation Learning. In CVPR
Jungkyoo Shin, Bumsoo Kim, and Eunwoo Kim. 2025 · 2025
Closest in time.
Missing Target-Relevant Information Prediction with World Model for Accurate Zero-Shot Composed Image Retrieval. In CVPR
Yuanmin Tang, Jing Yu, Keke Gai, Jiamin Zhuang, Gang Xiong, Gaopeng Gou, and Qi Wu. 2025 · 2025
Closest in time.
Clip-cid: Efficient clip distillation via cluster-instance discrimination. In AAAI , Vol. 39. 21974–21982
Kaicheng Yang, Tiancheng Gu, Xiang An, Haiqiang Jiang, Xiangzi Dai, Ziyong Feng, Weidong Cai, and Jiankang Deng. 2025 · 2025
Closest in time.