Fetching the paper…
Reading the bibliography…
Contrastive Language-Image Pretraining (CLIP) has been widely used for crossmodal information retrieval and multimodal understanding tasks.
Unsupervised cross-lingual representation learning at scale, 2020
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov · 1911
Earlier work this paper cites.
Docvqa: A dataset for vqa on document images, 2021b
Minesh Mathew, Dimosthenis Karatzas, and C. V. Jawahar · 2007
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier · 2014
Earlier work this paper cites.
Microsoft COCO Captions: Data Collection and Evaluation Server
Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C. Lawrence Zitnick · 2015
Earlier work this paper cites.
Bridge correlational neural networks for multilingual multimodal representation learning, 2016
Janarthanan Rajendran, Mitesh M. Khapra, Sarath Chandar, and Balaraman Ravindran · 2016
Earlier work this paper cites.
Representation Learning with Contrastive Predictive Coding
A. Van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Earlier work this paper cites.
Learning dense representations for entity retrieval
Daniel Gillick, Sayali Kulkarni, Larry Lansing, Alessandro Presta, Jason Baldridge, Eugene Ie, and Diego Garcia-Olano · 2019
Earlier work this paper cites.
Decoupled weight decay regularization, 2019
Ilya Loshchilov and Frank Hutter · 2019
Earlier work this paper cites.
Towards zero-shot cross-lingual image retrieval, 2020
Pranav Aggarwal and Ajinkya Kale · 2020
Earlier work this paper cites.
A simple framework for contrastive learning of visual representations
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton · 2020
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models, 2021
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2021
Earlier work this paper cites.
Scaling up visual and vision-language representation learning with noisy text supervision
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig · 2021
Earlier work this paper cites.
Align before fuse: Vision and language representation learning with momentum distillation
Junnan Li, Ramprasaath Selvaraju, Akhilesh Gotmare, Shafiq Joty, Caiming Xiong, and Steven Chu Hong Hoi · 2021
Earlier work this paper cites.
RocketQA: An optimized training approach to dense passage retrieval for open-domain question answering
Yingqi Qu, Yuchen Ding, Jing Liu, Kai Liu, Ruiyang Ren, Wayne Xin Zhao, Daxiang Dong, Hua Wu, and Haifeng Wang · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision, 2021
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Earlier work this paper cites.
Wit: Wikipedia-based image text dataset for multimodal multilingual machine learning
Krishna Srinivasan, Karthik Raman, Jiecao Chen, Michael Bendersky, and Marc Najork · 2021
Cited alongside, same era.
Understanding the behaviour of contrastive loss
Feng Wang and Huaping Liu · 2021
Cited alongside, same era.
Cross-lingual and multilingual clip
Fredrik Carlsson, Philipp Eisen, Faton Rekathati, and Magnus Sahlgren · 2022
Cited alongside, same era.
No language left behind: Scaling human-centered machine translation
Marta R Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, et al · 2022
Cited alongside, same era.
Flashattention: Fast and memory-efficient exact attention with io-awareness, 2022
Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher Ré · 2022
Gpt-4v(ision) system card, 2023
OpenAI · 2023
Later among the works it cites.
Roformer: Enhanced transformer with rotary position embedding, 2023
Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, and Yunfeng Liu · 2023
Later among the works it cites.
Eva-clip: Improved training techniques for clip at scale, 2023
Quan Sun, Yuxin Fang, Ledell Wu, Xinlong Wang, and Yue Cao · 2023
Later among the works it cites.
Sigmoid loss for language image pre-training, 2023
Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer · 2023
Later among the works it cites.
Getting vit in shape: Scaling laws for compute-optimal model design, 2024
Ibrahim Alabdulmohsin, Xiaohua Zhai, Alexander Kolesnikov, and Lucas Beyer · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
xformers: A modular and hackable transformer modelling library
Benjamin Lefaudeux, Francisco Massa, Diana Liskovich, Wenhan Xiong, Vittorio Caggiano, Sean Naren, Min Xu, Jieru Hu, Marta Tintore, Susan Zhang, Patrick Labatut, Daniel Haziza, Luca Wehrstedt, Jeremy Reizenstein, and Grigory Sizov · 2022
Cited alongside, same era.
Weixin Liang, Yuhui Zhang, Yongchan Kwon, Serena Yeung, and James Zou · 2022
Cited alongside, same era.
Crossmodal-3600: A Massively Multilingual Multimodal Evaluation Dataset
Ashish Thapliyal, Jordi Pont-Tuset, Xi Chen, and Radu Soricut · 2022
Cited alongside, same era.
Lit: Zero-shot transfer with locked-image text tuning
Xiaohua Zhai, Xiao Wang, Basil Mustafa, Andreas Steiner, Daniel Keysers, Alexander Kolesnikov, and Lucas Beyer · 2022
Cited alongside, same era.
Contrastive learning of medical visual representations from paired images and text
Yuhao Zhang, Hang Jiang, Yasuhide Miura, Christopher D Manning, and Curtis P Langlotz · 2022
Cited alongside, same era.
Towards complex document understanding by discrete reasoning
Fengbin Zhu, Wenqiang Lei, Fuli Feng, Chao Wang, Haozhou Zhang, and Tat-Seng Chua · 2022
Cited alongside, same era.
Datacomp: In search of the next generation of multimodal datasets, 2023
Samir Yitzhak Gadre, Gabriel Ilharco, Alex Fang, Jonathan Hayase, Georgios Smyrnis, Thao Nguyen, Ryan Marten, Mitchell Wortsman, Dhruba Ghosh, Jieyu Zhang, Eyal Orgad, Rahim Entezari, Giannis Daras, Sarah Pratt, Vivek Ramanujan, Yonatan Bitton, Kalyani Marathe, Stephen Mussmann, Richard Vencu, Mehdi Cherti, Ranjay Krishna, Pang Wei Koh, Olga Saukh, Alexander Ratner, Shuran Song, Hannaneh Hajishirzi, Ali Farhadi, Romain Beaumont, Sewoong Oh, Alex Dimakis, Jenia Jitsev, Yair Carmon, Vaishaal Shankar, and Ludwig Schmidt · 2023
Cited alongside, same era.
Closest in time.
Jianlv Chen, Shitao Xiao, Peitian Zhang, Kun Luo, Defu Lian, and Zheng Liu · 2024
Closest in time.
Mitigate the gap: Investigating approaches for improving cross-modal alignment in clip
Sedigheh Eslami and Gerard de Melo · 2024
Closest in time.
Colpali: Efficient document retrieval with vision language models, 2024
Manuel Faysse, Hugues Sibille, Tony Wu, Bilel Omrani, Gautier Viaud, Céline Hudelot, and Pierre Colombo · 2024
Closest in time.
Jina clip: Your clip model is also your text retriever, 2024
Andreas Koukounas, Georgios Mastrapas, Michael Günther, Bo Wang, Scott Martens, Isabelle Mohr, Saba Sturua, Mohammad Kalim Akram, Joan Fontanals Martínez, Saahil Ognawala, Susana Guzman, Maximilian Werk, Nan Wang, and Han Xiao · 2024
Closest in time.
Matryoshka representation learning, 2024
Aditya Kusupati, Gantavya Bhatt, Aniket Rege, Matthew Wallingford, Aditya Sinha, Vivek Ramanujan, William Howard-Snyder, Kaifeng Chen, Sham Kakade, Prateek Jain, and Ali Farhadi · 2024
Closest in time.
Multimodal ArXiv: A dataset for improving scientific comprehension of large vision-language models
Lei Li, Yuqi Wang, Runxin Xu, Peiyi Wang, Xiachong Feng, Lingpeng Kong, and Qi Liu · 2024
Closest in time.
Mm-embed: Universal multimodal retrieval with multimodal llms, 2024
Sheng-Chieh Lin, Chankyu Lee, Mohammad Shoeybi, Jimmy Lin, Bryan Catanzaro, and Wei Ping · 2024
Closest in time.
Simon Schrodi, David T Hoffmann, Max Argus, Volker Fischer, and Thomas Brox · 2024
Closest in time.
jina-embeddings-v3: Multilingual embeddings with task lora, 2024
Saba Sturua, Isabelle Mohr, Mohammad Kalim Akram, Michael Günther, Bo Wang, Markus Krimmel, Feng Wang, Georgios Mastrapas, Andreas Koukounas, Nan Wang, and Han Xiao · 2024
Closest in time.
Multilingual e5 text embeddings: A technical report
Liang Wang, Nan Yang, Xiaolong Huang, Linjun Yang, Rangan Majumder, and Furu Wei · 2024
Closest in time.