Fetching the paper…
Reading the bibliography…
Contrastive Language-Image Pretraining (CLIP) is widely used to train models to align images and texts in a common embedding space by mapping them to fixed-sized vectors.
Framing Image Description as a Ranking Task: Data, Models and Evaluation Metrics
Hodosh, M., Young, P., and Hockenmaier, J · 2013
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Young, P., Lai, A., Hodosh, M., and Hockenmaier, J · 2014
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Bowman, S., Angeli, G., Potts, C., and Manning, C. D · 2015
Earlier work this paper cites.
Microsoft COCO Captions: Data Collection and Evaluation Server
Chen, X., Fang, H., Lin, T., Vedantam, R., Gupta, S., Dollár, P., and Zitnick, C. L · 2015
Earlier work this paper cites.
MS MARCO: A Human Generated MAchine Reading COmprehension Dataset
Bajaj, P., Campos, D., Craswell, N., Deng, L., Gao, J., Liu, X., Majumder, R., McNamara, A., Mitra, B., Nguyen, T., et al · 2016
Earlier work this paper cites.
YFCC100M: The New Data in Multimedia Research
Thomee, B., Shamma, D. A., Friedland, G., Elizalde, B., Ni, K., Poland, D., Borth, D., and Li, L · 2016
Earlier work this paper cites.
Fixing Weight Decay Regularization in Adam
Loshchilov, I. and Hutter, F · 2017
Earlier work this paper cites.
Representation Learning with Contrastive Predictive Coding
Van den Oord, A., Li, Y., and Vinyals, O · 2018
Earlier work this paper cites.
HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering
Yang, Z., Qi, P., Zhang, S., Bengio, Y., Cohen, W. W., Salakhutdinov, R., and Manning, C. D · 2018
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J., Chang, M., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
Natural Questions: a Benchmark for Question Answering Research
Kwiatkowski, T., Palomaki, J., Redfield, O., Collins, M., Parikh, A., Alberti, C., Epstein, D., Polosukhin, I., Kelcey, M., Devlin, J., Lee, K., Toutanova, K. N., Jones, L., Chang, M.-W., Dai, A., Uszkoreit, J., Le, Q., and Petrov, S · 2019
Earlier work this paper cites.
Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
Reimers, N. and Gurevych, I · 2019
Earlier work this paper cites.
OpenCLIP (0.1)
Ilharco, G., Wortsman, M., Wightman, R., Gordon, C., Carlini, N., Taori, R., Dave, A., Shankar, V., Namkoong, H., Miller, J., Hajishirzi, H., Farhadi, A., and Schmidt, L · 2021
Earlier work this paper cites.
Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation
Press, O., Smith, N. A., and Lewis, M · 2021
Cited alongside, same era.
Learning Transferable Visual Models From Natural Language Supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I · 2021
Cited alongside, same era.
LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs
Schuhmann, C., Vencu, R., Beaumont, R., Kaczmarczyk, R., Mullis, C., Katta, A., Coombes, T., Jitsev, J., and Komatsuzaki, A · 2021
Cited alongside, same era.
Understanding the Behaviour of Contrastive Loss
Wang, F. and Liu, H · 2021
Cited alongside, same era.
Large Dual Encoders Are Generalizable Retrievers
Ni, J., Qu, C., Lu, J., Dai, Z., Ábrego, G. H., Ma, J., Zhao, V. Y., Luan, Y., Hall, K. B., Chang, M., and Yang, Y · 2022
Cited alongside, same era.
Jina Embeddings 2: 8192-Token General-Purpose Text Embeddings for Long Documents
Günther, M., Ong, J., Mohr, I., Abdessalem, A., Abel, T., Akram, M. K., Guzman, S., Mastrapas, G., Sturua, S., Wang, B., Werk, M., Wang, N., and Xiao, H · 2023
Later among the works it cites.
Three Towers: Flexible Contrastive Learning with Pretrained Image Models
Kossen, J., Collier, M., Mustafa, B., Wang, X., Zhai, X., Beyer, L., Steiner, A., Berent, J., Jenatton, R., and Kokiopoulou, E · 2023
Later among the works it cites.
MTEB: Massive Text Embedding Benchmark
Muennighoff, N., Tazi, N., Magne, L., and Reimers, N · 2023
Later among the works it cites.
OpenAI · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
LAION-5B: An open large-scale dataset for training next generation image-text models
Schuhmann, C., Beaumont, R., Vencu, R., Gordon, C. W., Wightman, R., Cherti, M., Coombes, T., Katta, A., Mullis, C., Wortsman, M., Schramowski, P., Kundurthy, S. R., Crowson, K., Schmidt, L., Kaczmarczyk, R., and Jitsev, J · 2022
Cited alongside, same era.
Text Embeddings by Weakly-Supervised Contrastive Pre-training
Wang, L., Yang, N., Huang, X., Jiao, B., Yang, L., Jiang, D., Majumder, R., and Wei, F · 2022
Cited alongside, same era.
LiT: Zero-Shot Transfer with Locked-image text Tuning
Zhai, X., Wang, X., Mustafa, B., Steiner, A., Keysers, D., Kolesnikov, A., and Beyer, L · 2022
Cited alongside, same era.
ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
Chen, L., Li, J., Dong, X., Zhang, P., He, C., Wang, J., Zhao, F., and Lin, D · 2023
Cited alongside, same era.
Reproducible Scaling Laws for Contrastive Language-Image Learning
Cherti, M., Beaumont, R., Wightman, R., Wortsman, M., Ilharco, G., Gordon, C., Schuhmann, C., Schmidt, L., and Jitsev, J · 2023
Cited alongside, same era.
EVA-02: A Visual Representation for Neon Genesis
Fang, Y., Sun, Q., Wang, X., Huang, T., Wang, X., and Cao, Y · 2023
Cited alongside, same era.
Jina Embeddings: A Novel Set of High-Performance Sentence Embedding Models
Günther, M., Mastrapas, G., Wang, B., Xiao, H., and Geuter, J · 2023
Cited alongside, same era.
Sun, Q., Fang, Y., Wu, L., Wang, X., and Cao, Y · 2023
Later among the works it cites.
Sigmoid Loss for Language Image Pre-Training
Zhai, X., Mustafa, B., Kolesnikov, A., and Beyer, L · 2023
Later among the works it cites.
Retrieving Multimodal Information for Augmented Generation: A Survey
Zhao, R., Chen, H., Wang, W., Jiao, F., Long, D. X., Qin, C., Ding, B., Guo, X., Li, M., Li, X., and Joty, S · 2023
Later among the works it cites.
Chen, J., Xiao, S., Zhang, P., Luo, K., Lian, D., and Liu, Z · 2024
Closest in time.
Multi-Task Contrastive Learning for 8192-Token Bilingual Text Embeddings
Mohr, I., Krimmel, M., Sturua, S., Akram, M. K., Koukounas, A., Günther, M., Mastrapas, G., Ravishankar, V., Martínez, J. F., Wang, F., Liu, Q., Yu, Z., Fu, J., Ognawala, S., Guzman, S., Wang, B., Werk, M., Wang, N., and Xiao, H · 2024
Closest in time.
DINOv2: Learning Robust Visual Features without Supervision
Oquab, M., Darcet, T., Moutakanni, T., Vo, H., Szafraniec, M., Khalidov, V., Fernandez, P., Haziza, D., Massa, F., El-Nouby, A., Assran, M., Ballas, N., Galuba, W., Howes, R., Huang, P.-Y., Li, S.-W., Misra, I., Rabbat, M., Sharma, V., Synnaeve, G., Xu, H., Jegou, H., Mairal, J., Labatut, P., Joulin, A., and Bojanowski, P · 2024
Closest in time.
EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters
Sun, Q., Wang, J., Yu, Q., Cui, Y., Zhang, F., Zhang, X., and Wang, X · 2024
Closest in time.
Long-CLIP: Unlocking the Long-Text Capability of CLIP
Zhang, B., Zhang, P., Dong, X., Zang, Y., and Wang, J · 2024
Closest in time.