Fetching the paper…
Reading the bibliography…
We introduce SigLIP 2, a family of new multilingual vision-language encoders that build on the success of the original SigLIP.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
The pascal visual object classes (voc) challenge
M. Everingham, L. Van Gool, C. K. Williams, J. Winn, and A. Zisserman · 2010
Earlier work this paper cites.
ReferItGame: Referring to objects in photographs of natural scenes
S. Kazemzadeh, V. Ordonez, M. Matten, and T. Berg · 2014
Earlier work this paper cites.
Microsoft COCO: common objects in context
T. Lin, M. Maire, S. J. Belongie, L. D. Bourdev, R. B. Girshick, J. Hays, P. Perona, D. Ramanan, P. Doll’a r, and C. L. Zitnick · 2014
Earlier work this paper cites.
The role of context for object detection and semantic segmentation in the wild
R. Mottaghi, X. Chen, X. Liu, N.-G. Cho, S.-W. Lee, S. Fidler, R. Urtasun, and A. Yuille · 2014
Earlier work this paper cites.
Modeling context in referring expressions
L. Yu, P. Poirson, S. Yang, A. C. Berg, and T. L. Berg · 2016
Earlier work this paper cites.
Fixing weight decay regularization in adam
I. Loshchilov, F. Hutter, et al · 2017
Earlier work this paper cites.
Scene parsing through ade20k dataset
B. Zhou, H. Zhao, X. Puig, S. Fidler, A. Barriuso, and A. Torralba · 2017
Earlier work this paper cites.
Coco-stuff: Thing and stuff classes in context
H. Caesar, J. Uijlings, and V. Ferrari · 2018
Earlier work this paper cites.
Objectnet: A large-scale bias-controlled dataset for pushing the limits of object recognition models
A. Barbu, D. Mayo, J. Alverio, W. Luo, C. Wang, D. Gutfreund, J. Tenenbaum, and B. Katz · 2019
Earlier work this paper cites.
Lvis: A dataset for large vocabulary instance segmentation
A. Gupta, P. Dollar, and R. Girshick · 2019
Earlier work this paper cites.
Do imagenet classifiers generalize to imagenet?
B. Recht, R. Roelofs, L. Schmidt, and V. Shankar · 2019
Earlier work this paper cites.
Semantic understanding of scenes through the ade20k dataset
B. Zhou, H. Zhao, X. Puig, T. Xiao, S. Fidler, A. Barriuso, and A. Torralba · 2019
Earlier work this paper cites.
L. Beyer, O. J. Hénaff, A. Kolesnikov, X. Zhai, and A. v. d. Oord · 2020
Earlier work this paper cites.
TextCaps: A dataset for image captioning with reading comprehension
O. Sidorov, R. Hu, M. Rohrbach, and A. Singh · 2020
Earlier work this paper cites.
Google landmarks dataset v2-a large-scale benchmark for instance-level recognition and retrieval
T. Weyand, A. Araujo, B. Cao, and J. Sim · 2020
Earlier work this paper cites.
Emerging properties in self-supervised vision transformers
M. Caron, H. Touvron, I. Misra, H. Jégou, J. Mairal, P. Bojanowski, and A. Joulin · 2021
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby · 2021
Earlier work this paper cites.
Scicap: Generating captions for scientific figures
T.-Y. Hsu, C. L. Giles, and T.-H. Huang · 2021
Earlier work this paper cites.
OpenCLIP, 2021
G. Ilharco, M. Wortsman, R. Wightman, C. Gordon, N. Carlini, R. Taori, A. Dave, V. Shankar, H. Namkoong, J. Miller, H. Hajishirzi, A. Farhadi, and L. Schmidt · 2021
Earlier work this paper cites.
Scaling up visual and vision-language representation learning with noisy text supervision
C. Jia, Y. Yang, Y. Xia, Y. Chen, Z. Parekh, H. Pham, Q. V. Le, Y. Sung, Z. Li, and T. Duerig · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever · 2021
Earlier work this paper cites.
Vision transformers for dense prediction
R. Ranftl, A. Bochkovskiy, and V. Koltun · 2021
Earlier work this paper cites.
Screen2words: Automatic mobile ui summarization with multimodal learning
B. Wang, G. Li, X. Zhou, Z. Chen, T. Grossman, and Y. Li · 2021
Cited alongside, same era.
PaLI: A jointly-scaled multilingual language-image model
X. Chen, X. Wang, S. Changpinyo, A. J. Piergiovanni, P. Padlewski, D. Salz, S. Goodman, A. Grycner, B. Mustafa, L. Beyer, A. Kolesnikov, J. Puigcerver, N. Ding, K. Rong, H. Akbari, G. Mishra, L. Xue, A. Thapliyal, J. Bradbury, W. Kuo, M. Seyedhosseini, C. Jia, B. K. Ayan, C. Riquelme, A. Steiner, A. Angelova, X. Zhai, N. Houlsby, and R. Soricut · 2022
Cited alongside, same era.
Decoupling zero-shot semantic segmentation
J. Ding, N. Xue, G.-S. Xia, and D. Dai · 2022
Cited alongside, same era.
Simple open-vocabulary object detection
M. Minderer, A. Gritsenko, A. Stone, M. Neumann, D. Weissenborn, A. Dosovitskiy, A. Mahendran, A. Arnab, M. Dehghani, Z. Shen, et al · 2022
Cited alongside, same era.
Prioritized training on points that are learnable, worth learning, and not yet learnt
S. Mindermann, J. M. Brauner, M. T. Razzak, M. Sharma, A. Kirsch, W. Xu, B. Höltgen, A. N. Gomez, A. Morisot, S. Farquhar, et al · 2022
Pytorch FSDP: experiences on scaling fully sharded data parallel
Y. Zhao, A. Gu, R. Varma, L. Luo, C. Huang, M. Xu, L. Wright, H. Shojanazeri, M. Ott, S. Shleifer, A. Desmaison, C. Balioglu, P. Damania, B. Nguyen, G. Chauhan, Y. Hao, A. Mathews, and S. Li · 2023
Later among the works it cites.
Clip the bias: How useful is balancing data in multimodal learning?
I. Alabdulmohsin, X. Wang, A. P. Steiner, P. Goyal, A. D’Amour, and X. Zhai · 2024
Later among the works it cites.
PaliGemma: A versatile 3B VLM for transfer
L. Beyer, A. Steiner, A. S. Pinto, A. Kolesnikov, X. Wang, D. Salz, M. Neumann, I. Alabdulmohsin, M. Tschannen, E. Bugliarello, T. Unterthiner, D. Keysers, S. Koppula, F. Liu, A. Grycner, A. Gritsenko, N. Houlsby, M. Kumar, K. Rong, J. Eisenschlos, R. Kabra, M. Bauer, M. Bošnjak, X. Chen, M. Minderer, P. Voigtlaender, I. Bica, I. Balazevic, J. Puigcerver, P. Papalampidi, O. Henaff, X. Xiong, R. Soricut, J. Harmsen, and X. Zhai · 2024
Later among the works it cites.
Cat-seg: Cost aggregation for open-vocabulary semantic segmentation
S. Cho, H. Shin, S. Hong, A. Arnab, P. H. Seo, and S. Kim · 2024
Later among the works it cites.
Patch n’pack: NaViT, a vision transformer for any aspect ratio and resolution
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
SLIP: Self-supervision meets language-image pre-training
N. Mu, A. Kirillov, D. Wagner, and S. Xie · 2022
Cited alongside, same era.
The dollar street dataset: Images representing the geographic and socioeconomic diversity of the world
W. A. G. Rojas, S. Diamos, K. R. Kini, D. Kanter, V. J. Reddi, and C. Coleman · 2022
Cited alongside, same era.
Crossmodal-3600: A massively multilingual multimodal evaluation dataset
A. V. Thapliyal, J. Pont Tuset, X. Chen, and R. Soricut · 2022
Cited alongside, same era.
SimVLM: Simple visual language model pretraining with weak supervision
Z. Wang, J. Yu, A. W. Yu, Z. Dai, Y. Tsvetkov, and Y. Cao · 2022
Cited alongside, same era.
CoCa: Contrastive captioners are image-text foundation models
J. Yu, Z. Wang, V. Vasudevan, L. Yeung, M. Seyedhosseini, and Y. Wu · 2022
Cited alongside, same era.
Image BERT pre-training with online tokenizer
J. Zhou, C. Wei, H. Wang, W. Shen, C. Xie, A. Yuille, and T. Kong · 2022
Cited alongside, same era.
Getting vit in shape: Scaling laws for compute-optimal model design
I. Alabdulmohsin, X. Zhai, A. Kolesnikov, and L. Beyer · 2023
Cited alongside, same era.
M. Dehghani, B. Mustafa, J. Djolonga, J. Heek, M. Minderer, M. Caron, A. Steiner, J. Puigcerver, R. Geirhos, I. M. Alabdulmohsin, et al · 2024
Later among the works it cites.
Data curation via joint example selection further accelerates multimodal learning
T. Evans, N. Parthasarathy, H. Merzic, and O. J. Henaff · 2024
Later among the works it cites.
Data filtering networks
A. Fang, A. M. Jose, A. Jain, L. Schmidt, A. T. Toshev, and V. Shankar · 2024
Later among the works it cites.
Multimodal autoregressive pre-training of large vision encoders
E. Fini, M. Shukor, X. Li, P. Dufter, M. Klein, D. Haldimann, S. Aitharaju, V. G. T. da Costa, L. Béthune, Z. Gan, A. T. Toshev, M. Eichner, M. Nabi, Y. Yang, J. M. Susskind, and A. El-Nouby · 2024
Later among the works it cites.
Datacomp: In search of the next generation of multimodal datasets
S. Y. Gadre, G. Ilharco, A. Fang, J. Hayase, G. Smyrnis, T. Nguyen, R. Marten, M. Wortsman, D. Ghosh, J. Zhang, et al · 2024
Later among the works it cites.
Introduction to Cloud TPU
Google Cloud · 2024
Later among the works it cites.
VeCLIP: Improving clip training via visual-enriched captions
Z. Lai, H. Zhang, B. Zhang, W. Wu, H. Bai, A. Timofeev, X. Du, Z. Gan, J. Shan, C.-N. Chuah, Y. Yang, and M. Cao · 2024
Later among the works it cites.
MM1: methods, analysis & insights from multimodal LLM pre-training
B. McKinzie, Z. Gan, J. Fauconnier, S. Dodge, B. Zhang, P. Dufter, D. Shah, X. Du, F. Peng, F. Weers, A. Belyi, H. Zhang, K. Singh, D. Kang, A. Jain, H. Hè, M. Schwarzer, T. Gunter, X. Kong, A. Zhang, J. Wang, C. Wang, N. Du, T. Lei, S. Wiseman, G. Yin, M. Lee, Z. Wang, R. Pang, P. Grasch, A. Toshev, and Y. Yang · 2024
Later among the works it cites.
SILC: Improving vision language pretraining with self-distillation
M. F. Naeem, Y. Xian, X. Zhai, L. Hoyer, L. Van Gool, and F. Tombari · 2024
Later among the works it cites.
Improving multimodal datasets with image captioning
T. Nguyen, S. Y. Gadre, G. Ilharco, S. Oh, and L. Schmidt · 2024
Later among the works it cites.
Dinov2: Learning robust visual features without supervision
M. Oquab, T. Darcet, T. Moutakanni, H. V. Vo, M. Szafraniec, V. Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby, et al · 2024
Later among the works it cites.
No filter: Cultural and socioeconomic diversityin contrastive vision-language models
A. Pouget, L. Beyer, E. Bugliarello, X. Wang, A. P. Steiner, X. Zhai, and I. Alabdulmohsin · 2024
Later among the works it cites.
Geode: a geographically diverse evaluation dataset for object recognition
V. V. Ramaswamy, S. Y. Lin, D. Zhao, A. Adcock, L. van der Maaten, D. Ghadiyaram, and O. Russakovsky · 2024
Later among the works it cites.
Paligemma 2: A family of versatile vlms for transfer
A. Steiner, A. S. Pinto, M. Tschannen, D. Keysers, X. Wang, Y. Bitton, A. Gritsenko, M. Minderer, A. Sherbondy, S. Long, et al · 2024
Later among the works it cites.
Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
S. Tong, E. Brown, P. Wu, S. Woo, M. Middepogu, S. C. Akula, J. Yang, S. Yang, A. Iyer, X. Pan, A. Wang, R. Fergus, Y. LeCun, and S. Xie · 2024
Later among the works it cites.
Active data curation effectively distills large-scale multimodal models
V. Udandarao, N. Parthasarathy, M. F. Naeem, T. Evans, S. Albanie, F. Tombari, Y. Xian, A. Tonioni, and O. J. Hénaff · 2024
Later among the works it cites.
LocCa: Visual pretraining with location-aware captioners
B. Wan, M. Tschannen, Y. Xian, F. Pavetic, I. Alabdulmohsin, X. Wang, A. S. Pinto, A. Steiner, L. Beyer, and X. Zhai · 2024
Later among the works it cites.
Demystifying clip data
H. Xu, S. Xie, X. Tan, P.-Y. Huang, R. Howes, V. Sharma, S.-W. Li, G. Ghosh, L. Zettlemoyer, and C. Feichtenhofer · 2024
Later among the works it cites.
TIPS: Text-image pretraining with spatial awareness
K.-K. Maninis, K. Chen, S. Ghosh, A. Karpur, K. Chen, Y. Xia, B. Cao, D. Salz, G. Han, J. Dlabal, et al · 2025
Closest in time.