Fetching the paper…
Reading the bibliography…
Pre-trained multi-modal vision-language models (VLMs) are becoming increasingly popular due to their exceptional performance on downstream vision applications, particularly in the few- and zero-shot settings.
Measuring dataset granularity, 2019
Yin Cui, Zeqi Gu, Dhruv Mahajan, Laurens van der Maaten, Serge Belongie, and Ser-Nam Lim · 1912
Earlier work this paper cites.
The use of multiple measurements in taxonomic problems
Ronald A Fisher · 1936
Earlier work this paper cites.
Silhouettes: A graphical aid to the interpretation and validation of cluster analysis
Peter J. Rousseeuw · 1987
Earlier work this paper cites.
Predicting neural network accuracy from weights, 2020
Thomas Unterthiner, Daniel Keysers, Sylvain Gelly, Olivier Bousquet, and Ilya Tolstikhin · 2002
Earlier work this paper cites.
The PASCAL Visual Object Classes Challenge 2007 (VOC2007) Results
M. Everingham, L. Van Gool, C. K. I. Williams, J. Winn, and A. Zisserman · 2007
Earlier work this paper cites.
Automated flower classification over a large number of classes
Maria-Elena Nilsback and Andrew Zisserman · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images, 2009
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Mnist handwritten digit database
Yann LeCun, Corinna Cortes, and CJ Burges · 2010
Earlier work this paper cites.
Sun database: Large-scale scene recognition from abbey to zoo
J. Xiao, J. Hays, K. A. Ehinger, A. Oliva, and A. Torralba · 2010
Earlier work this paper cites.
An analysis of single-layer networks in unsupervised feature learning
Adam Coates, Andrew Ng, and Honglak Lee · 2011
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Bo Wu, and Andrew Y. Ng · 2011
Earlier work this paper cites.
The german traffic sign recognition benchmark: a multi-class classification competition
Johannes Stallkamp, Marc Schlipsing, Jan Salmen, and Christian Igel · 2011
Earlier work this paper cites.
Cats and dogs
Omkar M. Parkhi, Andrea Vedaldi, Andrew Zisserman, and C. V. Jawahar · 2012
Earlier work this paper cites.
Challenges in representation learning: Facial expression recognition challenge, 2013
Yoshua Bengio Dumitru Ian Goodfellow, Will Cukierski · 2013
Earlier work this paper cites.
Vision meets robotics: The kitti dataset
Andreas Geiger, Philip Lenz, Christoph Stiller, and Raquel Urtasun · 2013
Earlier work this paper cites.
3d object representations for fine-grained categorization
Jonathan Krause, Michael Stark, Jia Deng, and Li Fei-Fei · 2013
Earlier work this paper cites.
Describing textures in the wild
M. Cimpoi, S. Maji, I. Kokkinos, S. Mohamed, and A. Vedaldi · 2014
Earlier work this paper cites.
Kaggle diabetic retinopathy detection, 2015
Kaggle and EyePacs · 2015
Cited alongside, same era.
Estimating accuracy from unlabeled data: A bayesian approach
Emmanouil Antonios Platanios, Avinava Dubey, and Tom Mitchell · 2016
Cited alongside, same era.
Remote sensing image scene classification: Benchmark and state of the art
Gong Cheng, Junwei Han, and Xiaoqiang Lu · 2017
Cited alongside, same era.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Justin Johnson, Bharath Hariharan, Laurens Van Der Maaten, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick · 2017
Cited alongside, same era.
Estimating accuracy from unlabeled data: A probabilistic logic approach
Emmanouil Platanios, Hoifung Poon, Tom M Mitchell, and Eric J Horvitz · 2017
Cited alongside, same era.
Pactran: Pac-bayesian metrics for estimating the transferability of pretrained models to classification tasks
Nan Ding, Xi Chen, Tomer Levinboim, Soravit Changpinyo, and Radu Soricut · 2022
Later among the works it cites.
Domino: Discovering systematic errors with cross-modal embeddings
Sabri Eyuboglu, Maya Varma, Khaled Kamal Saab, Jean-Benoit Delbrouck, Christopher Lee-Messer, Jared Dunnmon, James Zou, and Christopher Re · 2022
Later among the works it cites.
Data determines distributional robustness in contrastive language image pre-training (clip), 2022
Alex Fang, Gabriel Ilharco, Mitchell Wortsman, Yuhao Wan, Vaishaal Shankar, Achal Dave, and Ludwig Schmidt · 2022
Later among the works it cites.
Towards better selective classification, 2022
Leo Feng, Mohamed Osama Ahmed, Hossein Hajimirsadeghi, and Amir Abdi · 2022
Later among the works it cites.
Distilling model failures as directions in latent space, 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Florian Scheidegger, Roxana Istrate, Giovanni Mariani, Luca Benini, Costas Bekas, and Cristiano Malossi · 2018
Cited alongside, same era.
Rotation equivariant CNNs for digital pathology, June 2018
Bastiaan S Veeling, Jasper Linmans, Jim Winkens, Taco Cohen, and Max Welling · 2018
Cited alongside, same era.
Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification
Patrick Helber, Benjamin Bischke, Andreas Dengel, and Damian Borth · 2019
Cited alongside, same era.
A large-scale study of representation learning with the visual task adaptation benchmark, 2020
Xiaohua Zhai, Joan Puigcerver, Alexander Kolesnikov, Pierre Ruyssen, Carlos Riquelme, Mario Lucic, Josip Djolonga, Andre Susano Pinto, Maxim Neumann, Alexey Dosovitskiy, Lucas Beyer, Olivier Bachem, Michael Tschannen, Marcin Michalski, Olivier Bousquet, Sylvain Gelly, and Neil Houlsby · 2020
Cited alongside, same era.
Detecting errors and estimating accuracy on unlabeled data with self-training ensembles
Jiefeng Chen, Frederick Liu, Besim Avci, Xi Wu, Yingyu Liang, and Somesh Jha · 2021
Cited alongside, same era.
Are labels always necessary for classifier accuracy evaluation?
Weijian Deng and Liang Zheng · 2021
Cited alongside, same era.
Understanding dataset difficulty with 𝒱 \mathcal{V} -usable information, 2021
Kawin Ethayarajh, Yejin Choi, and Swabha Swayamdipta · 2021
Cited alongside, same era.
Saachi Jain, Hannah Lawrence, Ankur Moitra, and Aleksander Madry · 2022
Later among the works it cites.
Assessing generalization of SGD via disagreement
Yiding Jiang, Vaishnavh Nagarajan, Christina Baek, and J Zico Kolter · 2022
Later among the works it cites.
Mind the gap: Understanding the modality gap in multi-modal contrastive representation learning
Weixin Liang, Yuhui Zhang, Yongchan Kwon, Serena Yeung, and James Zou · 2022
Later among the works it cites.
Transferability estimation using bhattacharyya class separability
Michal Pándy, Andrea Agostinelli, Jasper Uijlings, Vittorio Ferrari, and Thomas Mensink · 2022
Later among the works it cites.
Transferability estimation using bhattacharyya class separability
Michal Pándy, Andrea Agostinelli, Jasper Uijlings, Vittorio Ferrari, and Thomas Mensink · 2022
Later among the works it cites.
LAION-5b: An open large-scale dataset for training next generation image-text models
Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade W Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, Patrick Schramowski, Srivatsa R Kundurthy, Katherine Crowson, Ludwig Schmidt, Robert Kaczmarczyk, and Jenia Jitsev · 2022
Later among the works it cites.
Chris van der Lee, Thiago Castro Ferreira, Chris Emmery, Travis Wiltshire, and Emiel Krahmer · 2022
Later among the works it cites.
Assessing the value of transfer learning metrics for rf domain adaptation, 2022
Lauren J. Wong, Sean McPherson, and Alan J. Michaels · 2022
Later among the works it cites.
Gpt-4 technical report, 2023
OpenAI · 2023
Closest in time.
What can we learn by predicting accuracy?
Olivier Risser-Maroix and Benjamin Chamand · 2023
Closest in time.
Llama: Open and efficient foundation language models, 2023
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample · 2023
Closest in time.
Image as a foreign language: BEiT pretraining for vision and vision-language tasks
Wenhui Wang, Hangbo Bao, Li Dong, Johan Bjorck, Zhiliang Peng, Qiang Liu, Kriti Aggarwal, Owais Khan Mohammed, Saksham Singhal, Subhojit Som, and Furu Wei · 2023
Closest in time.
Diagnosing and rectifying vision models using language
Yuhui Zhang, Jeff Z HaoChen, Shih-Cheng Huang, Kuan-Chieh Wang, James Zou, and Serena Yeung · 2023
Closest in time.