Fetching the paper…
Reading the bibliography…
Pre-training vision-language models with contrastive objectives has shown promising results that are both scalable to large uncurated datasets and transferable to many downstream applications.
Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories
L. Fei-Fei, R. Fergus, and P. Perona · 2004
Earlier work this paper cites.
Automated flower classification over a large number of classes
M.-E. Nilsback and A. Zisserman · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky, G. Hinton, et al · 2009
Earlier work this paper cites.
Cats and dogs
O. M. Parkhi, A. Vedaldi, A. Zisserman, and C. Jawahar · 2012
Earlier work this paper cites.
3d object representations for fine-grained categorization
J. Krause, M. Stark, J. Deng, and L. Fei-Fei · 2013
Earlier work this paper cites.
Fine-grained visual classification of aircraft
S. Maji, E. Rahtu, J. Kannala, M. Blaschko, and A. Vedaldi · 2013
Earlier work this paper cites.
Food-101–mining discriminative components with random forests
L. Bossard, M. Guillaumin, and L. V. Gool · 2014
Earlier work this paper cites.
Describing textures in the wild
M. Cimpoi, S. Maji, I. Kokkinos, S. Mohamed, and A. Vedaldi · 2014
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server
X. Chen, H. Fang, T.-Y. Lin, R. Vedantam, S. Gupta, P. Dollár, and C. L. Zitnick · 2015
Earlier work this paper cites.
Unsupervised visual representation learning by context prediction
C. Doersch, A. Gupta, and A. A. Efros · 2015
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
B. A. Plummer, L. Wang, C. M. Cervantes, J. C. Caicedo, J. Hockenmaier, and S. Lazebnik · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, et al · 2015
Earlier work this paper cites.
Facenet: A unified embedding for face recognition and clustering
F. Schroff, D. Kalenichenko, and J. Philbin · 2015
Earlier work this paper cites.
Going deeper with convolutions
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2015
Earlier work this paper cites.
Unsupervised learning of visual representations by solving jigsaw puzzles
M. Noroozi and P. Favaro · 2016
Earlier work this paper cites.
Context encoders: Feature learning by inpainting
D. Pathak, P. Krahenbuhl, J. Donahue, T. Darrell, and A. A. Efros · 2016
Cited alongside, same era.
Improved deep metric learning with multi-class n-pair loss objective
K. Sohn · 2016
Cited alongside, same era.
Yfcc100m: The new data in multimedia research
B. Thomee, D. A. Shamma, G. Friedland, B. Elizalde, K. Ni, D. Poland, D. Borth, and L.-J. Li · 2016
Cited alongside, same era.
Sun database: Exploring a large collection of scene categories
J. Xiao, K. A. Ehinger, J. Hays, A. Torralba, and A. Oliva · 2016
Cited alongside, same era.
Colorful image colorization
R. Zhang, P. Isola, and A. A. Efros · 2016
Cited alongside, same era.
Multi-task self-supervised visual learning
C. Doersch and A. Zisserman · 2017
Cited alongside, same era.
Eda: Easy data augmentation techniques for boosting performance on text classification tasks
J. Wei and K. Zou · 2019
Later among the works it cites.
A simple framework for contrastive learning of visual representations
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton · 2020
Later among the works it cites.
Randaugment: Practical automated data augmentation with a reduced search space
E. D. Cubuk, B. Zoph, J. Shlens, and Q. V. Le · 2020
Later among the works it cites.
Supervised contrastive learning
P. Khosla, P. Teterwak, C. Wang, A. Sarna, Y. Tian, P. Isola, A. Maschinot, C. Liu, and D. Krishnan · 2020
Later among the works it cites.
End-to-end learning of visual representations from uncurated instructional videos
A. Miech, J.-B. Alayrac, L. Smaira, I. Laptev, J. Sivic, and A. Zisserman · 2020
Later among the works it cites.
Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Representation learning by learning to count
M. Noroozi, H. Pirsiavash, and P. Favaro · 2017
Cited alongside, same era.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Cited alongside, same era.
Split-brain autoencoders: Unsupervised learning by cross-channel prediction
R. Zhang, P. Isola, and A. A. Efros · 2017
Cited alongside, same era.
Unsupervised representation learning by predicting image rotations
S. Gidaris, P. Singh, and N. Komodakis · 2018
Cited alongside, same era.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
P. Sharma, N. Ding, S. Goodman, and R. Soricut · 2018
Cited alongside, same era.
Representation learning with contrastive predictive coding
A. Van den Oord, Y. Li, and O. Vinyals · 2018
Cited alongside, same era.
S. Changpinyo, P. Sharma, N. Ding, and R. Soricut · 2021
Later among the works it cites.
Exploring simple siamese representation learning
X. Chen and K. He · 2021
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby · 2021
Later among the works it cites.
Scaling up visual and vision-language representation learning with noisy text supervision
C. Jia, Y. Yang, Y. Xia, Y.-T. Chen, Z. Parekh, H. Pham, Q. Le, Y.-H. Sung, Z. Li, and T. Duerig · 2021
Later among the works it cites.
Diffusionclip: Text-guided image manipulation using diffusion models
G. Kim and J. C. Ye · 2021
Later among the works it cites.
Slip: Self-supervision meets language-image pre-training
N. Mu, A. Kirillov, D. Wagner, and S. Xie · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Later among the works it cites.
Zero-shot text-to-image generation
A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. Voss, A. Radford, M. Chen, and I. Sutskever · 2021
Later among the works it cites.
Y. Cui, L. Zhao, F. Liang, Y. Li, and J. Shao · 2022
Closest in time.
Supervision exists everywhere: A data efficient contrastive language-image pre-training paradigm
Y. Li, F. Liang, L. Zhao, Y. Cui, W. Ouyang, J. Shao, F. Yu, and J. Yan · 2022
Closest in time.