Fetching the paper…
Reading the bibliography…
The development of CLIP [Radford et al., 2021] has sparked a debate on whether language supervision can result in vision models with more transferable representations than traditional image-only methods.
“Logic and Conversation”
H.. Grice · 1975
Earlier work this paper cites.
“Learning visual representations using images with captions”
A. Quattoni, M. Collins and T. Darrell · 2007
Earlier work this paper cites.
“ImageNet: A large-scale hierarchical image database”
J. Deng et al · 2009
Earlier work this paper cites.
“Multimodal learning with deep boltzmann machines”
N. Srivastava and R.. Salakhutdinov · 2012
Earlier work this paper cites.
“Devise: A deep visual-semantic embedding model”
A. Frome et al · 2013
Earlier work this paper cites.
“Framing image description as a ranking task: Data, models and evaluation metrics”
M. Hodosh, P. Young and J. Hockenmaier · 2013
Earlier work this paper cites.
“Analyzing the performance of multilayer neural networks for object recognition”
P. Agrawal, R. Girshick and J. Malik · 2014
Earlier work this paper cites.
“Return of the devil in the details: Delving deep into convolutional nets”
K. Chatfield, K. Simonyan, A. Vedaldi and A. Zisserman · 2014
Earlier work this paper cites.
“DeCAF: A deep convolutional activation feature for generic visual recognition”
J. Donahue et al · 2014
Earlier work this paper cites.
“Microsoft coco: Common objects in context”
T. Lin et al · 2014
Earlier work this paper cites.
“CNN Features off-the-shelf: an Astounding Baseline for Recognition.”
A.. Razavian, H. Azizpour, J. Sullivan and S. Carlsson · 2014
Earlier work this paper cites.
“How transferable are features in deep neural networks?”
J. Yosinski, J. Clune, Y. Bengio and H. Lipson · 2014
Earlier work this paper cites.
“Factors of transferability for a generic convnet representation”
H. Azizpour et al · 2015
Earlier work this paper cites.
“Microsoft coco captions: Data collection and evaluation server”
X. Chen et al · 2015
Earlier work this paper cites.
“ImageNet Large Scale Visual Recognition Challenge”
O. Russakovsky et al · 2015
Earlier work this paper cites.
“Best practices for fine-tuning visual classifiers to new domains”
B. Chu et al · 2016
Earlier work this paper cites.
“What makes ImageNet good for transfer learning?”
M. Huh, P. Agrawal and A.. Efros · 2016
Earlier work this paper cites.
“Deep Residual Learning for Image Recognition”
K. He, X. Zhang, S. Ren and J. Sun · 2016
Earlier work this paper cites.
“YFCC100M: The new data in multimedia research”
B. Thomee et al · 2016
Earlier work this paper cites.
“Bag of Tricks for Efficient Text Classification”
A. Joulin, E. Grave, P. Bojanowski and T. Mikolov · 2017
Earlier work this paper cites.
A. Vaswani et al · 2017
Earlier work this paper cites.
“Multimodal machine learning: A survey and taxonomy”
T. Baltrušaitis, C. Ahuja and L. Morency · 2018
Earlier work this paper cites.
“Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning”
P. Sharma, N. Ding, S. Goodman and R. Soricut · 2018
Earlier work this paper cites.
“Unsupervised feature learning via non-parametric instance discrimination”
Z. Wu, Y. Xiong, S.. Yu and D. Lin · 2018
Cited alongside, same era.
“ObjectNet: A large-scale bias-controlled dataset for pushing the limits of object recognition models”
A. Barbu et al · 2019
Cited alongside, same era.
“PyTorch Lightning”, 2019
W. Falcon and the PyTorch · 2019
Cited alongside, same era.
“Deep multimodal representation learning: A survey”
W. Guo, J. Wang and S. Wang · 2019
Cited alongside, same era.
“Do better imagenet models transfer better?”
S. Kornblith, J. Shlens and Q.. Le · 2019
Cited alongside, same era.
“Do ImageNet Classifiers Generalize to ImageNet?”
B. Recht, R. Roelofs, L. Schmidt and V. Shankar · 2019
Cited alongside, same era.
“Does language help generalization in vision models?”
B. Devillers, B. Choksi, R. Bielawski and R. VanRullen · 2021
Later among the works it cites.
“How well do self-supervised models transfer?”
L. Ericsson, H. Gouk and T.. Hospedales · 2021
Later among the works it cites.
“The many faces of robustness: A critical analysis of out-of-distribution generalization”
D. Hendrycks et al · 2021
Later among the works it cites.
“Provable guarantees for self-supervised deep learning with spectral contrastive loss”
J.. HaoChen, C. Wei, A. Gaidon and T. Ma · 2021
Later among the works it cites.
“Natural adversarial examples”
D. Hendrycks et al · 2021
Later among the works it cites.
“OpenCLIP”, 2021
G. Ilharco et al · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Learning robust global representations by penalizing local predictive power”
H. Wang, S. Ge, E.. Xing and Z.. Lipton · 2019
Cited alongside, same era.
“A large-scale study of representation learning with the visual task adaptation benchmark”
X. Zhai et al · 2019
Cited alongside, same era.
“Language Models are Few-Shot Learners”
T.. Brown et al · 2020
Cited alongside, same era.
“A simple framework for contrastive learning of visual representations”
T. Chen, S. Kornblith, M. Norouzi and G. Hinton · 2020
Cited alongside, same era.
“Big self-supervised models are strong semi-supervised learners”
T. Chen et al · 2020
Cited alongside, same era.
“Unsupervised learning of visual features by contrasting cluster assignments”
M. Caron et al · 2020
Cited alongside, same era.
E. Kreiss, N.. Goodman and C. Potts · 2021
Later among the works it cites.
“SLIP: Self-supervision meets Language-Image Pre-training”
N. Mu, A. Kirillov, D. Wagner and S. Xie · 2021
Later among the works it cites.
“Representation Learning via Invariant Causal Mechanisms”
J. Mitrovic et al · 2021
Later among the works it cites.
“Learning transferable visual models from natural language supervision”
A. Radford et al · 2021
Later among the works it cites.
“LAION-400M: Open dataset of clip-filtered 400 million image-text pairs”
C. Schuhmann et al · 2021
Later among the works it cites.
“Divide and contrast: Self-supervised learning from uncurated data”
Y. Tian, O.. Henaff and A. van Oord · 2021
Later among the works it cites.
“Self-supervised Learning from a Multi-view Perspective”
Y.. Tsai, Y. Wu, R.. Salakhutdinov and L. Morency · 2021
Later among the works it cites.
“GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model”, https://github.com/kingoflolz/mesh-transformer-jax , 2021
B. Wang and A. Komatsuzaki · 2021
Later among the works it cites.
“When does contrastive visual representation learning work?”
Elijah Cole et al · 2022
Closest in time.
“Image Captioning as an Assistive Technology: Lessons Learned from VizWiz 2020 Challenge”
P. Dognin et al · 2022
Closest in time.
“Data Determines Distributional Robustness in Contrastive Language Image Pre-training (CLIP)”
A. Fang et al · 2022
Closest in time.
J. Li, D. Li, C. Xiong and S. Hoi · 2022
Closest in time.
“Supervision exists everywhere: A data efficient contrastive language-image pre-training paradigm”
Y. Li et al · 2022
Closest in time.
“Optimal Representations for Covariate Shift”
Y. Ruan, Y. Dubois and C.. Maddison · 2022
Closest in time.
“On Mutual Information in Contrastive Learning for Visual Representations”
M. Wu et al · 2022
Closest in time.
“FILIP: Fine-grained Interactive Language-Image Pre-Training”
L. Yao et al · 2022
Closest in time.
“LiT: Zero-Shot Transfer with Locked-image Text Tuning”
X. Zhai et al · 2022
Closest in time.