Fetching the paper…
Reading the bibliography…
Recent advances in contrastive representation learning over paired image-text data have led to models such as CLIP that achieve state-of-the-art performance for zero-shot classification and distributional robustness.
Learning a similarity metric discriminatively, with application to face verification
S. Chopra, R. Hadsell, and Y. LeCun · 2005
Earlier work this paper cites.
Dimensionality reduction by learning an invariant mapping
R. Hadsell, S. Chopra, and Y. LeCun · 2006
Earlier work this paper cites.
Learning visual representations using images with captions
A. Quattoni, M. Collins, and T. Darrell · 2007
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky, G. Hinton, et al · 2009
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
M. Gutmann and A. Hyvärinen · 2010
Earlier work this paper cites.
Representation learning: A review and new perspectives
Y. Bengio, A. Courville, and P. Vincent · 2013
Earlier work this paper cites.
Generative adversarial nets
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2014
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server
X. Chen, H. Fang, T.-Y. Lin, R. Vedantam, S. Gupta, P. Dollár, and C. L. Zitnick · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
A. Karpathy and L. Fei-Fei · 2015
Earlier work this paper cites.
Deep learning
Y. LeCun, Y. Bengio, and G. Hinton · 2015
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
B. A. Plummer, L. Wang, C. M. Cervantes, J. C. Caicedo, J. Hockenmaier, and S. Lazebnik · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, et al · 2015
Earlier work this paper cites.
Facenet: A unified embedding for face recognition and clustering
F. Schroff, D. Kalenichenko, and J. Philbin · 2015
Earlier work this paper cites.
Deep metric learning via lifted structured feature embedding
H. Oh Song, Y. Xiang, S. Jegelka, and S. Savarese · 2016
Earlier work this paper cites.
Improved deep metric learning with multi-class n-pair loss objective
K. Sohn · 2016
Earlier work this paper cites.
See, hear, and read: Deep aligned representations
Y. Aytar, C. Vondrick, and A. Torralba · 2017
Earlier work this paper cites.
Dualgan: Unsupervised dual learning for image-to-image translation
Z. Yi, H. Zhang, P. Tan, and M. Gong · 2017
Earlier work this paper cites.
Unpaired image-to-image translation using cycle-consistent adversarial networks
J.-Y. Zhu, T. Park, P. Isola, and A. A. Efros · 2017
Earlier work this paper cites.
Multimodal machine learning: A survey and taxonomy
T. Baltrušaitis, C. Ahuja, and L.-P. Morency · 2018
Earlier work this paper cites.
Stargan: Unified generative adversarial networks for multi-domain image-to-image translation
Y. Choi, M. Choi, M. Kim, J.-W. Ha, S. Kim, and J. Choo · 2018
Earlier work this paper cites.
Flow-gan: Combining maximum likelihood and adversarial learning in generative models
A. Grover, M. Dhar, and S. Ermon · 2018
Earlier work this paper cites.
Representation learning with contrastive predictive coding
A. v. d. Oord, Y. Li, and O. Vinyals · 2018
Cited alongside, same era.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
P. Sharma, N. Ding, S. Goodman, and R. Soricut · 2018
Cited alongside, same era.
Multimodal generative models for scalable weakly-supervised learning
M. Wu and N. Goodman · 2018
Cited alongside, same era.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al · 2019
Cited alongside, same era.
Do imagenet classifiers generalize to imagenet?
B. Recht, R. Roelofs, L. Schmidt, and V. Shankar · 2019
Cited alongside, same era.
Variational mixture-of-experts autoencoders for multi-modal deep generative models
Supervision exists everywhere: A data efficient contrastive language-image pre-training paradigm
Y. Li, F. Liang, L. Zhao, Y. Cui, W. Ouyang, J. Shao, F. Yu, and J. Yan · 2021
Later among the works it cites.
Pretrained transformers as universal computation engines
K. Lu, A. Grover, P. Abbeel, and I. Mordatch · 2021
Later among the works it cites.
Hybrid contrastive learning of tri-modal representation for multimodal sentiment analysis
S. Mai, Y. Zeng, S. Zheng, and H. Hu · 2021
Later among the works it cites.
Accuracy on the line: on the strong correlation between out-of-distribution and in-distribution generalization
J. P. Miller, R. Taori, A. Raghunathan, S. Sagawa, P. W. Koh, V. Shankar, P. Liang, Y. Carmon, and L. Schmidt · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Shi, B. Paige, P. Torr, et al · 2019
Cited alongside, same era.
Learning robust global representations by penalizing local predictive power
H. Wang, S. Ge, Z. Lipton, and E. P. Xing · 2019
Cited alongside, same era.
A simple framework for contrastive learning of visual representations
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton · 2020
Cited alongside, same era.
Bootstrap your own latent-a new approach to self-supervised learning
J.-B. Grill, F. Strub, F. Altché, C. Tallec, P. Richemond, E. Buchatskaya, C. Doersch, B. Avila Pires, Z. Guo, M. Gheshlaghi Azar, et al · 2020
Cited alongside, same era.
Alignflow: Cycle consistent learning from multiple domains via normalizing flows
A. Grover, C. Chute, R. Shu, Z. Cao, and S. Ermon · 2020
Cited alongside, same era.
Momentum contrast for unsupervised visual representation learning
K. He, H. Fan, Y. Wu, S. Xie, and R. Girshick · 2020
Cited alongside, same era.
Understanding contrastive representation learning through alignment and uniformity on the hypersphere
T. Wang and P. Isola · 2020
Cited alongside, same era.
N. Mu, A. Kirillov, D. Wagner, and S. Xie · 2021
Later among the works it cites.
Glide: Towards photorealistic image generation and editing with text-guided diffusion models
A. Nichol, P. Dhariwal, A. Ramesh, P. Shyam, P. Mishkin, B. McGrew, I. Sutskever, and M. Chen · 2021
Later among the works it cites.
Pervasive label errors in test sets destabilize machine learning benchmarks
C. G. Northcutt, A. Athalye, and J. Mueller · 2021
Later among the works it cites.
Combined scaling for zero-shot transfer learning
H. Pham, Z. Dai, G. Ghiasi, H. Liu, A. W. Yu, M.-T. Luong, M. Tan, and Q. V. Le · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Later among the works it cites.
Zero-shot text-to-image generation
A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. Voss, A. Radford, M. Chen, and I. Sutskever · 2021
Later among the works it cites.
Clip-forge: Towards zero-shot text-to-shape generation
A. Sanghi, H. Chu, J. G. Lambourne, Y. Wang, C.-Y. Cheng, and M. Fumero · 2021
Later among the works it cites.
Flava: A foundational language and vision alignment model
A. Singh, R. Hu, V. Goswami, G. Couairon, W. Galuba, M. Rohrbach, and D. Kiela · 2021
Later among the works it cites.
A fistful of words: Learning transferable visual models from bag-of-words supervision
A. Tejankar, B. Wu, S. Xie, M. Khabsa, H. Pirsiavash, and H. Firooz · 2021
Later among the works it cites.
Vlmo: Unified vision-language pre-training with mixture-of-modality-experts
W. Wang, H. Bao, L. Dong, and F. Wei · 2021
Later among the works it cites.
Multimodal contrastive training for visual representation learning
X. Yuan, Z. Lin, J. Kuen, J. Zhang, Y. Wang, M. Maire, A. Kale, and B. Faieta · 2021
Later among the works it cites.
Barlow twins: Self-supervised learning via redundancy reduction
J. Zbontar, L. Jing, I. Misra, Y. LeCun, and S. Deny · 2021
Later among the works it cites.
Merlot: Multimodal neural script knowledge models
R. Zellers, X. Lu, J. Hessel, Y. Yu, J. S. Park, J. Cao, A. Farhadi, and Y. Choi · 2021
Later among the works it cites.
Vinvl: Revisiting visual representations in vision-language models
P. Zhang, X. Li, X. Hu, J. Yang, L. Zhang, L. Wang, Y. Choi, and J. Gao · 2021
Later among the works it cites.
Dall-eval: Probing the reasoning skills and social biases of text-to-image generative transformers
J. Cho, A. Zala, and M. Bansal · 2022
Closest in time.
Vqgan-clip: Open domain image generation and editing with natural language guidance
K. Crowson, S. R. Biderman, D. Kornis, D. Stander, E. Hallahan, L. Castricato, and E. Raff · 2022
Closest in time.
Multi-modal alignment using representation codebook
J. Duan, L. Chen, S. Tran, J. Yang, Y. Xu, B. Zeng, C. Tao, and T. Chilimbi · 2022
Closest in time.
Hierarchical text-conditional image generation with clip latents
A. Ramesh, P. Dhariwal, A. Nichol, C. Chu, and M. Chen · 2022
Closest in time.