Fetching the paper…
Reading the bibliography…
Recently, contrastive learning approaches (e.g., CLIP (Radford et al., 2021)) have received huge success in multimodal learning, where the model tries to minimize the distance between the representations of different views (e.g., image and its caption) of the same data point while keeping the representations of different data points away from each other.
A Theoretical Analysis of Contrastive Unsupervised Representation Learning, February 2019
Sanjeev Arora, Hrishikesh Khandeparkar, Mikhail Khodak, Orestis Plevrakis, and Nikunj Saunshi · 1902
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach, 2019
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 1907
Earlier work this paper cites.
Learning video representations from textual web supervision, 2020
Jonathan C. Stroud, Zhichao Lu, Chen Sun, Jia Deng, Rahul Sukthankar, Cordelia Schmid, and David A. Ross · 2007
Earlier work this paper cites.
Contrastive learning, multi-view redundancy, and linear models, April 2021
Christopher Tosh, Akshay Krishnamurthy, and Daniel Hsu · 2008
Earlier work this paper cites.
Microsoft COCO: Common Objects in Context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick · 2014
Earlier work this paper cites.
Deep residual learning for image recognition, 2015
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Fine-grained image classification via combining vision and language
X. He and Y. Peng · 2017
Earlier work this paper cites.
A mean field view of the landscape of two-layer neural networks
Song Mei, Andrea Montanari, and Phan-Minh Nguyen · 2018
Earlier work this paper cites.
Bootstrap your own latent - a new approach to self-supervised learning
Jean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec, Pierre Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Guo, Mohammad Gheshlaghi Azar, Bilal Piot, koray kavukcuoglu, Remi Munos, and Michal Valko · 2020
Earlier work this paper cites.
Momentum contrast for unsupervised visual representation learning
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick · 2020
Cited alongside, same era.
Debiased Contrastive Learning
Ching-Yao Chuang, Joshua Robinson, Yen-Chen Lin, Antonio Torralba, and Stefanie Jegelka · 2020
Cited alongside, same era.
Understanding Contrastive Representation Learning through Alignment and Uniformity on the Hypersphere
Tongzhou Wang and Phillip Isola · 2020
Cited alongside, same era.
Learning Transferable Visual Models From Natural Language Supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Cited alongside, same era.
Zero-shot text-to-image generation
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever · 2021
Cited alongside, same era.
Understanding self-supervised learning dynamics without contrastive pairs
Yuandong Tian, Xinlei Chen, and Surya Ganguli · 2021
Later among the works it cites.
Toward Understanding the Feature Learning Process of Self-supervised Contrastive Learning
Zixin Wen and Yuanzhi Li · 2021
Later among the works it cites.
What Makes Multi-Modal Learning Better than Single (Provably)
Yu Huang, Chenzhuang Du, Zihui Xue, Xuanyao Chen, Hang Zhao, and Longbo Huang · 2021
Later among the works it cites.
Understanding Negative Samples in Instance Discriminative Self-supervised Representation Learning
Kento Nozawa and Issei Sato · 2021
Later among the works it cites.
Provable Guarantees for Self-Supervised Deep Learning with Spectral Contrastive Loss
Jeff Z. HaoChen, Colin Wei, Adrien Gaidon, and Tengyu Ma · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A simple baseline for zero-shot semantic segmentation with pre-trained vision-language model, 2021
Mengde Xu, Zheng Zhang, Fangyun Wei, Yutong Lin, Yue Cao, Han Hu, and Xiang Bai · 2021
Cited alongside, same era.
Scaling up visual and vision-language representation learning with noisy text supervision
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig · 2021
Cited alongside, same era.
Exploring simple siamese representation learning
Xinlei Chen and Kaiming He · 2021
Cited alongside, same era.
Clip-td: Clip targeted distillation for vision-language tasks, 2022a
Zhecan Wang, Noel Codella, Yen-Chun Chen, Luowei Zhou, Jianwei Yang, Xiyang Dai, Bin Xiao, Haoxuan You, Shih-Fu Chang, and Lu Yuan
Cited in the paper.
Chaos is a Ladder: A New Theoretical Understanding of Contrastive Learning via Augmentation Overlap
Yifei Wang, Qi Zhang, Yisen Wang, Jiansheng Yang, and Zhouchen Lin
Cited in the paper.
Predicting What You Already Know Helps: Provable Self-Supervised Learning
Jason D Lee, Qi Lei, Nikunj Saunshi, and JIACHENG ZHUO · 2021
Later among the works it cites.
Understanding Dimensional Collapse in Contrastive Self-supervised Learning
Li Jing, Pascal Vincent, Yann LeCun, and Yuandong Tian · 2022
Later among the works it cites.
Contrasting the landscape of contrastive and non-contrastive learning, March 2022
Ashwini Pokle, Jinjin Tian, Yuchen Li, and Andrej Risteski · 2022
Later among the works it cites.