Fetching the paper…
Reading the bibliography…
We show that Contrastive Learning (CL) under a broad family of loss functions (including InfoNCE) has a unified formulation of coordinate-wise optimization on the network parameter $\boldsymbol{\theta}$ and pairwise importance $\alpha$, where the \emph{max player} $\boldsymbol{\theta}$ learns representation for contrastiveness, and the \emph{min player} $\alpha$ puts more weights on pairs of distinct samples that share similar representations.
The use of multiple measurements in taxonomic problems
Fisher, R. A · 1936
Earlier work this paper cites.
Principal component analysis
Wold, S., Esbensen, K., and Geladi, P · 1987
Earlier work this paper cites.
Neural networks and principal component analysis: Learning from examples without local minima
Baldi, P. and Hornik, K · 1989
Earlier work this paper cites.
Dimensionality reduction by learning an invariant mapping
Hadsell, R., Chopra, S., and LeCun, Y · 2006
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A., Hinton, G., et al · 2009
Earlier work this paper cites.
Understanding self-supervised learning with dual deep networks
Tian, Y., Yu, L., Chen, X., and Ganguli, S · 2010
Earlier work this paper cites.
An analysis of single-layer networks in unsupervised feature learning
Coates, A., Ng, A., and Lee, H · 2011
Earlier work this paper cites.
Trace optimization and eigenproblems in dimension reduction methods
Kokiopoulou, E., Chen, J., and Saad, Y · 2011
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Saxe, A. M., McClelland, J. L., and Ganguli, S · 2014
Earlier work this paper cites.
Facenet: A unified embedding for face recognition and clustering
Schroff, F., Kalenichenko, D., and Philbin, J · 2015
Earlier work this paper cites.
Ba, J. L., Kiros, J. R., and Hinton, G. E · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Deep learning without poor local minima
Kawaguchi, K · 2016
Earlier work this paper cites.
Deep metric learning via lifted structured feature embedding
Oh Song, H., Xiang, Y., Jegelka, S., and Savarese, S · 2016
Earlier work this paper cites.
Improved deep metric learning with multi-class n-pair loss objective
Sohn, K · 2016
Earlier work this paper cites.
Identity matters in deep learning
Hardt, M. and Ma, T · 2017
Earlier work this paper cites.
Densely connected convolutional networks
Huang, G., Liu, Z., Van Der Maaten, L., and Weinberger, K. Q · 2017
Cited alongside, same era.
Learning and generalization in overparameterized neural networks, going beyond two layers
Allen-Zhu, Z., Li, Y., and Liang, Y · 2018
Cited alongside, same era.
On the optimization of deep networks: Implicit acceleration by overparameterization
Arora, S., Cohen, N., and Hazan, E · 2018
Cited alongside, same era.
Mutual information neural estimation
Belghazi, M. I., Baratin, A., Rajeshwar, S., Ozair, S., Bengio, Y., Courville, A., and Hjelm, D · 2018
Cited alongside, same era.
Deep clustering for unsupervised learning of visual features
Caron, M., Bojanowski, P., Joulin, A., and Douze, M · 2018
Cited alongside, same era.
Deep linear networks with arbitrary loss: All local minima are global
Bootstrap your own latent: A new approach to self-supervised learning
Grill, J.-B., Strub, F., Altché, F., Tallec, C., Richemond, P. H., Buchatskaya, E., Doersch, C., Pires, B. A., Guo, Z. D., Azar, M. G., et al · 2020
Later among the works it cites.
Momentum contrast for unsupervised visual representation learning
He, K., Fan, H., Wu, Y., Xie, S., and Girshick, R. B · 2020
Later among the works it cites.
Hard negative mixing for contrastive learning
Kalantidis, Y., Sariyildiz, M. B., Pion, N., Weinzaepfel, P., and Larlus, D · 2020
Later among the works it cites.
Supervised contrastive learning
Khosla, P., Teterwak, P., Wang, C., Sarna, A., Tian, Y., Isola, P., Maschinot, A., Liu, C., and Krishnan, D · 2020
Later among the works it cites.
Self-supervised learning of pretext-invariant representations
Misra, I. and Maaten, L. v. d · 2020
Later among the works it cites.
Student specialization in deep relu networks with finite width and input dimension
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Laurent, T. and Brecht, J · 2018
Cited alongside, same era.
Representation learning with contrastive predictive coding
Oord, A. v. d., Li, Y., and Vinyals, O · 2018
Cited alongside, same era.
Spurious local minima are common in two-layer relu neural networks
Safran, I. and Shamir, O · 2018
Cited alongside, same era.
A theoretical framework for deep locally connected relu network
Tian, Y · 2018
Cited alongside, same era.
Unsupervised feature learning via non-parametric instance discrimination
Wu, Z., Xiong, Y., Yu, S. X., and Lin, D · 2018
Cited alongside, same era.
Critical points of linear neural networks: Analytical forms and landscape properties
Zhou, Y. and Liang, Y · 2018
Cited alongside, same era.
A theoretical analysis of contrastive unsupervised representation learning
Arora, S., Khandeparkar, H., Khodak, M., Plevrakis, O., and Saunshi, N · 2019
Cited alongside, same era.
Tian, Y · 2020
Later among the works it cites.
Contrastive multiview coding
Tian, Y., Krishnan, D., and Isola, P · 2020
Later among the works it cites.
On layer normalization in the transformer architecture
Xiong, R., Yang, Y., He, D., Zheng, K., Zheng, S., Xing, C., Zhang, H., Lan, Y., Wang, L., and Liu, T · 2020
Later among the works it cites.
Provable guarantees for self-supervised deep learning with spectral contrastive loss
HaoChen, J. Z., Wei, C., Gaidon, A., and Ma, T · 2021
Later among the works it cites.
The power of contrast for feature learning: A theoretical analysis
Ji, W., Deng, Z., Nakada, R., Zou, J., and Zhang, L · 2021
Later among the works it cites.
Predicting what you already know helps: Provable self-supervised learning
Lee, J. D., Lei, Q., Saunshi, N., and Zhuo, J · 2021
Later among the works it cites.
Contrastive learning with hard negative samples
Robinson, J., Chuang, C.-Y., Sra, S., and Jegelka, S · 2021
Later among the works it cites.
Toward understanding the feature learning process of self-supervised contrastive learning
Wen, Z. and Li, Y · 2021
Later among the works it cites.
Decoupled contrastive learning
Yeh, C.-H., Hong, C.-Y., Hsu, Y.-C., Liu, T.-L., Chen, Y., and LeCun, Y · 2021
Later among the works it cites.
Barlow twins: Self-supervised learning via redundancy reduction
Zbontar, J., Jing, L., Misra, I., LeCun, Y., and Deny, S · 2021
Later among the works it cites.
Understanding dimensional collapse in contrastive self-supervised learning
Jing, L., Vincent, P., LeCun, Y., and Tian, Y · 2022
Closest in time.