Fetching the paper…
Reading the bibliography…
Despite the empirical successes of self-supervised learning (SSL) methods, it is unclear what characteristics of their representations lead to high downstream accuracies.
S. Straszewicz, “Over exposed points of closed point sets,” Fundamenta Mathematicae , vol. 24, pp. 139–143, 1935
1935
Earlier work this paper cites.
E. T. Jaynes, “Information theory and statistical mechanics,” Physical Review , vol. 106, no. 4, pp. 620–630, 1957
1957
Earlier work this paper cites.
T. M. Cover, “Geometrical and statistical properties of systems of linear inequalities with applications in pattern recognition,” IEEE Transactions on Electronic Computers , vol. EC-14, no. 3, pp. 326–334, Jun. 1965
1965
Earlier work this paper cites.
R. Sinkhorn and P. Knopp, “Concerning nonnegative matrices and doubly stochastic matrices,” Pacific Journal of Mathematics , vol. 21, no. 2, pp. 343–348, 1967
1967
Earlier work this paper cites.
B. Grünbaum, V. Klee, M. A. Perles, and G. C. Shephard, Convex polytopes . Springer, 1967, vol. 16
1967
Earlier work this paper cites.
T. M. Cover, “Capacity problems for linear machines,” Pattern recognition , pp. 283–289, 1968
1968
Earlier work this paper cites.
R. T. Rockafellar, Convex analysis . Princeton university press, 1970
1970
Earlier work this paper cites.
V. N. Vapnik and A. Y. Chervonenkis, “On uniform convergence of the frequencies of events to their probabilities,” Teoriya Veroyatnostei i ee Primeneniya , vol. 16, no. 2, pp. 264–279, 1971
1971
Earlier work this paper cites.
A. Brøndsted, An introduction to convex polytopes . Springer Science & Business Media, 1983, no. 90
1983
Earlier work this paper cites.
E. B. Baum, “On the capabilities of multilayer perceptrons,” Journal of Complexity , vol. 4, no. 3, pp. 193–215, 1988
1988
Earlier work this paper cites.
M. L. Eaton, “Group invariance applications in statistics,” Regional Conference Series in Probability and Statistics , vol. 1, pp. i–133, 1989
1989
Earlier work this paper cites.
G. J. Mitchison and R. M. Durbin, “Bounds on the learning capacity of some multi-layer networks,” Biological Cybernetics , vol. 60, no. 5, pp. 345–365, 1989
1989
Earlier work this paper cites.
E. D. Sontag, “Remarks on interpolation and recognition using neural nets,” in Advances in Neural Information Processing Systems (NeurIPS) , 1990
1990
Earlier work this paper cites.
P. Flajolet, D. Gardy, and L. Thimonier, “Birthday paradox, coupon collectors, caching algorithms and self-organizing search,” Discrete Applied Mathematics , vol. 39, no. 3, pp. 207–229, 1992
1992
Earlier work this paper cites.
A. Sakurai, “n-h-1 networks store no less n*h+1 examples, but sometimes no more,” in International Joint Conference on Neural Networks (IJCNN) , vol. 3, 1992
1992
Earlier work this paper cites.
T. Leen, “From data distributions to regularization in invariant learning,” in Advances in Neural Information Processing Systems (NeurIPS) , 1994
1994
Earlier work this paper cites.
C. Cortes and V. Vapnik, “Support-vector networks,” Machine Learning , vol. 20, pp. 273–297, 1995
1995
Earlier work this paper cites.
A. Kowalczyk, “Dense shattering and teaching dimensions for differentiable families (extended abstract),” in Conference on Learning Theory (COLT) , 1997
1997
Earlier work this paper cites.
——, “Shattering all sets of k points in ’general position’ requires (k - 1)/2 parameters,” Neural Computation , vol. 9, no. 2, pp. 337–348, 1997
1997
Earlier work this paper cites.
A. Kowalczyk, “Estimates of storage capacity of multilayer perceptron with threshold logic hidden units,” Neural Networks , vol. 10, no. 8, pp. 1417–1433, 1997
1997
Earlier work this paper cites.
O. Kallenberg, Foundations of modern probability . Springer, 1997
1997
Earlier work this paper cites.
C. J. C. Burges, “A tutorial on support vector machines for pattern recognition,” Data mining and knowledge discovery , vol. 2, no. 2, pp. 121–167, 1998
1998
Earlier work this paper cites.
C. Zalinescu, Convex Analysis in General Vector Spaces . World Scientific, 2002
2002
Earlier work this paper cites.
D. J. C. MacKay, Information theory, inference, and learning algorithms . Cambridge University Press, 2003
2003
Earlier work this paper cites.
S. Boyd and L. Vandenberghe, Convex Optimization . Cambridge University Press, 2004
2004
Earlier work this paper cites.
E. L. Lehmann and J. P. Romano, Testing Statistical Hypotheses, Third Edition , 3rd ed., ser. Springer texts in statistics. Springer, 2008
2008
Earlier work this paper cites.
W. Feller, An introduction to probability theory and its applications . John Wiley & Sons, 2008
2008
Earlier work this paper cites.
M.-E. Nilsback and A. Zisserman, “Automated flower classification over a large number of classes,” in Indian Conference on Computer Vision, Graphics & Image Processing (ICVGIP) , 2008
2008
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in Conference on Computer Vision and Pattern Recognition (CVPR) , 2009
2009
Earlier work this paper cites.
P. Berenbrink and T. Sauerwald, “The weighted coupon collector’s problem and applications,” in International Computing and Combinatorics Conference (COCOON) , 2009
2009
Earlier work this paper cites.
A. Krizhevsky, “Learning multiple layers of features from tiny images,” University of Toronto, Tech. Rep., 2009
2009
Earlier work this paper cites.
M. Gutmann and A. Hyvärinen, “Noise-contrastive estimation: A new estimation principle for unnormalized statistical models,” in Artificial Intelligence and Statistics (AISTATS) , 2010
2010
Earlier work this paper cites.
L. Narici and E. Beckenstein, Topological Vector Spaces , 2nd ed. Chapman and Hall/CRC, 2010
2010
Earlier work this paper cites.
S. Dutta and A. Goswami, “Mode estimation for discrete distributions,” Mathematical Methods of Statistics , vol. 19, no. 4, pp. 374–384, 2010
2010
Earlier work this paper cites.
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay, “Scikit-learn: Machine learning in Python,” Journal of Machine Learning Research (JMLR) , vol. 12, 2011
2011
Earlier work this paper cites.
G. M. Ziegler, Lectures on polytopes . Springer Science & Business Media, 2012, vol. 152
2012
Earlier work this paper cites.
O. M. Parkhi, A. Vedaldi, A. Zisserman, and C. V. Jawahar, “Cats and dogs,” in Conference on Computer Vision and Pattern Recognition (CVPR) , 2012
2012
Earlier work this paper cites.
A. Mnih and K. Kavukcuoglu, “Learning word embeddings efficiently with noise-contrastive estimation,” in Advances in Neural Information Processing Systems (NeurIPS) , 2013
2013
Earlier work this paper cites.
Y. Bengio, A. C. Courville, and P. Vincent, “Representation learning: A review and new perspectives,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 35, no. 8, pp. 1798–1828, 2013
2013
Earlier work this paper cites.
J. Krause, M. Stark, J. Deng, and L. Fei-Fei, “3d object representations for fine-grained categorization,” in ICCV Wokshop on 3D Representation and Recognition , 2013
2013
Earlier work this paper cites.
2013
Earlier work this paper cites.
S. Shalev-Shwartz and T. Zhang, “Accelerated proximal stochastic dual coordinate ascent for regularized loss minimization,” Mathematical Programming , pp. 1–41, 2014
2014
Earlier work this paper cites.
S. Shalev-Shwartz and S. Ben-David, Understanding Machine Learning: From Theory to Algorithms . Cambridge University Press, 2014
2014
Earlier work this paper cites.
L. Bossard, M. Guillaumin, and L. V. Gool, “Food-101 – mining discriminative components with random forests,” in European Conference on Computer Vision (ECCV) , 2014
2014
Earlier work this paper cites.
M. Cimpoi, S. Maji, I. Kokkinos, S. Mohamed, , and A. Vedaldi, “Describing textures in the wild,” in Conference on Computer Vision and Pattern Recognition (CVPR) , 2014
2014
Earlier work this paper cites.
cs231n, “Tinyimagenet,” 2015
2015
Cited alongside, same era.
S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in International Conference on Machine Learning (ICML) , 2015, pp. 448–456
2015
Cited alongside, same era.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in International Conference on Learning Representations (ICLR) , 2015
2015
Cited alongside, same era.
2016
Cited alongside, same era.
R. Zhang, P. Isola, and A. A. Efros, “Colorful image colorization,” in European Conference on Computer Vision (ECCV) , 2016
2016
Cited alongside, same era.
F. Graf, C. Hofer, M. Niethammer, and R. Kwitt, “Dissecting supervised contrastive learning,” in International Conference on Machine Learning (ICML) , 2021
2021
Later among the works it cites.
J. Zarka, F. Guth, and S. Mallat, “Separation and concentration in deep networks,” in International Conference on Learning Representations (ICLR) , 2021
2021
Later among the works it cites.
W. E and S. Wojtowytsch, “On the emergence of simplex symmetry in the final and penultimate layers of neural network classifiers,” in Mathematical and Scientific Machine Learning Conference (MSML) , 2021
2021
Later among the works it cites.
Z. Zhu, T. Ding, J. Zhou, X. Li, C. You, J. Sulam, and Q. Qu, “A geometric analysis of neural collapse with unconstrained features,” in Advances in Neural Information Processing Systems (NeurIPS) , 2021
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Yang, D. Parikh, and D. Batra, “Joint unsupervised learning of deep representations and image clusters,” in Conference on Computer Vision and Pattern Recognition (CVPR) , 2016
2016
Cited alongside, same era.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Computer Vision and Pattern Recognition (CVPR) , 2016, pp. 770–778
2016
Cited alongside, same era.
Z. Ma and M. Collins, “Noise contrastive estimation and negative sampling for conditional models: Consistency and statistical efficiency,” in Empirical Methods in Natural Language Processing , 2018
2018
Cited alongside, same era.
M. Caron, P. Bojanowski, A. Joulin, and M. Douze, “Deep clustering for unsupervised learning of visual features,” in European Conference on Computer Vision (ECCV) , 2018
2018
Cited alongside, same era.
Z. Wu, Y. Xiong, S. X. Yu, and D. Lin, “Unsupervised feature learning via non-parametric instance discrimination,” in Conference on Computer Vision and Pattern Recognition (CVPR) , 2018
2018
Cited alongside, same era.
L. McInnes, J. Healy, N. Saul, and L. Großberger, “UMAP: uniform manifold approximation and projection,” Journal of Open Source Software , vol. 3, no. 29, p. 861, 2018
2018
Cited alongside, same era.
N. Saunshi, O. Plevrakis, S. Arora, M. Khodak, and H. Khandeparkar, “A theoretical analysis of contrastive unsupervised representation learning,” in International Conference on Machine Learning (ICML) , 2019
2019
Cited alongside, same era.
M. Caron, H. Touvron, I. Misra, H. J’egou, J. Mairal, P. Bojanowski, and A. Joulin, “Emerging properties in self-supervised vision transformers,” in International Conference on Computer Vision (ICCV) , 2021
2021
Later among the works it cites.
X. Chen and K. He, “Exploring simple siamese representation learning,” in Conference on Computer Vision and Pattern Recognition (CVPR) , 2021
2021
Later among the works it cites.
Z. Fang, J. Wang, L. Wang, L. Zhang, Y. Yang, and Z. Liu, “SEED: Self-supervised distillation for visual representation,” in International Conference on Learning Representations (ICLR) , 2021
2021
Later among the works it cites.
A. Ermolov, A. Siarohin, E. Sangineto, and N. Sebe, “Whitening for self-supervised representation learning,” in International Conference on Machine Learning (ICML) , 2021
2021
Later among the works it cites.
J. Zbontar, L. Jing, I. Misra, Y. LeCun, and S. Deny, “Barlow Twins: Self-supervised learning via redundancy reduction,” in International Conference on Machine Learning (ICML) , 2021
2021
Later among the works it cites.
T. Hua, W. Wang, Z. Xue, S. Ren, Y. Wang, and H. Zhao, “On feature decorrelation in self-supervised learning,” in International Conference on Computer Vision (ICCV) , 2021
2021
Later among the works it cites.
T. Chen, C. Luo, and L. Li, “Intriguing properties of contrastive losses,” in Advances in Neural Information Processing Systems (NeurIPS) , 2021
2021
Later among the works it cites.
Y. H. Tsai, Y. Wu, R. R. Salakhutdinov, and L. Morency, “Self-supervised learning from a multi-view perspective,” in International Conference on Learning Representations (ICLR) , 2021
2021
Later among the works it cites.
J. Mitrovic, B. McWilliams, J. Walker, L. Buesing, and C. Blundell, “Representation learning via invariant causal mechanisms,” in International Conference on Learning Representations (ICLR) , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
C. Tosh, A. Krishnamurthy, and D. Hsu, “Contrastive estimation reveals topic posterior information to linear models,” Journal of Machine Learning Research (JMLR) , vol. 22, pp. 281:1–281:31, 2021
2021
Later among the works it cites.
K. Nozawa and I. Sato, “Understanding negative samples in instance discriminative self-supervised representation learning,” in Advances in Neural Information Processing Systems (NeurIPS) , 2021
2021
Later among the works it cites.
Y. Tian, X. Chen, and S. Ganguli, “Understanding self-supervised learning dynamics without contrastive pairs,” in International Conference on Machine Learning (ICML) , 2021
2021
Later among the works it cites.
P. Goyal, Q. Duval, J. Reizenstein, M. Leavitt, M. Xu, B. Lefaudeux, M. Singh, V. Reis, M. Caron, P. Bojanowski, A. Joulin, and I. Misra, “VISSL,” 2021
2021
Later among the works it cites.
S. Kornblith, T. Chen, H. Lee, and M. Norouzi, “Why do better loss functions lead to less transferable features?” in Advances in Neural Information Processing Systems (NeurIPS) , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
B. O’Neill, “The classical occupancy distribution: Computation and approximation,” The American Statistician , vol. 75, no. 4, pp. 364–375, 2021
2021
Later among the works it cites.
C. Fang, H. He, Q. Long, and W. J. Su, “Exploring deep neural networks via layer-peeled model: Minority collapse in imbalanced training,” Proceedings of the National Academy of Sciences , vol. 118, 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
F. Wang and H. Liu, “Understanding the behaviour of contrastive loss,” in Conference on Computer Vision and Pattern Recognition (CVPR) , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
A. Foster, R. Pukdee, and T. Rainforth, “Improving transformation invariance in contrastive representation learning,” in International Conference on Learning Representations (ICLR) , 2021
2021
Later among the works it cites.
J. von Kügelgen, Y. Sharma, L. Gresele, W. Brendel, B. Schölkopf, M. Besserve, and F. Locatello, “Self-supervised learning with data augmentations provably isolates content from style,” in NeurIPS , 2021
2021
Later among the works it cites.
J. Lu and S. Steinerberger, “Neural collapse under cross-entropy loss,” Applied and Computational Harmonic Analysis , vol. 59, pp. 224–241, 2022
2022
Closest in time.
W. Ji, Y. Lu, Y. Zhang, Z. Deng, and W. J. Su, “An unconstrained layer-peeled perspective on neural collapse,” in International Conference on Learning Representations (ICLR) , 2022
2022
Closest in time.
L. Jing, P. Vincent, Y. LeCun, and Y. Tian, “Understanding dimensional collapse in contrastive self-supervised learning,” in International Conference on Learning Representations (ICLR) , 2022
2022
Closest in time.
N. Saunshi, J. T. Ash, S. Goel, D. Misra, C. Zhang, S. Arora, S. M. Kakade, and A. Krishnamurthy, “Understanding contrastive learning requires incorporating inductive biases,” in International Conference on Machine Learning (ICML) , 2022
2022
Closest in time.
Y. Ruan, Y. Dubois, and C. J. Maddison, “Optimal representations for covariate shift,” in International Conference on Learning Representations (ICLR) , 2022
2022
Closest in time.
A. Xu and M. Raginsky, “Minimum excess risk in bayesian learning,” IEEE Transactions on Information Theory , 2022
2022
Closest in time.
J. T. Ash, S. Goel, A. Krishnamurthy, and D. Misra, “Investigating the role of negatives in contrastive representation learning,” in Artificial Intelligence and Statistics (AISTATS) , 2022
2022
Closest in time.
P. Awasthi, N. Dikkala, and P. Kamath, “Do more negative samples necessarily hurt in contrastive learning?” in International Conference on Machine Learning (ICML) , 2022
2022
Closest in time.
Y. Wang, Y. Zhang, Y. Wang, J. Yang, and Z. Lin, “Chaos is a ladder: A new theoretical understanding of contrastive learning via augmentation overlap,” in International Conference on Learning Representations (ICLR) , 2022
2022
Closest in time.
D. G. Mixon, H. Parshall, and J. Pi, “Neural collapse with unconstrained features,” Sampling Theory, Signal Processing, and Data Analysis , vol. 20, no. 2, pp. 1–13, 2022
2022
Closest in time.
C. Zhang, K. Zhang, C. Zhang, T. X. Pham, C. D. Yoo, and I. S. Kweon, “How does simsiam avoid collapse without negative samples? a unified understanding with self-supervised contrastive learning,” in ICLR , 2022
2022
Closest in time.
T. Galanti, A. György, and M. Hutter, “On the role of neural collapse in transfer learning,” in International Conference on Learning Representations (ICLR) , 2022
2022
Closest in time.
2022
Closest in time.
A. Bardes, J. Ponce, and Y. LeCun, “VICReg: Variance-invariance-covariance regularization for self-supervised learning,” in International Conference on Learning Representations (ICLR) , 2022
2022
Closest in time.
A. Pokle, J. Tian, Y. Li, and A. Risteski, “Contrasting the landscape of contrastive and non-contrastive learning,” in Artificial Intelligence and Statistics (AISTATS) , 2022
2022
Closest in time.
F.-F. Li, M. Andreeto, M. A. Ranzato, and P. Perona, “Caltech 101,” 2022
2022
Closest in time.