Fetching the paper…
Reading the bibliography…
Self-Supervised Learning (SSL) methods such as VICReg, Barlow Twins or W-MSE avoid collapse of their joint embedding architectures by constraining or regularizing the covariance matrix of their projector's output.
Eléments aléatoires dans un espace de banach
Mourier, E · 1953
Earlier work this paper cites.
On measures of dependence
Rényi, A · 1959
Earlier work this paper cites.
Misspecifications of the normal distribution
Melnick, E. L. and Tenenbein, A · 1982
Earlier work this paper cites.
Learning factorial codes by predictability minimization
Schmidhuber, J · 1992
Earlier work this paper cites.
Canonical correlation analysis when the data are curves
Leurgans, S. E., Moyeed, R. A., and Silverman, B. W · 1993
Earlier work this paper cites.
Independent component analysis, a new concept?
Comon, P · 1994
Earlier work this paper cites.
Elements of information theory
Cover, T. M · 1999
Earlier work this paper cites.
Source separation in post-nonlinear mixtures
Taleb, A. and Jutten, C · 1999
Earlier work this paper cites.
Independent component analysis: algorithms and applications
Hyvärinen, A. and Oja, E · 2000
Earlier work this paper cites.
On the influence of the kernel on the consistency of support vector machines
Steinwart, I · 2001
Earlier work this paper cites.
Kernel independent component analysis
Bach, F. R. and Jordan, M. I · 2002
Earlier work this paper cites.
Tsp speech database, 2002
Kabal, P · 2002
Earlier work this paper cites.
An approximation to the distribution of finite sample size mutual information estimates
Goebel, B., Dawy, Z., Hagenauer, J., and Mueller, J. C · 2005
Earlier work this paper cites.
Universal kernels
Micchelli, C. A., Xu, Y., and Zhang, H · 2006
Earlier work this paper cites.
Random features for large-scale kernel machines
Rahimi, A. and Recht, B · 2007
Earlier work this paper cites.
A hilbert space embedding for distributions
Smola, A., Gretton, A., Song, L., and Schölkopf, B · 2007
Earlier work this paper cites.
Visualizing higher-layer features of a deep network
Erhan, D., Bengio, Y., Courville, A., and Vincent, P · 2009
Earlier work this paper cites.
Regression by dependence minimization and its application to causal inference in additive noise models
Mooij, J., Janzing, D., Peters, J., and Schölkopf, B · 2009
Earlier work this paper cites.
Weight uncertainty in neural networks
Blundell, C., Cornebise, J., Kavukcuoglu, K., and Wierstra, D · 2015
Earlier work this paper cites.
Reducing overfitting in deep networks by decorrelating representations
Cogswell, M., Ahmed, F., Girshick, R., Zitnick, L., and Batra, D · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Learning independent features with adversarial nets for non-linear ica
Brakel, P. and Bengio, Y · 2017
Cited alongside, same era.
Relative error embeddings of the gaussian kernel distance
Chen, D. and Phillips, J. M · 2017
Cited alongside, same era.
Deep sets
Zaheer, M., Kottur, S., Ravanbakhsh, S., Poczos, B., Salakhutdinov, R. R., and Smola, A. J · 2017
Cited alongside, same era.
Sanity checks for saliency maps
Adebayo, J., Gilmer, J., Muelly, M., Goodfellow, I., Hardt, M., and Kim, B · 2018
Cited alongside, same era.
Decorrelated batch normalization
Huang, L., Yang, D., Lang, B., and Deng, J · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C · 2018
Cited alongside, same era.
Understanding contrastive representation learning through alignment and uniformity on the hypersphere
Wang, T. and Isola, P · 2020
Later among the works it cites.
Emerging properties in self-supervised vision transformers
Caron, M., Touvron, H., Misra, I., Jégou, H., Mairal, J., Bojanowski, P., and Joulin, A · 2021
Later among the works it cites.
Intriguing properties of contrastive losses
Chen, T., Luo, C., and Li, L · 2021
Later among the works it cites.
Exploring simple siamese representation learning
Chen, X. and He, K · 2021
Later among the works it cites.
Whitening for self-supervised representation learning
Ermolov, A., Siarohin, A., Sangineto, E., and Sebe, N · 2021
Later among the works it cites.
Provable guarantees for self-supervised deep learning with spectral contrastive loss
HaoChen, J. Z., Wei, C., Gaidon, A., and Ma, T · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kernel-based tests for joint independence
Pfister, N., Bühlmann, P., Schölkopf, B., and Peters, J · 2018
Cited alongside, same era.
On the approximation properties of random relu features
Sun, Y., Gilbert, A., and Tewari, A · 2018
Cited alongside, same era.
Rethinking the usage of batch normalization and dropout in the training of deep neural networks
Chen, G., Chen, P., Shi, Y., Hsieh, C.-Y., Liao, B., and Zhang, S · 2019
Cited alongside, same era.
Remap: Multi-layer entropy-guided pooling of dense cnn features for image retrieval
Husain, S. S. and Bober, M · 2019
Cited alongside, same era.
Similarity of neural network representations revisited
Kornblith, S., Norouzi, M., Lee, H., and Hinton, G · 2019
Cited alongside, same era.
Learning disentangled representation with pairwise independence
Li, Z., Tang, Y., Li, W., and He, Y · 2019
Cited alongside, same era.
Later among the works it cites.
Masked autoencoders are scalable vision learners
He, K., Chen, X., Xie, S., Li, Y., Dollár, P., and Girshick, R · 2021
Later among the works it cites.
On feature decorrelation in self-supervised learning
Hua, T., Wang, W., Xue, Z., Ren, S., Wang, Y., and Zhao, H · 2021
Later among the works it cites.
Towards the generalization of contrastive self-supervised learning
Huang, W., Yi, M., and Zhao, X · 2021
Later among the works it cites.
Understanding dimensional collapse in contrastive self-supervised learning
Jing, L., Vincent, P., LeCun, Y., and Tian, Y · 2021
Later among the works it cites.
Self-supervised learning with kernel dependence maximization
Li, Y., Pogodin, R., Sutherland, D. J., and Gretton, A · 2021
Later among the works it cites.
On disentangled representations learned from correlated data
Träuble, F., Creager, E., Kilbertus, N., Locatello, F., Dittadi, A., Goyal, A., Schölkopf, B., and Bauer, S · 2021
Later among the works it cites.
A note on connecting barlow twins with negative-sample-free contrastive learning, 2021
Tsai, Y.-H. H., Bai, S., Morency, L.-P., and Salakhutdinov, R · 2021
Later among the works it cites.
Understanding the behaviour of contrastive loss
Wang, F. and Liu, H · 2021
Later among the works it cites.
Resnet strikes back: An improved training procedure in timm
Wightman, R., Touvron, H., and Jégou, H · 2021
Later among the works it cites.
Barlow twins: Self-supervised learning via redundancy reduction
Zbontar, J., Jing, L., Misra, I., LeCun, Y., and Deny, S · 2021
Later among the works it cites.
Balestriero, R. and LeCun, Y · 2022
Closest in time.
Vicreg: Variance-invariance-covariance regularization for self-supervised learning
Bardes, A., Ponce, J., and LeCun, Y · 2022
Closest in time.
Guillotine regularization: Improving deep networks generalization by removing their head
Bordes, F., Balestriero, R., Garrido, Q., Bardes, A., and Vincent, P · 2022
Closest in time.
Toward a geometrical understanding of self-supervised contrastive learning
Cosentino, R., Sengupta, A., Avestimehr, S., Soltanolkotabi, M., Ortega, A., Willke, T., and Tepper, M · 2022
Closest in time.
On the duality between contrastive and non-contrastive self-supervised learning
Garrido, Q., Chen, Y., Bardes, A., Najman, L., and Lecun, Y · 2022
Closest in time.