Fetching the paper…
Reading the bibliography…
The vast majority of work in self-supervised learning, both theoretical and empirical (though mostly the latter), have largely focused on recovering good features for downstream tasks, with the definition of "good" often being intricately tied to the downstream task itself.
The expression of a tensor or a polyadic as a sum of products
F. L. Hitchcock · 1927
Earlier work this paper cites.
Foundations of the parafac procedure: Model and conditions for an explanatory factor analysis
R. Harshman · 1970
Earlier work this paper cites.
Three-way arrays: rank and uniqueness of trilinear decompositions, with application to arithmetic complexity and statistics
J. B. Kruskal · 1977
Earlier work this paper cites.
Multivariate normal mixtures: a fast consistent method of moments
B. G. Lindsay and P. Basak · 1993
Earlier work this paper cites.
A simple framework for contrastive learning of visual representations
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton · 2002
Earlier work this paper cites.
Optimally sparse representation in general (nonorthogonal) dictionaries via l1 minimization
D. L. Donoho and M. Elad · 2003
Earlier work this paper cites.
Contrastive estimation reveals topic posterior information to linear models
C. Tosh, A. Krishnamurthy, and D. Hsu · 2003
Earlier work this paper cites.
Sparse component analysis and blind source separation of underdetermined mixtures
P. Georgiev, F. Theis, and A. Cichocki · 2005
Earlier work this paper cites.
Learning nonsingular phylogenies and hidden markov models
E. Mossel and S. Roch · 2005
Earlier work this paper cites.
On the uniqueness of overcomplete dictionaries, and a practical way to retrieve them
M. Aharon, M. Elad, and A. M. Bruckstein · 2006
Earlier work this paper cites.
Stable signal recovery from incomplete and inaccurate measurements
E. J. Candes, J. K. Romberg, and T. Tao · 2006
Earlier work this paper cites.
Identifiability of parameters in latent structure models with many observed variables
E. S. Allman, C. Matias, and J. A. Rhodes · 2009
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
M. Gutmann and A. Hyvärinen · 2010
Earlier work this paper cites.
High-dimensional ising model selection using l1-regularized logistic regression
P. Ravikumar, M. J. Wainwright, and J. D. Lafferty · 2010
Earlier work this paper cites.
A method of moments for mixture models and hidden markov models
A. Anandkumar, D. Hsu, and S. M. Kakade · 2012
Earlier work this paper cites.
Exact recovery of sparsely-used dictionaries
D. A. Spielman, H. Wang, and J. Wright · 2012
Cited alongside, same era.
Efficient estimation of word representations in vector space
T. Mikolov, K. Chen, G. Corrado, and J. Dean · 2013
Cited alongside, same era.
Tensor decompositions for learning latent variable models
A. Anandkumar, R. Ge, D. Hsu, S. M. Kakade, and M. Telgarsky · 2014
Cited alongside, same era.
New algorithms for learning incoherent and overcomplete dictionaries
S. Arora, R. Ge, and A. Moitra · 2014
Cited alongside, same era.
Efficiently learning ising models on arbitrary graphs
G. Bresler · 2015
Cited alongside, same era.
Unsupervised visual representation learning by context prediction
C. Doersch, A. Gupta, and A. A. Efros · 2015
Data-efficient image recognition with contrastive predictive coding
O. J. Hénaff, A. Srinivas, J. De Fauw, A. Razavi, C. Doersch, S. Eslami, and A. v. d. Oord · 2019
Later among the works it cites.
A structural probe for finding syntax in word representations
J. Hewitt and C. D. Manning · 2019
Later among the works it cites.
Superglue: A stickier benchmark for general-purpose language understanding systems
A. Wang, Y. Pruksachatkun, N. Nangia, A. Singh, J. Michael, F. Hill, O. Levy, and S. Bowman · 2019
Later among the works it cites.
A large-scale study of representation learning with the visual task adaptation benchmark
X. Zhai, J. Puigcerver, A. Kolesnikov, P. Ruyssen, C. Riquelme, M. Lucic, J. Djolonga, A. S. Pinto, M. Neumann, A. Dosovitskiy, et al · 2019
Later among the works it cites.
Flow contrastive estimation of energy-based models
R. Gao, E. Nijkamp, D. P. Kingma, Z. Xu, A. M. Dai, and Y. N. Wu · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Unsupervised feature extraction by time-contrastive learning and nonlinear ica
A. Hyvarinen and H. Morioka · 2016
Cited alongside, same era.
Context encoders: Feature learning by inpainting
D. Pathak, P. Krahenbuhl, J. Donahue, T. Darrell, and A. A. Efros · 2016
Cited alongside, same era.
Interaction screening: Efficient and sample-optimal learning of ising models
M. Vuffray, S. Misra, A. Lokhov, and M. Chertkov · 2016
Cited alongside, same era.
Nonlinear ICA of Temporally Dependent Stationary Sources
A. Hyvarinen and H. Morioka · 2017
Cited alongside, same era.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Cited alongside, same era.
Later among the works it cites.
Predicting what you already know helps: Provable self-supervised learning, 2020
J. D. Lee, Q. Lei, N. Saunshi, and J. Zhuo · 2020
Later among the works it cites.
A mathematical exploration of why language models help solve downstream tasks
N. Saunshi, S. Malladi, and S. Arora · 2020
Later among the works it cites.
T. Wang and P. Isola · 2020
Later among the works it cites.
Provable guarantees for self-supervised deep learning with spectral contrastive loss
J. Z. HaoChen, C. Wei, A. Gaidon, and T. Ma · 2021
Later among the works it cites.
Masked autoencoders are scalable vision learners
K. He, X. Chen, S. Xie, Y. Li, P. Dollár, and R. Girshick · 2021
Later among the works it cites.
Compressive visual representations
K.-H. Lee, A. Arnab, S. Guadarrama, J. Canny, and I. Fischer · 2021
Later among the works it cites.
Self-supervised learning is more robust to dataset imbalance
H. Liu, J. Z. HaoChen, A. Gaidon, and T. Ma · 2021
Later among the works it cites.
Dabs: A domain-agnostic benchmark for self-supervised learning
A. Tamkin, V. Liu, R. Lu, D. Fein, C. Schultz, and N. Goodman · 2021
Later among the works it cites.
Why do pretrained language models help in downstream tasks? an analysis of head and prompt tuning
C. Wei, S. M. Xie, and T. Ma · 2021
Later among the works it cites.
Toward understanding the feature learning process of self-supervised contrastive learning
Z. Wen and Y. Li · 2021
Later among the works it cites.