Fetching the paper…
Reading the bibliography…
Self-supervised pretraining on unlabeled data followed by supervised fine-tuning on labeled data is a popular paradigm for learning from limited labeled examples.
Large batch optimization for deep learning: Training bert in 76 minutes
You, Y., Li, J., Reddi, S., Hseu, J., Kumar, S., Bhojanapalli, S., Song, X., Demmel, J., Keutzer, K., and Hsieh, C.-J. (2019) · 1904
Earlier work this paper cites.
Deep learning recommendation model for personalization and recommendation systems
Naumov, M., Mudigere, D., Shi, H.-J. M., Huang, J., Sundaraman, N., Park, J., Wang, X., Gupta, U., Wu, C.-J., Azzolini, A. G., et al. (2019) · 1906
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V. (2019) · 1907
Earlier work this paper cites.
A large-scale study of representation learning with the visual task adaptation benchmark
Zhai, X., Puigcerver, J., Kolesnikov, A., Ruyssen, P., Riquelme, C., Lucic, M., Djolonga, J., Pinto, A. S., Neumann, M., Dosovitskiy, A., et al. (2019) · 1910
Earlier work this paper cites.
Revisiting self-supervised visual representation learning
Kolesnikov, A., Zhai, X., and Beyer, L. (2019) · 1929
Earlier work this paper cites.
Pac learning from positive statistical queries
Denis, F. (1998) · 1998
Earlier work this paper cites.
The foundations of cost-sensitive learning
Elkan, C. (2001) · 2001
Earlier work this paper cites.
Fine-tuning pretrained language models: Weight initializations, data orders, and early stopping
Dodge, J., Ilharco, G., Schwartz, R., Farhadi, A., Hajishirzi, H., and Smith, N. (2020) · 2002
Earlier work this paper cites.
Few-shot learning via learning the representation, provably
Du, S. S., Hu, W., Kakade, S. M., Lee, J. D., and Lei, Q. (2020) · 2002
Earlier work this paper cites.
Partially supervised classification of text documents
Liu, B., Lee, W. S., Yu, P. S., and Li, X. (2002) · 2002
Earlier work this paper cites.
Implicit feedback for inferring user preference: a bibliography
Kelly, D. and Teevan, J. (2003) · 2003
Earlier work this paper cites.
Building text classifiers using positive and unlabeled examples
Liu, B., Dai, Y., Li, X., Lee, W. S., and Yu, P. S. (2003) · 2003
Earlier work this paper cites.
Supervised contrastive learning
Khosla, P., Teterwak, P., Wang, C., Sarna, A., Tian, Y., Isola, P., Maschinot, A., Liu, C., and Krishnan, D. (2020) · 2004
Earlier work this paper cites.
Assran, M., Ballas, N., Castrejon, L., and Rabbat, M. (2020) · 2006
Earlier work this paper cites.
Greedy layer-wise training of deep networks
Bengio, Y., Lamblin, P., Popovici, D., and Larochelle, H. (2006) · 2006
Earlier work this paper cites.
Bootstrap your own latent: A new approach to self-supervised learning
Grill, J.-B., Strub, F., Altch’e, F., Tallec, C., Richemond, P. H., Buchatskaya, E., Doersch, C., Pires, B. Á., Guo, Z. D., Azar, M. G., Piot, B., Kavukcuoglu, K., Munos, R., and Valko, M. (2020) · 2006
Earlier work this paper cites.
A fast learning algorithm for deep belief nets
Hinton, G. E., Osindero, S., and Teh, Y.-W. (2006) · 2006
Earlier work this paper cites.
Chuang, C.-Y., Robinson, J., Yen-Chen, L., Torralba, A., and Jegelka, S. (2020) · 2007
Earlier work this paper cites.
Learning classifiers from only positive and unlabeled data
Elkan, C. and Noto, K. (2008) · 2008
Earlier work this paper cites.
Semi-supervised novelty detection
Blanchard, G., Lee, G., and Scott, C. (2010) · 2010
Cited alongside, same era.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Gutmann, M. and Hyvärinen, A. (2010) · 2010
Cited alongside, same era.
Supervised contrastive learning for pre-trained language model fine-tuning
Gunel, B., Du, J., Conneau, A., and Stoyanov, V. (2020) · 2011
Cited alongside, same era.
Distributed representations of words and phrases and their compositionality
Mikolov, T., Sutskever, I., Chen, K., Corrado, G. S., and Dean, J. (2013) · 2013
Cited alongside, same era.
Recursive deep models for semantic compositionality over a sentiment treebank
Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A. Y., and Potts, C. (2013) · 2013
Cited alongside, same era.
Improving generalization via scalable neighborhood component analysis
Wu, Z., Efros, A. A., and Yu, S. X. (2018) · 2018
Later among the works it cites.
Generalized cross entropy loss for training deep neural networks with noisy labels
Zhang, Z. and Sabuncu, M. (2018) · 2018
Later among the works it cites.
Learning from positive and unlabeled data: A survey
Bekker, J. and Davis, J. (2020) · 2020
Later among the works it cites.
Momentum contrast for unsupervised visual representation learning
He, K., Fan, H., Wu, Y., Xie, S., and Girshick, R. (2020) · 2020
Later among the works it cites.
Dedpul: Difference-of-estimated-densities-based positive-unlabeled learning
Ivanov, D. (2020) · 2020
Later among the works it cites.
Contrastive multiview coding
Tian, Y., Krishnan, D., and Isola, P. (2020) · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Du Plessis, M. C., Niu, G., and Sugiyama, M. (2014) · 2014
Cited alongside, same era.
Striving for simplicity: The all convolutional net
Springenberg, J. T., Dosovitskiy, A., Brox, T., and Riedmiller, M. (2014) · 2014
Cited alongside, same era.
Semi-supervised sequence learning
Dai, A. M. and Le, Q. V. (2015) · 2015
Cited alongside, same era.
Pu learning for matrix completion
Hsieh, C.-J., Natarajan, N., and Dhillon, I. (2015) · 2015
Cited alongside, same era.
Skip-thought vectors
Kiros, R., Zhu, Y., Salakhutdinov, R. R., Zemel, R., Urtasun, R., Torralba, A., and Fidler, S. (2015) · 2015
Cited alongside, same era.
Class-prior estimation for learning from positive and unlabeled data
Christoffel, M., Niu, G., and Sugiyama, M. (2016) · 2016
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J. (2016) · 2016
Cited alongside, same era.
Emerging properties in self-supervised vision transformers
Caron, M., Touvron, H., Misra, I., J’egou, H., Mairal, J., Bojanowski, P., and Joulin, A. (2021) · 2021
Later among the works it cites.
Pu active learning for recommender systems
Chen, J.-L., Cai, J.-J., Jiang, Y., and Huang, S.-J. (2021) · 2021
Later among the works it cites.
Simcse: Simple contrastive learning of sentence embeddings
Gao, T., Yao, X., and Chen, D. (2021) · 2021
Later among the works it cites.
Mixture proportion estimation and pu learning: A modern approach
Garg, S., Wu, Y., Smola, A. J., Balakrishnan, S., and Lipton, Z. (2021) · 2021
Later among the works it cites.
Dissecting supervised constrastive learning
Graf, F., Hofer, C., Niethammer, M., and Kwitt, R. (2021) · 2021
Later among the works it cites.
Pulns: Positive-unlabeled learning with effective negative sample selector
Luo, C., Zhao, P., Chen, C., Qiao, B., Du, C., Zhang, H., Wu, W., Cai, S., He, B., Rajmohan, S., et al. (2021) · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I. (2021) · 2021
Later among the works it cites.
Barlow twins: Self-supervised learning via redundancy reduction
Zbontar, J., Jing, L., Misra, I., LeCun, Y., and Deny, S. (2021) · 2021
Later among the works it cites.
Neighborhood contrastive learning for novel class discovery
Zhong, Z., Fini, E., Roy, S., Luo, Z., Ricci, E., and Sebe, N. (2021) · 2021
Later among the works it cites.
Dual contrastive learning: Text classification via label-aware data augmentation
Chen, Q., Zhang, R., Zheng, Y., and Mao, Y. (2022) · 2022
Closest in time.
Fine-tuning can distort pretrained features and underperform out-of-distribution
Kumar, A., Raghunathan, A., Jones, R., Ma, T., and Liang, P. (2022) · 2022
Closest in time.
Mixture proportion estimation via kernel embeddings of distributions
Ramaswamy, H., Scott, C., and Tewari, A. (2016) · 2060
Closest in time.