Fetching the paper…
Reading the bibliography…
Recently, pseudo label based semi-supervised learning has achieved great success in many fields.
The nature of statistical learning theory, 1996
Sain, S. R · 1996
Earlier work this paper cites.
Error bounds for transductive learning via compression and clustering
Derbeko, P., El-Yaniv, R., and Meir, R · 2003
Earlier work this paper cites.
Semi-supervised learning by entropy minimization
Grandvalet, Y. and Bengio, Y · 2004
Earlier work this paper cites.
Sequence transduction with recurrent neural networks
Graves, A · 2012
Earlier work this paper cites.
Pseudo-label: The simple and efficient semi-supervised learning method for deep neural networks
Lee, D.-H. et al · 2013
Earlier work this paper cites.
Semi-supervised learning with deep generative models
Kingma, D. P., Mohamed, S., Jimenez Rezende, D., and Welling, M · 2014
Earlier work this paper cites.
Scheduled sampling for sequence prediction with recurrent neural networks
Bengio, S., Vinyals, O., Jaitly, N., and Shazeer, N · 2015
Earlier work this paper cites.
Norm-based capacity control in neural networks
Neyshabur, B., Tomioka, R., and Srebro, N · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Temporal ensembling for semi-supervised learning
Laine, S. and Aila, T · 2016
Earlier work this paper cites.
Exploring generalization in deep learning
Neyshabur, B., Bhojanapalli, S., McAllester, D., and Srebro, N · 2017
Cited alongside, same era.
Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results
Tarvainen, A. and Valpola, H · 2017
Cited alongside, same era.
State-of-the-art speech recognition with sequence-to-sequence models
Chiu, C.-C., Sainath, T. N., Wu, Y., Prabhavalkar, R., Nguyen, P., Chen, Z., Kannan, A., Weiss, R. J., Rao, K., Gonina, E., et al · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Cited alongside, same era.
A priori estimates of the population risk for two-layer neural networks
Ma, C., Wu, L., et al · 2018
Cited alongside, same era.
Bridging theory and algorithm for domain adaptation
Zhang, Y., Liu, T., Long, M., and Jordan, M · 2019
Later among the works it cites.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Baevski, A., Zhou, Y., Mohamed, A., and Auli, M · 2020
Later among the works it cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Later among the works it cites.
Self-training avoids using spurious features under domain shift
Chen, Y., Wei, C., Kumar, A., and Ma, T · 2020
Later among the works it cites.
Early-learning regularization prevents memorization of noisy labels
Liu, S., Niles-Weed, J., Razavian, N., and Fernandez-Granda, C · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A dirt-t approach to unsupervised domain adaptation
Shu, R., Bui, H. H., Narui, H., and Ermon, S · 2018
Cited alongside, same era.
Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks
Arora, S., Du, S., Hu, W., Li, Z., and Wang, R · 2019
Cited alongside, same era.
The speechtransformer for large-scale mandarin chinese speech recognition
Li, J., Wang, X., Li, Y., et al · 2019
Cited alongside, same era.
Roberta: A robustly optimized bert pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 2019
Cited alongside, same era.
A priori estimates of the population risk for residual networks
Weinan, E., Ma, C., and Wang, Q · 2019
Cited alongside, same era.
Oymak, S. and Gulcu, T. C · 2020
Later among the works it cites.
Fixmatch: Simplifying semi-supervised learning with consistency and confidence
Sohn, K., Berthelot, D., Carlini, N., Zhang, Z., Zhang, H., Raffel, C. A., Cubuk, E. D., Kurakin, A., and Li, C.-L · 2020
Later among the works it cites.
Theoretical analysis of self-training with deep networks on unlabeled data
Wei, C., Shen, K., Chen, Y., and Ma, T · 2020
Later among the works it cites.
Self-training with noisy student improves imagenet classification
Xie, Q., Luong, M.-T., Hovy, E., and Le, Q. V · 2020
Later among the works it cites.
Ratt: Leveraging unlabeled data to guarantee generalization
Garg, S., Balakrishnan, S., Kolter, Z., and Lipton, Z · 2021
Later among the works it cites.