Fetching the paper…
Reading the bibliography…
Recent success of large-scale pre-trained language models crucially hinge on fine-tuning them on large amounts of labeled data for the downstream task, that are typically expensive to acquire.
Roberta: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 1907
Earlier work this paper cites.
Semi-supervised learning for text classification by layer partitioning
Alexander Hanbo Li and Abhinav Sethy · 1911
Earlier work this paper cites.
Probability of error of some adaptive pattern-recognition machines
H. J. Scudder III · 1965
Earlier work this paper cites.
Elementary applied statistics
Linton G Freeman · 1965
Earlier work this paper cites.
A mathematical theory of communication
Claude E. Shannon · 2001
Earlier work this paper cites.
Curriculum learning
Yoshua Bengio, Jérôme Louradour, Ronan Collobert, and Jason Weston · 2009
Earlier work this paper cites.
Semi-supervised learning
Olivier Chapelle, Bernhard Schlkopf, and Alexander Zien · 2010
Earlier work this paper cites.
Self-paced learning for latent variable models
M. P. Kumar, Benjamin Packer, and Daphne Koller · 2010
Earlier work this paper cites.
Bayesian active learning for classification and preference learning
Neil Houlsby, Ferenc Huszar, Zoubin Ghahramani, and Máté Lengyel · 2011
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts · 2011
Earlier work this paper cites.
Practical variational inference for neural networks
Alex Graves · 2011
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Y. Ng, and Christopher Potts · 2013
Earlier work this paper cites.
Hidden factors and hidden topics: understanding rating dimensions with review text
Julian J. McAuley and Jure Leskovec · 2013
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey E. Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
Learning with pseudo-ensembles
Philip Bachman, Ouais Alsharif, and Doina Precup · 2014
Earlier work this paper cites.
Semi-supervised learning with deep generative models
Diederik P. Kingma, Shakir Mohamed, Danilo Jimenez Rezende, and Max Welling · 2014
Earlier work this paper cites.
Bayesian convolutional neural networks with bernoulli approximate variational inference
Yarin Gal and Zoubin Ghahramani · 2015
Cited alongside, same era.
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Jake Zhao, and Yann LeCun · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Cited alongside, same era.
Semi-supervised sequence learning
Andrew M. Dai and Quoc V. Le · 2015
Cited alongside, same era.
Explaining and harnessing adversarial examples
Ian J. Goodfellow, Jonathon Shlens, and Christian Szegedy · 2015
Cited alongside, same era.
Adversarial training methods for semi-supervised text classification
Takeru Miyato, Andrew M. Dai, and Ian J. Goodfellow · 2017
Later among the works it cites.
Temporal ensembling for semi-supervised learning
Samuli Laine and Timo Aila · 2017
Later among the works it cites.
Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results
Antti Tarvainen and Harri Valpola · 2017
Later among the works it cites.
Self-paced co-training
Fan Ma, Deyu Meng, Qi Xie, Zina Li, and Xuanyi Dong · 2017
Later among the works it cites.
Active bias: Training more accurate neural networks by emphasizing high variance samples
Haw-Shiuan Chang, Erik G. Learned-Miller, and Andrew McCallum · 2017
Later among the works it cites.
Learning adversarial networks for semi-supervised text classification via policy gradient
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Semi-supervised learning with ladder networks
Antti Rasmus, Mathias Berglund, Mikko Honkala, Harri Valpola, and Tapani Raiko · 2015
Cited alongside, same era.
Weight uncertainty in neural network
Charles Blundell, Julien Cornebise, Koray Kavukcuoglu, and Daan Wierstra · 2015
Cited alongside, same era.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Yarin Gal and Zoubin Ghahramani · 2016
Cited alongside, same era.
Regularization with stochastic transformations and perturbations for deep semi-supervised learning
Mehdi Sajjadi, Mehran Javanmardi, and Tolga Tasdizen · 2016
Cited alongside, same era.
Language as a latent variable: Discrete generative models for sentence compression
Yishu Miao and Phil Blunsom · 2016
Cited alongside, same era.
Improving neural machine translation models with monolingual data
Rico Sennrich, Barry Haddow, and Alexandra Birch · 2016
Cited alongside, same era.
Training region-based object detectors with online hard example mining
Abhinav Shrivastava, Abhinav Gupta, and Ross B. Girshick · 2016
Cited alongside, same era.
Yan Li and Jieping Ye · 2018
Later among the works it cites.
Structvae: Tree-structured latent variable models for semi-supervised semantic parsing
Pengcheng Yin, Chunting Zhou, Junxian He, and Graham Neubig · 2018
Later among the works it cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Later among the works it cites.
Unsupervised data augmentation for consistency training, 2019
Qizhe Xie, Zihang Dai, Eduard Hovy, Minh-Thang Luong, and Quoc V. Le · 2019
Later among the works it cites.
Revisiting self-training for neural sequence generation, 2019
Junxian He, Jiatao Gu, Jiajun Shen, and Marc’Aurelio Ranzato · 2019
Later among the works it cites.
Variational pretraining for semi-supervised text classification
Suchin Gururangan, Tam Dang, Dallas Card, and Noah A. Smith · 2019
Later among the works it cites.
Delta-training: Simple semi-supervised text classification using pretrained word embeddings
Hwiyeol Jo and Ceyda Cinarel · 2019
Later among the works it cites.
Learning to self-train for semi-supervised few-shot classification
Xinzhe Li, Qianru Sun, Yaoyao Liu, Qin Zhou, Shibao Zheng, Tat-Seng Chua, and Bernt Schiele · 2019
Later among the works it cites.
Enhancing deep active learning using selective self-training for image classification
Emmeleia Panagiota Mastoropoulou · 2019
Later among the works it cites.