Fetching the paper…
Reading the bibliography…
Strong student models can learn from weaker teachers: when trained on the predictions of a weaker model, a strong pretrained student can learn to correct the weak model's errors and generalize to examples where the teacher is not confident, even when these examples are excluded from training.
On the uniform convergence of relative frequencies of events to their probabilities
Vapnik, V. (1971) · 1971
Earlier work this paper cites.
On the density of families of sets
Sauer, N. (1972) · 1972
Earlier work this paper cites.
A combinatorial problem; stability and order for models and theories in infinitary languages
Shelah, S. (1972) · 1972
Earlier work this paper cites.
Maximum likelihood estimation of observer error-rates using the em algorithm
Dawid, A. P. and Skene, A. M. (1979) · 1979
Earlier work this paper cites.
The complexity of testing whether a graph is a superconcentrator
Blum, M., Karp, R. M., Vornberger, O., Papadimitriu, C. H., and Yannakakis, M. (1981) · 1981
Earlier work this paper cites.
Combining labeled and unlabeled data with co-training
Blum, A. and Mitchell, T. (1998) · 1998
Earlier work this paper cites.
Introduction to statistical learning theory
Bousquet, O., Boucheron, S., and Lugosi, G. (2003) · 2003
Earlier work this paper cites.
Co-training and expansion: Towards bridging theory and practice
Balcan, M.-F., Blum, A., and Yang, K. (2004) · 2004
Earlier work this paper cites.
Detecting change in data streams
Kifer, D., Ben-David, S., and Gehrke, J. (2004) · 2004
Earlier work this paper cites.
Optimal aggregation of classifiers in statistical learning
Tsybakov, A. B. (2004) · 2004
Earlier work this paper cites.
Analysis of representations for domain adaptation
Ben-David, S., Blitzer, J., Crammer, K., and Pereira, F. (2006) · 2006
Earlier work this paper cites.
Model compression
Buciluǎ, C., Caruana, R., and Niculescu-Mizil, A. (2006) · 2006
Earlier work this paper cites.
Blocking conductance and mixing in random walks
Kannan, R., Lovász, L., and Montenegro, R. (2006) · 2006
Earlier work this paper cites.
Learning bounds for domain adaptation
Blitzer, J., Crammer, K., Kulesza, A., Pereira, F., and Wortman, J. (2007) · 2007
Earlier work this paper cites.
Iterative learning for reliable crowdsourcing systems
Karger, D., Oh, S., and Shah, D. (2011) · 2011
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Maas, A. L., Daly, R. E., Pham, P. T., Huang, D., Ng, A. Y., and Potts, C. (2011) · 2011
Earlier work this paper cites.
Learning with noisy labels
Natarajan, N., Dhillon, I. S., Ravikumar, P. K., and Tewari, A. (2013) · 2013
Earlier work this paper cites.
Improved cheeger’s inequality and analysis of local graph partitioning using vertex expansion and expansion profile
Kwok, T. C., Lau, L. C., and Lee, Y. T. (2016) · 2016
Earlier work this paper cites.
Data programming: Creating large training sets, quickly
Ratner, A. J., De Sa, C. M., Wu, S., Selsam, D., and Ré, C. (2016) · 2016
Earlier work this paper cites.
Snorkel: Rapid training data creation with weak supervision
Ratner, A., Bach, S. H., Ehrenberg, H., Fries, J., Wu, S., and Ré, C. (2017) · 2017
Earlier work this paper cites.
Learning from noisy singly-labeled data
Khetan, A., Lipton, Z. C., and Anandkumar, A. (2018) · 2018
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F. (2018) · 2018
Earlier work this paper cites.
Weakly-supervised neural text classification
Meng, Y., Shen, J., Zhang, C., and Han, J. (2018) · 2018
Earlier work this paper cites.
Snuba: Automating weak supervision to label training data
Varma, P. and Ré, C. (2018) · 2018
Cited alongside, same era.
Weakly supervised classification of aortic valve malformations using unlabeled cardiac mri sequences
Fries, J. A., Varma, P., Chen, V. S., Xiao, K., Tejeda, H., Saha, P., Dunnmon, J., Chubb, H., Maskatia, S., Fiterau, M., et al. (2019) · 2019
Cited alongside, same era.
Chexpert: A large chest radiograph dataset with uncertainty labels and expert comparison
Irvin, J., Rajpurkar, P., Ko, M., Yu, Y., Ciurea-Ilcus, S., Chute, C., Marklund, H., Haghgoo, B., Ball, R., Shpanskaya, K., et al. (2019) · 2019
Cited alongside, same era.
Mimic-cxr, a de-identified publicly available database of chest radiographs with free-text reports
Johnson, A. E., Pollard, T. J., Berkowitz, S. J., Greenbaum, N. R., Lungren, M. P., Deng, C.-y., Mark, R. G., and Horng, S. (2019) · 2019
Cited alongside, same era.
Training complex models with multi-task weak supervision
Ratner, A., Hancock, B., Dunnmon, J., Sala, F., Pandey, S., and Ré, C. (2019) · 2019
Cited alongside, same era.
Want to reduce labeling cost? gpt-3 can help
Wang, S., Liu, Y., Xu, Y., Zhu, C., and Zeng, M. (2021) · 2021
Later among the works it cites.
Fine-tuning pre-trained language model with weak supervision: A contrastive-regularized self-training approach
Yu, Y., Zuo, S., Jiang, H., Ren, W., Zhao, T., and Zhang, C. (2021) · 2021
Later among the works it cites.
WRENCH: A comprehensive benchmark for weak supervision
Zhang, J., Yu, Y., Li, Y., Wang, Y., Yang, Y., Yang, M., and Ratner, A. (2021) · 2021
Later among the works it cites.
Large language models are few-shot clinical information extractors
Agrawal, M., Hegselmann, S., Lang, H., Kim, Y., and Sontag, D. (2022) · 2022
Later among the works it cites.
Shoring up the foundations: Fusing model embeddings and weak supervision
Chen, M. F., Fu, D. Y., Adila, D., Zhang, M., Sala, F., Fatahalian, K., and Ré, C. (2022) · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sentence-BERT: Sentence embeddings using Siamese BERT-networks
Reimers, N. and Gurevych, I. (2019) · 2019
Cited alongside, same era.
Learning dependency structures for weak supervision models
Varma, P., Sala, F., He, A., Ratner, A., and Ré, C. (2019) · 2019
Cited alongside, same era.
A clinical text classification paradigm using weak supervision and deep representation
Wang, Y., Sohn, S., Liu, S., Shen, F., Wang, L., Atkinson, E. J., Amin, S., and Liu, H. (2019) · 2019
Cited alongside, same era.
Improved sample complexities for deep neural networks and robust classification via an all-layer margin
Wei, C. and Ma, T. (2019) · 2019
Cited alongside, same era.
Self-training avoids using spurious features under domain shift
Chen, Y., Wei, C., Kumar, A., and Ma, T. (2020) · 2020
Cited alongside, same era.
Fast and three-rious: Speeding up weak supervision with triplet methods
Fu, D., Chen, M., Sala, F., Hooper, S., Fatahalian, K., and Ré, C. (2020) · 2020
Cited alongside, same era.
Understanding self-training for gradual domain adaptation
Kumar, A., Ma, T., and Liang, P. (2020) · 2020
Cited alongside, same era.
Eisenstein, J., Andor, D., Bohnet, B., Collins, M., and Mimno, D. (2022) · 2022
Later among the works it cites.
Self-training converts weak learners to strong learners in mixture models
Frei, S., Zou, D., Chen, Z., and Gu, Q. (2022) · 2022
Later among the works it cites.
Beyond separability: Analyzing the linear transferability of contrastive representations to related subpopulations
HaoChen, J. Z., Wei, C., Kumar, A., and Ma, T. (2022) · 2022
Later among the works it cites.
Label propagation with weak supervision
Pukdee, R., Sam, D., Ravikumar, P. K., and Balcan, N. (2022) · 2022
Later among the works it cites.
Machine Learning from Weak Supervision: An Empirical Risk Minimization Approach
Sugiyama, M., Bao, H., Ishida, T., Lu, N., Sakai, T., and Niu, G. (2022) · 2022
Later among the works it cites.
W2f: A weakly-supervised to fully-supervised framework for object detection
Zhang, Y., Bai, Y., Ding, M., Li, Y., and Ghanem, B. (2018) · 2022
Later among the works it cites.
Generalization on the unseen, logic reasoning and degree curriculum
Abbe, E., Bengio, S., Lotfi, A., and Rizk, K. (2023) · 2023
Later among the works it cites.
Weak-to-strong generalization: Eliciting strong capabilities with weak supervision
Burns, C., Izmailov, P., Kirchner, J. H., Baker, B., Gao, L., Aschenbrenner, L., Chen, Y., Ecoffet, A., Joglekar, M., Leike, J., et al. (2023) · 2023
Later among the works it cites.
Unified risk analysis for weakly supervised learning
Chiang, C.-K. and Sugiyama, M. (2023) · 2023
Later among the works it cites.
Is GPT-3 a good data annotator?
Ding, B., Qin, C., Liu, L., Chia, Y. K., Li, B., Joty, S., and Bing, L. (2023) · 2023
Later among the works it cites.
Annollm: Making large language models to be better crowdsourced annotators
He, X., Lin, Z., Gong, Y., Jin, A., Zhang, H., Lin, C., Jiao, J., Yiu, S. M., Duan, N., Chen, W., et al. (2023) · 2023
Later among the works it cites.
Kuzman, T., Mozetic, I., and Ljubešic, N. (2023) · 2023
Later among the works it cites.
Higher-order cheeger inequality for partitioning with buffers
Makarychev, K., Makarychev, Y., Shan, L., and Vijayaraghavan, A. (2023) · 2023
Later among the works it cites.
Losses over labels: Weakly supervised learning via direct loss construction
Sam, D. and Kolter, J. Z. (2023) · 2023
Later among the works it cites.
Prevalence and prevention of large language model use in crowd work
Veselovsky, V., Ribeiro, M. H., Cozzolino, P., Gordon, A., Rothschild, D., and West, R. (2023) · 2023
Later among the works it cites.
Self-instruct: Aligning language models with self-generated instructions
Wang, Y., Kordi, Y., Mishra, S., Liu, A., Smith, N. A., Khashabi, D., and Hajishirzi, H. (2023) · 2023
Later among the works it cites.
MTEB: Massive text embedding benchmark
Muennighoff, N., Tazi, N., Magne, L., and Reimers, N. (2023) · 2037
Closest in time.