Fetching the paper…
Reading the bibliography…
The classic teacher-student model in machine learning posits that a strong teacher supervises a weak student to improve the student's capabilities.
Model compression
Cristian Buciluǎ, Rich Caruana, and Alexandru Niculescu-Mizil · 2006
Earlier work this paper cites.
Introduction to the non-asymptotic analysis of random matrices
Roman Vershynin · 2010
Earlier work this paper cites.
Pseudo-label: The simple and efficient semi-supervised learning method for deep neural networks
Dong-Hyun Lee et al · 2013
Earlier work this paper cites.
Surprising asymptotic conical structure in critical sample eigen-directions
Dan Shen, Haipeng Shen, Hongtu Zhu, and JS Marron · 2013
Earlier work this paper cites.
Asymptotics of empirical eigen-structure for ultra-high dimensional spiked covariance model
Jianqing Fan and Weichen Wang · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
Understanding intermediate layers using linear classifier probes
Guillaume Alain and Yoshua Bengio · 2016
Earlier work this paper cites.
Parameter-efficient transfer learning for nlp
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly · 2019
Earlier work this paper cites.
Towards understanding knowledge distillation
Mary Phuong and Christoph Lampert · 2019
Earlier work this paper cites.
Pseudo-labeling and confirmation bias in deep semi-supervised learning
Eric Arazo, Diego Ortego, Paul Albert, Noel E O’Connor, and Kevin McGuinness · 2020
Earlier work this paper cites.
Benign overfitting in linear regression
Peter L Bartlett, Philip M Long, Gábor Lugosi, and Alexander Tsigler · 2020
Earlier work this paper cites.
Two models of double descent for weak features
Mikhail Belkin, Daniel Hsu, and Ji Xu · 2020
Earlier work this paper cites.
Knowledge distillation in wide neural networks: Risk bound, data efficiency and imperfect teacher
Guangda Ji and Zhanxing Zhu · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Earlier work this paper cites.
Self-distillation amplifies regularization in hilbert space
Hossein Mobahi, Mehrdad Farajtabar, and Peter Bartlett · 2020
Earlier work this paper cites.
Harmless interpolation of noisy data in regression
Vidya Muthukumar, Kailas Vodrahalli, Vignesh Subramanian, and Anant Sahai · 2020
Earlier work this paper cites.
Curriculum labeling: Revisiting pseudo-labeling for semi-supervised learning
Paola Cascante-Bonilla, Fuwen Tan, Yanjun Qi, and Vicente Ordonez · 2021
Earlier work this paper cites.
Finite-sample analysis of interpolating linear classifiers in the overparameterized regime
Niladri S Chatterji and Philip M Long · 2021
Earlier work this paper cites.
Synthetic data in machine learning for medicine and healthcare
Richard J Chen, Ming Y Lu, Tiffany Y Chen, Drew FK Williamson, and Faisal Mahmood · 2021
Earlier work this paper cites.
Knowledge distillation: A survey
Jianping Gou, Baosheng Yu, Stephen J Maybank, and Dacheng Tao · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2021
Earlier work this paper cites.
Characterizing the implicit bias via a primal-dual analysis
Ziwei Ji and Matus Telgarsky · 2021
Earlier work this paper cites.
The power of scale for parameter-efficient prompt tuning
Brian Lester, Rami Al-Rfou, and Noah Constant · 2021
Earlier work this paper cites.
Classification vs regression in overparameterized regimes: Does the loss function matter?
Vidya Muthukumar, Adhyyan Narang, Vignesh Subramanian, Mikhail Belkin, Daniel Hsu, and Anant Sahai · 2021
Earlier work this paper cites.
Synthetic data for deep learning
Sergey I Nikolenko · 2021
Cited alongside, same era.
Orthant probabilities and the attainment of maxima on a vertex of a simplex
Damián Pinasco, Ezequiel Smucler, and Ignacio Zalduendo · 2021
Cited alongside, same era.
Mamshad Nayeem Rizve, Kevin Duarte, Yogesh S Rawat, and Mubarak Shah · 2021
Cited alongside, same era.
Can you learn an algorithm? generalizing from easy to hard problems with recurrent networks
Avi Schwarzschild, Eitan Borgnia, Arjun Gupta, Furong Huang, Uzi Vishkin, Micah Goldblum, and Tom Goldstein · 2021
Cited alongside, same era.
Does knowledge distillation really work?
Samuel Stanton, Pavel Izmailov, Polina Kirichenko, Alexander A Alemi, and Andrew G Wilson · 2021
Cited alongside, same era.
What knowledge gets distilled in knowledge distillation?
Utkarsh Ojha, Yuheng Li, Anirudh Sundara Rajan, Yingyu Liang, and Yong Jae Lee · 2023
Later among the works it cites.
Knowledge distillation performs partial variance reduction
Mher Safaryan, Alexandra Peste, and Dan Alistarh · 2023
Later among the works it cites.
Random teachers are good teachers
Felix Sarnthein, Gregor Bachmann, Sotiris Anagnostidis, and Thomas Hofmann · 2023
Later among the works it cites.
Improving knowledge distillation via head and tail categories
Liuchi Xu, Jin Ren, Zhenhua Huang, Weishi Zheng, and Yunwen Chen · 2023
Later among the works it cites.
Towards the fundamental limits of knowledge transfer over finite domains, 2023
Qingyue Zhao and Banghua Zhu · 2023
Later among the works it cites.
Representation engineering: A top-down approach to ai transparency
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ke Wang and Christos Thrampoulidis · 2021
Cited alongside, same era.
Benign Overfitting in Multiclass Classification: All Roads Lead to Interpolation
Ke Wang, Vidya Muthukumar, and Christos Thrampoulidis · 2021
Cited alongside, same era.
Finetuned language models are zero-shot learners
Jason Wei, Maarten Bosma, Vincent Y Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le · 2021
Cited alongside, same era.
Flexmatch: Boosting semi-supervised learning with curriculum pseudo labeling
Bowen Zhang, Yidong Wang, Wenxin Hou, Hao Wu, Jindong Wang, Manabu Okumura, and Takahiro Shinozaki · 2021
Cited alongside, same era.
Survey on synthetic data generation, evaluation methods and gans
Alvaro Figueira and Bruno Vaz · 2022
Cited alongside, same era.
Training compute-optimal large language models
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al · 2022
Cited alongside, same era.
Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning
Haokun Liu, Derek Tam, Mohammed Muqeeth, Jay Mohta, Tenghao Huang, Mohit Bansal, and Colin A Raffel · 2022
Cited alongside, same era.
Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, et al · 2023
Later among the works it cites.
Practical insights into knowledge distillation for pre-trained models, 2024
Norah Alballa and Marco Canini · 2024
Closest in time.
Quantifying the gain in weak-to-strong generalization
Moses Charikar, Chirag Pabbaraju, and Kirankumar Shiragur · 2024
Closest in time.
Qlora: Efficient finetuning of quantized llms
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer · 2024
Closest in time.
Vision superalignment: Weak-to-strong generalization for vision foundation models
Jianyuan Guo, Hanting Chen, Chengcheng Wang, Kai Han, Chang Xu, and Yunhe Wang · 2024
Closest in time.
The unreasonable effectiveness of easy training data for hard tasks
Peter Hase, Mohit Bansal, Peter Clark, and Sarah Wiegreffe · 2024
Closest in time.
Weakly-supervised concealed object segmentation with sam-based pseudo labeling and multi-scale feature grouping
Chunming He, Kai Li, Yachao Zhang, Guoxia Xu, Longxiang Tang, Yulun Zhang, Zhenhua Guo, and Xiu Li · 2024
Closest in time.
On dark knowledge for distilling generators
Chi Hong, Robert Birke, Pin-Yu Chen, and Lydia Y. Chen · 2024
Closest in time.
Aligner: Achieving efficient alignment through weak-to-strong correction
Jiaming Ji, Boyuan Chen, Hantao Lou, Donghai Hong, Borong Zhang, Xuehai Pan, Juntao Dai, and Yaodong Yang · 2024
Closest in time.
Theoretical analysis of weak-to-strong generalization
Hunter Lang, David Sontag, and Aravindan Vijayaraghavan · 2024
Closest in time.
Best practices and lessons learned on synthetic data for language models
Ruibo Liu, Jerry Wei, Fangyu Liu, Chenglei Si, Yanzhe Zhang, Jinmeng Rao, Steven Zheng, Daiyi Peng, Diyi Yang, Denny Zhou, et al · 2024
Closest in time.
Co-supervised learning: Improving weak-to-strong generalization with hierarchical mixture of experts
Yuejiang Liu and Alexandre Alahi · 2024
Closest in time.
A statistical framework for weak-to-strong generalization
Seamus Somerstep, Felipe Maia Polo, Moulinath Banerjee, Ya’acov Ritov, Mikhail Yurochkin, and Yuekai Sun · 2024
Closest in time.
Your weak llm is secretly a strong teacher for alignment
Leitian Tao and Yixuan Li · 2024
Closest in time.
Implicit bias of next-token prediction
Christos Thrampoulidis · 2024
Closest in time.
Precise asymptotic generalization for multiclass classification with overparameterized linear models
David Wu and Anant Sahai · 2024
Closest in time.
Super (ficial)-alignment: Strong models may deceive weak models in weak-to-strong generalization
Wenkai Yang, Shiqi Shen, Guangyao Shen, Zhi Gong, and Yankai Lin · 2024
Closest in time.
Learning from biased soft labels
Hua Yuan, Yu Shi, Ning Xu, Xu Yang, Xin Geng, and Yong Rui · 2024
Closest in time.
Transcendence: Generative models can outperform the experts that train them
Edwin Zhang, Vincent Zhu, Naomi Saphra, Anat Kleiman, Benjamin L Edelman, Milind Tambe, Sham M Kakade, and Eran Malach · 2024
Closest in time.