Fetching the paper…
Reading the bibliography…
A growing number of machine learning scenarios rely on knowledge distillation where one uses the output of a surrogate model as labels to supervise the training of a target model.
Surprises in high-dimensional ridgeless least squares interpolation, 2020
Trevor Hastie, Andrea Montanari, Saharon Rosset, and Ryan J. Tibshirani · 1903
Earlier work this paper cites.
Andrea Montanari, Feng Ruan, Youngtak Sohn, and Jun Yan · 1911
Earlier work this paper cites.
On milman’s inequality and random subspaces which escape through a mesh in rn
Y. Gordon · 1988
Earlier work this paper cites.
Scaling laws for neural language models, 2020
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2001
Earlier work this paper cites.
Why distillation helps: a statistical perspective, 2020
Aditya Krishna Menon, Ankit Singh Rawat, Sashank J. Reddi, Seungyeon Kim, and Sanjiv Kumar · 2005
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky and Geoffrey Hinton · 2009
Earlier work this paper cites.
The smallest singular value of a random rectangular matrix, 2009
Mark Rudelson and Roman Vershynin · 2009
Earlier work this paper cites.
Benign overfitting in ridge regression, 2022
A. Tsigler and P. L. Bartlett · 2009
Earlier work this paper cites.
Theoretical analysis of self-training with deep networks on unlabeled data, 2022b
Colin Wei, Kendrick Shen, Yining Chen, and Tengyu Ma · 2010
Earlier work this paper cites.
Precise high-dimensional asymptotics for quantifying heterogeneous transfers, 2023
Fan Yang, Hongyang R. Zhang, Sen Wu, Christopher Ré, and Weijie J. Su · 2010
Earlier work this paper cites.
Deep residual learning for image recognition, 2015
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network, 2015
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
Regularized linear regression: A precise analysis of the estimation error
Christos Thrampoulidis, Samet Oymak, and Babak Hassibi · 2015
Earlier work this paper cites.
Deep learning scaling is predictable, empirically, 2017
Joel Hestness, Sharan Narang, Newsha Ardalani, Gregory Diamos, Heewoo Jun, Hassan Kianinejad, Md. Mostofa Ali Patwary, Yang Yang, and Yanqi Zhou · 2017
Earlier work this paper cites.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal · 2019
Earlier work this paper cites.
Benign overfitting in linear regression
Peter L. Bartlett, Philip M. Long, Gábor Lugosi, and Alexander Tsigler · 2020
Earlier work this paper cites.
Self-distillation amplifies regularization in hilbert space
Hossein Mobahi, Mehrdad Farajtabar, and Peter Bartlett · 2020
Earlier work this paper cites.
Asymptotic learning curves of kernel methods: empirical data versus teacher–student paradigm
Stefano Spigler, Mario Geiger, and Matthieu Wyart · 2020
Cited alongside, same era.
On the optimal weighted ℓ _ 2 \ell\_2 regularization in overparameterized linear regression
Denny Wu and Ji Xu · 2020
Cited alongside, same era.
Provable benefits of overparameterization in model compression: From double descent to pruning neural networks
Xiangyu Chang, Yingcong Li, Samet Oymak, and Christos Thrampoulidis · 2021
Cited alongside, same era.
A theoretical characterization of semi-supervised learning with self-training for gaussian mixture models
Samet Oymak and Talha Cihad Gulcu · 2021
Cited alongside, same era.
Asymptotics of ridge (less) regression under general source condition
Dominic Richards, Jaouad Mourtada, and Lorenzo Rosasco · 2021
Cited alongside, same era.
Towards a statistical theory of data selection under weak supervision, 2023
Germain Kolossov, Andrea Montanari, and Pulkit Tandon · 2023
Later among the works it cites.
On student-teacher deviations in distillation: does it pay to disobey?
Vaishnavh Nagarajan, Aditya K Menon, Srinadh Bhojanapalli, Hossein Mobahi, and Sanjiv Kumar · 2023
Later among the works it cites.
The eigenlearning framework: A conservation law perspective on kernel ridge regression and wide neural networks
James B Simon, Madeline Dickens, Dhruva Karkada, and Michael Deweese · 2023
Later among the works it cites.
Self-instruct: Aligning language models with self-generated instructions
Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A Smith, Daniel Khashabi, and Hannaneh Hajishirzi · 2023
Later among the works it cites.
Phi-3 technical report: A highly capable language model locally on your phone, 2024
Marah Abdin, Sam Ade Jacobs, Ammar Ahmad Awan, Jyoti Aneja, Ahmed Awadallah, Hany Awadalla, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Harkirat Behl, et al · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Generalization error rates in kernel regression: the crossover from the noiseless to noisy regime
Hugo Cui, Bruno Loureiro, Florent Krzakala, and Lenka Zdeborová · 2022
Cited alongside, same era.
Self-training converts weak learners to strong learners in mixture models
Spencer Frei, Difan Zou, Zixiang Chen, and Quanquan Gu · 2022
Cited alongside, same era.
Learning curves of generic features maps for realistic datasets with a teacher-student model*
Bruno Loureiro, Cédric Gerbelot, Hugo Cui, Sebastian Goldt, Florent Krzakala, Marc Mézard, and Lenka Zdeborová · 2022
Cited alongside, same era.
A solvable model of neural scaling laws, 2022
Alexander Maloney, Daniel A. Roberts, and James Sully · 2022
Cited alongside, same era.
The generalization error of random features regression: Precise asymptotics and the double descent curve
Song Mei and Andrea Montanari · 2022
Cited alongside, same era.
The interpolation phase transition in neural networks: Memorization and generalization under lazy training
Andrea Montanari and Yiqiao Zhong · 2022
Cited alongside, same era.
Beyond neural scaling laws: beating power law scaling via data pruning
Ben Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli, and Ari Morcos · 2022
Cited alongside, same era.
Closest in time.
Explaining neural scaling laws
Yasaman Bahri, Ethan Dyer, Jared Kaplan, Jaehoon Lee, and Utkarsh Sharma · 2024
Closest in time.
Quantifying the gain in weak-to-strong generalization, 2024
Moses Charikar, Chirag Pabbaraju, and Kirankumar Shiragur · 2024
Closest in time.
Dimension free ridge regression, 2024
Chen Cheng and Andrea Montanari · 2024
Closest in time.
Understanding inverse scaling and emergence in multitask representation learning
Muhammed E. Ildiz, Zhe Zhao, and Samet Oymak · 2024
Closest in time.
Scaling laws for learning with real and surrogate data, 2024
Ayush Jain, Andrea Montanari, and Eren Sasoglu · 2024
Closest in time.
Theoretical analysis of weak-to-strong generalization, 2024
Hunter Lang, David Sontag, and Aravindan Vijayaraghavan · 2024
Closest in time.
Minimum-norm interpolation under covariate shift, 2024
Neil Mallinar, Austin Zane, Spencer Frei, and Bin Yu · 2024
Closest in time.
4+3 phases of compute-optimal neural scaling laws, 2024
Elliot Paquette, Courtney Paquette, Lechao Xiao, and Jeffrey Pennington · 2024
Closest in time.
Optimal ridge regularization for out-of-distribution prediction, 2024
Pratik Patil, Jin-Hong Du, and Ryan J. Tibshirani · 2024
Closest in time.
James B. Simon, Dhruva Karkada, Nikhil Ghosh, and Mikhail Belkin · 2024
Closest in time.
Generalization error of min-norm interpolators in transfer learning, 2024
Yanke Song, Sohom Bhattacharya, and Pragya Sur · 2024
Closest in time.