Fetching the paper…
Reading the bibliography…
In recent years the empirical success of transfer learning with neural networks has stimulated an increasing interest in obtaining a theoretical understanding of its core properties.
Some inequalities for gaussian processes and applications
Yehoram Gordon · 1985
Earlier work this paper cites.
Spin glass theory and beyond: An Introduction to the Replica Method and Its Applications , volume 9
Marc Mézard, Giorgio Parisi, and Miguel Virasoro · 1987
Earlier work this paper cites.
Phase diagram of coupled glassy systems: A mean-field study
Silvio Franz and Giorgio Parisi · 1997
Earlier work this paper cites.
Statistical mechanics of learning
Andreas Engel and Christian Van den Broeck · 2001
Earlier work this paper cites.
The organization of behavior: A neuropsychological theory
Donald Olding Hebb · 2005
Earlier work this paper cites.
Information, physics, and computation
Marc Mezard and Andrea Montanari · 2009
Earlier work this paper cites.
Entropy landscape of solutions in the binary perceptron problem
Haiping Huang, KY Michael Wong, and Yoshiyuki Kabashima · 2013
Earlier work this paper cites.
Learning both weights and connections for efficient neural network
Song Han, Jeff Pool, John Tran, and William Dally · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeffrey Dean · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2015
Earlier work this paper cites.
Learning may need only a few bits of synaptic precision
Carlo Baldassi, Federica Gerace, Carlo Lucibello, Luca Saglietti, and Riccardo Zecchina · 2016
Earlier work this paper cites.
Deep compression: Compressing deep neural network with pruning, trained quantization and huffman coding
Song Han, Huizi Mao, and William J. Dally · 2016
Earlier work this paper cites.
Sequence-level knowledge distillation
Yoon Kim and Alexander M. Rush · 2016
Cited alongside, same era.
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna · 2016
Cited alongside, same era.
Patient-driven privacy control through generalized distillation
Z Berkay Celik, David Lopez-Paz, and Patrick McDaniel · 2017
Cited alongside, same era.
Learning efficient object detection models with knowledge distillation
Guobin Chen, Wongun Choi, Xiang Yu, Tony Han, and Manmohan Chandraker · 2017
Cited alongside, same era.
A gift from knowledge distillation: Fast optimization, network minimization and transfer learning
Junho Yim, Donggyu Joo, Jihoon Bae, and Junmo Kim · 2017
Cited alongside, same era.
Visual relationship detection with internal and external linguistic knowledge distillation
Optimal errors and phase transitions in high-dimensional generalized linear models
Jean Barbier, Florent Krzakala, Nicolas Macris, Léo Miolane, and Lenka Zdeborová · 2019
Later among the works it cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Jonathan Frankle and Michael Carbin · 2019
Later among the works it cites.
Modelling the influence of data structure on learning in neural networks
Sebastian Goldt, Marc Mézard, Florent Krzakala, and Lenka Zdeborová · 2019
Later among the works it cites.
Surprises in high-dimensional ridgeless least squares interpolation
Trevor Hastie, Andrea Montanari, Saharon Rosset, and Ryan J Tibshirani · 2019
Later among the works it cites.
The generalization error of random features regression: Precise asymptotics and double descent curve
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ruichi Yu, Ang Li, Vlad I Morariu, and Larry S Davis · 2017
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Cited alongside, same era.
Large scale distributed neural network training through online distillation
Rohan Anil, Gabriel Pereyra, Alexandre Passos, Róbert Ormándi, George E. Dahl, and Geoffrey E. Hinton · 2018
Cited alongside, same era.
The committee machine: Computational to statistical gaps in learning a two-layers neural network
Benjamin Aubin, Antoine Maillard, Florent Krzakala, Nicolas Macris, Lenka Zdeborová, et al · 2018
Cited alongside, same era.
Role of synaptic stochasticity in training low-precision neural networks
Carlo Baldassi, Federica Gerace, Hilbert J Kappen, Carlo Lucibello, Luca Saglietti, Enzo Tartaglione, and Riccardo Zecchina · 2018
Cited alongside, same era.
Darkrank: Accelerating deep metric learning via cross sample similarities transfer
Yuntao Chen, Naiyan Wang, and Zhaoxiang Zhang · 2018
Cited alongside, same era.
Born again neural networks
Tommaso Furlanello, Zachary Lipton, Michael Tschannen, Laurent Itti, and Anima Anandkumar · 2018
Cited alongside, same era.
Song Mei and Andrea Montanari · 2019
Later among the works it cites.
When does label smoothing help?
Rafael Müller, Simon Kornblith, and Geoffrey E Hinton · 2019
Later among the works it cites.
Towards understanding knowledge distillation
Mary Phuong and Christoph Lampert · 2019
Later among the works it cites.
A modern maximum-likelihood theory for high-dimensional logistic regression
Pragya Sur and Emmanuel J Candès · 2019
Later among the works it cites.
Revisit knowledge distillation: a teacher-free framework
Li Yuan, Francis EH Tay, Guilin Li, Tao Wang, and Jiashi Feng · 2019
Later among the works it cites.
Generalisation error in learning with random features and the hidden manifold model
Federica Gerace, Bruno Loureiro, Florent Krzakala, Marc Mezard, and Lenka Zdeborova · 2020
Closest in time.
The role of regularization in classification of high-dimensional noisy Gaussian mixture
Francesca Mignacco, Florent Krzakala, Yue Lu, Pierfrancesco Urbani, and Lenka Zdeborova · 2020
Closest in time.
On the unreasonable effectiveness of knowledge distillation: Analysis in the kernel regime
Arman Rahbar, Ashkan Panahi, Chiranjib Bhattacharyya, Devdatt P. Dubhashi, and Morteza Haghir Chehreghani · 2020
Closest in time.
Understanding and improving knowledge distillation
Jiaxi Tang, Rakesh Shivanna, Zhe Zhao, Dong Lin, Anima Singh, Ed H Chi, and Sagar Jain · 2020
Closest in time.