Fetching the paper…
Reading the bibliography…
We formally study how ensemble of deep learning models can improve test accuracy, and how the superior performance of ensemble can be distilled into a single model using knowledge distillation.
Sanjeev Arora, Simon S. Du, Wei Hu, Zhiyuan Li, and Ruosong Wang · 1901
Earlier work this paper cites.
Can SGD Learn Recurrent Neural Networks with Provable Generalization?
Zeyuan Allen-Zhu and Yuanzhi Li · 1902
Earlier work this paper cites.
On exact computation with an infinitely wide neural net
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, Ruslan Salakhutdinov, and Ruosong Wang · 1904
Earlier work this paper cites.
What Can ResNet Learn Efficiently, Going Beyond Kernels?
Zeyuan Allen-Zhu and Yuanzhi Li · 1905
Earlier work this paper cites.
Bootstrapping regression models
David A Freedman et al · 1981
Earlier work this paper cites.
Neural network ensembles
Lars Kai Hansen and Peter Salamon · 1990
Earlier work this paper cites.
When networks disagree: Ensemble methods for hybrid neural networks
Michael P Perrone and Leon N Cooper · 1992
Earlier work this paper cites.
Neural network ensembles, cross validation, and active learning
Anders Krogh and Jesper Vedelsby · 1994
Earlier work this paper cites.
Bagging predictors
Leo Breiman · 1996
Earlier work this paper cites.
A decision-theoretic generalization of on-line learning and an application to boosting
Yoav Freund and Robert E Schapire · 1997
Earlier work this paper cites.
Wrappers for feature subset selection
Ron Kohavi, George H John, et al · 1997
Earlier work this paper cites.
The random subspace method for constructing decision forests
Tin Kam Ho · 1998
Earlier work this paper cites.
On combining classifiers
Josef Kittler, Mohamad Hatef, Robert PW Duin, and Jiri Matas · 1998
Earlier work this paper cites.
Boosting the margin: A new explanation for the effectiveness of voting methods
Robert E Schapire, Yoav Freund, Peter Bartlett, Wee Sun Lee, et al · 1998
Earlier work this paper cites.
A short introduction to boosting
Yoav Freund, Robert Schapire, and Naoki Abe · 1999
Earlier work this paper cites.
Popular ensemble methods: An empirical study
David Opitz and Richard Maclin · 1999
Earlier work this paper cites.
Feature selection for ensembles
David W Opitz · 1999
Earlier work this paper cites.
Ensemble methods in machine learning
Thomas G Dietterich · 2000
Earlier work this paper cites.
Additive logistic regression: a statistical view of boosting (with discussion and a rejoinder by the authors)
Jerome Friedman, Trevor Hastie, Robert Tibshirani, et al · 2000
Earlier work this paper cites.
Greedy function approximation: a gradient boosting machine
Jerome H Friedman · 2001
Earlier work this paper cites.
Ensembling neural networks: many could be better than all
Zhi-Hua Zhou, Jianxin Wu, and Wei Tang · 2002
Earlier work this paper cites.
Attribute bagging: improving accuracy of classifier ensembles by using random feature subsets
Robert Bryll, Ricardo Gutierrez-Osuna, and Francis Quek · 2003
Earlier work this paper cites.
Feature selection for ensembles: A hierarchical multi-objective genetic algorithm approach
Luiz S Oliveira, Robert Sabourin, Flávio Bortolozzi, and Ching Y Suen · 2003
Earlier work this paper cites.
Bias-variance analysis of support vector machines for the development of svm-based ensemble methods
Giorgio Valentini and Thomas G Dietterich · 2004
Earlier work this paper cites.
Diversity in search strategies for ensemble feature selection
Alexey Tsymbal, Mykola Pechenizkiy, and Pádraig Cunningham · 2005
Earlier work this paper cites.
An experimental bias-variance analysis of svm ensembles based on resampling techniques
Giorgio Valentini · 2005
Earlier work this paper cites.
Ensemble based systems in decision making
Robi Polikar · 2006
Earlier work this paper cites.
Rotation forest: A new classifier ensemble method
Juan José Rodriguez, Ludmila I Kuncheva, and Carlos J Alonso · 2006
Earlier work this paper cites.
Dynamic weighted majority: An ensemble method for drifting concepts
J Zico Kolter and Marcus A Maloof · 2007
Cited alongside, same era.
Data mining with decision trees: theory and applications , volume 69
Lior Rokach and Oded Z Maimon · 2008
Cited alongside, same era.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Cited alongside, same era.
On feature selection, bias-variance, and bagging
M Arthur Munson and Rich Caruana · 2009
Cited alongside, same era.
A review on ensembles for the class imbalance problem: bagging-, boosting-, and hybrid-based approaches
Mikel Galar, Alberto Fernandez, Edurne Barrenechea, Humberto Bustince, and Francisco Herrera · 2011
Cited alongside, same era.
Semantic road segmentation via multi-scale ensembles of learned features
Jose M Alvarez, Yann LeCun, Theo Gevers, and Antonio M Lopez · 2012
Theoretical insights into the optimization landscape of over-parameterized shallow neural networks
Mahdi Soltanolkotabi, Adel Javanmard, and Jason D Lee · 2017
Later among the works it cites.
Yuandong Tian · 2017
Later among the works it cites.
Recovery guarantees for one-hidden-layer neural networks
Kai Zhong, Zhao Song, Prateek Jain, Peter L Bartlett, and Inderjit S Dhillon · 2017
Later among the works it cites.
Learning two layer rectified neural networks in polynomial time
Ainesh Bakshi, Rajesh Jayaram, and David P Woodruff · 2018
Later among the works it cites.
Feature selection in machine learning: A new perspective
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Stochastic gradient descent optimizes over-parameterized deep relu networks
Difan Zou, Yuan Cao, Dongruo Zhou, and Quanquan Gu · 2012
Cited alongside, same era.
Fast decorrelated neural network ensembles with random weights
Monther Alhamdoosh and Dianhui Wang · 2014
Cited alongside, same era.
Combining pattern classifiers: methods and algorithms
Ludmila I Kuncheva · 2014
Cited alongside, same era.
Comparison and anti-concentration bounds for maxima of gaussian random vectors
Victor Chernozhukov, Denis Chetverikov, and Kengo Kato · 2015
Cited alongside, same era.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Cited alongside, same era.
Distilling knowledge from ensembles of neural networks for speech recognition
Yevgen Chebotar and Austin Waters · 2016
Cited alongside, same era.
Jie Cai, Jiawei Luo, Shulin Wang, and Sheng Yang · 2018
Later among the works it cites.
Tommaso Furlanello, Zachary C Lipton, Michael Tschannen, Laurent Itti, and Anima Anandkumar · 2018
Later among the works it cites.
Learning two-layer neural networks with symmetric inputs
Rong Ge, Rohith Kuditipudi, Zhize Li, and Xiang Wang · 2018
Later among the works it cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Later among the works it cites.
Knowledge distillation by on-the-fly native ensemble
Xu Lan, Xiatian Zhu, and Shaogang Gong · 2018
Later among the works it cites.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Yuanzhi Li and Yingyu Liang · 2018
Later among the works it cites.
Algorithmic regularization in over-parameterized matrix sensing and neural networks with quadratic activations
Yuanzhi Li, Tengyu Ma, and Hongyang Zhang · 2018
Later among the works it cites.
Polynomial convergence of gradient descent for training one-hidden-layer neural networks
Santosh Vempala and John Wilmes · 2018
Later among the works it cites.
Ensembles for feature selection: A review and future trends
Verónica Bolón-Canedo and Amparo Alonso-Betanzos · 2019
Later among the works it cites.
Generalization bounds of stochastic gradient descent for wide and deep neural networks
Yuan Cao and Quanquan Gu · 2019
Later among the works it cites.
Linearized two-layers neural networks in high dimension
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2019
Later among the works it cites.
Finite depth and width corrections to the neural tangent kernel
Boris Hanin and Mihai Nica · 2019
Later among the works it cites.
Yuanzhi Li, Colin Wei, and Tengyu Ma · 2019
Later among the works it cites.
Xiaodong Liu, Pengcheng He, Weizhu Chen, and Jianfeng Gao · 2019
Later among the works it cites.
A high-bias, low-variance introduction to machine learning for physicists
Pankaj Mehta, Marin Bukov, Ching-Hao Wang, Alexandre GR Day, Clint Richardson, Charles K Fisher, and David J Schwab · 2019
Later among the works it cites.
Samet Oymak and Mahdi Soltanolkotabi · 2019
Later among the works it cites.
Greg Yang · 2019
Later among the works it cites.
On the power and limitations of random features for understanding neural networks
Gilad Yehudai and Ohad Shamir · 2019
Later among the works it cites.
Be your own teacher: Improve the performance of convolutional neural networks via self distillation
Linfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen, Chenglong Bao, and Kaisheng Ma · 2019
Later among the works it cites.
Backward feature correction: How deep learning performs deep learning
Zeyuan Allen-Zhu and Yuanzhi Li · 2020
Closest in time.
When can wasserstein gans minimize wasserstein distance?
Yuanzhi Li and Zehao Dou · 2020
Closest in time.
Learning over-parametrized two-layer relu neural networks beyond ntk
Yuanzhi Li, Tengyu Ma, and Hongyang R Zhang · 2020
Closest in time.
Self-distillation amplifies regularization in hilbert space
Hossein Mobahi, Mehrdad Farajtabar, and Peter L Bartlett · 2020
Closest in time.
Neural kernels without tangents
Vaishaal Shankar, Alex Fang, Wenshuo Guo, Sara Fridovich-Keil, Ludwig Schmidt, Jonathan Ragan-Kelley, and Benjamin Recht · 2020
Closest in time.