Fetching the paper…
Reading the bibliography…
With the motive of training all the parameters of a neural network, we study why and when one can achieve this by iteratively creating, training, and combining randomly selected subnetworks.
On the convergence of fedavg on non-iid data, 2019
Xiang Li, Kaixuan Huang, Wenhao Yang, Shusen Wang, and Zhihua Zhang · 1907
Earlier work this paper cites.
First analysis of local gd on heterogeneous data, 2019a
Ahmed Khaled, Konstantin Mishchenko, and Peter Richtárik · 1909
Earlier work this paper cites.
Tighter theory for local sgd on identical and heterogeneous data, 2019b
Ahmed Khaled, Konstantin Mishchenko, and Peter Richtárik · 1909
Earlier work this paper cites.
On the convergence of local descent methods in federated learning, 2019
Farzin Haddadpour and Mehrdad Mahdavi · 1910
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Efficient large-scale distributed training of conditional maximum entropy models
Ryan Mcdonald, Mehryar Mohri, Nathan Silberman, Dan Walker, and Gideon S Mann · 2009
Earlier work this paper cites.
Parallelized stochastic gradient descent
Martin Zinkevich, Markus Weimer, Lihong Li, and Alex J Smola · 2010
Earlier work this paper cites.
Regularization of neural networks using dropconnect
Li Wan, Matthew Zeiler, Sixin Zhang, Yann Le Cun, and Rob Fergus · 2013
Earlier work this paper cites.
Fast dropout training
Sida Wang and Christopher Manning · 2013
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
Dimmwitted: A study of main-memory statistical analytics
Ce Zhang and Christopher Ré · 2014
Earlier work this paper cites.
Binaryconnect: Training deep neural networks with binary weights during propagations
Matthieu Courbariaux, Yoshua Bengio, and Jean-Pierre David · 2015
Earlier work this paper cites.
Deep residual learning for image recognition, 2015
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Variational dropout and the local reparameterization trick
Durk P Kingma, Tim Salimans, and Max Welling · 2015
Earlier work this paper cites.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Yarin Gal and Zoubin Ghahramani · 2016
Earlier work this paper cites.
Deep learning , volume 1
Ian Goodfellow, Yoshua Bengio, Aaron Courville, and Yoshua Bengio · 2016
Earlier work this paper cites.
Parallel SGD: When does averaging help?
Jian Zhang, Christopher De Sa, Ioannis Mitliagkas, and Christopher Ré · 2016
Earlier work this paper cites.
A closer look at memorization in deep networks
Devansh Arpit, Stanisław Jastrzębski, Nicolas Ballas, David Krueger, Emmanuel Bengio, Maxinder S Kanwal, Tegan Maharaj, Asja Fischer, Aaron Courville, Yoshua Bengio, et al · 2017
Earlier work this paper cites.
Spectrally-normalized margin bounds for neural networks
Peter L Bartlett, Dylan J Foster, and Matus J Telgarsky · 2017
Earlier work this paper cites.
Gintare Karolina Dziugaite and Daniel M Roy · 2017
Earlier work this paper cites.
Generalization in deep learning
Kenji Kawaguchi, Leslie Pack Kaelbling, and Yoshua Bengio · 2017
Earlier work this paper cites.
Exploring generalization in deep learning
Behnam Neyshabur, Srinadh Bhojanapalli, David Mcallester, and Nati Srebro · 2017
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean · 2017
Earlier work this paper cites.
Learning and generalization in overparameterized neural networks, going beyond two layers, 2018
Zeyuan Allen-Zhu, Yuanzhi Li, and Yingyu Liang · 2018
Earlier work this paper cites.
Stronger generalization bounds for deep nets via a compression approach
Sanjeev Arora, Rong Ge, Behnam Neyshabur, and Yi Zhang · 2018
Earlier work this paper cites.
Overfitting or perfect fitting? risk bounds for classification and regression rules that interpolate
Mikhail Belkin, Daniel J Hsu, and Partha Mitra · 2018
Earlier work this paper cites.
Gradient descent provably optimizes over-parameterized neural networks
Simon S Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2018
Earlier work this paper cites.
Size-independent sample complexity of neural networks
Noah Golowich, Alexander Rakhlin, and Ohad Shamir · 2018
Cited alongside, same era.
Neural tangent kernel: convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Cited alongside, same era.
The power of interpolation: Understanding the effectiveness of SGD in modern over-parametrized learning
Siyuan Ma, Raef Bassily, and Mikhail Belkin · 2018
Cited alongside, same era.
A PAC-Bayesian approach to spectrally-normalized margin bounds for neural networks
Behnam Neyshabur, Srinadh Bhojanapalli, and Nathan Srebro · 2018
Cited alongside, same era.
Theoretical insights into the optimization landscape of over-parameterized shallow neural networks
Mahdi Soltanolkotabi, Adel Javanmard, and Jason D Lee · 2018
Cited alongside, same era.
The implicit bias of gradient descent on separable data
The implicit regularization of ordinary least squares ensembles
Daniel LeJeune, Hamid Javadi, and Richard Baraniuk · 2020
Later among the works it cites.
Gshard: Scaling giant models with conditional computation and automatic sharding
Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen · 2020
Later among the works it cites.
Gradient descent with early stopping is provably robust to label noise for overparameterized neural networks
Mingchen Li, Mahdi Soltanolkotabi, and Samet Oymak · 2020
Later among the works it cites.
A mean field analysis of deep ResNet and beyond: Towards provably optimization via overparameterization from depth
Yiping Lu, Chao Ma, Yulong Lu, Jianfeng Lu, and Lexing Ying · 2020
Later among the works it cites.
On convergence and generalization of dropout training, 2020
Poorya Mianjy and Raman Arora · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro · 2018
Cited alongside, same era.
Local sgd converges fast and communicates little, 2018
Sebastian U. Stich · 2018
Cited alongside, same era.
An empirical study of example forgetting during deep neural network learning
Mariya Toneva, Alessandro Sordoni, Remi Tachet des Combes, Adam Trischler, Yoshua Bengio, and Geoffrey J Gordon · 2018
Cited alongside, same era.
Slimmable neural networks
Jiahui Yu, Linjie Yang, Ning Xu, Jianchao Yang, and Thomas Huang · 2018
Cited alongside, same era.
Non-vacuous generalization bounds at the ImageNet scale: a PAC-Bayesian compression approach
Wenda Zhou, Victor Veitch, Morgane Austern, Ryan P Adams, and Peter Orbanz · 2018
Cited alongside, same era.
Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks, 2019
Sanjeev Arora, Simon S. Du, Wei Hu, Zhiyuan Li, and Ruosong Wang · 2019
Cited alongside, same era.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal · 2019
Cited alongside, same era.
Toward moderate overparameterization: Global convergence guarantees for training shallow neural networks
Samet Oymak and Mahdi Soltanolkotabi · 2020
Later among the works it cites.
On generalization of adaptive methods for over-parameterized linear regression
Vatsal Shah, Soumya Basu, Anastasios Kyrillidis, and Sujay Sanghavi · 2020
Later among the works it cites.
Quadratic suffices for over-parametrization via matrix chernoff bound, 2020
Zhao Song and Xin Yang · 2020
Later among the works it cites.
Distributed learning of deep neural networks using independent subnet training, 2020
Binhang Yuan, Cameron R. Wolfe, Chen Dun, Yuxin Tang, Anastasios Kyrillidis, and Christopher M. Jermaine · 2020
Later among the works it cites.
Gradient descent optimizes over-parameterized deep ReLU networks
Difan Zou, Yuan Cao, Dongruo Zhou, and Quanquan Gu · 2020
Later among the works it cites.
Deep learning: a statistical viewpoint
Peter L Bartlett, Andrea Montanari, and Alexander Rakhlin · 2021
Closest in time.
Mikhail Belkin · 2021
Closest in time.
Few-shot image classification: Just use a library of pre-trained feature extractors and a simple classifier
Arkabandhu Chowdhury, Mingchao Jiang, Swarat Chaudhuri, and Chris Jermaine · 2021
Closest in time.
Local sgd optimizes overparameterized neural networks in polynomial time
Yuyang Deng and Mehrdad Mahdavi · 2021
Closest in time.
Resist: Layer-wise decomposition of resnets for distributed training, 2021
Chen Dun, Cameron R. Wolfe, Christopher M. Jermaine, and Anastasios Kyrillidis · 2021
Closest in time.
Modeling from features: a mean-field framework for over-parameterized deep neural networks
Cong Fang, Jason Lee, Pengkun Yang, and Tong Zhang · 2021
Closest in time.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
William Fedus, Barret Zoph, and Noam Shazeer · 2021
Closest in time.
Fl-ntk: A neural tangent kernel-based framework for federated learning convergence analysis, 2021
Baihe Huang, Xiaoxiao Li, Zhao Song, and Xin Yang · 2021
Closest in time.
The flip side of the reweighted coin: Duality of adaptive dropout and regularization
Daniel LeJeune, Hamid Javadi, and Richard Baraniuk · 2021
Closest in time.
Simultaneous training of partially masked neural networks
Amirkeivan Mohtashami, Martin Jaggi, and Sebastian U Stich · 2021
Closest in time.
Quynh Nguyen · 2021
Closest in time.
Subquadratic overparameterization for shallow neural networks, 2021
Chaehwan Song, Ali Ramezani-Kebrya, Thomas Pethick, Armin Eftekhari, and Volkan Cevher · 2021
Closest in time.
Pufferfish: Communication-efficient models at no extra cost
Hongyi Wang, Saurabh Agarwal, and Dimitris Papailiopoulos · 2021
Closest in time.
GIST: Distributed training for large-scale graph convolutional networks
Cameron R Wolfe, Jingkang Yang, Arindam Chowdhury, Chen Dun, Artun Bayer, Santiago Segarra, and Anastasios Kyrillidis · 2021
Closest in time.
Minipatch learning as implicit ridge-like regularization
Tianyi Yao, Daniel LeJeune, Hamid Javadi, Richard G Baraniuk, and Genevera I Allen · 2021
Closest in time.
Understanding deep learning (still) requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2021
Closest in time.