Fetching the paper…
Reading the bibliography…
In this paper, we propose a new covering technique localized for the trajectories of SGD.
Wassily Hoeffding, Probability Inequalities for Sums of Bounded Random Variables , Journal of the American Statistical Association (1963)
1963
Earlier work this paper cites.
Leon Bottou and Yoshua Bengio, Convergence Properties of the K-Means Algorithms , Advances in Neural Information Processing Systems, 1994
1994
Earlier work this paper cites.
Olivier Bousquet and André Elisseeff, Stability and Generalization , Journal of Machine Learning Research (2002)
2002
Earlier work this paper cites.
Yann Guermeur, Combining Discriminant Models with New Multi-Class SVMs , Pattern Analysis & Applications (2002)
2002
Earlier work this paper cites.
David J.C. MacKay, Information Theory, Inference and Learning Algorithms , Cambridge University Press, 2003
2003
Earlier work this paper cites.
Tong Zhang, Statistical Analysis of Some Multi-Category Large Margin Classification Methods , Journal of Machine Learning Research (2004)
2004
Earlier work this paper cites.
András Antos, Improved Minimax Bounds on the Test and Training Distortion of Empirically Designed Vector Quantizers , IEEE Transactions on Information Theory (2005)
2005
Earlier work this paper cites.
Shai Shalev-Shwartz, Ohad Shamir, Nathan Srebro, and Karthik Sridharan, Stochastic Convex Optimization , Conference on Learning Theory, 2009
2009
Earlier work this paper cites.
David Sculley, Web-Scale K-Means Clustering , International Conference on World Wide Web, 2010
2010
Earlier work this paper cites.
Robert Jenssen, Marius Kloft, Alexander Zien, Sören Sonnenburg, and Klaus-Robert Müller, A Scatter-Based Prototype Framework and Multi-Class Extension of Support Vector Machines , PloS one (2012)
2012
Earlier work this paper cites.
Simon J.D. Prince, Computer vision: models, learning, and inference , Cambridge University Press, 2012
2012
Earlier work this paper cites.
Corinna Cortes, Mehryar Mohri, and Afshin Rostamizadeh, Multi-Class Classification with Maximum Margin Multiple Kernel , International Conference on Machine Learning, 2013
2013
Earlier work this paper cites.
Clément Levrard, Fast rates for empirical vector quantization , Electronic Journal of Statistics (2013)
2013
Earlier work this paper cites.
Kenneth Falconer, Fractal geometry: mathematical foundations and applications , Wiley, 2014
2014
Earlier work this paper cites.
Shai Shalev-Shwartz and Shai Ben-David, Understanding Machine Learning: From Theory to Algorithms , Cambridge university press, 2014
2014
Earlier work this paper cites.
Amit Daniely, Sivan Sabato, Shai Ben-David, and Shai Shalev-Shwartz, Multiclass Learnability and the ERM Principle , Journal of Machine Learning Research (2015)
2015
Earlier work this paper cites.
2015
Cited alongside, same era.
Yunwen Lei, Ürün Dogan, Alexander Binder, and Marius Kloft, Multi-class SVMs: From Tighter Data-Dependent Generalization Bounds to Novel Algorithms , Advances in Neural Information Processing Systems, 2015
2015
Cited alongside, same era.
Yu Nesterov, Universal gradient methods for convex optimization problems , Mathematical Programming (2015)
2015
Cited alongside, same era.
Matthew Thorpe, Florian Theil, Adam M Johansen, and Neil Cade, Convergence of the k-Means Minimization Problem using Γ \Gamma -Convergence , SIAM Journal on Applied Mathematics (2015)
2015
Cited alongside, same era.
Aymeric Dieuleveut, Alain Durmus, and Francis Bach, Bridging the Gap between Constant Step Size Stochastic Gradient Descent and Markov Chains , The Annals of Statistics (2020)
2020
Later among the works it cites.
Jian Li, Xuanyuan Luo, and Mingda Qiao, On Generalization Error Bounds of Noisy Gradient Methods for Non-Convex Learning , International Conference on Learning Representations, 2020
2020
Later among the works it cites.
Yunwen Lei and Yiming Ying, Fine-Grained Analysis of Stability and Generalization for Stochastic Gradient Descent , International Conference on Machine Learning, 2020
2020
Later among the works it cites.
Umut Şimşekli, Ozan Sener, George Deligiannidis, and Murat A. Erdogdu, Hausdorff Dimension, Heavy Tails, and Generalization in Neural Networks , Advances in Neural Information Processing Systems, 2020
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
Cheng Tang and Claire Monteleoni, On Lloyd’s Algorithm: New Theoretical Insights for Clustering in Practice , International Conference on Artificial Intelligence and Statistics, 2016
2016
Cited alongside, same era.
Ben London, A PAC-Bayesian analysis of randomized learning with application to stochastic gradient descent , Advances in Neural Information Processing Systems (2017)
2017
Cited alongside, same era.
Maxim Raginsky, Alexander Rakhlin, and Matus Telgarsky, Non-convex learning via Stochastic Gradient Langevin Dynamics: a nonasymptotic analysis , Conference on Learning Theory, 2017
2017
Cited alongside, same era.
Murat A Erdogdu, Lester Mackey, and Ohad Shamir, Global non-convex optimization with discretized diffusions , Advances in Neural Information Processing Systems 31
2018
Cited alongside, same era.
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro, The Implicit Bias of Gradient Descent on Separable Data , The Journal of Machine Learning Research (2018)
2018
Cited alongside, same era.
Roman Vershynin, High-dimensional probability: An introduction with applications in data science , Cambridge University Press, 2018
2018
Cited alongside, same era.
Vitaly Feldman and Jan Vondrak, High probability generalization bounds for uniformly stable algorithms with nearly optimal rate , Conference on Learning Theory, 2019
2019
Cited alongside, same era.
Idan Amir, Tomer Koren, and Roi Livni, SGD Generalizes Better Than GD (And Regularization Doesn’t Help) , Conference on Learning Theory, 2021
2021
Later among the works it cites.
Alexander Camuto, George Deligiannidis, Murat A. Erdogdu, Mert Gürbüzbalaban, Umut Şimşekli, and Lingjiong Zhu, Fractal Structure and Generalization Properties of Stochastic Optimization Algorithms , Advances in Neural Information Processing Systems, 2021
2021
Later among the works it cites.
Nikita Doikov and Yurii Nesterov, Minimizing uniformly convex functions by cubic regularization of newton method , Journal of Optimization Theory and Applications (2021)
2021
Later among the works it cites.
Murat A Erdogdu and Rasa Hosseinzadeh, On the convergence of langevin monte carlo: The interplay between tail growth and smoothness , Conference on Learning Theory, PMLR, 2021, pp. 1776–1822
2021
Later among the works it cites.
Farzan Farnia and Asuman Ozdaglar, Train simultaneously, generalize better: Stability of gradient-based minimax learners , International Conference on Machine Learning, 2021
2021
Later among the works it cites.
Shaojie Li and Yong Liu, Sharper Generalization Bounds for Clustering , International Conference on Machine Learning, 2021
2021
Later among the works it cites.
Yunwen Lei, Mingrui Liu, and Yiming Ying, Generalization Guarantee of SGD for Pairwise Learning , International Conference on Artificial Intelligence and Statistics, 2021
2021
Later among the works it cites.
Sejun Park, Jaeho Lee, Chulhee Yun, and Jinwoo Shin, Provable Memorization via Deep Neural Networks using Sub-linear Parameters , Conference on Learning Theory, 2021
2021
Later among the works it cites.
Lu Yu, Krishnakumar Balasubramanian, Stanislav Volgushev, and Murat A. Erdogdu, An Analysis of Constant Step Size SGD in the Non-convex Regime: Asymptotic Normality and Bias , Advances in Neural Information Processing Systems, 2021
2021
Later among the works it cites.
Chulhee Yun, Shankar Krishnan, and Hossein Mobahi, A Unifying View on Implicit Bias in Training Linear Neural Networks , International Conference on Learning Representations, 2021
2021
Later among the works it cites.
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals, Understanding deep learning (still) requires rethinking generalization , Communications of the ACM (2021)
2021
Later among the works it cites.