Fetching the paper…
Reading the bibliography…
For linear classifiers, the relationship between (normalized) output margin and generalization is captured in a clear and simple bound -- a large output margin implies good generalization.
The sizes of compact subsets of hilbert space and continuity of gaussian processes
Richard M Dudley · 1967
Earlier work this paper cites.
A training algorithm for optimal margin classifiers
Bernhard E Boser, Isabelle M Guyon, and Vladimir N Vapnik · 1992
Earlier work this paper cites.
Support-vector networks
Corinna Cortes and Vladimir Vapnik · 1995
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results
Peter L Bartlett and Shahar Mendelson · 2002
Earlier work this paper cites.
Concentration inequalities and empirical processes theory applied to the analysis of learning algorithms
Olivier Bousquet · 2002
Earlier work this paper cites.
Empirical margin distributions and bounding the generalization error of combined classifiers
Vladimir Koltchinskii, Dmitry Panchenko, et al · 2002
Earlier work this paper cites.
Boosting as a regularized path to a maximum margin classifier
Saharon Rosset, Ji Zhu, and Trevor Hastie · 2004
Earlier work this paper cites.
Kernel methods in machine learning
Thomas Hofmann, Bernhard Schölkopf, and Alexander J Smola · 2008
Earlier work this paper cites.
On the complexity of linear prediction: Risk bounds, margin bounds, and regularization
Sham M Kakade, Karthik Sridharan, and Ambuj Tewari · 2009
Earlier work this paper cites.
Optimistic rates for learning with a smooth loss
Nathan Srebro, Karthik Sridharan, and Ambuj Tewari · 2010
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Deep learning face representation by joint identification-verification
Yi Sun, Yuheng Chen, Xiaogang Wang, and Xiaoou Tang · 2014
Earlier work this paper cites.
Train faster, generalize better: Stability of stochastic gradient descent
Moritz Hardt, Benjamin Recht, and Yoram Singer · 2015
Earlier work this paper cites.
Norm-based capacity control in neural networks
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2015
Earlier work this paper cites.
Entropy-sgd: Biasing gradient descent into wide valleys
Pratik Chaudhari, Anna Choromanska, Stefano Soatto, Yann LeCun, Carlo Baldassi, Christian Borgs, Jennifer Chayes, Levent Sagun, and Riccardo Zecchina · 2016
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Earlier work this paper cites.
A discriminative feature learning approach for deep face recognition
Yandong Wen, Kaipeng Zhang, Zhifeng Li, and Yu Qiao · 2016
Earlier work this paper cites.
Sergey Zagoruyko and Nikos Komodakis · 2016
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Earlier work this paper cites.
Spectrally-normalized margin bounds for neural networks
Peter L Bartlett, Dylan J Foster, and Matus J Telgarsky · 2017
Earlier work this paper cites.
Gintare Karolina Dziugaite and Daniel M Roy · 2017
Earlier work this paper cites.
Size-independent sample complexity of neural networks
Noah Golowich, Alexander Rakhlin, and Ohad Shamir · 2017
Cited alongside, same era.
Implicit regularization in matrix factorization
Suriya Gunasekar, Blake E Woodworth, Srinadh Bhojanapalli, Behnam Neyshabur, and Nati Srebro · 2017
Cited alongside, same era.
Yuanzhi Li, Tengyu Ma, and Hongyang Zhang · 2017
Cited alongside, same era.
Soft-margin softmax for deep classification
Xuezhi Liang, Xiaobo Wang, Zhen Lei, Shengcai Liao, and Stan Z Li · 2017
Cited alongside, same era.
Towards deep learning models resistant to adversarial attacks
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu · 2017
Robustness may be at odds with accuracy
Dimitris Tsipras, Shibani Santurkar, Logan Engstrom, Alexander Turner, and Aleksander Madry · 2018
Later among the works it cites.
Manifold mixup: Encouraging meaningful on-manifold interpolation as a regularizer
Vikas Verma, Alex Lamb, Christopher Beckham, Aaron Courville, Ioannis Mitliagkis, and Yoshua Bengio · 2018
Later among the works it cites.
Regularization Matters: Generalization and Optimization of Neural Nets v.s. their Induced Kernel
Colin Wei, Jason D. Lee, Qiang Liu, and Tengyu Ma · 2018
Later among the works it cites.
Hessian-based analysis of large batch training and robustness to adversaries
Zhewei Yao, Amir Gholami, Qi Lei, Kurt Keutzer, and Michael W Mahoney · 2018
Later among the works it cites.
Rademacher complexity for adversarially robust generalization
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Robust large margin deep neural networks
Jure Sokolić, Raja Giryes, Guillermo Sapiro, and Miguel RD Rodrigues · 2017
Cited alongside, same era.
Stronger generalization bounds for deep nets via a compression approach
Sanjeev Arora, Rong Ge, Behnam Neyshabur, and Yi Zhang · 2018
Cited alongside, same era.
To understand deep learning we need to understand kernel learning
Mikhail Belkin, Siyuan Ma, and Soumik Mandal · 2018
Cited alongside, same era.
Large margin deep networks for classification
Gamaleldin Elsayed, Dilip Krishnan, Hossein Mobahi, Kevin Regan, and Samy Bengio · 2018
Cited alongside, same era.
Generalizable adversarial training via spectral normalization
Farzan Farnia, Jesse M Zhang, and David Tse · 2018
Cited alongside, same era.
Implicit bias of gradient descent on linear convolutional networks
Suriya Gunasekar, Jason D Lee, Daniel Soudry, and Nati Srebro · 2018
Cited alongside, same era.
On the relation between the sharpest directions of dnn loss and the sgd step length
Stanislaw Jastrzebski, Zachary Kenton, Nicolas Ballas, Asja Fischer, Yoshua Bengio, and Amos Storkey · 2018
Cited alongside, same era.
Dong Yin, Kannan Ramchandran, and Peter Bartlett · 2018
Later among the works it cites.
Benign overfitting in linear regression
Peter L Bartlett, Philip M Long, Gábor Lugosi, and Alexander Tsigler · 2019
Closest in time.
Learning Imbalanced Datasets with Label-Distribution-Aware Margin Loss
Kaidi Cao, Colin Wei, Adrien Gaidon, Nikos Arechiga, and Tengyu Ma · 2019
Closest in time.
Robustness (python library), 2019
Logan Engstrom, Andrew Ilyas, Shibani Santurkar, and Dimitris Tsipras · 2019
Closest in time.
Surprises in high-dimensional ridgeless least squares interpolation
Trevor Hastie, Andrea Montanari, Saharon Rosset, and Ryan J Tibshirani · 2019
Closest in time.
A hessian based complexity measure for deep networks
Hamid Javadi, Randall Balestriero, and Richard Baraniuk · 2019
Closest in time.
Size-free generalization bounds for convolutional neural networks
Philip M Long and Hanie Sedghi · 2019
Closest in time.
Gradient descent maximizes the margin of homogeneous neural networks
Kaifeng Lyu and Jian Li · 2019
Closest in time.
Vc classes are adversarially robustly learnable, but only improperly
Omar Montasser, Steve Hanneke, and Nathan Srebro · 2019
Closest in time.
Lexicographic and depth-sensitive margins in homogeneous and non-homogeneous deep models
Mor Shpigel Nacson, Suriya Gunasekar, Jason D Lee, Nathan Srebro, and Daniel Soudry · 2019
Closest in time.
Deterministic pac-bayesian generalization bounds for deep networks via generalizing noise-resilience
Vaishnavh Nagarajan and J Zico Kolter · 2019
Closest in time.
Adversarial training can hurt generalization
Aditi Raghunathan, Sang Michael Xie, Fanny Yang, John C Duchi, and Percy Liang · 2019
Closest in time.
Data-dependent sample complexity of deep neural networks via lipschitz augmentation
Colin Wei and Tengyu Ma · 2019
Closest in time.
Adversarial examples improve image recognition
Cihang Xie, Mingxing Tan, Boqing Gong, Jiang Wang, Alan Yuille, and Quoc V Le · 2019
Closest in time.
Adversarial margin maximization networks
Ziang Yan, Yiwen Guo, and Changshui Zhang · 2019
Closest in time.
Theoretically principled trade-off between robustness and accuracy
Hongyang Zhang, Yaodong Yu, Jiantao Jiao, Eric Xing, Laurent El Ghaoui, and Michael Jordan · 2019
Closest in time.