Fetching the paper…
Reading the bibliography…
One popular hypothesis of neural network generalization is that the flat local minima of loss surface in parameter space leads to good generalization.
Improving the convergence of back-propagation learning with second order methods
Sue Becker, Yann Le Cun, et al · 1988
Earlier work this paper cites.
Double backpropagation increasing generalization performance
Harris Drucker and Yann Le Cun · 1991
Earlier work this paper cites.
Comparison theorems in Riemannian geometry , volume 365
Jeff Cheeger and David G Ebin · 2008
Earlier work this paper cites.
Deep learning via hessian-free optimization
James Martens · 2010
Earlier work this paper cites.
Intriguing properties of neural networks
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus · 2013
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Christian Szegedy Ian J. Goodfellow, Jonathon Shlens · 2014
Earlier work this paper cites.
Qualitatively characterizing neural network optimization problems
Ian J Goodfellow, Oriol Vinyals, and Andrew M Saxe · 2015
Earlier work this paper cites.
Deep learning without poor local minima
Kenji Kawaguchi · 2016
Earlier work this paper cites.
Adversarial examples in the physical world
Alexey Kurakin, Ian Goodfellow, and Samy Bengio · 2016
Cited alongside, same era.
Unifying adversarial training algorithms with flexible deep data gradient regularization
II Ororbia, G Alexander, C Lee Giles, and Daniel Kifer · 2016
Cited alongside, same era.
The limitations of deep learning in adversarial settings
Nicolas Papernot, Patrick McDaniel, Somesh Jha, Matt Fredrikson, Z Berkay Celik, and Ananthram Swami · 2016
Cited alongside, same era.
Distillation as a defense to adversarial perturbations against deep neural networks
Nicolas Papernot, Patrick McDaniel, Xi Wu, Somesh Jha, and Ananthram Swami · 2016
Cited alongside, same era.
Towards evaluating the robustness of neural networks
Nicholas Carlini and David Wagner · 2017
Cited alongside, same era.
An empirical analysis of deep network loss surfaces
Daniel Jiwoong Im, Michael Tao, and Kristin Branson · 2017
Later among the works it cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2017
Later among the works it cites.
Visualizing the loss landscape of neural nets
Hao Li, Zheng Xu, Gavin Taylor, and Tom Goldstein · 2017
Later among the works it cites.
mixup: Beyond empirical risk minimization
Hongyi Zhang, Moustapha Cisse, Yann N Dauphin, and David Lopez-Paz · 2017
Later among the works it cites.
Negative eigenvalues of the hessian in deep neural networks
Guillaume Alain, Nicolas Le Roux, and Pierre-Antoine Manzagol · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Entropy-sgd: Biasing gradient descent into wide valleys
P Chaudhari, Anna Choromanska, S Soatto, Yann LeCun, C Baldassi, C Borgs, J Chayes, Levent Sagun, and R Zecchina · 2017
Cited alongside, same era.
Parseval networks: Improving robustness to adversarial examples
Moustapha Cisse, Piotr Bojanowski, Edouard Grave, Yann Dauphin, and Nicolas Usunier · 2017
Cited alongside, same era.
Sharp minima can generalize for deep nets
Laurent Dinh, Razvan Pascanu, Samy Bengio, and Yoshua Bengio · 2017
Cited alongside, same era.
Closest in time.
Harini Kannan, Alexey Kurakin, and Ian Goodfellow · 2018
Closest in time.
Towards deep learning models resistant to adversarial attacks
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu · 2018
Closest in time.
Improving the adversarial robustness and interpretability of deep neural networks by regularizing their input gradients
Andrew Slavin Ross and Finale Doshi-Velez · 2018
Closest in time.