Fetching the paper…
Reading the bibliography…
Recently, Sharpness-Aware Minimization (SAM), which connects the geometry of the loss landscape and generalization, has demonstrated significant performance boosts on training large-scale models such as vision transformers.
Adaptive estimation of a quadratic functional by model selection
Beatrice Laurent and Pascal Massart · 2000
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Optimizing neural networks with kronecker-factored approximate curvature
James Martens and Roger Grosse · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Firecaffe: near-linear acceleration of deep neural network training on compute clusters
Forrest N Iandola, Matthew W Moskewicz, Khalid Ashraf, and Kurt Keutzer · 2016
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Earlier work this paper cites.
Sergey Zagoruyko and Nikos Komodakis · 2016
Earlier work this paper cites.
Extremely large minibatch sgd: Training resnet-50 on imagenet in 15 minutes
Takuya Akiba, Shuji Suzuki, and Keisuke Fukuda · 2017
Earlier work this paper cites.
Valeriu Codreanu, Damian Podareanu, and Vikram Saletore · 2017
Earlier work this paper cites.
Adabatch: Adaptive batch sizes for training deep neural networks
Aditya Devarakonda, Maxim Naumov, and Michael Garland · 2017
Earlier work this paper cites.
Sharp minima can generalize for deep nets
Laurent Dinh, Razvan Pascanu, Samy Bengio, and Yoshua Bengio · 2017
Earlier work this paper cites.
Gintare Karolina Dziugaite and Daniel M Roy · 2017
Earlier work this paper cites.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Earlier work this paper cites.
Elad Hoffer, Itay Hubara, and Daniel Soudry · 2017
Earlier work this paper cites.
Three factors influencing minima in sgd
Stanislaw Jastrzebski, Zachary Kenton, Devansh Arpit, Nicolas Ballas, Asja Fischer, Yoshua Bengio, and Amos Storkey · 2017
Earlier work this paper cites.
Visualizing the loss landscape of neural nets
Hao Li, Zheng Xu, Gavin Taylor, Christoph Studer, and Tom Goldstein · 2017
Cited alongside, same era.
Scaling distributed machine learning with system and algorithm co-design
Mu Li · 2017
Cited alongside, same era.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Cited alongside, same era.
A bayesian perspective on generalization and stochastic gradient descent
Samuel L Smith and Quoc V Le · 2017
Cited alongside, same era.
Don’t decay the learning rate, increase the batch size
Samuel L Smith, Pieter-Jan Kindermans, Chris Ying, and Quoc V Le · 2017
Entropy-sgd: Biasing gradient descent into wide valleys
Pratik Chaudhari, Anna Choromanska, Stefano Soatto, Yann LeCun, Carlo Baldassi, Christian Borgs, Jennifer Chayes, Levent Sagun, and Riccardo Zecchina · 2019
Later among the works it cites.
Asymmetric valleys: Beyond sharp and flat local minima
Haowei He, Gao Huang, and Yang Yuan · 2019
Later among the works it cites.
Peter Mattson, Christine Cheng, Cody Coleman, Greg Diamos, Paulius Micikevicius, David Patterson, Hanlin Tang, Gu-Yeon Wei, Peter Bailis, Victor Bittorf, et al · 2019
Later among the works it cites.
On the convergence of adam and beyond
Sashank J Reddi, Satyen Kale, and Sanjiv Kumar · 2019
Later among the works it cites.
Yet another accelerated sgd: Resnet-50 training on imagenet in 74.7 seconds
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
mixup: Beyond empirical risk minimization
Hongyi Zhang, Moustapha Cisse, Yann N Dauphin, and David Lopez-Paz · 2017
Cited alongside, same era.
JAX: composable transformations of Python+NumPy programs, 2018
James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, and Qiao Zhang · 2018
Cited alongside, same era.
Averaging weights leads to wider optima and better generalization
Pavel Izmailov, Dmitrii Podoprikhin, Timur Garipov, Dmitry Vetrov, and Andrew Gordon Wilson · 2018
Cited alongside, same era.
Xianyan Jia, Shutao Song, Wei He, Yangzihao Wang, Haidong Rong, Feihu Zhou, Liqiang Xie, Zhenyu Guo, Yuanzhou Yang, Liwei Yu, et al · 2018
Cited alongside, same era.
On the local minima of the empirical risk
Chi Jin, Lydia T Liu, Rong Ge, and Michael I Jordan · 2018
Cited alongside, same era.
Second-order optimization method for large mini-batch: Training resnet-50 on imagenet in 35 epochs
Kazuki Osawa, Yohei Tsuji, Yuichiro Ueno, Akira Naruse, Rio Yokota, and Satoshi Matsuoka · 2018
Cited alongside, same era.
Adaptive methods for nonconvex optimization
S Reddi, Manzil Zaheer, Devendra Sachan, Satyen Kale, and Sanjiv Kumar · 2018
Cited alongside, same era.
Masafumi Yamazaki, Akihiko Kasagi, Akihiro Tabuchi, Takumi Honda, Masahiro Miwa, Naoto Fukumoto, Tsuguchika Tabaru, Atsushi Ike, and Kohta Nakashima · 2019
Later among the works it cites.
Positively scale-invariant flatness of relu neural networks
Mingyang Yi, Qi Meng, Wei Chen, Zhi-ming Ma, and Tie-Yan Liu · 2019
Later among the works it cites.
Large batch optimization for deep learning: Training bert in 76 minutes
Yang You, Jing Li, Sashank Reddi, Jonathan Hseu, Sanjiv Kumar, Srinadh Bhojanapalli, Xiaodan Song, James Demmel, Kurt Keutzer, and Cho-Jui Hsieh · 2019
Later among the works it cites.
Why adam beats sgd for attention models
Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit, Seungyeon Kim, Sashank J Reddi, Sanjiv Kumar, and Suvrit Sra · 2019
Later among the works it cites.
Randaugment: Practical automated data augmentation with a reduced search space
Ekin D Cubuk, Barret Zoph, Jonathon Shlens, and Quoc V Le · 2020
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Later among the works it cites.
Sharpness-aware minimization for efficiently improving generalization
Pierre Foret, Ariel Kleiner, Hossein Mobahi, and Behnam Neyshabur · 2020
Later among the works it cites.
Normalized flat minima: Exploring scale invariant definition of flat minima for neural networks using pac-bayesian analysis
Yusuke Tsuzuku, Issei Sato, and Masashi Sugiyama · 2020
Later among the works it cites.
When vision transformers outperform resnets without pretraining or strong data augmentations
Xiangning Chen, Cho-Jui Hsieh, and Boqing Gong · 2021
Later among the works it cites.
Catastrophic fisher explosion: Early phase fisher matrix impacts generalization
Stanislaw Jastrzebski, Devansh Arpit, Oliver Astrand, Giancarlo B Kerg, Huan Wang, Caiming Xiong, Richard Socher, Kyunghyun Cho, and Krzysztof J Geras · 2021
Later among the works it cites.
Asam: Adaptive sharpness-aware minimization for scale-invariant learning of deep neural networks
Jungmin Kwon, Jeongseop Kim, Hyunseo Park, and In Kwon Choi · 2021
Later among the works it cites.
Concurrent adversarial learning for large-batch training
Yong Liu, Xiangning Chen, Minhao Cheng, Cho-Jui Hsieh, and Yang You · 2021
Later among the works it cites.