Fetching the paper…
Reading the bibliography…
In this work, we propose an Implicit Regularization Enhancement (IRE) framework to accelerate the discovery of flat solutions in deep learning, thereby improving generalization and convergence.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 1907
Earlier work this paper cites.
On the variance of the adaptive learning rate and beyond
Liyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen, Xiaodong Liu, Jianfeng Gao, and Jiawei Han · 1908
Earlier work this paper cites.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
A method of solving a convex programming problem with convergence rate o ( 1 / k 2 ) o(1/k^{2})
Yurii Nesterov · 1983
Earlier work this paper cites.
Learning representations by back-propagating errors
David E Rumelhart, Geoffrey E Hinton, and Ronald J Williams · 1986
Earlier work this paper cites.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images, 2009
Alex Krizhevsky and Geoffrey Hinton · 2009
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Understanding the exploding gradient problem
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio · 2012
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman · 2012
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2014
Earlier work this paper cites.
Optimizing neural networks with Kronecker-factored approximate curvature
James Martens and Roger Grosse · 2015
Earlier work this paper cites.
A Kronecker-factored approximate Fisher matrix for convolution layers
Roger Grosse and James Martens · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Earlier work this paper cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Earlier work this paper cites.
Sergey Zagoruyko and Nikos Komodakis · 2016
Earlier work this paper cites.
Accurate, large minibatch SGD: Training ImageNet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Earlier work this paper cites.
Elad Hoffer, Itay Hubara, and Daniel Soudry · 2017
Earlier work this paper cites.
Three factors influencing minima in SGD
Stanisław Jastrzębski, Zachary Kenton, Devansh Arpit, Nicolas Ballas, Asja Fischer, Yoshua Bengio, and Amos Storkey · 2017
Earlier work this paper cites.
Stochastic modified equations and adaptive stochastic gradient algorithms
Qianxiao Li, Cheng Tai, and E Weinan · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Towards understanding generalization of deep learning: Perspective of loss landscapes
Lei Wu, Zhanxing Zhu, and Weinan E · 2017
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Earlier work this paper cites.
The loss landscape of overparameterized neural networks
Yaim Cooper · 2018
Earlier work this paper cites.
Essentially no barriers in neural network energy landscape
Felix Draxler, Kambis Veschgini, Manfred Salmhofer, and Fred Hamprecht · 2018
Earlier work this paper cites.
Gradient descent provably optimizes over-parameterized neural networks
Simon S Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2018
Earlier work this paper cites.
Loss surfaces, mode connectivity, and fast ensembling of dnns
Timur Garipov, Pavel Izmailov, Dmitrii Podoprikhin, Dmitry P Vetrov, and Andrew G Wilson · 2018
Earlier work this paper cites.
Fast approximate natural gradient descent in a Kronecker-factored eigenbasis
Thomas George, César Laurent, Xavier Bouthillier, Nicolas Ballas, and Pascal Vincent · 2018
Earlier work this paper cites.
Implicit bias of gradient descent on linear convolutional networks
Suriya Gunasekar, Jason D Lee, Daniel Soudry, and Nati Srebro · 2018
Cited alongside, same era.
The implicit bias of gradient descent on separable data
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro · 2018
Cited alongside, same era.
High-dimensional probability: An introduction with applications in data science , volume 47
Roman Vershynin · 2018
Cited alongside, same era.
How SGD selects the global minima in over-parameterized learning: A dynamical stability perspective
Lei Wu, Chao Ma, and Weinan E · 2018
Cited alongside, same era.
Transformer-XL: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V Le, and Ruslan Salakhutdinov · 2019
Cited alongside, same era.
On linear stability of SGD and input-smoothness of neural networks
Chao Ma and Lexing Ying · 2021
Later among the works it cites.
The implicit bias of minima stability: A view from function space
Rotem Mulayoff, Tomer Michaeli, and Daniel Soudry · 2021
Later among the works it cites.
Implicit bias of SGD for diagonal linear networks: a provable benefit of stochasticity
Scott Pesme, Loucas Pillaud-Vivien, and Nicolas Flammarion · 2021
Later among the works it cites.
Understanding gradient descent on the edge of stability in deep learning
Sanjeev Arora, Zhiyuan Li, and Abhishek Panigrahi · 2022
Later among the works it cites.
What happens after SGD reaches zero loss?–a mathematical framework
Zhiyuan Li, Tianhao Wang, and Sanjeev Arora · 2022
Later among the works it cites.
Towards efficient and scalable sharpness-aware minimization
Yong Liu, Siqi Mai, Xiangning Chen, Cho-Jui Hsieh, and Yang You · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gradient descent finds global minima of deep neural networks
Simon Du, Jason Lee, Haochuan Li, Liwei Wang, and Xiyu Zhai · 2019
Cited alongside, same era.
Openwebtext corpus
Aaron Gokaslan and Vanya Cohen · 2019
Cited alongside, same era.
Fantastic generalization measures and where to find them
Yiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan, and Samy Bengio · 2019
Cited alongside, same era.
Stochastic modified equations and dynamics of stochastic gradient algorithms I: Mathematical foundations
Qianxiao Li, Cheng Tai, and E Weinan · 2019
Cited alongside, same era.
Adaptive gradient methods with dynamic bound of learning rate
Liangchen Luo, Yuanhao Xiong, Yan Liu, and Xu Sun · 2019
Cited alongside, same era.
Gradient descent maximizes the margin of homogeneous neural networks
Kaifeng Lyu and Jian Li · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Cited alongside, same era.
Later among the works it cites.
Understanding the generalization benefit of normalization layers: Sharpness reduction
Kaifeng Lyu, Zhiyuan Li, and Sanjeev Arora · 2022
Later among the works it cites.
Beyond the quadratic approximation: The multiscale structure of neural network loss landscapes
Chao Ma, Daniel Kunin, Lei Wu, and Lexing Ying · 2022
Later among the works it cites.
Make sharpness-aware minimization stronger: A sparsified perturbation approach
Peng Mi, Li Shen, Tianhe Ren, Yiyi Zhou, Xiaoshuai Sun, Rongrong Ji, and Dacheng Tao · 2022
Later among the works it cites.
Implicit bias of the step size in linear diagonal neural networks
Mor Shpigel Nacson, Kavya Ravichandran, Nathan Srebro, and Daniel Soudry · 2022
Later among the works it cites.
Rethinking sharpness-aware minimization as variational inference
Szilvia Ujváry, Zsigmond Telek, Anna Kerekes, Anna Mészáros, and Ferenc Huszár · 2022
Later among the works it cites.
On margin maximization in linear and ReLU networks
Gal Vardi, Ohad Shamir, and Nati Srebro · 2022
Later among the works it cites.
On accelerated perceptrons and beyond
Guanghui Wang, Rafael Hanashiro, Etash Guha, and Jacob Abernethy · 2022
Later among the works it cites.
The alignment property of SGD noise and how it helps select flat minima: A stability analysis
Lei Wu, Mingze Wang, and Weijie J Su · 2022
Later among the works it cites.
(S)GD over diagonal linear networks: Implicit regularisation, large stepsizes and edge of stability
Mathieu Even, Scott Pesme, Suriya Gunasekar, and Nicolas Flammarion · 2023
Later among the works it cites.
The inductive bias of flatness regularization for deep matrix factorization
Khashayar Gatmiry, Zhiyuan Li, Ching-Yao Chuang, Sashank Reddi, Tengyu Ma, and Stefanie Jegelka · 2023
Later among the works it cites.
The MiniPile challenge for data-efficient language models
Jean Kaddour · 2023
Later among the works it cites.
The asymmetric maximum margin bias of quasi-homogeneous neural networks
Daniel Kunin, Atsushi Yamamura, Chao Ma, and Surya Ganguli · 2023
Later among the works it cites.
Same pre-training loss, better downstream: Implicit bias matters for language models
Hong Liu, Sang Michael Xie, Zhiyuan Li, and Tengyu Ma · 2023
Later among the works it cites.
Saddle-to-saddle dynamics in diagonal linear networks
Scott Pesme and Nicolas Flammarion · 2023
Later among the works it cites.
LLaMA: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Later among the works it cites.
On the implicit bias in deep-learning algorithms
Gal Vardi · 2023
Later among the works it cites.
Understanding multi-phase optimization dynamics and rich nonlinear behaviors of ReLU networks
Mingze Wang and Chao Ma · 2023
Later among the works it cites.
The noise geometry of stochastic gradient descent: A quantitative and analytical characterization
Mingze Wang and Lei Wu · 2023
Later among the works it cites.
The implicit regularization of dynamical stability in stochastic gradient descent
Lei Wu and Weijie J Su · 2023
Later among the works it cites.
Decentralized SGD and average-direction SAM are asymptotically equivalent
Tongtian Zhu, Fengxiang He, Kaixuan Chen, Mingli Song, and Dacheng Tao · 2023
Later among the works it cites.
Symbolic discovery of optimization algorithms
Xiangning Chen, Chen Liang, Da Huang, Esteban Real, Kaiyuan Wang, Hieu Pham, Xuanyi Dong, Thang Luong, Cho-Jui Hsieh, Yifeng Lu, et al · 2024
Closest in time.
(S) GD over diagonal linear networks: Implicit bias, large stepsizes and edge of stability
Mathieu Even, Scott Pesme, Suriya Gunasekar, and Nicolas Flammarion · 2024
Closest in time.
Sophia: A scalable stochastic second-order optimizer for language model pre-training
Hong Liu, Zhiyuan Li, David Hall, Percy Liang, and Tengyu Ma · 2024
Closest in time.
Normalization layers are all that sharpness-aware minimization needs
Maximilian Mueller, Tiffany Vlaar, David Rolnick, and Matthias Hein · 2024
Closest in time.
Roformer: Enhanced transformer with rotary position embedding
Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu · 2024
Closest in time.
Achieving margin maximization exponentially fast via progressive norm rescaling
Mingze Wang, Zeping Min, and Lei Wu · 2024
Closest in time.