Fetching the paper…
Reading the bibliography…
Recently, flat-minima optimizers, which seek to find parameters in low-loss neighborhoods, have been shown to improve a neural network's generalization performance over stochastic and adaptive gradient-based optimizers.
Roberta: A robustly optimized bert pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 1907
Earlier work this paper cites.
Praktische verfahren der gleichungsauflösung
Mises, R. and Pollaczek-Geiringer, H · 1929
Earlier work this paper cites.
Calculation of gauss quadrature rules
Golub, G. H. and Welsch, J. H · 1969
Earlier work this paper cites.
Learning Representations by Back-Propagating Errors , pp. 696–699
Rumelhart, D. E., Hinton, G. E., and Williams, R. J · 1988
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
Polyak, B. T. and Juditsky, A. B · 1992
Earlier work this paper cites.
Flat minima
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Actor-critic algorithms
Konda, V. and Tsitsiklis, J · 1999
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A · 2009
Earlier work this paper cites.
Revisiting "qualitatively characterizing neural network optimization problems"
Frankle, J · 2012
Earlier work this paper cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Dauphin, Y. N., Pascanu, R., Gülçehre, Ç., Cho, K., Ganguli, S., and Bengio, Y · 2014
Earlier work this paper cites.
Qualitatively characterizing neural network optimization problems
Goodfellow, I. J. and Vinyals, O · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Complex embeddings for simple link prediction
Trouillon, T., Welbl, J., Riedel, S., Gaussier, É., and Bouchard, G · 2016
Earlier work this paper cites.
Wide residual networks
Zagoruyko, S. and Komodakis, N · 2016
Earlier work this paper cites.
Entropy-sgd: Biasing gradient descent into wide valleys
Chaudhari, P., Choromanska, A., Soatto, S., LeCun, Y., Baldassi, C., Borgs, C., Chayes, J. T., Sagun, L., and Zecchina, R · 2017
Earlier work this paper cites.
Improved regularization of convolutional neural networks with cutout
Devries, T. and Taylor, G. W · 2017
Earlier work this paper cites.
Computing nonvacuous generalization bounds for deep (stochastic) neural networks with many more parameters than training data
Dziugaite, G. K. and Roy, D. M · 2017
Earlier work this paper cites.
Inductive representation learning on large graphs
Hamilton, W. L., Ying, R., and Leskovec, J · 2017
Earlier work this paper cites.
Deep pyramidal residual networks
Han, D., Kim, J., and Kim, J · 2017
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., and Hochreiter, S · 2017
Earlier work this paper cites.
Snapshot ensembles: Train 1, get M for free
Huang, G., Li, Y., Pleiss, G., Liu, Z., Hopcroft, J. E., and Weinberger, K. Q · 2017
Earlier work this paper cites.
Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension
Joshi, M., Choi, E., Weld, D. S., and Zettlemoyer, L · 2017
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P · 2017
Earlier work this paper cites.
Semi-supervised classification with graph convolutional networks
Kipf, T. N. and Welling, M · 2017
Earlier work this paper cites.
Simple and scalable predictive uncertainty estimation using deep ensembles
Lakshminarayanan, B., Pritzel, A., and Blundell, C · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Optimization methods for large-scale machine learning
Bottou, L., Curtis, F. E., and Nocedal, J · 2018
Earlier work this paper cites.
Essentially no barriers in neural network energy landscape
Draxler, F., Veschgini, K., Salmhofer, M., and Hamprecht, F. A · 2018
Earlier work this paper cites.
Bilevel programming for hyperparameter optimization and meta-learning
Franceschi, L., Frasconi, P., Salzo, S., Grazzi, R., and Pontil, M · 2018
Earlier work this paper cites.
Loss surfaces, mode connectivity, and fast ensembling of dnns
Garipov, T., Izmailov, P., Podoprikhin, D., Vetrov, D. P., and Wilson, A. G · 2018
Earlier work this paper cites.
Averaging weights leads to wider optima and better generalization
Izmailov, P., Podoprikhin, D., Garipov, T., Vetrov, D. P., and Wilson, A. G · 2018
Earlier work this paper cites.
An alternative view: When does sgd escape local minima?
Kleinberg, B., Li, Y., and Yuan, Y · 2018
Earlier work this paper cites.
Canonical tensor decomposition for knowledge base completion
Lacroix, T., Usunier, N., and Obozinski, G · 2018
Cited alongside, same era.
Visualizing the loss landscape of neural nets
Li, H., Xu, Z., Taylor, G., Studer, C., and Goldstein, T · 2018
Cited alongside, same era.
Improving stability in deep reinforcement learning with weight averaging
Nikishin, E., Izmailov, P., Athiwaratkun, B., Podoprikhin, D., Garipov, T., Shvechikov, P., Vetrov, D., and Wilson, A. G · 2018
Cited alongside, same era.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S · 2018
Cited alongside, same era.
Unsupervised feature learning via non-parametric instance discrimination
Wu, Z., Xiong, Y., Yu, S. X., and Lin, D · 2018
Cited alongside, same era.
There are many consistent explanations of unlabeled data: Why you should average, 2019
What is being transferred in transfer learning?
Neyshabur, B., Sedghi, H., and Zhang, C · 2020
Later among the works it cites.
Training sensitivity in graph isomorphism network
Rahman, M. K · 2020
Later among the works it cites.
Sankar, A. R., Khasbage, Y., Vigneswaran, R., and Balasubramanian, V. N · 2020
Later among the works it cites.
Transformers: State-of-the-art natural language processing
Wolf, T., Chaumond, J., Debut, L., Sanh, V., Delangue, C., Moi, A., Cistac, P., Funtowicz, M., Davison, J., Shleifer, S., et al · 2020
Later among the works it cites.
Pyhessian: Neural networks through the lens of the hessian
Yao, Z., Gholami, A., Keutzer, K., and Mahoney, M. W · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Athiwaratkun, B., Finzi, M., Izmailov, P., and Wilson, A. G · 2019
Cited alongside, same era.
Cluster-gcn: An efficient algorithm for training deep and large graph convolutional networks
Chiang, W., Liu, X., Si, S., Li, Y., Bengio, S., and Hsieh, C · 2019
Cited alongside, same era.
Autoaugment: Learning augmentation strategies from data
Cubuk, E. D., Zoph, B., Mane, D., Vasudevan, V., and Le, Q. V · 2019
Cited alongside, same era.
HAWQ: hessian aware quantization of neural networks with mixed-precision
Dong, Z., Yao, Z., Gholami, A., Mahoney, M. W., and Keutzer, K · 2019
Cited alongside, same era.
Large scale structure of neural network loss landscapes
Fort, S. and Jastrzebski, S · 2019
Cited alongside, same era.
An investigation into neural net optimization via hessian eigenvalue density
Ghorbani, B., Krishnan, S., and Xiao, Y · 2019
Cited alongside, same era.
A closer look at deep learning heuristics: Learning rate restarts, warmup and distillation
Gotmare, A., Keskar, N. S., Xiong, C., and Socher, R · 2019
Cited alongside, same era.
Zhou, P., Feng, J., Ma, C., Xiong, C., Hoi, S. C. H., et al · 2020
Later among the works it cites.
Sharpness-aware minimization improves language model generalization, 2021
Bahri, D., Mobahi, H., and Tay, Y · 2021
Later among the works it cites.
SWAD: Domain generalization by seeking flat minima
Cha, J., Chun, S., Lee, K., Cho, H.-C., Park, S., Lee, Y., and Park, S · 2021
Later among the works it cites.
Exploring simple siamese representation learning
Chen, X. and He, K · 2021
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2021
Later among the works it cites.
Sharpness-aware minimization for efficiently improving generalization
Foret, P., Kleiner, A., Mobahi, H., and Neyshabur, B · 2021
Later among the works it cites.
Stochastic training is not necessary for generalization
Geiping, J., Goldblum, M., Pope, P. E., Moeller, M., and Goldstein, T · 2021
Later among the works it cites.
Diversity is all you need to improve bayesian model averaging
Grewal, Y. and Bui, T. D · 2021
Later among the works it cites.
Leveraging passage retrieval with generative models for open domain question answering
Izacard, G. and Grave, É · 2021
Later among the works it cites.
Causal effect inference for structured treatments
Kaddour, J., Zhu, Y., Liu, Q., Kusner, M., and Silva, R · 2021
Later among the works it cites.
Characterizing possible failure modes in physics-informed neural networks
Krishnapriyan, A. S., Gholami, A., Zhe, S., Kirby, R. M., and Mahoney, M. W · 2021
Later among the works it cites.
Asam: Adaptive sharpness-aware minimization for scale-invariant learning of deep neural networks
Kwon, J., Kim, J., Park, H., and Choi, I. K · 2021
Later among the works it cites.
Relative flatness and generalization
Petzka, H., Kamp, M., Adilova, L., Sminchisescu, C., and Boley, M · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I · 2021
Later among the works it cites.
Relating adversarially robust generalization to flat minima
Stutz, D., Hein, M., and Schiele, B · 2021
Later among the works it cites.
Mlp-mixer: An all-mlp architecture for vision
Tolstikhin, I. O., Houlsby, N., Kolesnikov, A., Beyer, L., Zhai, X., Unterthiner, T., Yung, J., Steiner, A., Keysers, D., Uszkoreit, J., Lucic, M., and Dosovitskiy, A · 2021
Later among the works it cites.
Going deeper with image transformers
Touvron, H., Cord, M., Sablayrolles, A., Synnaeve, G., and Jégou, H · 2021
Later among the works it cites.
Taxonomizing local versus global structure in neural network loss landscapes
Yang, Y., Hodgkinson, L., Theisen, R., Zou, J., Gonzalez, J. E., Ramchandran, K., and Mahoney, M. W · 2021
Later among the works it cites.
Barlow twins: Self-supervised learning via redundancy reduction
Zbontar, J., Jing, L., Misra, I., LeCun, Y., and Deny, S · 2021
Later among the works it cites.
Towards understanding sharpness-aware minimization
Andriushchenko, M. and Flammarion, N · 2022
Closest in time.
Low-pass filtering sgd for recovering flat optima in the deep learning optimization landscape, 2022
Bisla, D., Wang, J., and Choromanska, A · 2022
Closest in time.
Stochastic weight averaging revisited
Guo, H., Jin, J., and Liu, B · 2022
Closest in time.
Stop wasting my time! saving days of imagenet and bert training with latest weight averaging
Kaddour, J · 2022
Closest in time.
Causal machine learning: A survey and open problems
Kaddour, J., Lynch, A., Liu, Q., Kusner, M. J., and Silva, R · 2022
Closest in time.
Hierarchical text-conditional image generation with clip latents
Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M · 2022
Closest in time.
Adversarial weight perturbation improves generalization in graph neural networks, 2022
Wu, Y., Bojchevski, A., and Huang, H · 2022
Closest in time.
Zhao, Y., Zhang, H., and Hu, X · 2022
Closest in time.
Surrogate gap minimization improves sharpness-aware training
Zhuang, J., Gong, B., Yuan, L., Cui, Y., Adam, H., Dvornek, N. C., sekhar tatikonda, s Duncan, J., and Liu, T · 2022
Closest in time.