Random gradient-free minimization of convex functions
Nesterov, Y. and Spokoiny, V. (2017) · 2017
Later among the works it cites.
Exploring generalization in deep learning
Neyshabur, B., Bhojanapalli, S., McAllester, D., and Srebro, N. (2017) · 2017
Later among the works it cites.
Information-theoretic analysis of generalization capability of learning algorithms
Xu, A. and Raginsky, M. (2017) · 2017
Later among the works it cites.
Continual learning through synaptic intelligence
Zenke, F., Poole, B., and Ganguli, S. (2017) · 2017
Later among the works it cites.
Understanding deep learning requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O. (2017) · 2017
Later among the works it cites.
Memory aware synapses: Learning what (not) to forget
Aljundi, R., Babiloni, F., Elhoseiny, M., Rohrbach, M., and Tuytelaars, T. (2018) · 2018
Later among the works it cites.
On the optimization of deep networks: Implicit acceleration by overparameterization
Arora, S., Cohen, N., and Hazan, E. (2018) · 2018
Later among the works it cites.
Co-teaching: Robust training of deep neural networks with extremely noisy labels
Han, B., Yao, Q., Yu, X., Niu, G., Xu, M., Hu, W., Tsang, I., and Sugiyama, M. (2018) · 2018
Later among the works it cites.
Fast and scalable bayesian deep learning by weight-perturbation in adam
Khan, M., Nielsen, D., Tangkaratt, V., Lin, W., Gal, Y., and Srivastava, A. (2018) · 2018
Later among the works it cites.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Li, Y. and Liang, Y. (2018) · 2018
Later among the works it cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F. (2018) · 2018
Later among the works it cites.
Mobilenetv2: Inverted residuals and linear bottlenecks
Sandler, M., Howard, A., Zhu, M., Zhmoginov, A., and Chen, L.-C. (2018) · 2018
Later among the works it cites.
Smoothout: Smoothing out sharp minima to improve generalization in deep learning
Original
Wen, W., Wang, Y., Yan, F., Xu, C., Wu, C., Chen, Y., and Li, H. (2018) · 2018
Later among the works it cites.
Improving the antinoise ability of dnns via a bio-inspired noise adaptive activation function rand softplus
Chen, Y., Mai, Y., Xiao, J., and Zhang, L. (2019) · 2019
Later among the works it cites.
Control batch size and learning rate to generalize well: Theoretical and empirical evidence
He, F., Liu, T., and Tao, D. (2019) · 2019
Later among the works it cites.
Calculating the mutual information between two spike trains
Houghton, C. (2019) · 2019
Later among the works it cites.
Effect of depth and width on local minima in deep learning
Kawaguchi, K., Huang, J., and Kaelbling, L. P. (2019) · 2019
Later among the works it cites.
Continual lifelong learning with neural networks: A review
Parisi, G. I., Kemker, R., Part, J. L., Kanan, C., and Wermter, S. (2019) · 2019
Later among the works it cites.
Toward understanding the importance of noise in training neural networks
Zhou, M., Liu, T., Li, Y., Lin, D., Zhou, E., and Zhao, T. (2019) · 2019
Later among the works it cites.
The anisotropic noise in stochastic gradient descent: Its behavior of escaping from sharp minima and regularization effects
Zhu, Z., Wu, J., Yu, B., Wu, L., and Ma, J. (2019) · 2019
Later among the works it cites.
Being bayesian, even just a bit, fixes overconfidence in relu networks
Kristiadi, A., Hein, M., and Hennig, P. (2020) · 2020
Closest in time.
A theoretical analysis of catastrophic forgetting through the ntk overlap matrix
Doan, T., Bennani, M. A., Mazoure, B., Rabusseau, G., and Alquier, P. (2021) · 2021
Closest in time.