Measuring the intrinsic dimension of objective landscapes
Original
Li, C., Farkhoor, H., Liu, R., and Yosinski, J · 2018
Later among the works it cites.
The full spectrum of deepnet hessians at scale: Dynamics with sgd training and sample size
Original
Papyan, V · 2018
Later among the works it cites.
Fast generalization error bound of deep learning from a kernel perspective
Suzuki, T · 2018
Later among the works it cites.
Non-vacuous generalization bounds at the imagenet scale: a pac-bayesian compression approach
Zhou, W., Veitch, V., Austern, M., Adams, R. P., and Orbanz, P · 2018
Later among the works it cites.
Benign overfitting in linear regression
Original
Bartlett, P. L., Long, P. M., Lugosi, G., and Tsigler, A · 2019
Later among the works it cites.
Pyro: Deep universal probabilistic programming
Bingham, E., Chen, J. P., Jankowiak, M., Obermeyer, F., Pradhan, N., Karaletsos, T., Singh, R., Szerlip, P., Horsfall, P., and Goodman, N. D · 2019
Later among the works it cites.
Entropy-SGD: Biasing gradient descent into wide valleys
Chaudhari, P., Choromanska, A., Soatto, S., LeCun, Y., Baldassi, C., Borgs, C., Chayes, J., Sagun, L., and Zecchina, R · 2019
Later among the works it cites.
Semi-flat minima and saddle points by embedding neural networks to overparameterization
Fukumizu, K., Yamaguch, S., Mototake, Y.-i., and Tanaka, M · 2019
Later among the works it cites.
Surprises in High-Dimensional Ridgeless Least Squares Interpolation
Original
Hastie, T., Montanari, A., Rosset, S., and Tibshirani, R. J · 2019
Later among the works it cites.
Understanding generalization through visualizations
Original
Huang, W. R., Emam, Z., Goldblum, M., Fowl, L., Terry, J. K., Huang, F., and Goldstein, T · 2019
Later among the works it cites.
Subspace inference for bayesian deep learning
Izmailov, P., Maddox, W. J., Kirichenko, P., Garipov, T., Vetrov, D., and Wilson, A. G · 2019
Later among the works it cites.
Fantastic generalization measures and where to find them
Original
Jiang, Y., Neyshabur, B., Mobahi, H., Krishnan, D., and Bengio, S · 2019
Later among the works it cites.
Wide neural networks of any depth evolve as linear models under gradient descent
Lee, J., Xiao, L., Schoenholz, S. S., Bahri, Y., Sohl-Dickstein, J., and Pennington, J · 2019
Later among the works it cites.
On Linearizing Neural Networks for Transfer Learning
Maddox, W. J., Tang, S., Moreno, P. G., Wilson, A. G., and Damianou, A · 2019
Later among the works it cites.
Detecting extrapolation with local ensembles
Original
Madras, D., Atwood, J., and D’Amour, A · 2019
Later among the works it cites.
The generalization error of random features regression: Precise asymptotics and double descent curve
Original
Mei, S. and Montanari, A · 2019
Later among the works it cites.
Harmless interpolation of noisy data in regression
Original
Muthukumar, V., Vodrahalli, K., and Sahai, A · 2019
Later among the works it cites.
Deep double descent: Where bigger models and more data hurt
Original
Nakkiran, P., Kaplun, G., Bansal, Y., Yang, T., Barak, B., and Sutskever, I · 2019
Later among the works it cites.
Composable effects for flexible and accelerated probabilistic programming in numpyro
Original
Phan, D., Pradhan, N., and Jankowiak, M · 2019
Later among the works it cites.
Bayesian deep learning and a probabilistic perspective of generalization
Original
Wilson, A. G. and Izmailov, P · 2020
Closest in time.