When does label smoothing help?
Original
Müller, R., Kornblith, S., and Hinton, G · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S · 2019
Later among the works it cites.
Exploring randomly wired neural networks for image recognition
Xie, S., Kirillov, A., Girshick, R., and He, K · 2019
Later among the works it cites.
Underspecification presents challenges for credibility in modern machine learning
Original
D’Amour, A., Heller, K., Moldovan, D., Adlam, B., Alipanahi, B., Beutel, A., Chen, C., Deaton, J., Eisenstein, J., Hoffman, M. D., et al · 2020
Later among the works it cites.
Sharpness-aware minimization for efficiently improving generalization
Original
Foret, P., Kleiner, A., Mobahi, H., and Neyshabur, B · 2020
Later among the works it cites.
Deep learning versus kernel learning: an empirical study of loss landscape geometry and the time evolution of the neural tangent kernel
Original
Fort, S., Dziugaite, G. K., Paul, M., Kharaghani, S., Roy, D. M., and Ganguli, S · 2020
Later among the works it cites.
Revisiting” qualitatively characterizing neural network optimization problems”
Original
Frankle, J · 2020
Later among the works it cites.
Linear mode connectivity and the lottery ticket hypothesis
Frankle, J., Dziugaite, G. K., Roy, D., and Carbin, M · 2020
Later among the works it cites.
Subspace inference for bayesian deep learning
Izmailov, P., Maddox, W. J., Kirichenko, P., Garipov, T., Vetrov, D., and Wilson, A. G · 2020
Later among the works it cites.
Pruning neural networks without any data by iteratively conserving synaptic flow
Original
Tanaka, H., Kunin, D., Yamins, D. L., and Ganguli, S · 2020
Later among the works it cites.
Measuring robustness to natural distribution shifts in image classification
Taori, R., Dave, A., Shankar, V., Carlini, N., Recht, B., and Schmidt, L · 2020
Later among the works it cites.
Batchensemble: an alternative approach to efficient ensemble and lifelong learning
Original
Wen, Y., Tran, D., and Ba, J · 2020
Later among the works it cites.
Cyclical stochastic gradient mcmc for bayesian deep learning
Zhang, R., Li, C., Zhang, J., Chen, C., and Wilson, A. G · 2020
Later among the works it cites.
Neural networks with late-phase weights
Oswald, J. V., Kobayashi, S., Sacramento, J., Meulemans, A., Henning, C., and Grewe, B. F · 2021
Closest in time.