Fetching the paper…
Reading the bibliography…
Grokking, or delayed generalization, is a phenomenon where generalization in a deep neural network (DNN) occurs long after achieving near zero training error.
Robustness and generalization
Xu, H. and Mannor, S · 2012
Earlier work this paper cites.
On the number of linear regions of deep neural networks
Montufar, G. F., Pascanu, R., Cho, K., and Bengio, Y · 2014
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
Madry, A., Makelov, A., Schmidt, L., Tsipras, D., and Vladu, A · 2017
Earlier work this paper cites.
On the expressive power of deep neural networks
Raghu, M., Poole, B., Kleinberg, J., Ganguli, S., and Dickstein, J. S · 2017
Earlier work this paper cites.
Handbook of discrete and computational geometry
Toth, C. D., O’Rourke, J., and Goodman, J. E · 2017
Earlier work this paper cites.
A spline theory of deep networks
Balestriero, R. and Baraniuk, R · 2018
Earlier work this paper cites.
Sensitivity and generalization in neural networks: an empirical study
Novak, R., Bahri, Y., Abolafia, D. A., Pennington, J., and Sohl-Dickstein, J · 2018
Earlier work this paper cites.
Robustness may be at odds with accuracy
Tsipras, D., Santurkar, S., Engstrom, L., Turner, A., and Madry, A · 2018
Earlier work this paper cites.
Nearly-tight vc-dimension and pseudodimension bounds for piecewise linear neural networks
Bartlett, P. L., Harvey, N., Liaw, C., and Mehrabian, A · 2019
Earlier work this paper cites.
Complexity of linear regions in deep networks
Hanin, B. and Rolnick, D · 2019
Earlier work this paper cites.
Adversarial examples are not bugs, they are features
Ilyas, A., Santurkar, S., Tsipras, D., Engstrom, L., Tran, B., and Madry, A · 2019
Earlier work this paper cites.
Implicit regularization in over-parameterized neural networks
Kubo, M., Banno, R., Manabe, H., and Minoji, M · 2019
Cited alongside, same era.
Adversarial robustness through local linearization
Qin, C., Martens, J., Gowal, S., Krishnan, D., Dvijotham, K., Fawzi, A., De, S., Stanforth, R., and Kohli, P · 2019
Cited alongside, same era.
Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacks
Croce, F. and Hein, M · 2020
Cited alongside, same era.
Dropout vs. batch normalization: an empirical study of their impact to deep learning
Garbin, C., Zhu, X., and Marques, O · 2020
Cited alongside, same era.
Zoom in: An introduction to circuits
Olah, C., Cammarata, N., Schubert, L., Goh, G., Petrov, M., and Carter, S · 2020
Cited alongside, same era.
Polarity sampling: Quality and diversity control of pre-trained generative networks via singular values
Humayun, A. I., Balestriero, R., and Baraniuk, R · 2022
Later among the works it cites.
Test sample accuracy scales with training sample density in neural networks
Ji, X., Pascanu, R., Hjelm, R. D., Lakshminarayanan, B., and Vedaldi, A · 2022
Later among the works it cites.
Why robust generalization in deep learning is difficult: Perspective of expressive power
Li, B., Jin, J., Zhong, H., Hopcroft, J., and Wang, L · 2022
Later among the works it cites.
Omnigrok: Grokking beyond algorithmic data
Liu, Z., Michaud, E. J., and Tegmark, M · 2022
Later among the works it cites.
Grokking: Generalization beyond overfitting on small algorithmic datasets
Power, A., Burda, Y., Edwards, H., Babuschkin, I., and Misra, V · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Papyan, V., Han, X., and Donoho, D. L · 2020
Cited alongside, same era.
Comparing wiener complexity with eccentric complexity
Xu, K., Ilić, A., Iršič, V., Klavžar, S., and Li, H · 2021
Cited alongside, same era.
Max-affine spline insights into deep network pruning
You, H., Balestriero, R., Lu, Z., Kou, Y., Shi, H., Zhang, S., Wu, S., Lin, Y., and Baraniuk, R · 2021
Cited alongside, same era.
Towards understanding sharpness-aware minimization
Andriushchenko, M. and Flammarion, N · 2022
Cited alongside, same era.
Balestriero, R. and Baraniuk, R. G · 2022
Cited alongside, same era.
Hidden progress in deep learning: Sgd learns parities near the computational limit
Barak, B., Edelman, B., Goel, S., Kakade, S., Malach, E., and Zhang, C · 2022
Cited alongside, same era.
Are all linear regions created equal?
Gamba, M., Chmielewski-Anders, A., Sullivan, J., Azizpour, H., and Bjorkman, M · 2022
Cited alongside, same era.
Police: Provably optimal linear constraint enforcement for deep neural networks
Balestriero, R. and LeCun, Y · 2023
Later among the works it cites.
Provable instance specific robustness via linear constraints
Humayun, A. I., Casco-Rodriguez, J., Balestriero, R., and Baraniuk, R · 2023
Later among the works it cites.
Progress measures for grokking via mechanistic interpretability
Nanda, N., Chan, L., Lieberum, T., Smith, J., and Steinhardt, J · 2023
Later among the works it cites.
A blessing of dimensionality in membership inference through regularization
Tan, J., LeJeune, D., Mason, B., Javadi, H., and Baraniuk, R. G · 2023
Later among the works it cites.
Explaining grokking through circuit efficiency
Varma, V., Shah, R., Kenton, Z., Kramár, J., and Kumar, R · 2023
Later among the works it cites.
Benign overfitting and grokking in relu networks for xor cluster data
Xu, Z., Wang, Y., Frei, S., Vardi, G., and Hu, W · 2023
Later among the works it cites.