Fetching the paper…
Reading the bibliography…
We identify and formalize a fundamental gradient descent phenomenon resulting in a learning proclivity in over-parameterized neural networks.
Arora, S., Du, S. S., Hu, W., Li, Z., and Wang, R. (2019c) · 1901
Earlier work this paper cites.
Cross-entropy loss and low-rank features have responsibility for adversarial examples
Nar, K., Ocal, O., Sastry, S. S., and Ramchandran, K. (2019) · 1901
Earlier work this paper cites.
Frequency principle: Fourier analysis sheds light on deep neural networks
Xu, Z.-Q. J., Zhang, Y., Luo, T., Xiao, Y., and Ma, Z. (2019a) · 1901
Earlier work this paper cites.
Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference
McCoy, R. T., Pavlick, E., and Linzen, T. (2019) · 1902
Earlier work this paper cites.
Benchmarking neural network robustness to common corruptions and perturbations
Hendrycks, D. and Dietterich, T. (2019) · 1903
Earlier work this paper cites.
Learning robust representations by projecting superficial statistics out
Wang, H., He, Z., Lipton, Z. C., and Xing, E. P. (2019) · 1903
Earlier work this paper cites.
Approximating cnns with bag-of-local-features models works surprisingly well on imagenet
Brendel, W. and Bethge, M. (2019) · 1904
Earlier work this paper cites.
Sgd on neural networks learns functions of increasing complexity
Nakkiran, P., Kaplun, G., Kalimeris, D., Yang, T., Edelman, B. L., Zhang, F., and Barak, B. (2019) · 1905
Earlier work this paper cites.
Generalization guarantees for neural networks via harnessing the low-rank structure of the jacobian
Oymak, S., Fabian, Z., Li, M., and Soltanolkotabi, M. (2019) · 1906
Earlier work this paper cites.
Arjovsky, M., Bottou, L., Gulrajani, I., and Lopez-Paz, D. (2019) · 1907
Earlier work this paper cites.
Probing neural network comprehension of natural language arguments
Niven, T. and Kao, H.-Y. (2019) · 1907
Earlier work this paper cites.
A fine-grained spectral perspective on neural networks
Yang, G. and Salman, H. (2019) · 1907
Earlier work this paper cites.
The generalization error of random features regression: Precise asymptotics and double descent curve
Mei, S. and Montanari, A. (2019) · 1908
Earlier work this paper cites.
Dynamics of deep neural networks and neural tangent hierarchy
Huang, J. and Yau, H.-T. (2019) · 1909
Earlier work this paper cites.
Connections between support vector machines, wasserstein distance and gradient-penalty gans
Jolicoeur-Martineau, A. and Mitliagkas, I. (2019) · 1910
Earlier work this paper cites.
Persistency of excitation for robustness of neural networks
Nar, K. and Sastry, S. S. (2019) · 1911
Earlier work this paper cites.
Clever Hans:(the horse of Mr. Von Osten.) a contribution to experimental animal and human psychology
Pfungst, O. (1911) · 1911
Earlier work this paper cites.
Sagawa, S., Koh, P. W., Hashimoto, T. B., and Liang, P. (2019) · 1911
Earlier work this paper cites.
Towards understanding the spectral bias of deep learning
Cao, Y., Fang, Z., Wu, Y., Zhou, D.-X., and Gu, Q. (2019) · 1912
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R. (2014) · 1958
Earlier work this paper cites.
Geometrical and statistical properties of systems of linear inequalities with applications in pattern recognition
Cover, T. M. (1965) · 1965
Earlier work this paper cites.
On the existence of maximum likelihood estimates in logistic regression models
Albert, A. and Anderson, J. A. (1984) · 1984
Earlier work this paper cites.
A simple weight decay can improve generalization
Krogh, A. and Hertz, J. A. (1992) · 1992
Earlier work this paper cites.
The lack of a priori distinctions between learning algorithms
Wolpert, D. H. (1996) · 1996
Earlier work this paper cites.
Statistical learning theory wiley
Vapnik, V. and Vapnik, V. (1998) · 1998
Earlier work this paper cites.
Probabilistic kernel regression models
Jaakkola, T. S. and Haussler, D. (1999) · 1999
Earlier work this paper cites.
Uniqueness of the svm solution
Burges, C. J. and Crisp, D. J. (2000) · 2000
Earlier work this paper cites.
Invariant risk minimization games
Ahuja, K., Shanmugam, K., Varshney, K., and Dhurandhar, A. (2020b) · 2002
Earlier work this paper cites.
A generalized neural tangent kernel analysis for two-layer neural networks
Chen, Z., Cao, Y., Gu, Q., and Zhang, T. (2020) · 2002
Earlier work this paper cites.
Out-of-distribution generalization via risk extrapolation (rex)
Krueger, D., Caballero, E., Jacobsen, J.-H., Zhang, A., Binas, J., Priol, R. L., and Courville, A. (2020) · 2003
Earlier work this paper cites.
Shortcut learning in deep neural networks
Geirhos, R., Jacobsen, J.-H., Michaelis, C., Zemel, R., Brendel, W., Bethge, M., and Wichmann, F. A. (2020) · 2004
Earlier work this paper cites.
Logistic regression and boosting for labeled bags of instances
Xu, X. and Frank, E. (2004) · 2004
Earlier work this paper cites.
What shapes feature representations? exploring datasets, architectures, and training
Hermann, K. L. and Lampinen, A. K. (2020) · 2006
Earlier work this paper cites.
Evaluation of neural architectures trained with square loss vs cross-entropy in classification tasks
Hui, L. and Belkin, M. (2020) · 2006
Earlier work this paper cites.
The pitfalls of simplicity bias in neural networks
Shah, H., Tamuly, K., Raghunathan, A., Jain, P., and Netrapalli, P. (2020) · 2006
Earlier work this paper cites.
Learning from failure: Training debiased classifier from biased classifier
Nam, J., Cha, H., Ahn, S., Lee, J., and Shin, J. (2020) · 2007
Cited alongside, same era.
Implicit regularization in deep learning: A view from function space
Baratin, A., George, T., Laurent, C., Hjelm, R. D., Lajoie, G., Vincent, P., and Lacoste-Julien, S. (2020) · 2008
Cited alongside, same era.
A dual coordinate descent method for large-scale linear svm
Hsieh, C.-J., Chang, K.-W., Lin, C.-J., Keerthi, S. S., and Sundararajan, S. (2008) · 2008
Cited alongside, same era.
Learning multiple layers of features from tiny images
Krizhevsky, A., Hinton, G., et al. (2009) · 2009
Cited alongside, same era.
Learning explanations that are hard to vary
Parascandolo, G., Neitz, A., Orvieto, A., Gresele, L., and Schölkopf, B. (2020) · 2009
Recognition in terra incognita
Beery, S., Van Horn, G., and Perona, P. (2018) · 2018
Later among the works it cites.
A note on lazy training in supervised differentiable programming
Chizat, L. and Bach, F. (2018) · 2018
Later among the works it cites.
On the learning dynamics of deep neural networks
Combes, R. T. d., Pezeshki, M., Shabanian, S., Courville, A., and Bengio, Y. (2018) · 2018
Later among the works it cites.
Implicit bias of gradient descent on linear convolutional networks
Gunasekar, S., Lee, J. D., Soudry, D., and Srebro, N. (2018) · 2018
Later among the works it cites.
Annotation artifacts in natural language inference data
Gururangan, S., Swayamdipta, S., Levy, O., Schwartz, R., Bowman, S. R., and Smith, N. A. (2018) · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Linear regression games: Convergence guarantees to approximate out-of-distribution solutions
Ahuja, K., Shanmugam, K., and Dhurandhar, A. (2020a) · 2010
Cited alongside, same era.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Saxe, A. M., McClelland, J. L., and Ganguli, S. (2013) · 2013
Cited alongside, same era.
Intriguing properties of neural networks
Szegedy, C., Zaremba, W., Sutskever, I., Bruna, J., Erhan, D., Goodfellow, I., and Fergus, R. (2013) · 2013
Cited alongside, same era.
Explaining and harnessing adversarial examples
Goodfellow, I. J., Shlens, J., and Szegedy, C. (2014) · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J. (2014) · 2014
Cited alongside, same era.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Neyshabur, B., Tomioka, R., and Srebro, N. (2014) · 2014
Cited alongside, same era.
Understanding machine learning: From theory to algorithms
Shalev-Shwartz, S. and Ben-David, S. (2014) · 2014
Cited alongside, same era.
Black-box adversarial attacks with limited queries and information
Ilyas, A., Engstrom, L., Athalye, A., and Lin, J. (2018) · 2018
Later among the works it cites.
Excessive invariance causes adversarial vulnerability
Jacobsen, J.-H., Behrmann, J., Zemel, R., and Bethge, M. (2018) · 2018
Later among the works it cites.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C. (2018) · 2018
Later among the works it cites.
An analytic theory of generalization dynamics and transfer learning in deep linear networks
Lampinen, A. K. and Ganguli, S. (2018) · 2018
Later among the works it cites.
A simple unified framework for detecting out-of-distribution samples and adversarial attacks
Lee, K., Lee, K., Lee, H., and Shin, J. (2018) · 2018
Later among the works it cites.
Implicit regularization in nonconvex statistical estimation: Gradient descent converges linearly for phase retrieval and matrix completion
Ma, C., Wang, K., Chi, Y., and Chen, Y. (2018) · 2018
Later among the works it cites.
Rosenfeld, A., Zemel, R., and Tsotsos, J. K. (2018) · 2018
Later among the works it cites.
The implicit bias of gradient descent on separable data
Soudry, D., Hoffer, E., Nacson, M. S., Gunasekar, S., and Srebro, N. (2018) · 2018
Later among the works it cites.
Deep learning generalizes because the parameter-function map is biased towards simple functions
Valle-Pérez, G., Camargo, C. Q., and Louis, A. A. (2018) · 2018
Later among the works it cites.
Confounding variables can degrade generalization performance of radiological deep learning models
Zech, J. R., Badgeley, M. A., Liu, M., Costa, A. B., Titano, J. J., and Oermann, E. K. (2018) · 2018
Later among the works it cites.
On the inductive bias of neural tangent kernels
Bietti, A. and Mairal, J. (2019) · 2019
Later among the works it cites.
Generalization bounds of stochastic gradient descent for wide and deep neural networks
Cao, Y. and Gu, Q. (2019) · 2019
Later among the works it cites.
Implicit regularization of discrete gradient dynamics in linear neural networks
Gidel, G., Bach, F., and Lacoste-Julien, S. (2019) · 2019
Later among the works it cites.
Dynamics of stochastic gradient descent for two-layer neural networks in the teacher-student setup
Goldt, S., Advani, M., Saxe, A. M., Krzakala, F., and Zdeborová, L. (2019) · 2019
Later among the works it cites.
Adversarial examples are not bugs, they are features
Ilyas, A., Santurkar, S., Tsipras, D., Engstrom, L., Tran, B., and Madry, A. (2019) · 2019
Later among the works it cites.
The implicit bias of gradient descent on nonseparable data
Ji, Z. and Telgarsky, M. (2019) · 2019
Later among the works it cites.
Unmasking clever hans predictors and assessing what machines really learn
Lapuschkin, S., Wäldchen, S., Binder, A., Montavon, G., Samek, W., and Müller, K.-R. (2019) · 2019
Later among the works it cites.
Wide neural networks of any depth evolve as linear models under gradient descent
Lee, J., Xiao, L., Schoenholz, S., Bahri, Y., Novak, R., Sohl-Dickstein, J., and Pennington, J. (2019) · 2019
Later among the works it cites.
On the spectral bias of neural networks
Rahaman, N., Baratin, A., Arpit, D., Draxler, F., Lin, M., Hamprecht, F., Bengio, Y., and Courville, A. (2019) · 2019
Later among the works it cites.
The convergence rate of neural networks for learned functions of different frequencies
Ronen, B., Jacobs, D., Kasten, Y., and Kritchman, S. (2019) · 2019
Later among the works it cites.
A mathematical theory of semantic development in deep neural networks
Saxe, A. M., McClelland, J. L., and Ganguli, S. (2019) · 2019
Later among the works it cites.
Gradient descent for one-hidden-layer neural networks: Polynomial convergence and sq lower bounds
Vempala, S. and Wilmes, J. (2019) · 2019
Later among the works it cites.
High-dimensional dynamics of generalization error in neural networks
Advani, M. S., Saxe, A. M., and Sompolinsky, H. (2020) · 2020
Closest in time.
Nngeometry: Easy and fast fisher information matrices and neural tangent kernels in pytorch
George, T. (2020) · 2020
Closest in time.
Why do deep residual networks generalize better than deep feedforward networks?—a neural tangent kernel perspective
Huang, K., Wang, Y., Tao, M., and Zhao, T. (2020) · 2020
Closest in time.
Hidden stratification causes clinically meaningful failures in machine learning for medical imaging
Oakden-Rayner, L., Dunnmon, J., Carneiro, G., and Ré, C. (2020) · 2020
Closest in time.
Gradient descent optimizes over-parameterized deep relu networks
Zou, D., Cao, Y., Zhou, D., and Gu, Q. (2020) · 2020
Closest in time.
Machine learning for covid-19 diagnosis: Promising, but still too flawed
Roberts, M. (2021) · 2021
Closest in time.
When and why pinns fail to train: A neural tangent kernel perspective
Wang, S., Yu, X., and Perdikaris, P. (2021) · 2021
Closest in time.