Fetching the paper…
Reading the bibliography…
In this work, we analyze the role of the network architecture in shaping the inductive bias of deep classifiers.
T. M. Mitchell, “The Need for Biases in Learning Generalizations,” tech. rep., Rutgers University, 1980
1980
Earlier work this paper cites.
USA: Johns Hopkins University Press, 1996
G. H. Golub and C. F. Van Loan, Matrix Computations (3rd Ed.) · 1996
Earlier work this paper cites.
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,” Proceedings of the IEEE
1998
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L. J. Li, L. Kai, and F. F. Li, “ImageNet: A large-scale hierarchical image database,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
2009
Earlier work this paper cites.
J. S. Bergstra, R. Bardenet, Y. Bengio, and B. Kégl, “Algorithms for Hyper-Parameter Optimization,” in Advances in Neural Information Processing Systems (NeurIPS)
2011
Earlier work this paper cites.
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis, “Human-level control through deep reinforcement learning,” Nature
2015
Earlier work this paper cites.
K. Simonyan and A. Zisserman, “Very deep convolutional networks for Large-Scale Image Recognition,” in International Conference on Learning Representations, (ICLR)
2015
Earlier work this paper cites.
S. Ioffe and C. Szegedy, “Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift,” in Proceedings of the 32nd International Conference on Machine Learning (ICML)
2015
Earlier work this paper cites.
S. Mallat, “Understanding deep convolutional networks,” Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
2016
Earlier work this paper cites.
M. Defferrard, X. Bresson, and P. Vandergheynst, “Convolutional Neural Networks on Graphs with Fast Localized Spectral Filtering,” in Advances in Neural Information Processing Systems (NeurIPS)
2016
Earlier work this paper cites.
B. Poole, S. Lahiri, M. Raghu, J. Sohl-Dickstein, and S. Ganguli, “Exponential expressivity in deep neural networks through transient chaos,” in Advances in Neural Information Processing Systems 29 (NeurIPS)
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
G. Huang, Z. Liu, L. Van Der Maaten, and K. Q. Weinberger, “Densely connected convolutional networks,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
2017
Earlier work this paper cites.
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals, “Understanding deep learning requires rethinking generalization,” in International Conference on Learning Representations (ICLR)
2017
Cited alongside, same era.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin, “Attention is All you Need,” in Advances in Neural Information Processing Systems (NeurIPS)
2017
Cited alongside, same era.
S. S. Schoenholz, J. Gilmer, S. Ganguli, and J. Sohl-Dickstein, “Deep information propagation,” in International Conference on Learning Representations (ICLR)
2017
Cited alongside, same era.
B. Zoph and Q. V. Le, “Neural architecture search with reinforcement learning,” in International Conference on Learning Representations, (ICLR)
2017
Cited alongside, same era.
Pearson, 4 edition ed., 2017
R. C. Gonzalez and R. E. Woods, Digital Image Processing · 2017
A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, and S. R. Bowman, “GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding,” in International Conference on Learning Representations (ICLR)
2019
Later among the works it cites.
N. Rahaman, A. Baratin, D. Arpit, F. Draxler, M. Lin, F. Hamprecht, Y. Bengio, and A. Courville, “On the Spectral Bias of Neural Networks,” in Proceedings of the 36th International Conference on Machine Learning (ICML)
2019
Later among the works it cites.
A. Bietti and J. Mairal, “On the Inductive Bias of Neural Tangent Kernels,” in Advances in Neural Information Processing Systems (NeurIPS)
2019
Later among the works it cites.
E. Abbe and C. Sandon, “Provable limitations of deep learning,” arXiv:1812.06369
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
2018
Cited alongside, same era.
D. Soudry, E. Hoffer, M. S. Nacson, S. Gunasekar, and N. Srebro, “The Implicit Bias of Gradient Descent on Separable Data,” in International Conference on Learning Representations (ICLR)
2018
Cited alongside, same era.
S. Gunasekar, J. Lee, D. Soudry, and N. Srebro, “Characterizing Implicit Bias in Terms of Optimization Geometry,” in Proceedings of the 35th International Conference on Machine Learning (ICML)
2018
Cited alongside, same era.
P. Chaudhari and S. Soatto, “Stochastic gradient descent performs variational inference, converges to limit cycles for deep networks,” in International Conference on Learning Representations (ICLR)
2018
Cited alongside, same era.
M. Nye and A. Saxe, “Are Efficient Deep Representations Learnable?,” in International Conference on Learning Representations, (ICLR)
2018
Cited alongside, same era.
A. Brutzkus, A. Globerson, E. Malach, and S. Shalev-Shwartz, “SGD Learns Over-parameterized Networks that Provably Generalize on Linearly Separable Data,” in International Conference on Learning Representations (ICLR)
2018
Cited alongside, same era.
J. Pennington, S. Schoenholz, and S. Ganguli, “The emergence of spectral universality in deep networks,” in International Conference on Artificial Intelligence and Statistics (AISTATS)
2018
Cited alongside, same era.
S. d'Ascoli, L. Sagun, G. Biroli, and J. Bruna, “Finding the Needle in the Haystack with Convolutions: on the benefits of architectural bias,” in Advances in Neural Information Processing Systems (NeurIPS)
2019
Later among the works it cites.
R. Zhang, “Making Convolutional Networks Shift-Invariant Again,” in Proceedings of the 36th International Conference on Machine Learning (ICML)
2019
Later among the works it cites.
D. Yin, R. G. Lopes, J. Shlens, E. D. Cubuk, and J. Gilmer, “A Fourier Perspective on Model Robustness in Computer Vision,” in Advances in Neural Information Processing Systems (NeurIPS)
2019
Later among the works it cites.
H. Wang, X. Wu, Z. Huang, and E. P. Xing, “High Frequency Component Helps Explain the Generalization of Convolutional Neural Networks,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
2019
Later among the works it cites.
B. Ghorbani, S. Krishnan, and Y. Xiao, “An Investigation into Neural Net Optimization via Hessian Eigenvalue Density,” in Proceedings of the 36th International Conference on Machine Learning (ICML)
2019
Later among the works it cites.
2019
Later among the works it cites.
J.-B. Cordonnier, A. Loukas, and M. Jaggi, “On the relationship between self-attention and convolutional layers,” in International Conference on Learning Representations (ICLR)
2020
Closest in time.
G. Ortiz-Jimenez, A. Modas, S.-M. Moosavi-Dezfooli, and P. Frossard, “Hold me tight! Influence of discriminative features on deep network boundaries,” in Advances in Neural Information Processing Systems (NeurIPS)
2020
Closest in time.
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei, “Language Models are Few-Shot Learners,” in Advances in Neural Information Processing Systems (NeurIPS)
2020
Closest in time.