Fetching the paper…
Reading the bibliography…
In this work we revisit the most fundamental building block in deep learning, the multi-layer perceptron (MLP), and study the limits of its performance on vision tasks.
The perceptron: a probabilistic model for information storage and organization in the brain
Rosenblatt, F. (1958) · 1958
Earlier work this paper cites.
Cybernetic Predicting Devices
Ivakhnenko, A., Lapa, V., and ENGINEERING., P. U. L. I. S. O. E. (1965) · 1965
Earlier work this paper cites.
A theory of adaptive pattern classifiers
Amari, S. (1967) · 1967
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L. (2009) · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A. (2009) · 2009
Earlier work this paper cites.
An analysis of single-layer networks in unsupervised feature learning
Coates, A., Ng, A., and Lee, H. (2011) · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E. (2012) · 2012
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Neyshabur, B., Tomioka, R., and Srebro, N. (2014) · 2014
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Saxe, A. M., McClelland, J. L., and Ganguli, S. (2014) · 2014
Earlier work this paper cites.
Tiny imagenet visual recognition challenge
Le, Y. and Yang, X. S. (2015) · 2015
Earlier work this paper cites.
How far can we go without convolution: Improving fully-connected networks
Lin, Z., Memisevic, R., and Konda, K. R. (2015) · 2015
Earlier work this paper cites.
Layer normalization
Ba, J. L., Kiros, J. R., and Hinton, G. E. (2016) · 2016
Earlier work this paper cites.
Convolution by evolution: Differentiable pattern producing networks
Fernando, C., Banarse, D., Reynolds, M., Besse, F., Pfau, D., Jaderberg, M., Lanctot, M., and Wierstra, D. (2016) · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J. (2015) · 2016
Earlier work this paper cites.
Exponential expressivity in deep neural networks through transient chaos
Poole, B., Lahiri, S., Raghu, M., Sohl-Dickstein, J., and Ganguli, S. (2016) · 2016
Earlier work this paper cites.
Globally optimal gradient descent for a ConvNet with Gaussian inputs
Brutzkus, A. and Globerson, A. (2017) · 2017
Earlier work this paper cites.
A downsampled variant of imagenet as an alternative to the cifar datasets
Chrabaszcz, P., Loshchilov, I., and Hutter, F. (2017) · 2017
Earlier work this paper cites.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Goyal, P., Dollár, P., Girshick, R., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K. (2017) · 2017
Earlier work this paper cites.
Deep learning scaling is predictable, empirically
Hestness, J., Narang, S., Ardalani, N., Diamos, G., Jun, H., Kianinejad, H., Patwary, M. M. A., Yang, Y., and Zhou, Y. (2017) · 2017
Earlier work this paper cites.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
Hoffer, E., Hubara, I., and Soudry, D. (2017) · 2017
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P. (2017) · 2017
Earlier work this paper cites.
Convergence analysis of two-layer neural networks with relu activation
Li, Y. and Yuan, Y. (2017) · 2017
Earlier work this paper cites.
Deep information propagation
Schoenholz, S. S., Gilmer, J., Ganguli, S., and Sohl-Dickstein, J. (2017) · 2017
Earlier work this paper cites.
Do deep convolutional nets really need to be deep and convolutional?
Urban, G., Geras, K. J., Kahou, S. E., Aslan, O., Wang, S., Mohamed, A., Philipose, M., Richardson, M., and Caruana, R. (2017) · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. u., and Polosukhin, I. (2017) · 2017
Earlier work this paper cites.
Large batch training of convolutional networks
You, Y., Gitman, I., and Ginsburg, B. (2017) · 2017
Earlier work this paper cites.
On the optimization of deep networks: Implicit acceleration by overparameterization
Arora, S., Cohen, N., and Hazan, E. (2018) · 2018
Earlier work this paper cites.
Relational inductive biases, deep learning, and graph networks
Battaglia, P. W., Hamrick, J. B., Bapst, V., Sanchez-Gonzalez, A., Zambaldi, V., Malinowski, M., Tacchetti, A., Raposo, D., Santoro, A., Faulkner, R., Gulcehre, C., Song, F., Ballard, A., Gilmer, J., Dahl, G., Vaswani, A., Allen, K., Nash, C., Langston, V., Dyer, C., Heess, N., Wierstra, D., Kohli, P., Botvinick, M., Vinyals, O., Li, Y., and Pascanu, R. (2018) · 2018
Cited alongside, same era.
Implicit bias of gradient descent on linear convolutional networks
Gunasekar, S., Lee, J. D., Soudry, D., and Srebro, N. (2018) · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C. (2018) · 2018
Cited alongside, same era.
A mean field view of the landscape of two-layer neural networks
Mei, S., Montanari, A., and Nguyen, P.-M. (2018) · 2018
Cited alongside, same era.
Scalable training of artificial neural networks with adaptive sparse connectivity inspired by network science
Mocanu, D., Mocanu, E., Stone, P., Nguyen, P., Gibescu, M., and Liotta, A. (2018) · 2018
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N. (2021) · 2021
Later among the works it cites.
Pay attention to MLPs
Liu, H., Dai, Z., So, D., and Le, Q. V. (2021) · 2021
Later among the works it cites.
The generalization error of random features regression: Precise asymptotics and the double descent curve
Mei, S. and Montanari, A. (2021) · 2021
Later among the works it cites.
Nerf: Representing scenes as neural radiance fields for view synthesis
Mildenhall, B., Srinivasan, P. P., Tancik, M., Barron, J. T., Ramamoorthi, R., and Ng, R. (2021) · 2021
Later among the works it cites.
What kinds of functions do deep neural networks learn? insights from variational spline theory
Parhi, R. and Nowak, R. D. (2021) · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ai and compute
OpenAI (2018) · 2018
Cited alongside, same era.
The implicit bias of gradient descent on separable data
Soudry, D., Hoffer, E., and Srebro, N. (2018) · 2018
Cited alongside, same era.
mixup: Beyond empirical risk minimization
Zhang, H., Cisse, M., Dauphin, Y. N., and Lopez-Paz, D. (2018) · 2018
Cited alongside, same era.
Finding the needle in the haystack with convolutions: on the benefits of architectural bias
d'Ascoli, S., Sagun, L., Biroli, G., and Bruna, J. (2019) · 2019
Cited alongside, same era.
Gradient descent provably optimizes over-parameterized neural networks
Du, S. S., Zhai, X., Poczos, B., and Singh, A. (2019) · 2019
Cited alongside, same era.
Beyond human-level accuracy: Computational challenges in deep learning
Hestness, J., Ardalani, N., and Diamos, G. (2019) · 2019
Cited alongside, same era.
The role of over-parametrization in generalization of neural networks
Neyshabur, B., Li, Z., Bhojanapalli, S., LeCun, Y., and Srebro, N. (2019) · 2019
Cited alongside, same era.
Ridnik, T., Ben-Baruch, E., Noy, A., and Zelnik, L. (2021) · 2021
Later among the works it cites.
MLP-mixer: An all-MLP architecture for vision
Tolstikhin, I., Houlsby, N., Kolesnikov, A., Beyer, L., Zhai, X., Unterthiner, T., Yung, J., Steiner, A. P., Keysers, D., Uszkoreit, J., Lucic, M., and Dosovitskiy, A. (2021) · 2021
Later among the works it cites.
ResMLP: Feedforward networks for image classification with data-efficient training
Touvron, H., Bojanowski, P., Caron, M., Cord, M., El-Nouby, A., Grave, E., Izacard, G., Joulin, A., Synnaeve, G., Verbeek, J., and Jegou, H. (2021) · 2021
Later among the works it cites.
CycleMLP: A MLP-like architecture for dense prediction
Chen, S., Xie, E., GE, C., Chen, R., Liang, D., and Luo, P. (2022) · 2022
Later among the works it cites.
Hire-mlp: Vision mlp via hierarchical rearrangement
Guo, J., Tang, Y., Han, K., Chen, X., Wu, H., Xu, C., Xu, C., and Wang, Y. (2021) · 2022
Later among the works it cites.
Training compute-optimal large language models
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., de Las Casas, D., Hendricks, L. A., Welbl, J., Clark, A., Hennigan, T., Noland, E., Millican, K., van den Driessche, G., Damoc, B., Guy, A., Osindero, S., Simonyan, K., Elsen, E., Rae, J. W., Vinyals, O., and Sifre, L. (2022) · 2022
Later among the works it cites.
AS-MLP: An axial shifted MLP architecture for vision
Lian, D., Yu, Z., Sun, X., and Gao, S. (2022) · 2022
Later among the works it cites.
A convnet for the 2020s
Liu, Z., Mao, H., Wu, C.-Y., Feichtenhofer, C., Darrell, T., and Xie, S. (2022) · 2022
Later among the works it cites.
A solvable model of neural scaling laws
Maloney, A., Roberts, D. A., and Sully, J. (2022) · 2022
Later among the works it cites.
The Principles of Deep Learning Theory: An Effective Theory Approach to Understanding Neural Networks
Roberts, D. A., Yaida, S., and Hanin, B. (2022) · 2022
Later among the works it cites.
Patches are all you need?
Trockman, A. and Kolter, J. Z. (2022) · 2022
Later among the works it cites.
Metaformer is actually what you need for vision
Yu, W., Luo, M., Zhou, P., Si, C., Zhou, Y., Wang, X., Feng, J., and Yan, S. (2022) · 2022
Later among the works it cites.
Scaling vision transformers
Zhai, X., Kolesnikov, A., Houlsby, N., and Beyer, L. (2022) · 2022
Later among the works it cites.
The curious case of benign memorization
Anagnostidis, S., Bachmann, G., Noci, L., and Hofmann, T. (2023) · 2023
Closest in time.
Broken neural scaling laws
Caballero, E., Gupta, K., Rish, I., and Krueger, D. (2023) · 2023
Closest in time.
Symbolic discovery of optimization algorithms
Chen, X., Liang, C., Huang, D., Real, E., Wang, K., Liu, Y., Pham, H., Dong, X., Luong, T., Hsieh, C.-J., Lu, Y., and Le, Q. V. (2023) · 2023
Closest in time.
FFCV: Accelerating training by removing data bottlenecks
Leclerc, G., Ilyas, A., Engstrom, L., Park, S. M., Salman, H., and Madry, A. (2023) · 2023
Closest in time.
Gpt-4 technical report
OpenAI (2023) · 2023
Closest in time.
Linear neural network layers promote learning single- and multiple-index models
Parkinson, S., Ongie, G., and Willett, R. (2023) · 2023
Closest in time.
Vector-valued variation spaces and width bounds for dnns: Insights on weight decay regularization
Shenouda, J., Parhi, R., Lee, K., and Nowak, R. D. (2023) · 2023
Closest in time.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al. (2023) · 2023
Closest in time.