Fetching the paper…
Reading the bibliography…
In this paper, we investigate the Rademacher complexity of deep sparse neural networks, where each neuron receives a small number of inputs.
Arora, S., Du, S. S., Hu, W., Li, Z., and Wang, R · 1901
Earlier work this paper cites.
Alphastar: An evolutionary computation perspective, 2019
Arulkumaran, K., Cully, A., and Togelius, J · 1902
Earlier work this paper cites.
Why resnet works? residuals generalize, 2019
He, F., Liu, T., and Tao, D · 1904
Earlier work this paper cites.
Generalization bounds of stochastic gradient descent for wide and deep neural networks, 2019
Cao, Y. and Gu, Q · 1905
Earlier work this paper cites.
Data-dependent sample complexity of deep neural networks via lipschitz augmentation, 2019
Wei, C. and Ma, T · 1905
Earlier work this paper cites.
Approximation by superpositions of a sigmoidal function
Cybenko, G · 1989
Earlier work this paper cites.
Approximation capabilities of multilayer feedforward networks
Hornik, K · 1991
Earlier work this paper cites.
Probability in Banach spaces
Ledoux, M. and Talagrand, M · 1991
Earlier work this paper cites.
Neural networks for optimal approximation of smooth and analytic functions
Mhaskar, H. N · 1996
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P · 1998
Earlier work this paper cites.
Statistical Learning Theory
Vapnik, V. N · 1998
Earlier work this paper cites.
On best approximation by ridge functions
Maiorov, V · 1999
Earlier work this paper cites.
Lower bounds for approximation by mlp neural networks
Maiorov, V. and Pinkus, A · 1999
Earlier work this paper cites.
On the approximation of functional classes equipped with a uniform measure using ridge functions
Maiorov, V., Meir, R., and Ratsaby, J · 1999
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results
Bartlett, P. L. and Mendelson, S · 2001
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results
Bartlett, P. L. and Mendelson, S · 2002
Earlier work this paper cites.
MNIST handwritten digit database, 2010
LeCun, Y. and Cortes, C · 2010
Earlier work this paper cites.
The deep bootstrap framework: Good online learners are good offline generalizers, 2020
Nakkiran, P., Neyshabur, B., and Sedghi, H · 2010
Earlier work this paper cites.
Concentration Inequalities - A Nonasymptotic Theory of Independence
Boucheron, S., Lugosi, G., and Massart, P · 2013
Earlier work this paper cites.
Understanding Machine Learning: From Theory to Algorithms
Shalev-Shwartz, S. and Ben-David, S · 2014
Cited alongside, same era.
Norm-based capacity control in neural networks
Neyshabur, B., Tomioka, R., and Srebro, N · 2015
Cited alongside, same era.
Deep Residual Learning for Image Recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Cited alongside, same era.
Sgdr: Stochastic gradient descent with warm restarts, 2016
Loshchilov, I. and Hutter, F · 2016
Cited alongside, same era.
Salimans, T. and Kingma, D. P · 2016
Cited alongside, same era.
Mastering the game of Go with deep neural networks and tree search
Generalization bounds for neural networks via approximate description length
Daniely, A. and Granot, E · 2019
Later among the works it cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Later among the works it cites.
Deterministic PAC-bayesian generalization bounds for deep networks via generalizing noise-resilience
Nagarajan, V. and Kolter, Z · 2019
Later among the works it cites.
Non-vacuous generalization bounds at the imagenet scale: a PAC-bayesian compression approach
Zhou, W., Veitch, V., Austern, M., Adams, R. P., and Orbanz, P · 2019
Later among the works it cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., van den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., Dieleman, S., Grewe, D., Nham, J., Kalchbrenner, N., Sutskever, I., Lillicrap, T., Leach, M., Kavukcuoglu, K., Graepel, T., and Hassabis, D · 2016
Cited alongside, same era.
Spectrally-normalized margin bounds for neural networks
Bartlett, P. L., Foster, D. J., and Telgarsky, M · 2017
Cited alongside, same era.
Size-Independent Sample Complexity of Neural Networks
Golowich, N., Rakhlin, A., and Shamir, O · 2017
Cited alongside, same era.
Approximating continuous functions by relu nets of minimal width, 2017
Hanin, B. and Sellke, M · 2017
Cited alongside, same era.
Nearly-tight vc-dimension bounds for piecewise linear neural networks
Harvey, N., Liaw, C., and Mehrabian, A · 2017
Cited alongside, same era.
When and why are deep networks better than shallow ones?
Mhaskar, H., Liao, Q., and Poggio, T · 2017
Cited alongside, same era.
Exploring generalization in deep learning
Neyshabur, B., Bhojanapalli, S., Mcallester, D., and Srebro, N · 2017
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2020
Later among the works it cites.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Later among the works it cites.
Fantastic generalization measures and where to find them
Jiang, Y., Neyshabur, B., Mobahi, H., Krishnan, D., and Bengio, S · 2020
Later among the works it cites.
Generalization bounds for deep convolutional neural networks
Long, P. M. and Sedghi, H · 2020
Later among the works it cites.
Theoretical issues in deep networks
Poggio, T., Banburski, A., and Liao, Q · 2020
Later among the works it cites.
Fit without fear: remarkable mathematical phenomena of deep learning through the prism of interpolation
Belkin, M · 2021
Later among the works it cites.
Evaluating large language models trained on code, 2021
Chen, M., Tworek, J., Jun, H., Yuan, Q., de Oliveira Pinto, H. P., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., Ray, A., Puri, R., Krueger, G., Petrov, M., Khlaaf, H., Sastry, G., Mishkin, P., Chan, B., Gray, S., Ryder, N., Pavlov, M., Power, A., Kaiser, L., Bavarian, M., Winter, C., Tillet, P., Such, F. P., Cummings, D., Plappert, M., Chantzis, F., Barnes, E., Herbert-Voss, A., Guss, W. H., Nichol, A., Paino, A., Tezak, N., Tang, J., Babuschkin, I., Balaji, S., Jain, S., Saunders, W., Hesse, C., Carr, A. N., Leike, J., Achiam, J., Misra, V., Morikawa, E., Radford, A., Knight, M., Brundage, M., Murati, M., Mayer, K., Welinder, P., McGrew, B., Amodei, D., McCandlish, S., Sutskever, I., and Zaremba, W · 2021
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2021
Later among the works it cites.
Norm-based generalisation bounds for deep multi-class convolutional neural networks
Ledent, A., Mustafa, W., Lei, Y., and Kloft, M · 2021
Later among the works it cites.
Stability & generalisation of gradient descent for shallow neural networks without the neural tangent kernel
Richards, D. and Kuzborskij, I · 2021
Later among the works it cites.
Scaling vision transformers, 2021
Zhai, X., Kolesnikov, A., Houlsby, N., and Beyer, L · 2021
Later among the works it cites.
PAC-bayes compression bounds so tight that they can explain generalization
Lotfi, S., Finzi, M. A., Kapoor, S., Potapczynski, A., Goldblum, M., and Wilson, A. G · 2022
Later among the works it cites.
Foundations of deep learning: Compositional sparsity of computable functions
Poggio, T · 2022
Later among the works it cites.