Fetching the paper…
Reading the bibliography…
When training deep neural networks, a model's generalization error is often observed to follow a power scaling law dependent both on the model size and the data size.
Curvature measures
Herbert Federer · 1959
Earlier work this paper cites.
Nonlinear dimensionality reduction by locally linear embedding
Sam T Roweis and Lawrence K Saul · 2000
Earlier work this paper cites.
A global geometric framework for nonlinear dimensionality reduction
Joshua B Tenenbaum, Vin de Silva, and John C Langford · 2000
Earlier work this paper cites.
Maximum likelihood estimation of intrinsic dimension
Elizaveta Levina and Peter Bickel · 2004
Earlier work this paper cites.
Fast learning rates for plug-in classifiers
Jean-Yves Audibert and Alexandre B Tsybakov · 2007
Earlier work this paper cites.
Scikit-learn: Machine learning in Python
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay · 2011
Earlier work this paper cites.
An Introduction to Manifolds
Loring W. Tu · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Speech recognition with deep recurrent neural networks
Alex Graves, Abdel-rahman Mohamed, and Geoffrey Hinton · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael S. Bernstein, Alexander C. Berg, and Li Fei-Fei · 2014
Earlier work this paper cites.
Sphere Packings, Lattices and Groups , volume 290
John Conway and N. Sloane · 2016
Earlier work this paper cites.
Error bounds for approximations with deep relu networks
Dmitry Yarotsky · 2016
Earlier work this paper cites.
Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates
Shixiang Gu, Ethan Holly, Timothy Lillicrap, and Sergey Levine · 2017
Earlier work this paper cites.
Deep learning scaling is predictable, empirically
Joel Hestness, Sharan Narang, Newsha Ardalani, Gregory Diamos, Heewoo Jun, Hassan Kianinejad, Md Mostofa Ali Patwary, Yang Yang, and Yanqi Zhou · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Deep learning for healthcare: review, opportunities and challenges
Riccardo Miotto, Fei Wang, Shuang Wang, Xiaoqian Jiang, and Joel T Dudley · 2018
Earlier work this paper cites.
Estimating the reach of a manifold, 2019
Eddie Aamari, Jisu Kim, Frédéric Chazal, Bertrand Michel, Alessandro Rinaldo, and Larry Wasserman · 2019
Earlier work this paper cites.
Efficient approximation of deep relu networks for functions on low dimensional manifolds
Minshuo Chen, Haoming Jiang, Wenjing Liao, and Tuo Zhao · 2019
Cited alongside, same era.
Openwebtext corpus
Aaron Gokaslan, Vanya Cohen, Ellie Pavlick, and Stefanie Tellex · 2019
Cited alongside, same era.
Fast convergence rates of deep neural networks for classification, 2019
Yongdai Kim, Ilsang Ohn, and Dongha Kim · 2019
Cited alongside, same era.
Approximation and non-parametric estimation of resnet-type convolutional neural networks
Kenta Oono and Taiji Suzuki · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Cited alongside, same era.
A constructive prediction of the generalization error across scales
On deep generative models for approximation and estimation of distributions on manifolds
Biraj Dahal, Alexander Havrilla, Minshuo Chen, Tuo Zhao, and Wenjing Liao · 2022
Later among the works it cites.
Inductive biases and variable creation in self-attention mechanisms
Benjamin L Edelman, Surbhi Goel, Sham Kakade, and Cyril Zhang · 2022
Later among the works it cites.
Training compute-optimal large language models, 2022
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katie Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karen Simonyan, Erich Elsen, Jack W. Rae, Oriol Vinyals, and Laurent Sifre · 2022
Later among the works it cites.
Vision transformers provably learn spatial structure
Samy Jelassi, Michael Sander, and Yuanzhi Li · 2022
Later among the works it cites.
The stack: 3 tb of permissively licensed source code
Denis Kocetkov, Raymond Li, Loubna Ben Allal, Jia Li, Chenghao Mou, Carlos Muñoz Ferrandis, Yacine Jernite, Margaret Mitchell, Sean Hughes, Thomas Wolf, Dzmitry Bahdanau, Leandro von Werra, and Harm de Vries · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jonathan S Rosenfeld, Amir Rosenfeld, Yonatan Belinkov, and Nir Shavit · 2019
Cited alongside, same era.
Are transformers universal approximators of sequence-to-sequence functions?
Chulhee Yun, Srinadh Bhojanapalli, Ankit Singh Rawat, Sashank J. Reddi, and Sanjiv Kumar · 2019
Cited alongside, same era.
Intrinsic dimensionality explains the effectiveness of language model fine-tuning
Armen Aghajanyan, Luke Zettlemoyer, and Sonal Gupta · 2020
Cited alongside, same era.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Cited alongside, same era.
Wide neural networks of any depth evolve as linear models under gradient descent *
Jaehoon Lee, Lechao Xiao, Samuel S Schoenholz, Yasaman Bahri, Roman Novak, Jascha Sohl-Dickstein, and Jeffrey Pennington · 2020
Cited alongside, same era.
Adaptive approximation and generalization of deep neural network with intrinsic dimensionality
Ryumei Nakada and Masaaki Imaizumi · 2020
Cited alongside, same era.
Nonparametric regression using deep neural networks with ReLU activation function
Johannes Schmidt-Hieber · 2020
Cited alongside, same era.
Later among the works it cites.
Scaling laws from the data manifold dimension
Utkarsh Sharma and Jared Kaplan · 2022
Later among the works it cites.
Statistically meaningful approximation: a case study on approximating turing machines with transformers
Colin Wei, Yining Chen, and Tengyu Ma · 2022
Later among the works it cites.
An analysis of attention via the lens of exchangeability and latent variable models
Yufeng Zhang, Boyi Liu, Qi Cai, Lingxiao Wang, and Zhaoran Wang · 2022
Later among the works it cites.
Transformers as statisticians: Provable in-context learning with in-context algorithm selection, 2023
Yu Bai, Fan Chen, Huan Wang, Caiming Xiong, and Song Mei · 2023
Later among the works it cites.
Pythia: A suite for analyzing large language models across training and scaling, 2023
Stella Biderman, Hailey Schoelkopf, Quentin Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, Aviya Skowron, Lintang Sutawika, and Oskar van der Wal · 2023
Later among the works it cites.
Tinystories: How small can language models be and still speak coherent english?, 2023
Ronen Eldan and Yuanzhi Li · 2023
Later among the works it cites.
Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing
Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig · 2023
Later among the works it cites.
An intrinsic dimension perspective of transformers for sequential modeling, 2023
Zeping Min, Qian Ge, and Zhong Li · 2023
Later among the works it cites.
The shape of learning: Anisotropy and intrinsic dimensions in transformer-based models
Anton Razzhigaev, Matvey Mikhalchuk, Elizaveta Goncharova, Ivan Oseledets, Denis Dimitrov, and Andrey Kuznetsov · 2023
Later among the works it cites.
Approximation and estimation ability of transformers for sequence-to-sequence functions with infinite dimensional input
Shokichi Takakura and Taiji Suzuki · 2023
Later among the works it cites.
On statistical rates and provably efficient criteria of latent diffusion transformers (dits), 2024
Jerry Yao-Chieh Hu, Weimin Wu, Zhao Song, and Han Liu · 2024
Closest in time.
gzip predicts data-dependent scaling laws, 2024
Rohan Pandey · 2024
Closest in time.
Transformers, parallel computation, and logarithmic depth
Clayton Sanford, Daniel Hsu, and Matus Telgarsky · 2024
Closest in time.