Fetching the paper…
Reading the bibliography…
When data is plentiful, the loss achieved by well-trained neural networks scales as a power-law $L \propto N^{-\alpha}$ in the number of network parameters $N$.
Intrinsic dimension of data representations in deep neural networks, 2019, 1905.12784
Alessio Ansuini, Alessandro Laio, Jakob H. Macke, and Davide Zoccolan · 1905
Earlier work this paper cites.
Stefano Spigler, Mario Geiger, and Matthieu Wyart · 1905
Earlier work this paper cites.
Efficientnet: Rethinking model scaling for convolutional neural networks
Mingxing Tan and Quoc V. Le · 1905
Earlier work this paper cites.
A constructive prediction of the generalization error across scales, 2019, 1909.12673
Jonathan S. Rosenfeld, Amir Rosenfeld, Yonatan Belinkov, and Nir Shavit · 1909
Earlier work this paper cites.
Leveraging procedural generation to benchmark reinforcement learning, 2019, 1912.01588
Karl Cobbe, Christopher Hesse, Jacob Hilton, and John Schulman · 1912
Earlier work this paper cites.
Scaling laws for neural language models, 2020, 2001.08361
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2001
Earlier work this paper cites.
Maximum likelihood estimation of intrinsic dimension
Elizaveta Levina and Peter J Bickel · 2005
Earlier work this paper cites.
All of nonparametric statistics
Larry Wasserman · 2006
Earlier work this paper cites.
Local polynomial regression on unknown manifolds
Peter J Bickel, Bo Li, et al · 2007
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
MNIST handwritten digit database
Yann LeCun and Corinna Cortes · 2010
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Bo Wu, and Andrew Y Ng · 2011
Cited alongside, same era.
Analysis of a random forests model
GÊrard Biau · 2012
Cited alongside, same era.
API design for machine learning software: experiences from the scikit-learn project
Lars Buitinck, Gilles Louppe, Mathieu Blondel, Fabian Pedregosa, Andreas Mueller, Olivier Grisel, Vlad Niculae, Peter Prettenhofer, Alexandre Gramfort, Jaques Grobler, Robert Layton, Jake VanderPlas, Arnaud Joly, Brian Holt, and Gaël Varoquaux · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization, 2014, 1412.6980
Diederik P. Kingma and Jimmy Ba · 2014
Cited alongside, same era.
On the number of linear regions of deep neural networks
Guido F Montufar, Razvan Pascanu, Kyunghyun Cho, and Yoshua Bengio · 2014
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin · 2017
Later among the works it cites.
Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms
Han Xiao, Kashif Rasul, and Roland Vollgraf · 2017
Later among the works it cites.
Measuring the intrinsic dimension of objective landscapes, 2018, 1804.08838
Chunyuan Li, Heerad Farkhoor, Rosanne Liu, and Jason Yosinski · 2018
Later among the works it cites.
Generating wikipedia by summarizing long sequences
Peter J. Liu, Mohammad Saleh, Etienne Pot, Ben Goodrich, Ryan Sepassi, Lukasz Kaiser, and Noam Shazeer · 2018
Later among the works it cites.
Dimensionality-driven learning with noisy labels, 2018, 1806.02612
Xingjun Ma, Yisen Wang, Michael E. Houle, Shuo Zhou, Sarah M. Erfani, Shu-Tao Xia, Sudanthi Wijewickrema, and James Bailey · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015
Martín Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S. Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Ian Goodfellow, Andrew Harp, Geoffrey Irving, Michael Isard, Yangqing Jia, Rafal Jozefowicz, Lukasz Kaiser, Manjunath Kudlur, Josh Levenberg, Dandelion Mané, Rajat Monga, Sherry Moore, Derek Murray, Chris Olah, Mike Schuster, Jonathon Shlens, Benoit Steiner, Ilya Sutskever, Kunal Talwar, Paul Tucker, Vincent Vanhoucke, Vijay Vasudevan, Fernanda Viégas, Oriol Vinyals, Pete Warden, Martin Wattenberg, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng · 2015
Cited alongside, same era.
Efficient representation of low-dimensional manifolds using deep networks, 2016, 1602.04723
Ronen Basri and David Jacobs · 2016
Cited alongside, same era.
Intrinsic dimension estimation: Advances and open problems
Francesco Camastra and Antonino Staiano · 2016
Cited alongside, same era.
Estimating the intrinsic dimension of datasets by a minimal neighborhood information
Elena Facco, Maria d’Errico, Alex Rodriguez, and Alessandro Laio · 2017
Cited alongside, same era.
Deep learning scaling is predictable, empirically, 2017, 1712.00409
Joel Hestness, Sharan Narang, Newsha Ardalani, Gregory Diamos, Heewoo Jun, Hassan Kianinejad, Md. Mostofa Ali Patwary, Yang Yang, and Yanqi Zhou · 2017
Cited alongside, same era.
Later among the works it cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever · 2018
Later among the works it cites.
Beyond human-level accuracy: computational challenges in deep learning
Joel Hestness, Newsha Ardalani, and Gregory Diamos · 2019
Later among the works it cites.
Adversarial examples are not bugs, they are features
Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Logan Engstrom, Brandon Tran, and Aleksander Madry · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Later among the works it cites.
Zoom in: An introduction to circuits
Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov, and Shan Carter · 2020
Closest in time.