Fetching the paper…
Reading the bibliography…
The effectiveness of machine learning algorithms arises from being able to extract useful features from large amounts of data.
Greg Yang · 1902
Earlier work this paper cites.
The intrinsic dimensionality of signal collections
R. Bennett · 1969
Earlier work this paper cites.
Priors for infinite networks (tech. rep. no. crg-tr-94-1)
Radford M. Neal · 1994
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Mnist handwritten digit database
Yann LeCun, Corinna Cortes, and CJ Burges · 2010
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Bo Wu, and Andrew Y Ng · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Intrinsic dimension estimation: Advances and open problems
Francesco Camastra and Antonino Staiano · 2015
Earlier work this paper cites.
What makes imagenet good for transfer learning?
Minyoung Huh, Pulkit Agrawal, and Alexei A Efros · 2016
Earlier work this paper cites.
Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms, 2017
Han Xiao, Kashif Rasul, and Roland Vollgraf · 2017
Earlier work this paper cites.
Estimating the intrinsic dimension of datasets by a minimal neighborhood information
Elena Facco, Maria d’Errico, Alex Rodriguez, and Alessandro Laio · 2017
Earlier work this paper cites.
Deep learning scaling is predictable, empirically
Joel Hestness, Sharan Narang, Newsha Ardalani, Gregory F. Diamos, Heewoo Jun, Hassan Kianinejad, Md. Mostofa Ali Patwary, Yang Yang, and Yanqi Zhou · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Tongzhou Wang, Jun-Yan Zhu, Antonio Torralba, and Alexei A Efros · 2018
Earlier work this paper cites.
JAX: composable transformations of Python+NumPy programs, 2018
James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, and Skye Wanderman-Milne · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clement Hongler · 2018
Earlier work this paper cites.
Deep neural networks as gaussian processes
Jaehoon Lee, Yasaman Bahri, Roman Novak, Sam Schoenholz, Jeffrey Pennington, and Jascha Sohl-dickstein · 2018
Earlier work this paper cites.
Gaussian process behaviour in wide deep neural networks
Alexander G. de G. Matthews, Jiri Hron, Mark Rowland, Richard E. Turner, and Zoubin Ghahramani · 2018
Earlier work this paper cites.
Adversarial examples are not bugs, they are features, 2019
Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Logan Engstrom, Brandon Tran, and Aleksander Madry · 2019
Earlier work this paper cites.
Bayesian deep convolutional networks with many channels are gaussian processes
Roman Novak, Lechao Xiao, Jaehoon Lee, Yasaman Bahri, Greg Yang, Jiri Hron, Daniel A. Abolafia, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2019
Cited alongside, same era.
Deep convolutional networks as shallow gaussian processes
Adrià Garriga-Alonso, Laurence Aitchison, and Carl Edward Rasmussen · 2019
Cited alongside, same era.
On exact computation with an infinitely wide neural net
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, Russ R Salakhutdinov, and Ruosong Wang · 2019
Cited alongside, same era.
Enhanced convolutional neural tangent kernels
Zhiyuan Li, Ruosong Wang, Dingli Yu, Simon S Du, Wei Hu, Ruslan Salakhutdinov, and Sanjeev Arora · 2019
Cited alongside, same era.
Wide neural networks of any depth evolve as linear models under gradient descent
Jaehoon Lee, Lechao Xiao, Samuel S. Schoenholz, Yasaman Bahri, Roman Novak, Jascha Sohl-Dickstein, and Jeffrey Pennington · 2019
Cited alongside, same era.
Asymptotics of wide convolutional neural networks
Anders Andreassen and Ethan Dyer · 2020
Later among the works it cites.
Non-Gaussian processes and neural networks at finite widths
Sho Yaida · 2020
Later among the works it cites.
Infinite attention: NNGP and NTK for deep attention networks
Jiri Hron, Yasaman Bahri, Jascha Sohl-Dickstein, and Roman Novak · 2020
Later among the works it cites.
Disentangling trainability and generalization in deep learning
Lechao Xiao, Jeffrey Pennington, and Samuel S Schoenholz · 2020
Later among the works it cites.
The neural tangent kernel in high dimensions: Triple descent and a multi-scale theory of generalization
Ben Adlam and Jeffrey Pennington · 2020
Later among the works it cites.
The large learning rate phase of deep learning: the catapult mechanism
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Intrinsic dimension of data representations in deep neural networks
A. Ansuini, A. Laio, J. Macke, and D. Zoccolan · 2019
Cited alongside, same era.
Soft-label dataset distillation and text dataset distillation
Ilia Sucholutsky and Matthias Schonlau · 2019
Cited alongside, same era.
Autoaugment: Learning augmentation strategies from data
Ekin D. Cubuk, Barret Zoph, Dandelion Mane, Vijay Vasudevan, and Quoc V. Le · 2019
Cited alongside, same era.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Cited alongside, same era.
Flexible dataset distillation: Learn labels instead of images
Ondrej Bohdal, Yongxin Yang, and Timothy Hospedales · 2020
Cited alongside, same era.
What shapes feature representations? exploring datasets, architectures, and training
Katherine L Hermann and Andrew K Lampinen · 2020
Cited alongside, same era.
Aitor Lewkowycz, Yasaman Bahri, Ethan Dyer, Jascha Sohl-Dickstein, and Guy Gur-Ari · 2020
Later among the works it cites.
On the training dynamics of deep networks with l_2 regularization
Aitor Lewkowycz and Guy Gur-Ari · 2020
Later among the works it cites.
Towards nngp-guided neural architecture search
Daniel S Park, Jaehoon Lee, Daiyi Peng, Yuan Cao, and Jascha Sohl-Dickstein · 2020
Later among the works it cites.
A constructive prediction of the generalization error across scales
Jonathan S. Rosenfeld, Amir Rosenfeld, Yonatan Belinkov, and Nir Shavit · 2020
Later among the works it cites.
Scaling laws for neural language models, 2020
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Later among the works it cites.
On the infinite width limit of neural networks with a standard parameterization
Jascha Sohl-Dickstein, Roman Novak, Samuel S Schoenholz, and Jaehoon Lee · 2020
Later among the works it cites.
Dataset meta-learning from kernel ridge-regression
Timothy Nguyen, Zhourong Chen, and Jaehoon Lee · 2021
Closest in time.
Dataset condensation with differentiable siamese augmentation
Bo Zhao and Hakan Bilen · 2021
Closest in time.
Dataset condensation with gradient matching
Bo Zhao, Konda Reddy Mopuri, and Hakan Bilen · 2021
Closest in time.
Approximation and learning with deep convolutional models: a kernel perspective
Alberto Bietti · 2021
Closest in time.
Exploring the uncertainty properties of neural networks’ implicit priors in the infinite-width limit
Ben Adlam, Jaehoon Lee, Lechao Xiao, Jeffrey Pennington, and Jasper Snoek · 2021
Closest in time.
Neural architecture search on imagenet in four {gpu} hours: A theoretically inspired perspective
Wuyang Chen, Xinyu Gong, and Zhangyang Wang · 2021
Closest in time.
Scaling neural tangent kernels via sketching and random features
Amir Zandieh, Insu Han, Haim Avron, Neta Shoham, Chaewon Kim, and Jinwoo Shin · 2021
Closest in time.
Explaining neural scaling laws
Yasaman Bahri, Ethan Dyer, Jared Kaplan, Jaehoon Lee, and Utkarsh Sharma · 2021
Closest in time.