Fetching the paper…
Reading the bibliography…
A key factor in the success of deep neural networks is the ability to scale models to improve performance by varying the architecture depth and width.
Regression and anova with zero-one data: Measures of residual variation
Bradley Efron · 1978
Earlier work this paper cites.
Approximation by superpositions of a sigmoidal function
George Cybenko · 1989
Earlier work this paper cites.
Approximation capabilities of multilayer feedforward networks
Kurt Hornik · 1991
Earlier work this paper cites.
Approximation theory of the mlp model in neural networks
Allan Pinkus · 1999
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Algorithms for learning kernels based on centered alignment
Corinna Cortes, Mehryar Mohri, and Afshin Rostamizadeh · 2012
Earlier work this paper cites.
Feature selection via dependence maximization
Le Song, Alex Smola, Arthur Gretton, Justin Bedo, and Karsten Borgwardt · 2012
Earlier work this paper cites.
Do deep nets really need to be deep?
Jimmy Ba and Rich Caruana · 2014
Earlier work this paper cites.
Fitnets: Hints for thin deep nets
Adriana Romero, Nicolas Ballas, Samira Ebrahimi Kahou, Antoine Chassang, Carlo Gatta, and Yoshua Bengio · 2015
Earlier work this paper cites.
Representation benefits of deep feedforward networks
Matus Telgarsky · 2015
Earlier work this paper cites.
Understanding intermediate layers using linear classifier probes
Guillaume Alain and Yoshua Bengio · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Wide residual networks
Sergey Zagoruyko and Nikos Komodakis · 2016
Earlier work this paper cites.
Approximating continuous functions by relu nets of minimal width
Boris Hanin and Mark Sellke · 2017
Earlier work this paper cites.
The expressive power of neural networks: A view from the width
Zhou Lu, Hongming Pu, Feicheng Wang, Zhiqiang Hu, and Liwei Wang · 2017
Cited alongside, same era.
A closer look at deep learning heuristics: Learning rate restarts, warmup and distillation
Akhilesh Gotmare, Nitish Shirish Keskar, Caiming Xiong, and Richard Socher · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Cited alongside, same era.
Deep neural networks as gaussian processes
Jaehoon Lee, Jascha Sohl-dickstein, Jeffrey Pennington, Roman Novak, Sam Schoenholz, and Yasaman Bahri · 2018
Cited alongside, same era.
Resnet with one-neuron hidden layers is a universal approximator
Hongzhou Lin and Stefanie Jegelka · 2018
Cited alongside, same era.
Universality and individuality in neural dynamics across large populations of recurrent networks
Niru Maheswaranathan, Alex Williams, Matthew Golub, Surya Ganguli, and David Sussillo · 2019
Later among the works it cites.
Deep neuroethology of a virtual rodent
Josh Merel, Diego Aldarondo, Jesse Marshall, Yuval Tassa, Greg Wayne, and Bence Ölveczky · 2019
Later among the works it cites.
Transfusion: Understanding transfer learning for medical imaging
Maithra Raghu, Chiyuan Zhang, Jon Kleinberg, and Samy Bengio · 2019
Later among the works it cites.
Probing the state of the art: A critical look at visual representation evaluation
Cinjon Resnick, Zeping Zhan, and Joan Bruna · 2019
Later among the works it cites.
Comparison against task driven artificial neural networks reveals functional properties in mouse visual cortex
Jianghong Shi, Eric Shea-Brown, and Michael Buice · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alexander G de G Matthews, Mark Rowland, Jiri Hron, Richard E Turner, and Zoubin Ghahramani · 2018
Cited alongside, same era.
Insights on representational similarity in neural networks with canonical correlation
Ari Morcos, Maithra Raghu, and Samy Bengio · 2018
Cited alongside, same era.
Bayesian deep convolutional networks with many channels are gaussian processes
Roman Novak, Lechao Xiao, Jaehoon Lee, Yasaman Bahri, Greg Yang, Jiri Hron, Daniel A Abolafia, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2018
Cited alongside, same era.
Post selection inference with kernels
Makoto Yamada, Yuta Umezu, Kenji Fukumizu, and Ichiro Takeuchi · 2018
Cited alongside, same era.
Identifying and controlling important neurons in neural machine translation
Anthony Bau, Yonatan Belinkov, Hassan Sajjad, Nadir Durrani, Fahim Dalvi, and James Glass · 2019
Cited alongside, same era.
On lazy training in differentiable programming
Lenaic Chizat, Edouard Oyallon, and Francis Bach · 2019
Cited alongside, same era.
Selective brain damage: Measuring the disparate impact of model pruning
Sara Hooker, Aaron Courville, Yann Dauphin, and Andrea Frome · 2019
Cited alongside, same era.
Later among the works it cites.
Efficientnet: Rethinking model scaling for convolutional neural networks
Mingxing Tan and Quoc Le · 2019
Later among the works it cites.
Jessica AF Thompson, Yoshua Bengio, and Marc Schoenwiesner · 2019
Later among the works it cites.
Regularization matters: Generalization and optimization of neural nets vs their induced kernel
Colin Wei, Jason D Lee, Qiang Liu, and Tengyu Ma · 2019
Later among the works it cites.
Emerging cross-lingual structure in pretrained language models
Shijie Wu, Alexis Conneau, Haoran Li, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Later among the works it cites.
Let’s agree to agree: Neural networks share classification order on real datasets
Guy Hacohen and Daphna Weinshall · 2020
Closest in time.
What shapes feature representations? exploring datasets, architectures, and training
Katherine L Hermann and Andrew K Lampinen · 2020
Closest in time.
The large learning rate phase of deep learning: the catapult mechanism
Aitor Lewkowycz, Yasaman Bahri, Ethan Dyer, Jascha Sohl-Dickstein, and Guy Gur-Ari · 2020
Closest in time.
What is being transferred in transfer learning?
Behnam Neyshabur, Hanie Sedghi, and Chiyuan Zhang · 2020
Closest in time.
Similarity analysis of contextual word representation models
John M Wu, Yonatan Belinkov, Hassan Sajjad, Nadir Durrani, Fahim Dalvi, and James Glass · 2020
Closest in time.
Transferability of brain decoding using graph convolutional networks
Yu Zhang and Pierre Bellec · 2020
Closest in time.