Fetching the paper…
Reading the bibliography…
The deep learning literature is continuously updated with new architectures and training techniques.
On random graphs i
P Erdos and A Rényi · 1959
Earlier work this paper cites.
Backpropagation applied to handwritten zip code recognition
Yann LeCun, Bernhard Boser, John S Denker, Donnie Henderson, Richard E Howard, Wayne Hubbard, and Lawrence D Jackel · 1989
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Emergence of scaling in random networks
Albert-László Barabási and Réka Albert · 1999
Earlier work this paper cites.
Reducing the dimensionality of data with neural networks
Geoffrey E Hinton and Ruslan R Salakhutdinov · 2006
Earlier work this paper cites.
Greedy layer-wise training of deep networks
Yoshua Bengio, Pascal Lamblin, Dan Popovici, and Hugo Larochelle · 2007
Earlier work this paper cites.
Characterization of complex networks: A survey of measurements
L da F Costa, Francisco A Rodrigues, Gonzalo Travieso, and Paulino Ribeiro Villas Boas · 2007
Earlier work this paper cites.
Scale-free networks: a decade and beyond
Albert-László Barabási · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
Analyzing and modeling real-world phenomena with complex networks: a survey of applications
Luciano da Fontoura Costa, Osvaldo N Oliveira Jr, Gonzalo Travieso, Francisco Aparecido Rodrigues, Paulino Ribeiro Villas Boas, Lucas Antiqueira, Matheus Palhares Viana, and Luis Enrique Correa Rocha · 2011
Earlier work this paper cites.
Deep sparse rectifier neural networks
Xavier Glorot, Antoine Bordes, and Yoshua Bengio · 2011
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Andrew M Saxe, James L McClelland, and Surya Ganguli · 2013
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton · 2013
Earlier work this paper cites.
Random walk initialization for training very deep feedforward networks
David Sussillo and LF Abbott · 2014
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Adjusting for dropout variance in batch normalization and weight initialization
Dan Hendrycks and Kevin Gimpel · 2016
Earlier work this paper cites.
Impact of small-world network topology on the conventional artificial neural network for the diagnosis of diabetes
Okan Erkaymaz and Mahmut Ozer · 2016
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna · 2016
Cited alongside, same era.
Deep networks with stochastic depth
Gao Huang, Yu Sun, Zhuang Liu, Daniel Sedra, and Kilian Q Weinberger · 2016
Cited alongside, same era.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
Tim Salimans and Durk P Kingma · 2016
Cited alongside, same era.
Deep learning systems as complex networks
Alberto Testolin, Michele Piccolini, and Samir Suweis · 2020
Later among the works it cites.
Emergence of network motifs in deep neural networks
Matteo Zambra, Amos Maritan, and Alberto Testolin · 2020
Later among the works it cites.
Graph structure of neural networks
Jiaxuan You, Jure Leskovec, Kaiming He, and Saining Xie · 2020
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Later among the works it cites.
Randaugment: Practical automated data augmentation with a reduced search space
Ekin D Cubuk, Barret Zoph, Jonathon Shlens, and Quoc V Le · 2020
Later among the works it cites.
Random erasing data augmentation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Scott Gray, Alec Radford, and Diederik P Kingma · 2017
Cited alongside, same era.
Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms
Han Xiao, Kashif Rasul, and Roland Vollgraf · 2017
Cited alongside, same era.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Cited alongside, same era.
mixup: Beyond empirical risk minimization
Hongyi Zhang, Moustapha Cisse, Yann N Dauphin, and David Lopez-Paz · 2017
Cited alongside, same era.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Jonathan Frankle and Michael Carbin · 2018
Cited alongside, same era.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Yuanzhi Li and Yingyu Liang · 2018
Cited alongside, same era.
Deconstructing lottery tickets: Zeros, signs, and the supermask
Hattie Zhou, Janice Lan, Rosanne Liu, and Jason Yosinski · 2019
Cited alongside, same era.
Zhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li, and Yi Yang · 2020
Later among the works it cites.
David Picard · 2021
Later among the works it cites.
Resnet strikes back: An improved training procedure in timm
Ross Wightman, Hugo Touvron, and Hervé Jégou · 2021
Later among the works it cites.
Quantifying epistemic uncertainty in deep learning
Ziyi Huang, Henry Lam, and Haofeng Zhang · 2021
Later among the works it cites.
Effect of initial configuration of weights on training and function of artificial neural networks
Ricardo J Jesus, Mário L Antunes, Rui A da Costa, Sergey N Dorogovtsev, José FF Mendes, and Rui L Aguiar · 2021
Later among the works it cites.
A review on weight initialization strategies for neural networks
Meenal V Narkhede, Prashant P Bartakke, and Mukul S Sutaone · 2021
Later among the works it cites.
Revisiting weight initialization of deep neural networks
Maciej Skorski, Alessandro Temperoni, and Martin Theobald · 2021
Later among the works it cites.
Social interaction layers in complex networks for the dynamical epidemic modeling of covid-19 in brazil
Leonardo FS Scabini, Lucas C Ribas, Mariane B Neiva, Altamir GB Junior, Alex JF Farfán, and Odemir M Bruno · 2021
Later among the works it cites.
Structure and performance of fully connected neural networks: Emerging complex network properties
Leonardo FS Scabini and Odemir M Bruno · 2021
Later among the works it cites.
Characterizing learning dynamics of deep neural networks via complex networks
Emanuele La Malfa, Gabriele La Malfa, Giuseppe Nicosia, and Vito Latora · 2021
Later among the works it cites.
Vision transformer for small-size datasets
Seung Hoon Lee, Seunghyun Lee, and Byung Cheol Song · 2021
Later among the works it cites.
Sample-efficient neural architecture search by learning actions for monte carlo tree search
Linnan Wang, Saining Xie, Teng Li, Rodrigo Fonseca, and Yuandong Tian · 2021
Later among the works it cites.
A survey on vision transformer
Kai Han, Yunhe Wang, Hanting Chen, Xinghao Chen, Jianyuan Guo, Zhenhua Liu, Yehui Tang, An Xiao, Chunjing Xu, Yixing Xu, et al · 2022
Closest in time.
Asher Trockman and J Zico Kolter · 2022
Closest in time.