Fetching the paper…
Reading the bibliography…
Designing neural network architectures is a challenging task and knowing which specific layers of a model must be adapted to improve the performance is almost a mystery.
Making deep neural networks robust to label noise: A loss correction approach,
G. Patrini, A. Rozza, A. Krishna Menon, R. Nock, L. Qu, · 1952
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting,
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, R. Salakhutdinov, · 1958
Earlier work this paper cites.
Long short-term memory,
S. Hochreiter, J. Schmidhuber, · 1997
Earlier work this paper cites.
A fast learning algorithm for deep belief nets,
G. E. Hinton, S. Osindero, Y. W. Teh, · 2006
Earlier work this paper cites.
How many hidden layers and nodes?,
D. Stathakis, · 2009
Earlier work this paper cites.
A. Krizhevsky, V. Nair, G. Hinton, Cifar-10, 2009. URL: http://www.cs.toronto.edu/~kriz/cifar.html
2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks,
X. Glorot, Y. Bengio, · 2010
Earlier work this paper cites.
Y. LeCun, C. Cortes, MNIST handwritten digit database, http://yann.lecun.com/exdb/mnist/, 2010. URL: http://yann.lecun.com/exdb/mnist/
2010
Earlier work this paper cites.
Y. Netzer, T. Wang, A. Coates, A. Bissacco, B. Wu, A. Y. Ng, Reading digits in natural images with unsupervised feature learning, 2011
2011
Earlier work this paper cites.
Efficient backprop,
Y. A. LeCun, L. Bottou, G. B. Orr, K.-R. Müller, · 2012
Earlier work this paper cites.
Improving deep neural networks for lvcsr using rectified linear units and dropout,
G. E. Dahl, T. N. Sainath, G. E. Hinton, · 2013
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification,
K. He, X. Zhang, S. Ren, J. Sun, · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift,
S. Ioffe, C. Szegedy, · 2015
Cited alongside, same era.
Highway networks,
R. K. Srivastava, K. Greff, J. Schmidhuber, · 2015
Cited alongside, same era.
Deep learning in neural networks: An overview,
J. Schmidhuber, · 2015
Cited alongside, same era.
M. Abadi, A. Agarwal, P. Barham, E. Brevdo, Z. Chen, C. Citro, G. S. Corrado, A. Davis, J. Dean, M. Devin, S. Ghemawat, I. Goodfellow, A. Harp, G. Irving, M. Isard, Y. Jia, R. Jozefowicz, L. Kaiser, M. Kudlur, J. Levenberg, D. Mané, R. Monga, S. Moore, D. Murray, C. Olah, M. Schuster, J. Shlens, B. Steiner, I. Sutskever, K. Talwar, P. Tucker, V. Vanhoucke, V. Vasudevan, F. Viégas, O. Vinyals, P. Warden, M. Wattenberg, M. Wicke, Y. Yu, X. Zheng, TensorFlow: Large-scale machine learning on heterogeneous systems, 2015. URL: https://www.tensorflow.org/ , software available from tensorflow.org
2015
Cited alongside, same era.
2018
Later among the works it cites.
Mish: A self regularized non-monotonic neural activation function.,
D. Misra, · 2019
Later among the works it cites.
Efficientnet: Rethinking model scaling for convolutional neural networks,
M. Tan, Q. V. Le, · 2019
Later among the works it cites.
Sparse semi-autoencoders to solve the vanishing information problem in multi-layered neural networks,
R. Kamimura, H. Takeuchi, · 2019
Later among the works it cites.
J. Howard, imagenette, 2019. URL: https://github.com/fastai/imagenette/
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep residual learning for image recognition,
K. He, X. Zhang, S. Ren, J. Sun, · 2016
Cited alongside, same era.
Residual networks behave like ensembles of relatively shallow networks,
A. Veit, M. J. Wilber, S. Belongie, · 2016
Cited alongside, same era.
Data-dependent initializations of convolutional neural networks.,
P. Krähenbühl, C. Doersch, J. Donahue, T. Darrell, · 2016
Cited alongside, same era.
All you need is a good init.,
D. Mishkin, J. Matas, · 2016
Cited alongside, same era.
I. Goodfellow, Y. Bengio, A. Courville, Deep learning, MIT press, 2016
2016
Cited alongside, same era.
Identity mappings in deep residual networks,
K. He, X. Zhang, S. Ren, J. Sun, · 2016
Cited alongside, same era.
On the expressive power of deep neural networks,
M. Raghu, B. Poole, J. Kleinberg, S. Ganguli, J. Sohl-Dickstein, · 2017
Cited alongside, same era.
The shattered gradients problem: If resnets are the answer, then what is the question?,
D. Balduzzi, M. Frean, L. Leary, J. Lewis, K. W.-D. Ma, B. McWilliams, · 2017
Cited alongside, same era.
Lookahead optimizer: k steps forward, 1 step back,
M. Zhang, J. Lucas, J. Ba, G. E. Hinton, · 2019
Later among the works it cites.
On the variance of the adaptive learning rate and beyond,
L. Liu, H. Jiang, P. He, W. Chen, X. Liu, J. Gao, J. Han, · 2020
Later among the works it cites.
Norm-preservation: Why residual networks can become extremely deep?,
A. Zaeemzadeh, N. Rahnavard, M. Shah, · 2020
Later among the works it cites.
Limitation of capsule networks,
D. Peer, S. Stabinger, A. Rodríguez-Sánchez, · 2021
Closest in time.
Conflicting bundles: Adapting architectures towards the improved training of deep neural networks,
D. Peer, S. Stabinger, A. Rodríguez-Sánchez, · 2021
Closest in time.
conflicting_bundle.py - a python module to identify problematic layers in deep neural networks,
D. Peer, S. Stabinger, A. Rodríguez-Sánchez, · 2021
Closest in time.