Fetching the paper…
Reading the bibliography…
The results of training a neural network are heavily dependent on the architecture chosen; and even a modification of only its size, however small, typically involves restarting the training process.
Mathematical theory of probability and statistics , chapter VIII.9.3
von Mises, R · 1964
Earlier work this paper cites.
Dynamic node creation in backpropagation networks
Ash, T · 1989
Earlier work this paper cites.
Optimal brain damage
LeCun, Y., Denker, J. S., and Solla, S. A · 1989
Earlier work this paper cites.
Optimal brain surgeon and general network pruning
Hassibi, B., Stork, D. G., and Wolff, G. J · 1993
Earlier work this paper cites.
Natural gradient works efficiently in learning
Amari, S · 1998
Earlier work this paper cites.
Neurogenesis in the adult brain: Death of a dogma
Gross, C. G · 2000
Earlier work this paper cites.
Cifar-10 (canadian institute for advanced research)
Krizhevsky, A., Nair, V., and Hinton, G · 2009
Earlier work this paper cites.
Riemann Manifold Langevin and Hamiltonian Monte Carlo Methods
Girolami, M. and Calderhead, B · 2011
Earlier work this paper cites.
Scikit-learn: Machine learning in Python
Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., and Duchesnay, E · 2011
Earlier work this paper cites.
The mnist database of handwritten digit images for machine learning research
Deng, L · 2012
Earlier work this paper cites.
Efficient BackProp , pp. 9–48
LeCun, Y. A., Bottou, L., Orr, G. B., and Müller, K.-R · 2012
Earlier work this paper cites.
Functional neurogenesis in the adult hippocampus: Then and now
Vadodaria, K. C. and Jessberger, S · 2014
Earlier work this paper cites.
Tiny imagenet visual recognition challenge
Le, Y. and Yang, X. S · 2015
Earlier work this paper cites.
Optimizing neural networks with kronecker-factored approximate curvature
Martens, J. and Grosse, R. B · 2015
Cited alongside, same era.
Imagenet large scale visual recognition challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M. S., Berg, A. C., and Fei-Fei, L · 2015
Cited alongside, same era.
Going deeper with convolutions
Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S. E., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A · 2015
Cited alongside, same era.
Net2net: Accelerating learning via knowledge transfer
Chen, T., Goodfellow, I., and Shlens, J · 2016
Cited alongside, same era.
Rusu, A. A., Rabinowitz, N. C., Desjardins, G., Soyer, H., Kirkpatrick, J., Kavukcuoglu, K., Pascanu, R., and Hadsell, R · 2016
Cited alongside, same era.
Gpipe: Efficient training of giant neural networks using pipeline parallelism
Huang, Y., Cheng, Y., Bapna, A., Firat, O., Chen, D., Chen, M. X., Lee, H., Ngiam, J., Le, Q. V., Wu, Y., and Chen, Z · 2019
Later among the works it cites.
Splitting steepest descent for growing neural architectures
Wu, L., Wang, D., and Liu, Q · 2019
Later among the works it cites.
New insights and perspectives on the natural gradient method
Martens, J · 2020
Later among the works it cites.
Padé activation units: End-to-end learning of flexible activation functions in deep networks
Molina, A., Schramowski, P., and Kersting, K · 2020
Later among the works it cites.
Deep double descent: Where bigger models and more data hurt
Nakkiran, P., Kaplun, G., Bansal, Y., Yang, T., Barak, B., and Sutskever, I · 2020
Later among the works it cites.
Firefly neural architecture descent: a general approach for growing neural networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Draelos, T. J., Miner, N. E., Lamb, C. C., Cox, J. A., Vineyard, C. M., Carlson, K. D., Severa, W. M., James, C. D., and Aimone, J. B · 2017
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2017
Cited alongside, same era.
Opening the black box of deep neural networks via information
Shwartz-Ziv, R. and Tishby, N · 2017
Cited alongside, same era.
JAX: composable transformations of Python+NumPy programs
Bradbury, J., Frostig, R., Hawkins, P., Johnson, M. J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., VanderPlas, J., Wanderman-Milne, S., and Zhang, Q · 2018
Cited alongside, same era.
On the information bottleneck theory of deep learning
Saxe, A. M., Bansal, Y., Dapello, J., Advani, M., Kolchinsky, A., Tracey, B. D., and Cox, D. D · 2018
Cited alongside, same era.
Lifelong learning with dynamically expandable networks
Yoon, J., Yang, E., Lee, J., and Hwang, S. J · 2018
Cited alongside, same era.
Deep active learning with a neural architecture search
Geifman, Y. and El-Yaniv, R · 2019
Cited alongside, same era.
Wu, L., Liu, B., Stone, P., and Liu, Q · 2020
Later among the works it cites.
Knowledge-adaptation priors
Khan, M. E. and Swaroop, S · 2021
Later among the works it cites.
Gradmax: Growing neural networks using gradient information
Evci, U., van Merrienboer, B., Unterthiner, T., Pedregosa, F., and Vladymyrov, M · 2022
Later among the works it cites.
Biological underpinnings for lifelong learning machines
Kudithipudi, D., Aguilar-Simon, M., Babb, J., Bazhenov, M., Blackiston, D., Bongard, J., Brna, A., Chakravarthi Raja, S., Cheney, N., Clune, J., Daram, A., Fusi, S., Helfer, P., Kay, L., Ketz, N., Kira, Z., Kolouri, S., Krichmar, J., Kriegman, S., and Siegelmann, H · 2022
Later among the works it cites.
When, where, and how to add new neurons to anns
Maile, K., Rachelson, E., Luga, H., and Wilson, D. G · 2022
Later among the works it cites.
Flax: A neural network library and ecosystem for JAX
Heek, J., Levskaya, A., Oliver, A., Ritter, M., Rondepierre, B., Steiner, A., and van Zee, M · 2023
Closest in time.
A wholistic view of continual learning with deep neural networks: Forgotten lessons and the bridge to active and open world learning
Mundt, M., Hong, Y., Pliushch, I., and Ramesh, V · 2023
Closest in time.