Fetching the paper…
Reading the bibliography…
We analyze the learning dynamics of infinitely wide neural networks with a finite sized bottle-neck.
Handwritten digit recognition with a back-propagation network
LeCun, Y., Boser, B. E., Denker, J. S., Henderson, D., Howard, R. E., Hubbard, W. E., and Jackel, L. D · 1989
Earlier work this paper cites.
Bayesian learning for neural networks
Hinton, G. E. and Neal, R · 1995
Earlier work this paper cites.
On the optimization dynamics of wide hypernetworks
Littwin, E., Galanti, T., and Wolf, L · 2003
Earlier work this paper cites.
Tensor programs ii: Neural tangent kernel for any architecture
Yang, G · 2006
Earlier work this paper cites.
Continuous neural networks
Roux, N. L. and Bengio, Y · 2007
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A · 2009
Earlier work this paper cites.
Tensor programs iii: Neural matrix laws
Yang, G · 2009
Earlier work this paper cites.
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015
Abadi, M., Agarwal, A., Barham, P., Brevdo, E., Chen, Z., Citro, C., Corrado, G. S., Davis, A., Dean, J., Devin, M., Ghemawat, S., Goodfellow, I., Harp, A., Irving, G., Isard, M., Jia, Y., Jozefowicz, R., Kaiser, L., Kudlur, M., Levenberg, J., Mané, D., Monga, R., Moore, S., Murray, D., Olah, C., Schuster, M., Shlens, J., Steiner, B., Sutskever, I., Talwar, K., Tucker, P., Vanhoucke, V., Vasudevan, V., Viégas, F., Vinyals, O., Warden, P., Wattenberg, M., Wicke, M., Yu, Y., and Zheng, X · 2015
Earlier work this paper cites.
Steps toward deep kernel methods from infinite neural networks
Hazan, T. and Jaakkola, T · 2015
Cited alongside, same era.
Toward deeper understanding of neural networks: The power of initialization and a dual view on expressivity
Daniely, A., Frostig, R., and Singer, Y · 2016
Cited alongside, same era.
On the optimization of deep networks: Implicit acceleration by overparameterization
Arora, S., Cohen, N., and Hazan, E · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C · 2018
Cited alongside, same era.
Deep neural networks as gaussian processes
Lee, J., Bahri, Y., Novak, R., Schoenholz, S., Pennington, J., and Sohl-Dickstein, J · 2018
Addressing the loss-metric mismatch with adaptive loss alignment
Huang, C., Zhai, S., Talbott, W., Martin, M. B., Sun, S.-Y., Guestrin, C., and Susskind, J · 2019
Later among the works it cites.
Bayesian deep convolutional networks with many channels are gaussian processes
Novak, R., Xiao, L., Bahri, Y., Lee, J., Yang, G., Hron, J., Abolafia, D., Pennington, J., and Sohl-Dickstein, J · 2019
Later among the works it cites.
Yang, G · 2019
Later among the works it cites.
Wide neural networks with bottlenecks are deep gaussian processes
Agrawal, D., Papamarkou, T., and Hinkle, J · 2020
Later among the works it cites.
Why bigger is not always better: on finite and infinite neural networks
Aitchison, L · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Gaussian process behaviour in wide deep neural networks
Matthews, A., Rowland, M., Hron, J., Turner, R., and Ghahramani, Z · 2018
Cited alongside, same era.
On exact computation with an infinitely wide neural net
Arora, S., Du, S. S., Hu, W., Li, Z., Salakhutdinov, R., and Wang, R · 2019
Cited alongside, same era.
On random kernels of residual architectures
Littwin, E., Galanti, T., and Wolf, L
Cited in the paper.
Later among the works it cites.
Feature learning in infinite-width neural networks
Yang, G. and Hu, E. J · 2020
Later among the works it cites.
Tensor programs iib: Architectural universality of neural tangent kernel training dynamics
Yang, G. and Littwin, E · 2021
Closest in time.