Fetching the paper…
Reading the bibliography…
The aim of this paper is to introduce a new learning procedure for neural networks and to demonstrate that it works well enough on a few small problems to be worth further investigation.
Your classifier is secretly an energy based model and you should treat it like one
Grathwohl, W., Wang, K.-C., Jacobsen, J.-H., Duvenaud, D., Norouzi, M., and Swersky, K. (2019) · 1912
Earlier work this paper cites.
The perceptron: a probabilistic model for information storage and organization in the brain
Rosenblatt, F. (1958) · 1958
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R. (2014) · 1958
Earlier work this paper cites.
The function of dream sleep
Crick, F. and Mitchison, G. (1983) · 1983
Earlier work this paper cites.
Learning and relearning in Boltzmann machines
Hinton, G. and Sejnowski, T. (1986) · 1986
Earlier work this paper cites.
Learning representations by back-propagating errors
Rumelhart, D. E., Hinton, G. E., and Williams, R. J. (1986) · 1986
Earlier work this paper cites.
Feudal reinforcement learning
Dayan, P. and Hinton, G. E. (1992) · 1992
Earlier work this paper cites.
Weight perturbation: An optimal architecture and learning technique for analog vlsi feedforward and recurrent multilayer networks
Jabri, M. and Flower, B. (1992) · 1992
Earlier work this paper cites.
Predictive coding in the visual cortex: a functional interpretation of some extra-classical receptive-field effects
Rao, R. and Ballard, D. (1999) · 1999
Earlier work this paper cites.
Training products of experts by minimizing contrastive divergence
Hinton, G. E. (2002) · 2002
Earlier work this paper cites.
Extreme components analysis
Welling, M., Williams, C., and Agakov, F. (2003) · 2003
Earlier work this paper cites.
Greedy layer-wise training of deep networks
Bengio, Y., Lamblin, P., Popovici, D., and Larochelle, H. (2007) · 2006
Earlier work this paper cites.
Big self-supervised models are strong semi-supervised learners
Chen, T., Kornblith, S., Swersky, K., Norouzi, M., and Hinton, G. (2020b) · 2006
Earlier work this paper cites.
Bootstrap your own latent: A new approach to self-supervised learning
Grill, J.-B., Strub, F., Altché, F., Tallec, C., Richemond, P. H., Buchatskaya, E., Doersch, C., Pires, B. A., Guo, Z. D., Azar, M. G., Piot, B., Kavukcuoglu, K., Munos, R., and Valko, M. (2020) · 2006
Cited alongside, same era.
Training end-to-end analog neural networks with equilibrium propagation
Kendall, J., Pantone, R., Manickavasagam, K., Bengio, Y., and Scellier, B. (2020) · 2006
Cited alongside, same era.
Topographic product models applied to natural scene statistics
Osindero, S., Welling, M., and Hinton, G. E. (2006) · 2006
Cited alongside, same era.
Learning multiple layers of features from tiny images
Krizhevsky, A. and Hinton, G. (2009) · 2009
Cited alongside, same era.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Gutmann, M. and Hyvärinen, A. (2010) · 2010
Error forward-propagation: Reusing feedforward connections to propagate errors in deep learning
Kohan, A. A., Rietman, E. A., and Siegelmann, H. T. (2018) · 2018
Later among the works it cites.
Representation learning with contrastive predictive coding
van den Oord, A., Li, Y., and Vinyals, O. (2018) · 2018
Later among the works it cites.
Unsupervised feature learning via non-parametric instance discrimination
Wu, Z., Xiong, Y., Yu, S. X., and Lin, D. (2018) · 2018
Later among the works it cites.
Learning representations by maximizing mutual information across views
Bachman, P., Hjelm, R. D., and Buchwalter, W. (2019) · 2019
Later among the works it cites.
Putting an end to end-to-end: Gradient-isolated learning of representations
Löwe, S., O’Connor, P., and Veeling, B. (2019) · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Raghu, A., Raghu, M., Kornblith, S., Duvenaud, D., and Hinton, G. (2020) · 2011
Cited alongside, same era.
Normalization as a canonical neural computation
Carandini, M. and Heeger, D. J. (2013) · 2013
Cited alongside, same era.
Generative adversarial nets
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y. (2014) · 2014
Cited alongside, same era.
Distilling the knowledge in a neural network
Hinton, G., Vinyals, O., and Dean, J. (2014) · 2014
Cited alongside, same era.
synaptic feedback weights support error backpropagation for deep learning
Lillicrap, T., Cownden, D.and Tweed, D., and Akerman, C. (2016) · 2016
Cited alongside, same era.
Towards deep learning with segregated dendrites
Guerguiev, J., Lillicrap, T. P., and Richards, B. A. (2017) · 2017
Cited alongside, same era.
Regularizing neural networks by penalizing confident output distributions
Pereyra, G., Tucker, G., Chorowski, J., Kaiser, Ł., and Hinton, G. (2017) · 2017
Cited alongside, same era.
Dendritic solutions to the credit assignment problem
Richards, B. A. and Lillicrap, T. P. (2019) · 2019
Later among the works it cites.
Momentum contrast for unsupervised visual representation learning
He, K., Fan, H., Wu, Y., Xie, S., and Girshick, R. (2020) · 2020
Later among the works it cites.
Backpropagation and the brain
Lillicrap, T., Santoro, A., Marris, L., Akerman, C., and Hinton, G. E. (2020) · 2020
Later among the works it cites.
How to represent part-whole hierarchies in a neural network
Hinton, G. (2021) · 2021
Later among the works it cites.
Signal propagation: A framework for learning and inference in a forward pass
Kohan, A. A., Rietman, E. A., and Siegelmann, H. T. (2022) · 2022
Closest in time.
Generative trees: Adversarial and copycat
Nock, R. and Guillame-Bert, M. (2022) · 2022
Closest in time.
Scaling forward gradient with local losses
Ren, M., Kornblith, S., Liao, R., and Hinton, G. (2022) · 2022
Closest in time.