Fetching the paper…
Reading the bibliography…
Hypernetworks are neural networks that generate weights for another neural network.
L. Kozachenko and N. N. Leonenko, “Sample estimate of the entropy of a random vector,” Problemy Peredachi Informatsii
1987
Earlier work this paper cites.
G. E. Hinton and D. Van Camp, “Keeping the neural networks simple by minimizing the description length of the weights,” in Proceedings of the sixth annual conference on Computational learning theory
1993
Earlier work this paper cites.
J. Schmidhuber, “A ‘self-referential’weight matrix,” in ICANN’93
1993
Earlier work this paper cites.
PhD thesis, University of Toronto, 1995
R. M. Neal, Bayesian learning for neural networks · 1995
Earlier work this paper cites.
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,” Proceedings of the IEEE
1998
Earlier work this paper cites.
Springer series in statistics New York, 2001
J. Friedman, T. Hastie, and R. Tibshirani, The elements of statistical learning · 2001
Earlier work this paper cites.
A. Kraskov, H. Stögbauer, and P. Grassberger, “Estimating mutual information,” Physical review E
2004
Earlier work this paper cites.
J. Schmidhuber, “Learning to control fast-weight memories: An alternative to dynamic recurrent networks,” Learning
2008
Earlier work this paper cites.
K. Jarrett, K. Kavukcuoglu, Y. LeCun, et al
2009
Earlier work this paper cites.
K. O. Stanley, D. B. D’Ambrosio, and J. Gauci, “A hypercube-based encoding for evolving large-scale neural networks,” Artificial life
2009
Earlier work this paper cites.
X. Glorot and Y. Bengio, “Understanding the difficulty of training deep feedforward neural networks,” in Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics
2010
Earlier work this paper cites.
A. Graves, “Practical variational inference for neural networks,” in Advances in Neural Information Processing Systems
2011
Earlier work this paper cites.
X. Glorot, A. Bordes, and Y. Bengio, “Deep sparse rectifier neural networks,” in Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics
2011
Earlier work this paper cites.
M. Welling and Y. W. Teh, “Bayesian learning via stochastic gradient langevin dynamics,” in Proceedings of the 28th International Conference on Machine Learning (ICML-11)
2011
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in Advances in neural information processing systems
2012
Earlier work this paper cites.
D. P. Kingma and M. Welling, “Auto-encoding variational bayes,” arXiv preprint arXiv:1312.6114
2013
Earlier work this paper cites.
A. L. Maas, A. Y. Hannun, and A. Y. Ng, “Rectifier nonlinearities improve neural network acoustic models,” in Proc. ICML
2013
Earlier work this paper cites.
2013
Cited alongside, same era.
M. Denil, B. Shakibi, L. Dinh, M. A. Ranzato, and N. de Freitas, “Predicting parameters in deep learning,” in Advances in Neural Information Processing Systems 26
2013
Cited alongside, same era.
C. Marsh, “Introduction to continuous entropy,” 2013
2013
Cited alongside, same era.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” in Advances in Neural Information Processing Systems 27
2014
Cited alongside, same era.
Y. N. Dauphin, R. Pascanu, C. Gulcehre, K. Cho, S. Ganguli, and Y. Bengio, “Identifying and attacking the saddle point problem in high-dimensional non-convex optimization,” in Advances in Neural Information Processing Systems 27
2016
Later among the works it cites.
M. Andrychowicz, M. Denil, S. Gomez, M. W. Hoffman, D. Pfau, T. Schaul, and N. de Freitas, “Learning to learn by gradient descent by gradient descent,” in Advances in Neural Information Processing Systems
2016
Later among the works it cites.
K. Li and J. Malik, “Learning to optimize,” arXiv preprint arXiv:1606.01885
2016
Later among the works it cites.
L. Bertinetto, J. F. Henriques, J. Valmadre, P. Torr, and A. Vedaldi, “Learning feed-forward one-shot learners,” in Advances in Neural Information Processing Systems
2016
Later among the works it cites.
B. De Brabandere, X. Jia, T. Tuytelaars, and L. Van Gool, “Dynamic filter networks,” in Neural Information Processing Systems (NIPS)
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2014
Cited alongside, same era.
2014
Cited alongside, same era.
N. Srivastava, G. E. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “Dropout: a simple way to prevent neural networks from overfitting,” Journal of machine learning research
2014
Cited alongside, same era.
2014
Cited alongside, same era.
D. Kingma and J. Ba, “Adam: A method for stochastic optimization,” arXiv preprint arXiv:1412.6980
2014
Cited alongside, same era.
2015
Cited alongside, same era.
A. Choromanska, M. Henaff, M. Mathieu, G. B. Arous, and Y. LeCun, “The loss surfaces of multilayer networks,” in Artificial Intelligence and Statistics
2015
Cited alongside, same era.
K. He, X. Zhang, S. Ren, and J. Sun, “Delving deep into rectifiers: Surpassing human-level performance on imagenet classification,” in Proceedings of the IEEE international conference on computer vision
2015
Cited alongside, same era.
2016
Later among the works it cites.
2016
Later among the works it cites.
Y. Gal and Z. Ghahramani, “Dropout as a bayesian approximation: Representing model uncertainty in deep learning,” in international conference on machine learning
2016
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
PhD thesis, University of Amsterdam, 2017
D. P. Kingma, Variational inference & deep learning: A new synthesis · 2017
Later among the works it cites.
2017
Later among the works it cites.
https://distill.pub/2017/feature-visualization
C. Olah, A. Mordvintsev, and L. Schubert, “Feature visualization,” Distill · 2017
Later among the works it cites.
2018
Closest in time.
2018
Closest in time.
2018
Closest in time.