Fetching the paper…
Reading the bibliography…
Recent work on mode connectivity in the loss landscape of deep neural networks has demonstrated that the locus of (sub-)optimal weight vectors lies on continuous paths.
Sample estimate of the entropy of a random vector
Kozachenko, L. and Leonenko, N. N · 1987
Earlier work this paper cites.
Efficient estimations from a slowly convergent robbins-monro process
Ruppert, D · 1988
Earlier work this paper cites.
New stochastic approximation type procedures
Polyak, B. T · 1990
Earlier work this paper cites.
Keeping the neural networks simple by minimizing the description length of the weights
Hinton, G. E. and Van Camp, D · 1993
Earlier work this paper cites.
A ‘self-referential’weight matrix
Schmidhuber, J · 1993
Earlier work this paper cites.
Bayesian learning for neural networks
Neal, R. M · 1995
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P · 1998
Earlier work this paper cites.
The elements of statistical learning , volume 1
Friedman, J., Hastie, T., and Tibshirani, R · 2001
Earlier work this paper cites.
Estimating mutual information
Kraskov, A., Stögbauer, H., and Grassberger, P · 2004
Earlier work this paper cites.
Learning to control fast-weight memories: An alternative to dynamic recurrent networks
Schmidhuber, J · 2008
Earlier work this paper cites.
Neural architecture search with reinforcement learning
Zoph, B. and Le, Q. V · 2008
Earlier work this paper cites.
What is the best multi-stage architecture for object recognition?
Jarrett, K., Kavukcuoglu, K., LeCun, Y., et al · 2009
Earlier work this paper cites.
A hypercube-based encoding for evolving large-scale neural networks
Stanley, K. O., D’Ambrosio, D. B., and Gauci, J · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot, X. and Bengio, Y · 2010
Earlier work this paper cites.
Deep sparse rectifier neural networks
Glorot, X., Bordes, A., and Bengio, Y · 2011
Earlier work this paper cites.
Practical variational inference for neural networks
Graves, A · 2011
Earlier work this paper cites.
Bayesian learning via stochastic gradient langevin dynamics
Welling, M. and Teh, Y. W · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Earlier work this paper cites.
Predicting parameters in deep learning
Denil, M., Shakibi, B., Dinh, L., Ranzato, M. A., and de Freitas, N · 2013
Cited alongside, same era.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M · 2013
Cited alongside, same era.
Rectifier nonlinearities improve neural network acoustic models
Maas, A. L., Hannun, A. Y., and Ng, A. Y · 2013
Cited alongside, same era.
Introduction to continuous entropy
Marsh, C · 2013
Cited alongside, same era.
Intriguing properties of neural networks
Szegedy, C., Zaremba, W., Sutskever, I., Bruna, J., Erhan, D., Goodfellow, I., and Fergus, R · 2013
Cited alongside, same era.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Learning feed-forward one-shot learners
Bertinetto, L., Henriques, J. F., Valmadre, J., Torr, P., and Vedaldi, A · 2016
Later among the works it cites.
Dynamic filter networks
De Brabandere, B., Jia, X., Tuytelaars, T., and Van Gool, L · 2016
Later among the works it cites.
Topology and geometry of half-rectified network optimization
Freeman, C. D. and Bruna, J · 2016
Later among the works it cites.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Gal, Y. and Ghahramani, Z · 2016
Later among the works it cites.
Ha, D., Dai, A., and Le, Q. V · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dauphin, Y. N., Pascanu, R., Gulcehre, C., Cho, K., Ganguli, S., and Bengio, Y · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. and Ba, J · 2014
Cited alongside, same era.
The cifar-10 dataset
Krizhevsky, A., Nair, V., and Hinton, G · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A · 2014
Cited alongside, same era.
Dropout: a simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G. E., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R · 2014
Cited alongside, same era.
Weight uncertainty in neural network
Blundell, C., Cornebise, J., Kavukcuoglu, K., and Wierstra, D · 2015
Cited alongside, same era.
Improvement of the k-nn entropy estimator with applications in systems biology
Charzyńska, A. and Gambin, A · 2015
Cited alongside, same era.
Li, K. and Malik, J · 2016
Later among the works it cites.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
Salimans, T. and Kingma, D. P · 2016
Later among the works it cites.
Variational inference & deep learning: A new synthesis
Kingma, D. P · 2017
Later among the works it cites.
Krueger, D., Huang, C.-W., Islam, R., Turner, R., Lacoste, A., and Courville, A · 2017
Later among the works it cites.
Diversified texture synthesis with feed-forward networks
Li, Y., Fang, C., Yang, J., Wang, Z., Lu, X., and Yang, M.-H · 2017
Later among the works it cites.
Multiplicative normalizing flows for variational bayesian neural networks
Louizos, C. and Welling, M · 2017
Later among the works it cites.
Stochastic gradient descent as approximate bayesian inference
Mandt, S., Hoffman, M. D., and Blei, D. M · 2017
Later among the works it cites.
Feature visualization
Olah, C., Mordvintsev, A., and Schubert, L · 2017
Later among the works it cites.
Chang, O. and Lipson, H · 2018
Later among the works it cites.
Essentially no barriers in neural network energy landscape
Draxler, F., Veschgini, K., Salmhofer, M., and Hamprecht, F. A · 2018
Later among the works it cites.
Loss surfaces, mode connectivity, and fast ensembling of dnns
Garipov, T., Izmailov, P., Podoprikhin, D., Vetrov, D. P., and Wilson, A. G · 2018
Later among the works it cites.
Stochastic hyperparameter optimization through hypernetworks
Lorraine, J. and Duvenaud, D · 2018
Later among the works it cites.