——, “Where is the information in a deep neural network?” CoRR , vol. abs/1905.12213, 2019. [Online]. Available: http://arxiv.org/abs/1905.12213
Original
1905
Earlier work this paper cites.
A. Kolmogoroff, “Über die analytischen methoden in der wahrscheinlichkeitsrechnung,” Mathematische Annalen , vol. 104, no. 1, pp. 415–458, 1931
1931
Earlier work this paper cites.
K. Itô, On stochastic differential equations . American Mathematical Soc., 1951, vol. 4
1951
Earlier work this paper cites.
V. G. Boltyanskii, R. V. Gamkrelidze, and L. S. Pontryagin, “The theory of optimal processes. i. the maximum principle,” TRW SPACE TECHNOLOGY LABS LOS ANGELES CALIF, Tech. Rep., 1960
1960
Earlier work this paper cites.
R. E. Bellman and R. E. Kalaba, Selected papers on mathematical trends in control theory . Dover Publications, 1964
1964
Earlier work this paper cites.
Y. LeCun, B. E. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. E. Hubbard, and L. D. Jackel, “Handwritten digit recognition with a back-propagation network,” in Advances in neural information processing systems , 1990, pp. 396–404
1990
Earlier work this paper cites.
G. F. Franklin, J. D. Powell, A. Emami-Naeini, and J. D. Powell, Feedback control of dynamic systems . Addison-Wesley Reading, MA, 1994, vol. 3
1994
Earlier work this paper cites.
G. N. Milstein, Numerical integration of stochastic differential equations . Springer Science & Business Media, 1994, vol. 313
1994
Earlier work this paper cites.
D. P. Bertsekas, D. P. Bertsekas, D. P. Bertsekas, and D. P. Bertsekas, Dynamic programming and optimal control . Athena scientific Belmont, MA, 1995, vol. 1, no. 2
1995
Earlier work this paper cites.
C. K. Williams, “Computing with infinite networks,” in Advances in neural information processing systems , 1997, pp. 295–301
1997
Earlier work this paper cites.
R. Jordan, D. Kinderlehrer, and F. Otto, “The variational formulation of the fokker–planck equation,” SIAM journal on mathematical analysis , vol. 29, no. 1, pp. 1–17, 1998
1998
Earlier work this paper cites.
N. Qian, “On the momentum term in gradient descent learning algorithms,” Neural networks , vol. 12, no. 1, pp. 145–151, 1999
1999
Earlier work this paper cites.
S. P. Coraluppi and S. I. Marcus, “Risk-sensitive and minimax control of discrete-time, finite-state markov decision processes,” Automatica , vol. 35, no. 2, pp. 301–309, 1999
1999
Earlier work this paper cites.
B. Øksendal, “Stochastic differential equations,” in Stochastic differential equations . Springer, 2003, pp. 65–84
2003
Earlier work this paper cites.
A. Bovier, M. Eckhoff, V. Gayrard, and M. Klein, “Metastability in reversible diffusion processes i: Sharp asymptotics for capacities and exit times,” Journal of the European Mathematical Society , vol. 6, no. 4, pp. 399–424, 2004
2004
Earlier work this paper cites.
L. Ambrosio, N. Gigli, and G. Savaré, Gradient flows: in metric spaces and in the space of probability measures . Springer Science & Business Media, 2008
2008
Earlier work this paper cites.
T. Başar and P. Bernhard, H-infinity optimal control and related minimax design problems: a dynamic game approach . Springer Science & Business Media, 2008
2008
Earlier work this paper cites.
X. Glorot and Y. Bengio, “Understanding the difficulty of training deep feedforward neural networks,” in Proceedings of the thirteenth international conference on artificial intelligence and statistics , 2010, pp. 249–256
2010
Earlier work this paper cites.
E. Moulines and F. R. Bach, “Non-asymptotic analysis of stochastic approximation algorithms for machine learning,” in Advances in Neural Information Processing Systems , 2011, pp. 451–459
2011
Earlier work this paper cites.
W. Xu, “Towards optimal one pass large scale learning with averaged stochastic gradient descent,” arXiv preprint arXiv:1107.2490 , 2011
Original
2011
Earlier work this paper cites.
M. Welling and Y. W. Teh, “Bayesian learning via stochastic gradient langevin dynamics,” in Proceedings of the 28th international conference on machine learning (ICML-11) , 2011, pp. 681–688
2011
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in Advances in neural information processing systems , 2012, pp. 1097–1105
2012
Earlier work this paper cites.
N. Srivastava and R. R. Salakhutdinov, “Multimodal learning with deep boltzmann machines,” in Advances in neural information processing systems , 2012, pp. 2222–2230
2012
Earlier work this paper cites.
K. Simonyan, A. Vedaldi, and A. Zisserman, “Deep inside convolutional networks: Visualising image classification models and saliency maps,” arXiv preprint arXiv:1312.6034 , 2013
Original
2013
Earlier work this paper cites.
I. Sutskever, J. Martens, G. Dahl, and G. Hinton, “On the importance of initialization and momentum in deep learning,” in International conference on machine learning , 2013, pp. 1139–1147
2013
Earlier work this paper cites.
C. Szegedy, W. Zaremba, I. Sutskever, J. Bruna, D. Erhan, I. Goodfellow, and R. Fergus, “Intriguing properties of neural networks,” arXiv preprint arXiv:1312.6199 , 2013
Original
2013
Earlier work this paper cites.
I. J. Goodfellow, J. Shlens, and C. Szegedy, “Explaining and harnessing adversarial examples,” arXiv preprint arXiv:1412.6572 , 2014
Original
2014
Earlier work this paper cites.
Y. N. Dauphin, R. Pascanu, C. Gulcehre, K. Cho, S. Ganguli, and Y. Bengio, “Identifying and attacking the saddle point problem in high-dimensional non-convex optimization,” in Advances in neural information processing systems , 2014, pp. 2933–2941
2014
Earlier work this paper cites.
G. A. Pavliotis, Stochastic processes and applications: diffusion processes, the Fokker-Planck and Langevin equations . Springer, 2014, vol. 60
2014
Earlier work this paper cites.
M. Hardt, B. Recht, and Y. Singer, “Train faster, generalize better: Stability of stochastic gradient descent,” arXiv preprint arXiv:1509.01240 , 2015
Original
2015
Earlier work this paper cites.
N. Tishby and N. Zaslavsky, “Deep learning and the information bottleneck principle,” in 2015 IEEE Information Theory Workshop (ITW) . IEEE, 2015, pp. 1–5
2015
Earlier work this paper cites.
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot et al. , “Mastering the game of go with deep neural networks and tree search,” nature , vol. 529, no. 7587, p. 484, 2016
2016
Earlier work this paper cites.
S. Wang, W. Liu, J. Wu, L. Cao, Q. Meng, and P. J. Kennedy, “Training deep neural networks on imbalanced data sets,” in 2016 international joint conference on neural networks (IJCNN) . IEEE, 2016, pp. 4368–4374
2016
Earlier work this paper cites.
B. Poole, S. Lahiri, M. Raghu, J. Sohl-Dickstein, and S. Ganguli, “Exponential expressivity in deep neural networks through transient chaos,” in Advances in neural information processing systems , 2016, pp. 3360–3368
2016
Earlier work this paper cites.
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals, “Understanding deep learning requires rethinking generalization,” arXiv preprint arXiv:1611.03530 , 2016
Original
2016
Earlier work this paper cites.
N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, and P. T. P. Tang, “On large-batch training for deep learning: Generalization gap and sharp minima,” arXiv preprint arXiv:1609.04836 , 2016
Original
2016
Earlier work this paper cites.