Fetching the paper…
Reading the bibliography…
Whatever information a deep neural network has gleaned from training data is encoded in its weights.
Theory of statistical estimation
Fisher, R. A · 1925
Earlier work this paper cites.
A mathematical theory of communication
Shannon, C. E · 1948
Earlier work this paper cites.
Keeping the neural networks simple by minimizing the description length of the weights
Hinton, G. E. and Van Camp, D · 1993
Earlier work this paper cites.
Flat minima
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Mutual information, fisher information, and population coding
Brunel, N. and Nadal, J.-P · 1998
Earlier work this paper cites.
The information bottleneck method
Tishby, N., Pereira, F. C., and Bialek, W · 1999
Earlier work this paper cites.
Stability and generalization
Bousquet, O. and Elisseeff, A · 2002
Earlier work this paper cites.
Kramers’ law: Validity, derivations and generalisations
Berglund, N · 2011
Earlier work this paper cites.
Elements of information theory
Cover, T. M. and Thomas, J. A · 2012
Earlier work this paper cites.
beta-vae: Learning basic visual concepts with a constrained variational framework
Higgins, I., Matthey, L., Pal, A., Burgess, C., Glorot, X., Botvinick, M., Mohamed, S., and Lerchner, A · 2013
Earlier work this paper cites.
A pac-bayesian tutorial with a dropout bound
McAllester, D · 2013
Earlier work this paper cites.
New insights and perspectives on the natural gradient method
Martens, J · 2014
Earlier work this paper cites.
Striving for simplicity: The all convolutional net
Springenberg, J. T., Dosovitskiy, A., Brox, T., and Riedmiller, M · 2014
Earlier work this paper cites.
Train faster, generalize better: Stability of stochastic gradient descent
Hardt, M., Recht, B., and Singer, Y · 2015
Earlier work this paper cites.
Variational dropout and the local reparameterization trick
Kingma, D. P., Salimans, T., and Welling, M · 2015
Earlier work this paper cites.
Information geometry of the gaussian distribution in view of stochastic optimization
Malagò, L. and Pistone, G · 2015
Cited alongside, same era.
Deep learning and the information bottleneck principle
Tishby, N. and Zaslavsky, N · 2015
Cited alongside, same era.
Deep variational information bottleneck
Alemi, A. A., Fischer, I., Dillon, J. V., and Murphy, K · 2016
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P · 2016
Cited alongside, same era.
Understanding disentangling in β \beta -vae
Burgess, C. P., Higgins, I., Pal, A., Matthey, L., Watters, N., Desjardins, G., and Lerchner, A · 2017
Cited alongside, same era.
Optimization methods for large-scale machine learning
Bottou, L., Curtis, F. E., and Nocedal, J · 2018
Later among the works it cites.
Estimating information flow in neural networks
Goldfeld, Z., Berg, E. v. d., Greenewald, K., Melnyk, I., Nguyen, N., Kingsbury, B., and Polyanskiy, Y · 2018
Later among the works it cites.
Variational bayesian dropout: pitfalls and fixes
Hron, J., Matthews, A. G. d. G., and Ghahramani, Z · 2018
Later among the works it cites.
β \beta -bnn: A rate-distortion perspective on bayesian neural networks
Hu, S. X., Champs-sur Marne, F., Moreno, P. G., Lawrence, N., and Damianou, A · 2018
Later among the works it cites.
Visualizing the loss landscape of neural nets
Li, H., Xu, Z., Taylor, G., Studer, C., and Goldstein, T · 2018
Later among the works it cites.
Generalization error bounds for noisy, iterative algorithms
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Entropy-sgd: Biasing gradient descent into wide valleys
Chaudhari, P., Choromanska, A., Soatto, S., LeCun, Y., Baldassi, C., Borgs, C., Chayes, J., Sagun, L., and Zecchina, R · 2017
Cited alongside, same era.
Sharp minima can generalize for deep nets
Dinh, L., Pascanu, R., Bengio, S., and Bengio, Y · 2017
Cited alongside, same era.
Computing nonvacuous generalization bounds for deep (stochastic) neural networks with many more parameters than training data
Dziugaite, G. K. and Roy, D. M · 2017
Cited alongside, same era.
Stochastic modified equations and adaptive stochastic gradient algorithms
Li, Q., Tai, C., and E, W · 2017
Cited alongside, same era.
Variational dropout sparsifies deep neural networks
Molchanov, D., Ashukha, A., and Vetrov, D · 2017
Cited alongside, same era.
Opening the Black Box of Deep Neural Networks via Information
Shwartz-Ziv, R. and Tishby, N · 2017
Cited alongside, same era.
Information-theoretic analysis of generalization capability of learning algorithms
Xu, A. and Raginsky, M · 2017
Cited alongside, same era.
Pensia, A., Jog, V., and Loh, P.-L · 2018
Later among the works it cites.
Rezende, D. J. and Viola, F · 2018
Later among the works it cites.
On the information bottleneck theory of deep learning
Saxe, A. M., Bansal, Y., Dapello, J., Advani, M., Kolchinsky, A., Tracey, B. D., and Cox, D. D · 2018
Later among the works it cites.
The Information Complexity of Learning Tasks, their Structure and their Distance
Achille, A., Paolini, G., Mbeng, G., and Soatto, S · 2019
Closest in time.
Critical learning periods in deep networks
Achille, A., Rovere, M., and Soatto, S · 2019
Closest in time.
Quantitative central limit theorems for discrete stochastic processes
Cheng, X., Bartlett, P. L., and Jordan, M. I · 2019
Closest in time.
Information-theoretic generalization bounds for sgld via data-dependent estimates
Negrea, J., Haghifam, M., Dziugaite, G. K., Khisti, A., and Roy, D. M · 2019
Closest in time.
How much does your data exploration overfit? controlling bias via information usage
Russo, D. and Zou, J · 2019
Closest in time.
On the convex behavior of deep neural networks in relation to the layers’ width
Littwin, E. and Wolf, L · 2020
Closest in time.
Xie, Z., Sato, I., and Sugiyama, M · 2020
Closest in time.