Fetching the paper…
Reading the bibliography…
Information Theory (IT) has been used in Machine Learning (ML) from early days of this field.
M. A. RA Fisher, “On the mathematical foundations of theoretical statistics,” Phil. Trans. R. Soc. Lond. A
1922
Earlier work this paper cites.
B. O. Koopman, “On distributions admitting a sufficient statistic,” Transactions of the American Mathematical society
1936
Earlier work this paper cites.
C. E. Shannon, “A mathematical theory of communication,” Bell system technical journal
1948
Earlier work this paper cites.
E. L. Lehmann and H. Scheffé, “Completeness, Similar Regions, and Unbiased Estimation,” in Bulletin of the American Mathematical Society
1948
Earlier work this paper cites.
S. Kullback and R. A. Leibler, “On information and sufficiency,” The annals of mathematical statistics
1951
Earlier work this paper cites.
A. Kolmogorov, “On the Shannon theory of information transmission in the case of continuous signals,” IRE Transactions on Information Theory
1956
Earlier work this paper cites.
R. M. Dudley, “The sizes of compact subsets of Hilbert space and continuity of Gaussian processes,” Journal of Functional Analysis
1967
Earlier work this paper cites.
V. N. Vapnik and A. Y. Chervonenkis, “On the Uniform Convergence of Relative Frequencies of Events to Their Probabilities,” Theory of Probability and its Applications
1971
Earlier work this paper cites.
R. Blahut, “Computation of channel capacity and rate-distortion functions,” IEEE transactions on Information Theory
1972
Earlier work this paper cites.
S. Arimoto, “An algorithm for computing the capacity of arbitrary discrete memoryless channels,” IEEE Transactions on Information Theory
1972
Earlier work this paper cites.
1986
Earlier work this paper cites.
T. M. Cover, P. Gacs, and R. M. Gray, “Kolmogorov’s Contributions to Information Theory and Algorithmic Complexity,” The annals of probability
1989
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Flat minima,” Neural Computation
1997
Earlier work this paper cites.
N. Tishby, F. Pereira, and W. Bialek, “The information bottleneck method,” in Proceedings of the 37-th Annual Allerton Conference on Communication, Control and Computing
1999
Earlier work this paper cites.
N. Friedman, O. Mosenzon, N. Slonim, and N. Tishby, “Multivariate information bottleneck,” in Proceedings of the Seventeenth Conference on Uncertainty in Artificial Intelligence
2001
Earlier work this paper cites.
R. Gilad-Bachrach, A. Navot, and N. Tishby, “An information theoretic tradeoff between complexity and accuracy,” in Learning Theory and Kernel Machines
2003
Earlier work this paper cites.
G. Chechik, A. Globerson, N. Tishby, and Y. Weiss, “Information bottleneck for Gaussian variables,” Journal of machine learning research
2005
Earlier work this paper cites.
O. Shamir, S. Sabato, and N. Tishby, “Learning and generalization with the information bottleneck,” Theoretical Computer Science
2010
Earlier work this paper cites.
John Wiley & Sons, 2012
T. M. Cover and J. A. Thomas, Elements of Information Theory · 2012
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in Advances in Neural Information Processing Systems
2012
Cited alongside, same era.
D. P. Kingma and M. Welling, “Auto-encoding variational bayes,” arXiv preprint arXiv:1312.6114
2013
Cited alongside, same era.
Y. Bengio, L. Yao, G. Alain, and P. Vincent, “Generalized denoising auto-encoders as generative models,” in Advances in Neural Information Processing Systems
2013
Cited alongside, same era.
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “Dropout: A simple way to prevent neural networks from overfitting,” The Journal of Machine Learning Research
2014
Cited alongside, same era.
A. A. Alemi, I. Fischer, J. V. Dillon, and K. Murphy, “Deep Variational Information Bottleneck,” International Conference on Learning Representations
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
R. Vidal, J. Bruna, R. Giryes, and S. Soatto, “Mathematics of Deep Learning,” arXiv:1712.04741 [cs]
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2015
Cited alongside, same era.
2015
Cited alongside, same era.
2015
Cited alongside, same era.
2016
Cited alongside, same era.
2016
Cited alongside, same era.
2016
Cited alongside, same era.
2016
Cited alongside, same era.
2016
Cited alongside, same era.
2017
Later among the works it cites.
D. J. Im, S. Ahn, R. Memisevic, and Y. Bengio, “Denoising Criterion for Variational Auto-Encoding Framework.,” in AAAI
2017
Later among the works it cites.
2017
Later among the works it cites.
A. Xu and M. Raginsky, “Information-theoretic analysis of generalization capability of learning algorithms,” in Advances in Neural Information Processing Systems
2017
Later among the works it cites.
R. Bassily, S. Moran, I. Nachum, J. Shafer, and A. Yehudayoff, “Learners that Use Little Information,” in Algorithmic Learning Theory
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
A. M. Saxe, Y. Bansal, J. Dapello, M. Advani, A. Kolchinsky, B. D. Tracey, and D. D. Cox, “On the Information Bottleneck Theory of Deep Learning,” International Conference on Learning Representations
2018
Later among the works it cites.
A. Achille and S. Soatto, “Information Dropout: Learning Optimal Representations Through Noisy Computation,” IEEE Transactions on Pattern Analysis and Machine Intelligence
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
M. Vera, P. Piantanida, and L. R. Vega, “The Role of the Information Bottleneck in Representation Learning,” in 2018 IEEE International Symposium on Information Theory (ISIT)
2018
Later among the works it cites.