Fetching the paper…
Reading the bibliography…
In this theory paper, we investigate training deep neural networks (DNNs) for classification via minimizing the information bottleneck (IB) functional.
A. Rényi, “On the dimension and entropy of probability distributions,” Acta Mathematica Hungarica , vol. 10, no. 1-2, pp. 193–215, Mar. 1959
1959
Earlier work this paper cites.
I. Csiszár, “On the dimension and entropy of order α \alpha of the mixture of probability distributions,” Acta Mathematica Hungarica , vol. 13, no. 3-4, pp. 245–255, 1962
1962
Earlier work this paper cites.
T. M. Cover and J. A. Thomas, Elements of Information Theory , 1st ed. New York, NY: John Wiley & Sons, Inc., 1991
1991
Earlier work this paper cites.
B. R. Hunt and V. Y. Kaloshin, “How projections affect the dimension spectrum of fractal measures,” Nonlinearity , vol. 10, no. 5, p. 1031, 1997
1997
Earlier work this paper cites.
N. Tishby, F. C. Pereira, and W. Bialek, “The information bottleneck method,” in Proc. Allerton Conf. on Communication, Control, and Computing , Monticello, IL, Sep. 1999, pp. 368–377
1999
Earlier work this paper cites.
P. Mattila, M. Morán, and J.-M. Rey, “Dimension of a measure,” Studia Math , vol. 142, no. 3, pp. 219–233, 2000
2000
Earlier work this paper cites.
Y. Wu and S. Verdú, “Rényi information dimension: Fundamental limits of almost lossless analog compression,” IEEE Trans. Inf. Theory , vol. 56, no. 8, pp. 3721–3748, Aug. 2010
2011
Earlier work this paper cites.
H. Xu and S. Mannor, “Robustness and generalization,” Machine Learning , vol. 86, no. 3, pp. 391–423, Mar. 2012
2012
Earlier work this paper cites.
Y. Bengio, A. Courville, and P. Vincent, “Representation learning: A review and new perspectives,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 35, no. 8, pp. 1798–1828, Aug. 2013
2013
Earlier work this paper cites.
T. He, Y. Fan, Y. Qian, T. Tan, and K. Yu, “Reshaping deep neural network for fast decoding by node-pruning,” in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , May 2014, pp. 245–249
2014
Earlier work this paper cites.
N. Tishby and N. Zaslavsky, “Deep learning and the information bottleneck principle,” in Proc. IEEE Information Theory Workshop (ITW) , Jerusalem, Apr. 2015, pp. 1–5
2015
Earlier work this paper cites.
Y. Wu, S. Shamai (Shitz), and S. Verdú, “Information dimension and the degrees of freedom of the interference channel,” IEEE Trans. Inf. Theory , vol. 61, no. 1, pp. 256–279, Jan. 2015
2015
Earlier work this paper cites.
I. Goodfellow, Y. Bengio, and A. Courville, Deep Learning . MIT Press, 2016, www.deeplearningbook.org
2016
Earlier work this paper cites.
D. J. Strouse and D. J. Schwab, “The deterministic information bottleneck,” in Proc. of the Thirty-Second Conference on Uncertainty in Artificial Intelligence , Jun. 2016, pp. 696–705
2016
Earlier work this paper cites.
R. Liao, A. Schwing, R. Zemel, and R. Urtasun, “Learning deep parsimonious representations,” in Proc. Advances in Neural Information Processing Systems 29 , Barcelona, Dec. 2016, pp. 5076–5084
2016
Cited alongside, same era.
2017
Cited alongside, same era.
A. A. Alemi, I. Fischer, J. V. Dillon, and K. Murphy, “Deep variational information bottleneck,” in Proc. International Conference on Learning Representations (ICLR) , Toulon, Apr. 2017
2017
Cited alongside, same era.
H. Li, A. Kadav, I. Durdanovic, H. Samet, and H. P. Graf, “Pruning filters for efficient convnets,” in Proc. International Conference on Learning Representations (ICLR) , Toulon, Apr. 2017
2017
Cited alongside, same era.
P. K. Banerjee and G. Montufar, “The variational deficiency bottleneck,” arXiv:1810.11677v1 [cs.LG]
2018
Closest in time.
B. C. Geiger and G. Kubin, Information Loss in Deterministic Signal Processing Systems , ser. Understanding Complex Systems. Springer, 2018
2018
Closest in time.
R. Novak, Y. Bahri, D. A. Abolafia, J. Pennington, and J. Sohl-Dickstein, “Sensitivity and generalization in neural networks: an empirical study,” in Proc. International Conference on Learning Representations (ICLR) , Vancouver, May 2018
2018
Closest in time.
T. Zahavy, B. Kang, A. Sivak, J. Feng, H. Xu, and S. Mannor, “Ensemble robustness and generalization of stochastic deep learning algorithms,” in Proc. International Conference on Learning Representations (ICLR) , Vancouver, May 2018
2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
T. DeVries and G. W. Taylor, “Dataset augmentation in feature space,” in Proc. International Conference on Learning Representations (ICLR) , Toulon, Apr. 2017
2017
Cited alongside, same era.
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals, “Understanding deep learning requires rethinking generalization,” in Proc. International Conference on Learning Representations (ICLR) , Toulon, Apr. 2017
2017
Cited alongside, same era.
D. Arpit, S. Jastrzębski, N. Ballas, D. Krueger, E. Bengio, M. S. Kanwal, T. Maharaj, A. Fischer, A. Courville, Y. Bengio, and S. Lacoste-Julien, “A closer look at memorization in deep networks,” in Proc. 34th International Conference on Machine Learning (ICML) , Sydney, Aug. 2017, pp. 233–242
2017
Cited alongside, same era.
G. Pereyra, G. Tucker, J. Chorowski, L. Kaiser, and G. Hinton, “Regularizing neural networks by penalizing confident output distributions,” in Proc. International Conference on Learning Representations (ICLR) , Toulon, Apr. 2017
2017
Cited alongside, same era.
2017
Cited alongside, same era.
A. M. Saxe, Y. Bansal, J. Dapello, M. Advani, A. Kolchinsky, B. D. Tracey, and D. D. Cox, “On the information bottleneck theory of deep learning,” in Proc. International Conference on Learning Representations (ICLR) , Vancouver, May 2018
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Closest in time.
A. Achille and S. Soatto, “Emergence of invariance and disentanglement in deep representations,” Journal of Machine Learning Research , vol. 19, no. 50, pp. 1–34, 2018
2018
Closest in time.
J.-H. Jacobsen, A. W. Smeulders, and E. Oyallon, “i-revnet: Deep invertible networks,” in Proc. International Conference on Learning Representations (ICLR) , Vancouver, May 2018
2018
Closest in time.
A. Achille and S. Soatto, “Information dropout: Learning optimal representations through noisy computation,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 40, no. 12, pp. 2897–2905, Dec. 2018
2018
Closest in time.
B. Chang, L. Meng, E. Haber, L. Ruthotto, D. Begert, and E. Holtham, “Reversible architectures for arbitrarily deep residual neural networks,” in Proc. AAAI Conference on Artificial Intelligence , New Orleans, Feb. 2018
2018
Closest in time.
2018
Closest in time.
M. Vera, P. Piantanida, and L. R. Vega, “The role of the information bottleneck in representation learning,” in Proc. IEEE International Symposium on Information Theory (ISIT) , Vail, CO, Jun. 2018, pp. 1580–1584
2018
Closest in time.
D. Moyer, S. Gao, R. Brekelmans, A. Galstyan, and G. Ver Steeg, “Invariant representations without adversarial training,” in Proc. Advances in Neural Information Processing Systems 31 , Montreal, Dec. 2018, pp. 9102–9111
2018
Closest in time.
J. Frankle and M. Carbin, “The lottery ticket hypothesis: Finding sparse, trainable neural networks,” in Proc. International Conference on Learning Representations (ICLR) , New Orleans, May 2019
2019
Closest in time.
A. Kolchinsky, B. D. Tracey, and S. Van Kuyk, “Caveats for information bottleneck in deterministic scenarios,” in Proc. International Conference on Learning Representations (ICLR) , New Orleans, May 2019
2019
Closest in time.