Fetching the paper…
Reading the bibliography…
We review the current literature concerned with information plane analyses of neural network classifiers.
D. Messerschmitt, “Quantizing for maximum output entropy (corresp.),” IEEE Transactions on Information Theory , vol. 17, no. 5, pp. 612–612, Sep. 1971
1971
Earlier work this paper cites.
T. M. Cover and J. A. Thomas, Elements of Information Theory , 1st ed. New York, NY: John Wiley & Sons, Inc., 1991
1991
Earlier work this paper cites.
T. S. Han and S. Verdú, “Generalizing the Fano inequality,” IEEE Trans. Inf. Theory , vol. 40, no. 4, pp. 1247–1251, Jul. 1994
1994
Earlier work this paper cites.
A. Kraskov, H. Stögbauer, and P. Grassberger, “Estimating mutual information,” Physical Review E , vol. 69, no. 6, p. 066138, 2004
2004
Earlier work this paper cites.
M. Basirat, B. C. Geiger, and P. M. Roth, “A geometric perspective on information plane analysis,” Entropy , vol. 23, no. 6, p. 711, Jun. 2021, open-access
2009
Earlier work this paper cites.
D. P. Kingma and M. Welling, “Auto-encoding variational Bayes,” arXiv:1312.6114
2013
Earlier work this paper cites.
N. Tishby and N. Zaslavsky, “Deep learning and the information bottleneck principle,” in Proc. IEEE Information Theory Workshop (ITW) , Jerusalem, Apr. 2015, pp. 1–5
2015
Earlier work this paper cites.
2017
Earlier work this paper cites.
A. A. Alemi, I. Fischer, J. V. Dillon, and K. Murphy, “Deep variational information bottleneck,” in Proc. International Conference on Learning Representations (ICLR) , Toulon, Apr. 2017
2017
Earlier work this paper cites.
W. Samek, G. Montavon, A. V. andLars Kai Hansen, and K.-R. Müller, Eds., Explainable AI: Interpreting, Explaining and Visualizing Deep Learning , ser. Lecture Notes in Artificial Intelligence. Springer, 2017
2017
Earlier work this paper cites.
A. Kolchinsky and B. D. Tracey, “Estimating mixture entropy with pairwise distances,” Entropy , vol. 19, no. 7, p. 361, Jul. 2017
2017
Earlier work this paper cites.
T. Salimans, A. Karpathy, X. Chen, and D. P. Kingma, “PixelCNN++: Improving the PixelCNN with discretized logistic mixture likelihood and other modifications,” in Proc. Int. Conf. on Learning Representations (ICLR) , Toulon, Apr. 2017
2017
Earlier work this paper cites.
E. R. Balda, A. Behboodi, and R. Mathar, “An information theoretic view on learning of artificial neural networks,” in Proc. Int. Conf. on Signal Processing and Communication Systems (ICSPCS) , Dec. 2018, pp. 1–8
2018
Earlier work this paper cites.
——, “On the trajectory of stochastic gradient descent in the information plane,” 2018. [Online]. Available: openreview.net/forum?id=SkMON20ctX
2018
Earlier work this paper cites.
L. N. Darlow and A. Storkey, “What information does a resnet compress?” 2018. [Online]. Available: openreview.net/forum?id=HklbTjRcKX
2018
Earlier work this paper cites.
M. I. Belghazi, A. Baratin, S. Rajeshwar, S. Ozair, Y. Bengio, A. Courville, and D. Hjelm, “Mutual information neural estimation,” in Proc. Int. Conf. on Machine Learning (ICML) , Stockholm, Jul. 2018, pp. 531–540
2018
Earlier work this paper cites.
H. Fang, V. Wang, and M. Yamaguchi, “Dissecting deep learning networks – visualizing mutual information,” Entropy , vol. 20, no. 8, p. 823, Aug. 2018
2018
Earlier work this paper cites.
M. Gabrié, A. Manoel, C. Luneau, J. Barbier, N. Macris, F. Krzakala, and L. Zdeborová, “Entropy and mutual information in models of deep neural networks,” in Proc. Advances in Neural Information Processing Systems (NeurIPS) , Montreal, Dec. 2018, pp. 1821–1831
2018
Earlier work this paper cites.
A. M. Saxe, Y. Bansal, J. Dapello, M. Advani, A. Kolchinsky, B. D. Tracey, and D. D. Cox, “On the information bottleneck theory of deep learning,” in Proc. Int. Conf. on Learning Representations (ICLR) , Vancouver, May 2018
2018
Earlier work this paper cites.
M. Vera, P. Piantanida, and L. R. Vega, “The role of the information bottleneck in representation learning,” in Proc. IEEE Int. Sym. on Information Theory (ISIT) , Vail, CO, Jun. 2018, pp. 1580–1584
2018
Earlier work this paper cites.
H. Hafez-Kolahi and S. Kasaei, “Information bottleneck and its applications in deep learning,” Journal on Information Systems and Telecommunications , vol. 6, no. 3, pp. 2897–127, Jul. 2018
2018
Cited alongside, same era.
A. Achille and S. Soatto, “Information dropout: Learning optimal representations through noisy computation,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 40, no. 12, pp. 2897–2905, Dec. 2018
2018
Cited alongside, same era.
A. S. Morcos, D. G. Barrett, N. C. Rabinowitz, and M. Botvinick, “On the importance of single directions for generalization,” in Proc. Int. Conf. on Learning Representations (ICLR) , Vancouver, May 2018
2018
Cited alongside, same era.
A. Jacot, F. Gabriel, and C. Hongler, “Neural tangent kernel: convergence and generalization in neural networks,” in Proc. Advances in Neural Information Processing Systems (NeurIPS) , Montreal, Dec. 2018, pp. 8580–8589
2018
Cited alongside, same era.
2019
Later among the works it cites.
S. Yu and J. C. Príncipe, “Understanding autoencoders with information theoretic concepts,” Neural Networks , vol. 117, pp. 104–123, 2019
2019
Later among the works it cites.
J. Frankle and M. Carbin, “The lottery ticket hypothesis: Training pruned neural networks,” in Proc. Int. Conf. on Learning Representations (ICLR) , New Orleans, May 2019
2019
Later among the works it cites.
B. Davis, U. Bhatt, K. Bhardwaj, R. Marculescu, and J. Moura, “NIF: A framework for quantifying neural information flow in deep networks,” in Workshop on Network Interpretability for Deep Learning @ AAAI Conf. on Artificial Intelligence (AAAI) , Honolulu, Jan. 2019
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
G. Montavon, W. Samek, and K.-R. Müller, “Methods for interpreting and understanding deep neural networks,” Digital Signal Processing , vol. 73, pp. 1–15, 2018
2018
Cited alongside, same era.
A. Adadi and M. Berrada, “Peeking inside the black-box: A survey on explainable artificial intelligence (XAI),” IEEE Access , vol. 6, pp. 52 138–52 160, 2018
2018
Cited alongside, same era.
2018
Cited alongside, same era.
A. Achille and S. Soatto, “Emergence of invariance and disentanglement in deep representations,” Journal of Machine Learning Research , vol. 19, pp. 1–34, 2018
2018
Cited alongside, same era.
J.-H. Jacobsen, A. W. Smeulders, and E. Oyallon, “i-RevNet: Deep invertible networks,” in Proc. Int. Conf. on Learning Representations (ICLR) , Vancouver, May 2018
2018
Cited alongside, same era.
B. Chang, L. Meng, E. Haber, L. Ruthotto, D. Begert, and E. Holtham, “Reversible architectures for arbitrarily deep residual neural networks,” in Proc. AAAI Conf. on Artificial Intelligence , New Orleans, Feb. 2018
2018
Cited alongside, same era.
H. Cheng, D. Lian, S. Gao, and Y. Geng, “Evaluating capability of deep neural networks for image classification via information plane,” in Proc. European Conf. on Computer Vision (ECCV) , Munich, Sep. 2018, pp. 181–195
2018
Cited alongside, same era.
I. Chelombiev, C. J. Houghton, and C. O’Donnel, “Adaptive estimators show information compression in deep neural networks,” in Proc. Int. Conf. on Learning Representations (ICLR) , New Orleans, May 2019
2019
Cited alongside, same era.
A. Kolchinsky, B. D. Tracey, and D. H. Wolpert, “Nonlinear information bottleneck,” Entropy , vol. 21, no. 12, p. 1181, Dec. 2019
2019
Later among the works it cites.
B. Poole, S. Ozari, A. van den Oord, A. A. Alemi, and G. Tucker, “On variational bounds of mutual information,” in Proc. Int. Conf. on Machine Learning (ICML) , Long Beach, California, USA, Jun. 2019, pp. 5171–5180
2019
Later among the works it cites.
V. Abrol and J. Tanner, “Information-bottleneck under mean field initialization,” in Workshop Uncertainty & Robustness in Deep Learning at Int. Conf. on Machine Learning (ICML) , online, Jul. 2020
2020
Closest in time.
H. Jónsson, G. Cherubini, and E. Eleftheriou, “Convergence behavior of DNNs with mutual-information-based regularization,” Entropy , vol. 22, no. 7, p. 727, Jul. 2020
2020
Closest in time.
A. Kirsch, C. Lyle, and Y. Gal, “Scalable training with information bottleneck objectives,” in Workshop Uncertainty & Robustness in Deep Learning at Int. Conf. on Machine Learning (ICML) , online, Jul. 2020
2020
Closest in time.
2020
Closest in time.
S. Yu, K. Wickstrøm, R. Jenssen, and J. C. Príncipe, “Understanding convolutional neural networks with information theory: An initial exploration,” IEEE Trans. Neural Netw. Learn. Syst. , pp. 1–8, 2020
2020
Closest in time.
R. A. Amjad and B. C. Geiger, “Learning representations for neural network-based classification using the information bottleneck principle,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 42, no. 9, pp. 2225–2239, Sep. 2020
2020
Closest in time.
Z. Goldfeld and Y. Polyanskiy, “The information bottleneck problem and its applications in machine learning,” IEEE Journal on Selected Areas in Information Theory , vol. 1, no. 1, pp. 19–38, May 2020
2020
Closest in time.
X. Li, C. C. Cao, Y. Shi, W. Bai, H. Gao, L. Qiu, C. Wang, Y. Gao, S. Zhang, X. Xue, and L. Chen, “A survey of data-driven and knowledge-aware explainable AI,” IEEE Trans. Knowl. Data Eng. , pp. 1–1, 2020
2020
Closest in time.
N. I. Tapia and P. A. Estévez, “On the information plane of autoencoders,” in Proc. Int. Joint Conf. on Neural Networks (IJCNN) , virtual, Jul. 2020, pp. 1–8
2020
Closest in time.
A. Kirsch, C. Lyle, and Y. Gal, “Learning CIFAR-10 with a simple entropy estimator using information bottleneck objectives,” in Workshop Uncertainty & Robustness in Deep Learning at Int. Conf. on Machine Learning (ICML) , online, Jul. 2020
2020
Closest in time.
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
J. Li and D. Liu, “Information bottleneck theory on convolutional neural networks,” Neural Processing Letters , Feb. 2021
2021
Closest in time.