Fetching the paper…
Reading the bibliography…
The matrix-based Renyi's \alpha-entropy functional and its multivariate extension were recently developed in terms of the normalized eigenspectrum of a Hermitian matrix of the projected data in a reproducing kernel Hilbert space (RKHS).
A. Rényi, “On measures of entropy and information,” in Proc. of the 4th Berkeley Sympos. on Math. Statist. and Prob. , vol. 1, 1961, pp. 547–561
1961
Earlier work this paper cites.
M. Hellman and J. Raviv, “Probability of error, equivocation, and the chernoff bound,” IEEE Transactions on Information Theory , vol. 16, no. 4, pp. 368–372, 1970
1970
Earlier work this paper cites.
B. W. Silverman, Density estimation for statistics and data analysis . CRC press, 1986, vol. 26
1986
Earlier work this paper cites.
R. W. Yeung, “A new outlook on shannon’s information measures,” IEEE transactions on information theory , vol. 37, no. 3, pp. 466–474, 1991
1991
Earlier work this paper cites.
Y.-I. Moon, B. Rajagopalan, and U. Lall, “Estimation of mutual information using kernel density estimators,” Physical Review E , vol. 52, no. 3, p. 2318, 1995
1995
Earlier work this paper cites.
D. Koller and M. Sahami, “Toward optimal feature selection,” Stanford InfoLab, Tech. Rep., 1996
1996
Earlier work this paper cites.
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,” Proceedings of the IEEE , vol. 86, no. 11, pp. 2278–2324, 1998
1998
Earlier work this paper cites.
J. Shi and J. Malik, “Normalized cuts and image segmentation,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 22, no. 8, pp. 888–905, 2000
2000
Earlier work this paper cites.
N. Tishby, F. C. Pereira, and W. Bialek, “The information bottleneck method,” arXiv preprint physics/0004057 , 2000
2000
Earlier work this paper cites.
A. J. Bell, “The co-information lattice,” in Proceedings of the Fifth International Workshop on Independent Component Analysis and Blind Signal Separation: ICA , vol. 2003, 2003
2003
Earlier work this paper cites.
L. Paninski, “Estimation of entropy and mutual information,” Neural computation , vol. 15, no. 6, pp. 1191–1253, 2003
2003
Earlier work this paper cites.
A. Kraskov, H. Stögbauer, and P. Grassberger, “Estimating mutual information,” Physical review E , vol. 69, no. 6, p. 066138, 2004
2004
Earlier work this paper cites.
S. Yaramakala and D. Margaritis, “Speculative markov blanket discovery for optimal feature selection,” in ICDM , 2005, pp. 809–812
2005
Earlier work this paper cites.
R. Bhatia, “Infinitely divisible matrices,” The American Mathematical Monthly , vol. 113, no. 3, pp. 221–235, 2006
2006
Earlier work this paper cites.
A. Krizhevsky and G. Hinton, “Learning multiple layers of features from tiny images,” Citeseer, Tech. Rep., 2009
2009
Earlier work this paper cites.
R. Jenssen, “Kernel entropy component analysis,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 32, no. 5, pp. 847–860, 2009
2009
Earlier work this paper cites.
2010
Earlier work this paper cites.
J. C. Principe, Information theoretic learning: Renyi’s entropy and kernel perspectives . Springer Science & Business Media, 2010
2010
Earlier work this paper cites.
G. Brown, A. Pocock, M.-J. Zhao, and M. Luján, “Conditional likelihood maximisation: a unifying framework for information theoretic feature selection,” JMLR , vol. 13, no. Jan, pp. 27–66, 2012
2012
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in NeurIPS , 2012, pp. 1097–1105
2012
Cited alongside, same era.
T. M. Cover and J. A. Thomas, Elements of information theory . John Wiley & Sons, 2012
2012
Cited alongside, same era.
M. Müller-Lennert, F. Dupuis, O. Szehr, S. Fehr, and M. Tomamichel, “On quantum rényi entropies: A new generalization and some properties,” J. Math. Phys. , vol. 54, no. 12, p. 122203, 2013
2013
Cited alongside, same era.
D. Lopez-Paz, P. Hennig, and B. Schölkopf, “The randomized dependence coefficient,” in NeurIPS , 2013, pp. 1–9
2013
Cited alongside, same era.
N. Bertschinger, J. Rauh, E. Olbrich, J. Jost, and N. Ay, “Quantifying unique information,” Entropy , vol. 16, no. 4, pp. 2161–2183, 2014
2014
Cited alongside, same era.
T. Tax, P. A. Mediano, and M. Shanahan, “The partial information decomposition of generative neural network models,” Entropy , vol. 19, no. 9, p. 474, 2017
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
M. Thoma, “The hasyv2 dataset,” arXiv preprint arXiv:1701.08380 , 2017
2017
Later among the works it cites.
A. Kolchinsky and B. Tracey, “Estimating mixture entropy with pairwise distances,” Entropy , vol. 19, no. 7, p. 361, 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
V. Griffith and C. Koch, “Quantifying synergistic mutual information,” in Guided Self-Organization: Inception . Springer, 2014, pp. 159–190
2014
Cited alongside, same era.
N. Timme, W. Alford, B. Flecker, and J. M. Beggs, “Synergy, redundancy, and multivariate information measures: an experimentalists perspective,” J. Comput. Neurosci. , vol. 36, no. 2, pp. 119–140, 2014
2014
Cited alongside, same era.
N. X. Vinh, J. Chan, and J. Bailey, “Reconsidering mutual information based feature selection: A statistical significance view,” in AAAI , 2014
2014
Cited alongside, same era.
J. R. Vergara and P. A. Estévez, “A review of feature selection methods based on mutual information,” Neural computing and applications , vol. 24, no. 1, pp. 175–186, 2014
2014
Cited alongside, same era.
N. Tishby and N. Zaslavsky, “Deep learning and the information bottleneck principle,” in IEEE ITW , 2015, pp. 1–5
2015
Cited alongside, same era.
L. G. Sanchez Giraldo, M. Rao, and J. C. Principe, “Measures of entropy from data using infinitely divisible kernels,” IEEE Transactions on Information Theory , vol. 61, no. 1, pp. 535–548, 2015
2015
Cited alongside, same era.
K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” in ICLR , 2015
2015
Cited alongside, same era.
2017
Later among the works it cites.
J.-H. Luo, J. Wu, and W. Lin, “Thinet: A filter level pruning method for deep neural network compression,” in ICCV , 2017, pp. 5058–5066
2017
Later among the works it cites.
H. Li, A. Kadav, I. Durdanovic, H. Samet, and H. P. Graf, “Pruning filters for efficient convnets,” in ICLR , 2017
2017
Later among the works it cites.
A. Achille and S. Soatto, “Emergence of invariance and disentanglement in deep representations,” JMLR , vol. 19, no. 1, pp. 1947–1980, 2018
2018
Closest in time.
A. M. Saxe et al. , “On the information bottleneck theory of deep learning,” in ICLR , 2018
2018
Closest in time.
H. Mureşan and M. Oltean, “Fruit recognition from images using deep learning,” Acta Universitatis Sapientiae, Informatica , vol. 10, no. 1, pp. 26–42, 2018
2018
Closest in time.
I. Sason and S. Verdú, “Arimoto–rényi conditional entropy and bayesian m m -ary hypothesis testing,” IEEE Transactions on Information Theory , vol. 64, no. 1, pp. 4–25, 2018
2018
Closest in time.
2018
Closest in time.
2018
Closest in time.
S. Yu and J. C. Principe, “Understanding autoencoders with information theoretic concepts,” Neural Networks , vol. 117, pp. 104–123, 2019
2019
Closest in time.
S. Yu, L. G. Sanchez Giraldo, R. Jenssen, and J. C. Principe, “Multivariate extension of matrix-based renyi’s α \alpha -order entropy functional,” IEEE Transactions on Pattern Analysis and Machine Intelligence , 2019
2019
Closest in time.
S. Yu and J. C. Príncipe, “Simple stopping criteria for information theoretic feature selection,” Entropy , vol. 21, no. 1, p. 99, 2019
2019
Closest in time.
M. Noshad, Y. Zeng, and A. O. Hero, “Scalable mutual information estimation using dependence graphs,” in ICASSP , 2019, pp. 2962–2966
2019
Closest in time.