Fetching the paper…
Reading the bibliography…
We point out a limitation of the mutual information neural estimation (MINE) where the network fails to learn at the initial training phase, leading to slow convergence in the number of training iterations.
A mathematical theory of communication
C. E. Shannon · 1948
Earlier work this paper cites.
Coding theorems for a discrete source with a fidelity criterion
C. E. Shannon · 1959
Earlier work this paper cites.
Asymptotic evaluation of certain markov process expectations for large time. iv
M. D. Donsker and S. S. Varadhan · 1983
Earlier work this paper cites.
Common randomness in information theory and cryptography—Part I: Secret sharing
R. Ahlswede and I. Csiszár · 1993
Earlier work this paper cites.
Estimation of mutual information using kernel density estimators
Y.-I. Moon, B. Rajagopalan, and U. Lall · 1995
Earlier work this paper cites.
The mutual information: Detecting and evaluating dependencies between variables
R. Steuer, J. Kurths, C. O. Daub, J. Weise, and J. Selbig · 2002
Earlier work this paper cites.
An introduction to variable and feature selection
I. Guyon and A. Elisseeff · 2003
Earlier work this paper cites.
Estimation of entropy and mutual information
L. Paninski · 2003
Earlier work this paper cites.
Estimating mutual information
A. Kraskov, H. Stögbauer, and P. Grassberger · 2004
Earlier work this paper cites.
Estimating optimal feature subsets using efficient estimation of high-dimensional mutual information
T. W. Chow and D. Huang · 2005
Earlier work this paper cites.
Feature selection based on mutual information: Criteria of max-dependency, max-relevance, and min-redundancy
H. Peng, F. Long, and C. Ding · 2005
Cited alongside, same era.
Normalized mutual information feature selection
P. A. Estévez, M. Tesmer, C. A. Perez, and J. M. Zurada · 2009
Cited alongside, same era.
The Elements of Statistical Learning: Data mining, Inference, and Prediction, springer series in statistics, 2009
T. Hastie, R. Tibshirani, and J. Friedman · 2009
Cited alongside, same era.
Noise-contrastive estimation of unnormalized statistical models, with applications to natural image statistics
M. U. Gutmann and A. Hyvärinen · 2012
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Cited alongside, same era.
Sgd learns over-parameterized networks that provably generalize on linearly separable data
A. Brutzkus, A. Globerson, E. Malach, and S. Shalev-Shwartz · 2017
Later among the works it cites.
Mutual information neural estimation
M. I. Belghazi, A. Baratin, S. Rajeshwar, S. Ozair, Y. Bengio, A. Courville, and D. Hjelm · 2018
Later among the works it cites.
Binder 2.0-reproducible, interactive, sharable environments for science at scale
P. Jupyter, M. Bussonnier, J. Forde, J. Freeman, B. Granger, T. Head, C. Holdgraf, K. Kelley, G. Nalvarte, A. Osheroff, et al · 2018
Later among the works it cites.
Revisiting small batch training for deep neural networks
D. Masters and C. Luschi · 2018
Later among the works it cites.
Representation learning with contrastive predictive coding
A. v. d. Oord, Y. Li, and O. Vinyals · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
L. Dinh, D. Krueger, and Y. Bengio · 2014
Cited alongside, same era.
Generative adversarial nets
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2014
Cited alongside, same era.
Density estimation using real NVP
L. Dinh, J. Sohl-Dickstein, and S. Bengio · 2016
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, et al · 2016
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2016
Cited alongside, same era.
Later among the works it cites.
Deep contextualized word representations
M. E. Peters, M. Neumann, M. Iyyer, M. Gardner, C. Clark, K. Lee, and L. Zettlemoyer · 2018
Later among the works it cites.
On variational lower bounds of mutual information
B. Poole, S. Ozair, A. van den Oord, A. A. Alemi, and G. Tucker · 2018
Later among the works it cites.
A survey on deep learning: Algorithms, techniques, and applications
S. Pouyanfar, S. Sadiq, Y. Yan, H. Tian, Y. Tao, M. P. Reyes, M.-L. Shyu, S.-C. Chen, and S. Iyengar · 2018
Later among the works it cites.
Learning deep representations by mutual information estimation and maximization
R. D. Hjelm, A. Fedorov, S. Lavoie-Marchildon, K. Grewal, P. Bachman, A. Trischler, and Y. Bengio · 2019
Closest in time.