Fetching the paper…
Reading the bibliography…
Mutual Information (MI) has been widely used as a loss regularizer for training neural networks.
A mathematical theory of communication
Shannon, C. E · 1948
Earlier work this paper cites.
Remarks on some nonparametric estimates of a density function
Rosenblatt, M · 1956
Earlier work this paper cites.
On estimation of a probability density function and mode
Parzen, E · 1962
Earlier work this paper cites.
On the evaluation of an unknown probability density function, the direct estimation of the entropy from independent observations of a continuous random variable, and the distribution-free entropy test of goodness-of-fit
Tarasenko, F · 1968
Earlier work this paper cites.
A nonparametric estimation of the entropy for absolutely continuous distributions
Ahmad, I. and Lin, P.-E · 1976
Earlier work this paper cites.
Density-free convergence properties of various estimators of entropy
Györfi, L. and Van der Meulen, E. C · 1987
Earlier work this paper cites.
Sample estimate of the entropy of a random vector
Kozachenko, L. and Leonenko, N. N · 1987
Earlier work this paper cites.
Estimation of entropy and other functionals of a multivariate density
Joe, H · 1989
Earlier work this paper cites.
How to generate ordered maps by maximizing the mutual information between input and output signals
Linsker, R · 1989
Earlier work this paper cites.
Unsupervised classifiers, mutual information and 'phantom targets
Bridle, J., Heading, A., and MacKay, D · 1992
Earlier work this paper cites.
On the estimation of entropy
Hall, P. and Morton, S · 1993
Earlier work this paper cites.
Estimation of Integral Functionals of a Density
Birge, L. and Massart, P · 1995
Earlier work this paper cites.
Root-n consistent estimators of entropy for densities with unbounded support
Tsybakov, A. B. and Van der Meulen, E. C · 1996
Earlier work this paper cites.
Empirical entropy manipulation for real-world problems
Viola, P., Schraudolph, N. N., and Sejnowski, T. J · 1996
Earlier work this paper cites.
Nonparametric entropy estimation: An overview
Beirlant, J., Dudewicz, E. J., Györfi, L., and Van der Meulen, E. C · 1997
Earlier work this paper cites.
The IM algorithm: A variational approach to information maximization
Barber, D. and Agakov, F · 2003
Earlier work this paper cites.
Estimation of entropy and mutual information
Paninski, L · 2003
Earlier work this paper cites.
Estimating mutual information
Kraskov, A., Stögbauer, H., and Grassberger, P · 2004
Earlier work this paper cites.
Gradient-based manipulation of nonparametric entropy estimates
Schraudolph, N. N · 2004
Earlier work this paper cites.
Information theoretic learning
Principe, J. C., Xu, D., Fisher, J., and Haykin, S · 2006
Earlier work this paper cites.
Causality detection based on information-theoretic approaches in time series analysis
Hlaváčková-Schindler, K., Paluš, M., Vejmelka, M., and Bhattacharya, J · 2007
Earlier work this paper cites.
Approximating mutual information by maximum likelihood density ratio estimation
Suzuki, T., Sugiyama, M., Sese, J., and Kanamori, T · 2008
Earlier work this paper cites.
Information-theoretic methods
Torkkola, K · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A., Hinton, G., et al · 2009
Earlier work this paper cites.
Universal estimation of information measures for analog sources
Wang, Q., Kulkarni, S. R., and Verdú, S · 2009
Earlier work this paper cites.
Information measures in perspective
Ebrahimi, N., Soofi, E. S., and Soyer, R · 2010
Earlier work this paper cites.
Discriminative clustering by regularized information maximization
Krause, A., Perona, P., and Gomes, R · 2010
Earlier work this paper cites.
MNIST handwritten digit database
LeCun, Y. and Cortes, C · 2010
Cited alongside, same era.
Estimating divergence functionals and the likelihood ratio by convex risk minimization
Nguyen, X., Wainwright, M. J., and Jordan, M. I · 2010
Cited alongside, same era.
An analysis of single-layer networks in unsupervised feature learning
Coates, A., Ng, A., and Lee, H · 2011
Cited alongside, same era.
Reading digits in natural images with unsupervised feature learning
Netzer, Y., Wang, T., Coates, A., Bissacco, A., Wu, B., and Ng, A. Y · 2011
Cited alongside, same era.
Exponential concentration for mutual information estimation with application to forests
Liu, H., Wasserman, L., and Lafferty, J · 2012
Cited alongside, same era.
Ensemble estimators for multivariate entropy estimation
Sricharan, K., Wei, D., and Hero, A. O · 2013
EMI: exploration with mutual information
Kim, H., Kim, J., Jeong, Y., Levine, S., and Song, H. O · 2019
Later among the works it cites.
Roberta: A robustly optimized bert pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 2019
Later among the works it cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S · 2019
Later among the works it cites.
On variational bounds of mutual information
Poole, B., Ozair, S., Van Den Oord, A., Alemi, A., and Tucker, G · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Exponential concentration of a density functional estimator
Singh, S. and Poczos, B · 2014
Cited alongside, same era.
Nonparametric von mises estimators for entropies, divergences and mutual informations
Kandasamy, K., Krishnamurthy, A., Poczos, B., Wasserman, L., and robins, j. m · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Cited alongside, same era.
The winograd schema challenge: Evaluating progress in commonsense reasoning
Morgenstern, L. and Ortiz, C · 2015
Cited alongside, same era.
Deep variational information bottleneck
Alemi, A. A., Fischer, I., Dillon, J. V., and Murphy, K · 2016
Cited alongside, same era.
Demographic dialectal variation in social media: A case study of african-american english
Blodgett, S. L., Green, L., and O’Connor, B · 2016
Cited alongside, same era.
Model-based active exploration
Shyam, P., Jaśkowski, W., and Gomez, F · 2019
Later among the works it cites.
Empirical estimation of information measures: A literature guide
Verdú, S · 2019
Later among the works it cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R · 2019
Later among the works it cites.
Neural network acceptability judgments
Warstadt, A., Singh, A., and Bowman, S. R · 2019
Later among the works it cites.
Huggingface’s transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., et al · 2019
Later among the works it cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Yang, Z., Dai, Z., Yang, Y., Carbonell, J., Salakhutdinov, R., and Le, Q. V · 2019
Later among the works it cites.
InfoVAE: Information maximizing variational autoencoders
Zhao, S., Song, J., and Ermon, S · 2019
Later among the works it cites.
Fine-tuning pretrained language models: Weight initializations, data orders, and early stopping
Dodge, J., Ilharco, G., Schwartz, R., Farhadi, A., Hajishirzi, H., and Smith, N · 2020
Later among the works it cites.
Unsupervised multi-target domain adaptation: An information theoretic approach
Gholami, B., Sahu, P., Rudovic, O., Bousmalis, K., and Pavlovic, V · 2020
Later among the works it cites.
Mixout: Effective regularization to finetune large-scale pretrained language models
Lee, C., Cho, K., and Kang, W · 2020
Later among the works it cites.
Formal limitations on the measurement of mutual information
McAllester, D. and Stratos, K · 2020
Later among the works it cites.
Monte carlo gradient estimation in machine learning
Mohamed, S., Rosca, M., Figurnov, M., and Mnih, A · 2020
Later among the works it cites.
On the estimation of information measures of continuous distributions
Pichler, G., Piantanida, P., and Koliander, G · 2020
Later among the works it cites.
Null it out: Guarding protected attributes by iterative nullspace projection
Ravfogel, S., Elazar, Y., Gonen, H., Twiton, M., and Goldberg, Y · 2020
Later among the works it cites.
Understanding the limitations of variational mutual information estimators
Song, J. and Ermon, S · 2020
Later among the works it cites.
Sliced score matching: A scalable approach to density and score estimation
Song, Y., Garg, S., Shi, J., and Ermon, S · 2020
Later among the works it cites.
On mutual information maximization for representation learning
Tschannen, M., Djolonga, J., Rubenstein, P. K., Gelly, S., and Lucic, M · 2020
Later among the works it cites.
On the estimation of entropy for non-negative data
Chaubey, Y. P. and Vu, N. L · 2021
Later among the works it cites.
A novel estimator of mutual information for learning to disentangle textual representations
Colombo, P., Piantanida, P., and Clavel, C · 2021
Later among the works it cites.
Variational information bottleneck for effective low-resource fine-tuning
Mahabadi, R. K., Belinkov, Y., and Henderson, J · 2021
Later among the works it cites.
Ensemble estimation of generalized mutual information with applications to genomics
Moon, K. R., Sricharan, K., and Hero, A. O · 2021
Later among the works it cites.