Fetching the paper…
Reading the bibliography…
In this work we offer a framework for reasoning about a wide class of existing objectives in machine learning.
Information theory and statistical mechanics
Edwin T Jaynes · 1957
Earlier work this paper cites.
Thermal physics
Colin BP Finn · 1993
Earlier work this paper cites.
The information bottleneck method
N. Tishby, F.C. Pereira, and W. Biale · 1999
Earlier work this paper cites.
Predictability, complexity, and learning
William Bialek, Ilya Nemenman, and Naftali Tishby · 2001
Earlier work this paper cites.
Multivariate information bottleneck
Nir Friedman, Ori Mosenzon, Noam Slonim, and Naftali Tishby · 2001
Earlier work this paper cites.
Information projections revisited
Imre Csiszár and František Matúš · 2003
Earlier work this paper cites.
Estimating mutual information and multi–information in large networks
Noam Slonim, Gurinder S Atwal, Gasper Tkacik, and William Bialek · 2005
Earlier work this paper cites.
Statistical mechanics: entropy, order parameters, and complexity , volume 14
James Sethna · 2006
Earlier work this paper cites.
Edgeworth expansions and cumulants
Anirban DasGupta · 2008
Earlier work this paper cites.
Algebraic geometry and statistical learning theory , volume 25
Sumio Watanabe · 2009
Earlier work this paper cites.
Monte Carlo Methods for Rough Free Energy Landscapes: Population Annealing and Parallel Tempering
J. Machta and R. S. Ellis · 2011
Earlier work this paper cites.
Bayesian learning via stochastic gradient langevin dynamics
Max Welling and Yee W Teh · 2011
Earlier work this paper cites.
Elements of information theory
Thomas M Cover and Joy A Thomas · 2012
Earlier work this paper cites.
S. Still, D. A. Sivak, A. J. Bell, and G. E. Crooks · 2012
Cited alongside, same era.
Stochastic Gradient Hamiltonian Monte Carlo
T. Chen, E. B. Fox, and C. Guestrin · 2014
Cited alongside, same era.
Semi-Supervised Learning with Deep Generative Models
D. P. Kingma, D. J. Rezende, S. Mohamed, and M. Welling · 2014
Cited alongside, same era.
Auto-encoding variational Bayes
Diederik P Kingma and Max Welling · 2014
Cited alongside, same era.
Information bottleneck approach to predictive inference
Susanne Still · 2014
Cited alongside, same era.
Pratik Chaudhari and Stefano Soatto · 2017
Later among the works it cites.
β \beta -VAE: Learning Basic Visual Concepts with a Constrained Variational Framework
Irina Higgins, Loic Matthey, Arka Pal, Christopher Burgess, Xavier Glorot, Matthew Botvinick, Shakir Mohamed, and Alexander Lerchner · 2017
Later among the works it cites.
Snapshot Ensembles: Train 1, get M for free
G. Huang, Y. Li, G. Pleiss, Z. Liu, J. E. Hopcroft, and K. Q. Weinberger · 2017
Later among the works it cites.
Stochastic Gradient Descent as Approximate Bayesian Inference
S. Mandt, M. D. Hoffman, and D. M. Blei · 2017
Later among the works it cites.
A Bayesian Perspective on Generalization and Stochastic Gradient Descent
S. L. Smith and Q. V. Le · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
C. Blundell, J. Cornebise, K. Kavukcuoglu, and D. Wierstra · 2015
Cited alongside, same era.
A Complete Recipe for Stochastic Gradient MCMC
Y.-A. Ma, T. Chen, and E. B. Fox · 2015
Cited alongside, same era.
Gradient-based hyperparameter optimization through reversible learning
Dougal Maclaurin, David Duvenaud, and Ryan P. Adams · 2015
Cited alongside, same era.
Thermodynamics of information
Juan MR Parrondo, Jordan M Horowitz, and Takahiro Sagawa · 2015
Cited alongside, same era.
Deep variational information bottleneck
Alexander A Alemi, Ian Fischer, Joshua V Dillon, and Kevin Murphy · 2016
Cited alongside, same era.
X. Chen, D. P. Kingma, T. Salimans, Y. Duan, P. Dhariwal, J. Schulman, I. Sutskever, and P. Abbeel · 2016
Cited alongside, same era.
Emergence of Invariance and Disentangling in Deep Representations
A. Achille and S. Soatto · 2017
Cited alongside, same era.
Later among the works it cites.
Thermodynamic cost and benefit of data representations
S. Still · 2017
Later among the works it cites.
De Finetti, 2018
L Accardi · 2018
Closest in time.
Alexander A Alemi, Ben Poole, Joshua V Dillon, Rif A Saurous, and Kevin Murphy · 2018
Closest in time.
MINE: Mutual Information Neural Estimation
M. Ishmael Belghazi, A. Baratin, S. Rajeswar, S. Ozair, Y. Bengio, A. Courville, and R Devon Hjelm · 2018
Closest in time.
Representation Learning with Contrastive Predictive Coding
A. van den Oord, Y. Li, and O. Vinyals · 2018
Closest in time.
Mathematical theory of Bayesian statistics
Sumio Watanabe · 2018
Closest in time.
Energy-entropy competition and the effectiveness of stochastic gradient descent in machine learning
Y. Zhang, A. M. Saxe, M. S. Advani, and A. A. Lee · 2018
Closest in time.