Fetching the paper…
Reading the bibliography…
The Information Bottleneck principle offers both a mechanism to explain how deep neural networks train and generalize, as well as a regularized objective with which to train models.
A mathematical theory of communication
Claude E Shannon · 1948
Earlier work this paper cites.
Multivariate information transmission
William McGill · 1954
Earlier work this paper cites.
Emergence of invariance and disentanglement in deep representations
Alessandro Achille and Stefano Soatto · 1980
Earlier work this paper cites.
Accelerated learning in layered neural networks
Sara A. Solla, Esther Levin, and Michael Fleisher · 1988
Earlier work this paper cites.
Connectionist learning procedures
Geoffrey E Hinton · 1990
Earlier work this paper cites.
A new outlook on shannon’s information measures
Raymond W Yeung · 1991
Earlier work this paper cites.
Ying Zhang, Tao Xiang, Timothy M Hospedales, and Huchuan Lu · 1991
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
Boris T Polyak and Anatoli B Juditsky · 1992
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. Lecun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
The information bottleneck method
Naftali Tishby, Fernando C Pereira, and William Bialek · 2000
Earlier work this paper cites.
A renyi entropy convolution inequality with application
J-F Bercher and Christophe Vignat · 2002
Earlier work this paper cites.
Slurm: Simple linux utility for resource management
Morris A. Jette, Andy B. Yoo, and Mark Grondona · 2002
Earlier work this paper cites.
Conditional information bottleneck clustering
David Gondek and Thomas Hofmann · 2003
Earlier work this paper cites.
Information Theory, Inference, and Learning Algorithms
David J. C. MacKay · 2003
Earlier work this paper cites.
Estimating mutual information
Alexander Kraskov, Harald Stögbauer, and Peter Grassberger · 2004
Earlier work this paper cites.
Learning deep architectures for ai
Yoshua Bengio et al · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Learning and generalization with the information bottleneck
Ohad Shamir, Sivan Sabato, and Naftali Tishby · 2010
Earlier work this paper cites.
Bayesian active learning for classification and preference learning
Neil Houlsby, Ferenc Huszár, Zoubin Ghahramani, and Máté Lengyel · 2011
Earlier work this paper cites.
Elements of information theory
Thomas M Cover and Joy A Thomas · 2012
Cited alongside, same era.
An empirical investigation of catastrophic forgetting in gradient-based neural networks
Ian J Goodfellow, Mehdi Mirza, Da Xiao, Aaron Courville, and Yoshua Bengio · 2013
Cited alongside, same era.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2013
Cited alongside, same era.
Intriguing properties of neural networks
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Towards deep learning models resistant to adversarial attacks
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu · 2018
Later among the works it cites.
Formal limitations on the measurement of mutual information
David McAllester and Karl Stratos · 2018
Later among the works it cites.
Representation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Later among the works it cites.
Learning representations for neural network-based classification using the information bottleneck principle
Rana Ali Amjad and Bernhard Claus Geiger · 2019
Later among the works it cites.
The Conditional Entropy Bottleneck
Ian Fischer · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Cited alongside, same era.
Importance weighted autoencoders
Yuri Burda, Roger Grosse, and Ruslan Salakhutdinov · 2015
Cited alongside, same era.
Deep learning and the information bottleneck principle
Naftali Tishby and Noga Zaslavsky · 2015
Cited alongside, same era.
Deep variational information bottleneck
Alexander A Alemi, Ian Fischer, Joshua V Dillon, and Kevin Murphy · 2016
Cited alongside, same era.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Yarin Gal and Zoubin Ghahramani · 2016
Cited alongside, same era.
beta-vae: Learning basic visual concepts with a constrained variational framework
Irina Higgins, Loic Matthey, Arka Pal, Christopher Burgess, Xavier Glorot, Matthew Botvinick, Shakir Mohamed, and Alexander Lerchner · 2016
Cited alongside, same era.
Deepfool: a simple and accurate method to fool deep neural networks
Seyed-Mohsen Moosavi-Dezfooli, Alhussein Fawzi, and Pascal Frossard · 2016
Cited alongside, same era.
Partha Ghosh, Mehdi SM Sajjadi, Antonio Vergari, Michael Black, and Bernhard Schölkopf · 2019
Later among the works it cites.
Imagewang
Jeremy Howard · 2019
Later among the works it cites.
Multi-sample dropout for accelerated training and better generalization
Hiroshi Inoue · 2019
Later among the works it cites.
Batchbald: Efficient and diverse batch acquisition for deep bayesian active learning
Andreas Kirsch, Joost van Amersfoort, and Yarin Gal · 2019
Later among the works it cites.
Scalable mutual information estimation using dependence graphs
Morteza Noshad, Yu Zeng, and Alfred O Hero · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Later among the works it cites.
On variational bounds of mutual information
Ben Poole, Sherjil Ozair, Aaron van den Oord, Alexander A Alemi, and George Tucker · 2019
Later among the works it cites.
On the information bottleneck theory of deep learning
Andrew M Saxe, Yamini Bansal, Joel Dapello, Madhu Advani, Artemy Kolchinsky, Brendan D Tracey, and David D Cox · 2019
Later among the works it cites.
On mutual information maximization for representation learning
Michael Tschannen, Josip Djolonga, Paul K Rubenstein, Sylvain Gelly, and Mario Lucic · 2019
Later among the works it cites.
Efficient and scalable bayesian neural nets with rank-1 factors
Michael W Dusenberry, Ghassen Jerfel, Yeming Wen, Yi-an Ma, Jasper Snoek, Katherine Heller, Balaji Lakshminarayanan, and Dustin Tran · 2020
Closest in time.
The conditional entropy bottleneck
Ian Fischer · 2020
Closest in time.
Ceb improves model robustness
Ian Fischer and Alexander A. Alemi · 2020
Closest in time.
A theory of usable information under computational constraints
Yilun Xu, Shengjia Zhao, Jiaming Song, Russell Stewart, and Stefano Ermon · 2020
Closest in time.