Fetching the paper…
Reading the bibliography…
In the Information Bottleneck (IB), when tuning the relative strength between compression and prediction terms, how do the two terms behave, and what's their relationship with the dataset and the learned representation? In this paper, we set out to answer these questions by studying multiple phase transitions in the IB objective: $\text{IB}_\beta[p(z|x)] = I(X; Z) - \beta I(Y; Z)$ defined on the encoding distribution p(z|x) for input $X$, target $Y$ and representation $Z$, where sudden jumps of $dI(Y; Z)/d \beta$ and prediction accuracy are observed with increasing $\beta$.
Computation of channel capacity and rate-distortion functions
Richard Blahut · 1972
Earlier work this paper cites.
Emergence of invariance and disentanglement in deep representations
Alessandro Achille and Stefano Soatto · 1980
Earlier work this paper cites.
Relations between two sets of variates
Harold Hotelling · 1992
Earlier work this paper cites.
The information bottleneck method
Naftali Tishby, Fernando C Pereira, and William Bialek · 2000
Earlier work this paper cites.
Information bottleneck for gaussian variables
Gal Chechik, Amir Globerson, Naftali Tishby, and Yair Weiss · 2005
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky and Geoffrey Hinton · 2009
Earlier work this paper cites.
Learning and generalization with the information bottleneck
Ohad Shamir, Sivan Sabato, and Naftali Tishby · 2010
Earlier work this paper cites.
Meta-gaussian information bottleneck
Mélanie Rey and Volker Roth · 2012
Earlier work this paper cites.
Venkat Anantharam, Amin Gohari, Sudeep Kamath, and Chandra Nair · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2015
Cited alongside, same era.
Fisher information properties
Pablo Zegers · 2015
Cited alongside, same era.
Deep variational information bottleneck
Alexander A Alemi, Ian Fischer, Joshua V Dillon, and Kevin Murphy · 2016
Cited alongside, same era.
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Wide Residual Networks
The conditional entropy bottleneck, 2018
Ian Fischer · 2018
Later among the works it cites.
Xue Bin Peng, Angjoo Kanazawa, Sam Toyer, Pieter Abbeel, and Sergey Levine · 2018
Later among the works it cites.
Danilo Jimenez Rezende and Fabio Viola · 2018
Later among the works it cites.
Lecture: the information theory of deep neural networks: the statistical physics aspects
Naftali Tishby · 2018
Later among the works it cites.
Infobot: Transfer and exploration via the information bottleneck
Anirudh Goyal, Riashat Islam, Daniel Strouse, Zafarali Ahmed, Matthew Botvinick, Hugo Larochelle, Sergey Levine, and Yoshua Bengio · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Zagoruyko and N. Komodakis · 2016
Cited alongside, same era.
beta-vae: Learning basic visual concepts with a constrained variational framework
Irina Higgins, Loic Matthey, Arka Pal, Christopher Burgess, Xavier Glorot, Matthew Botvinick, Shakir Mohamed, and Alexander Lerchner · 2017
Cited alongside, same era.
The dynamics of differential learning i: Information-dynamics and task reachability
Alessandro Achille, Glen Mbeng, and Stefano Soatto · 2018
Cited alongside, same era.
Autoaugment: Learning augmentation policies from data
Ekin D Cubuk, Barret Zoph, Dandelion Mane, Vijay Vasudevan, and Quoc V Le · 2018
Cited alongside, same era.
Information dropout: Learning optimal representations through noisy computation
Alessandro Achille and Stefano Soatto
Cited in the paper.
The deterministic information bottleneck
DJ Strouse and David J Schwab
Cited in the paper.
The information bottleneck and geometric clustering
DJ Strouse and David J Schwab
Cited in the paper.
Later among the works it cites.
Dynamics-aware unsupervised discovery of skills
Archit Sharma, Shixiang Gu, Sergey Levine, Vikash Kumar, and Karol Hausman · 2019
Later among the works it cites.
Pareto-optimal data compression for binary classification tasks
Max Tegmark and Tailin Wu · 2019
Later among the works it cites.
Learnability for the information bottleneck
Tailin Wu, Ian Fischer, Isaac Chuang, and Max Tegmark · 2019
Later among the works it cites.