Fetching the paper…
Reading the bibliography…
Understanding the source of the superior generalization ability of NNs remains one of the most important problems in ML research.
Emergence of invariance and disentanglement in deep representations
Alessandro Achille and Stefano Soatto · 1980
Earlier work this paper cites.
Residuals and influence in regression
R Dennis Cook and Sanford Weisberg · 1982
Earlier work this paper cites.
Diffusion for global optimization in R n {R}^{n}
Tzuu-Shuh Chiang, Chii-Ruey Hwang, and Shuenn Jyi Sheu · 1987
Earlier work this paper cites.
Bootstrap methods: another look at the jackknife
Bradley Efron · 1992
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
A strong approximation theorem for stochastic recursive algorithms
Vivek S Borkar and Sanjoy K Mitter · 1999
Earlier work this paper cites.
The information bottleneck method
Naftali Tishby, Fernando C Pereira, and William Bialek · 2000
Earlier work this paper cites.
Elementary statistical physics
Charles Kittel · 2004
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Finding influential instances for distantly supervised relation extraction
Zifeng Wang, Rui Wen, Xi Chen, Shao-Lun Huang, Ningyu Zhang, and Yefeng Zheng · 2009
Earlier work this paper cites.
An analysis of single-layer networks in unsupervised feature learning
Adam Coates, Andrew Ng, and Honglak Lee · 2011
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Bo Wu, and Andrew Y Ng · 2011
Earlier work this paper cites.
Bayesian learning via stochastic gradient Langevin dynamics
Max Welling and Yee W Teh · 2011
Earlier work this paper cites.
Estimating uncertainty for massive data streams
Nicholas Chamandy, Omkar Muralidharan, Amir Najmi, and Siddartha Naidu · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2015
Cited alongside, same era.
High-Dimensional Statistics , chapter 1
Rigollet Philippe · 2015
Cited alongside, same era.
Deep learning and the information bottleneck principle
Naftali Tishby and Noga Zaslavsky · 2015
Cited alongside, same era.
Sharp minima can generalize for deep nets
Laurent Dinh, Razvan Pascanu, Samy Bengio, and Yoshua Bengio · 2017
Cited alongside, same era.
Understanding black-box predictions via influence functions
Pang Wei Koh and Percy Liang · 2017
Cited alongside, same era.
Non-convex learning via stochastic gradient Langevin dynamics: a nonasymptotic analysis
Maxim Raginsky, Alexander Rakhlin, and Matus Telgarsky · 2017
Cited alongside, same era.
An information-theoretic view for deep learning
Jingwei Zhang, Tongliang Liu, and Dacheng Tao · 2018
Later among the works it cites.
Estimating information flow in deep neural networks
Ziv Goldfeld, Ewout van den Berg, Kristjan Greenewald, Igor Melnyk, Nam Nguyen, Brian Kingsbury, and Yury Polyanskiy · 2019
Later among the works it cites.
InfoBot: transfer and exploration via the information bottleneck
Anirudh Goyal, Riashat Islam, DJ Strouse, Zafarali Ahmed, Hugo Larochelle, Matthew Botvinick, Yoshua Bengio, and Sergey Levine · 2019
Later among the works it cites.
Nonlinear information bottleneck
Artemy Kolchinsky, Brendan D Tracey, and David H Wolpert · 2019
Later among the works it cites.
Specializing word embeddings (for parsing) by information bottleneck
Xiang Lisa Li and Jason Eisner · 2019
Later among the works it cites.
Fisher-Rao metric, geometry, and complexity of neural networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Opening the black box of deep neural networks via information
Ravid Shwartz-Ziv and Naftali Tishby · 2017
Cited alongside, same era.
Information-theoretic analysis of generalization capability of learning algorithms
Aolin Xu and Maxim Raginsky · 2017
Cited alongside, same era.
Understanding disentangling in β \beta -VAE
Christopher P Burgess, Irina Higgins, Arka Pal, Loic Matthey, Nick Watters, Guillaume Desjardins, and Alexander Lerchner · 2018
Cited alongside, same era.
Compressing neural networks using the variational information bottleneck
Bin Dai, Chen Zhu, Baining Guo, and David Wipf · 2018
Cited alongside, same era.
Data-dependent PAC-Bayes priors via differential privacy
Gintare Karolina Dziugaite and Daniel M Roy · 2018
Cited alongside, same era.
Caveats for information bottleneck in deterministic scenarios
Artemy Kolchinsky, Brendan D Tracey, and Steven Van Kuyk · 2018
Cited alongside, same era.
Tengyuan Liang, Tomaso Poggio, Alexander Rakhlin, and James Stokes · 2019
Later among the works it cites.
Information-theoretic generalization bounds for SGLD via data-dependent estimates
Jeffrey Negrea, Mahdi Haghifam, Gintare Karolina Dziugaite, Ashish Khisti, and Daniel M Roy · 2019
Later among the works it cites.
PyTorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Later among the works it cites.
On the information bottleneck theory of deep learning
Andrew M Saxe, Yamini Bansal, Joel Dapello, Madhu Advani, Artemy Kolchinsky, Brendan D Tracey, and David D Cox · 2019
Later among the works it cites.
Deep multi-view information bottleneck
Qi Wang, Claire Boudreau, Qixing Luo, Pang-Ning Tan, and Jiayu Zhou · 2019
Later among the works it cites.
New insights and perspectives on the natural gradient method
James Martens · 2020
Later among the works it cites.
Disentangled information bottleneck
Ziqi Pan, Li Niu, Jianfu Zhang, and Liqing Zhang · 2020
Later among the works it cites.
Information in infinite ensembles of infinitely-wide neural networks
Ravid Shwartz-Ziv and Alexander A Alemi · 2020
Later among the works it cites.
Tailin Wu, Hongyu Ren, Pan Li, and Jure Leskovec · 2020
Later among the works it cites.
On the role of data in PAC-Bayes
Gintare Karolina Dziugaite, Kyle Hsu, Waseem Gharbieh, Gabriel Arpino, and Daniel Roy · 2021
Closest in time.