Fetching the paper…
Reading the bibliography…
Numerous deep learning algorithms have been inspired by and understood via the notion of information bottleneck, where unnecessary information is (often implicitly) minimized while task-relevant information is maximized.
The information bottleneck method
Tishby, N., Pereira, F. C., and Bialek, W · 1999
Earlier work this paper cites.
The nature of statistical learning theory
Vapnik, V · 1999
Earlier work this paper cites.
Document clustering using word clusters via the information bottleneck method
Slonim, N. and Tishby, N · 2000
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results
Bartlett, P. L. and Mendelson, S · 2002
Earlier work this paper cites.
Stability and generalization
Bousquet, O. and Elisseeff, A · 2002
Earlier work this paper cites.
Randomized algorithms
Hromkovič, J · 2004
Earlier work this paper cites.
Learning and generalization with the information bottleneck
Shamir, O., Sabato, S., and Tishby, N · 2010
Earlier work this paper cites.
Robustness and generalization
Xu, H. and Mannor, S · 2012
Earlier work this paper cites.
Constrained optimization and Lagrange multiplier methods
Bertsekas, D. P · 2014
Earlier work this paper cites.
Understanding machine learning: From theory to algorithms
Shalev-Shwartz, S. and Ben-David, S · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R · 2014
Earlier work this paper cites.
Variational dropout and the local reparameterization trick
Kingma, D. P., Salimans, T., and Welling, M · 2015
Earlier work this paper cites.
Deep learning and the information bottleneck principle
Tishby, N. and Zaslavsky, N · 2015
Earlier work this paper cites.
Deep variational information bottleneck
Alemi, A. A., Fischer, I., Dillon, J. V., and Murphy, K · 2016
Earlier work this paper cites.
Identity mappings in deep residual networks
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Estimating mixture entropy with pairwise distances
Kolchinsky, A. and Tracey, B. D · 2017
Earlier work this paper cites.
Stochastic gradient descent as approximate bayesian inference
Mandt, S., Hoffman, M. D., and Blei, D. M · 2017
Earlier work this paper cites.
Opening the black box of deep neural networks via information
Shwartz-Ziv, R. and Tishby, N · 2017
Earlier work this paper cites.
Information-theoretic analysis of generalization capability of learning algorithms
Xu, A. and Raginsky, M · 2017
Earlier work this paper cites.
Emergence of invariance and disentanglement in deep representations
Achille, A. and Soatto, S · 2018
Earlier work this paper cites.
Fixing a broken elbo
Alemi, A., Poole, B., Fischer, I., Dillon, J., Saurous, R. A., and Murphy, K · 2018
Earlier work this paper cites.
Learners that use little information
Bassily, R., Moran, S., Nachum, I., Shafer, J., and Yehudayoff, A · 2018
Cited alongside, same era.
The role of the information bottleneck in representation learning
Vera, M., Piantanida, P., and Vega, L. R · 2018
Cited alongside, same era.
Mitigating adversarial effects through randomization
Xie, C., Wang, J., Zhang, Z., Ren, Z., and Yuille, A · 2018
Cited alongside, same era.
Learning representations for neural network-based classification using the information bottleneck principle
Amjad, R. A. and Geiger, B. C · 2019
Cited alongside, same era.
Nearly-tight vc-dimension and pseudodimension bounds for piecewise linear neural networks
Bartlett, P. L., Harvey, N., Liaw, C., and Mehrabian, A · 2019
Cited alongside, same era.
Adaptive estimators show information compression in deep neural networks
Chelombiev, I., Houghton, C., and O’Donnell, C · 2019
Robustness certificates for sparse adversarial attacks by randomized ablation
Levine, A. and Feizi, S · 2020
Later among the works it cites.
Prevalence of neural collapse during the terminal phase of deep learning training
Papyan, V., Han, X., and Donoho, D. L · 2020
Later among the works it cites.
Randomization matters how to defend against strong adversarial attacks
Pinot, R., Ettedgui, R., Rizk, G., Chevaleyre, Y., and Atif, J · 2020
Later among the works it cites.
Reasoning about generalization via conditional mutual information
Steinke, T. and Zakynthinou, L · 2020
Later among the works it cites.
Burhanpurkar, M., Deng, Z., Dwork, C., and Zhang, L · 2021
Later among the works it cites.
Toward better generalization bounds with locally elastic stability
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Efficient and accurate estimation of lipschitz constants for deep neural networks
Fazlyab, M., Robey, A., Hassani, H., Morari, M., and Pappas, G · 2019
Cited alongside, same era.
Estimating information flow in deep neural networks
Goldfeld, Z., Van Den Berg, E., Greenewald, K., Melnyk, I., Nguyen, N., Kingsbury, B., and Polyanskiy, Y · 2019
Cited alongside, same era.
Lipschitz constant estimation of neural networks via sparse polynomial optimization
Latorre, F., Rolland, P., and Cevher, V · 2019
Cited alongside, same era.
A simple baseline for bayesian uncertainty in deep learning
Maddox, W. J., Izmailov, P., Garipov, T., Vetrov, D. P., and Wilson, A. G · 2019
Cited alongside, same era.
Uniform convergence may be unable to explain generalization in deep learning
Nagarajan, V. and Kolter, J. Z · 2019
Cited alongside, same era.
Theoretical evidence for adversarial robustness through randomization
Pinot, R., Meunier, L., Araujo, A., Kashima, H., Yger, F., Gouy-Pailler, C., and Atif, J · 2019
Cited alongside, same era.
Deng, Z., He, H., and Su, W · 2021
Later among the works it cites.
On the role of data in pac-bayes bounds
Dziugaite, G. K., Hsu, K., Gharbieh, W., Arpino, G., and Roy, D · 2021
Later among the works it cites.
Robust predictable control
Eysenbach, B., Salakhutdinov, R. R., and Levine, S · 2021
Later among the works it cites.
Compressive visual representations
Lee, K.-H., Arnab, A., Guadarrama, S., Canny, J., and Fischer, I · 2021
Later among the works it cites.
Discrete-valued neural communication
Liu, D., Lamb, A., Kawaguchi, K., Goyal, A., Sun, C., Mozer, M. C., and Bengio, Y · 2021
Later among the works it cites.
Training robust neural networks using lipschitz bounds
Pauli, P., Koch, A., Berberich, J., Kohler, P., and Allgöwer, F · 2021
Later among the works it cites.
Galloway, A., Golubeva, A., Salem, M., Nica, M., Ioannou, Y., and Taylor, G. W · 2022
Later among the works it cites.
Inductive biases for deep learning of higher-level cognition
Goyal, A. and Bengio, Y · 2022
Later among the works it cites.
When do extended physics-informed neural networks (xpinns) improve generalization?
Hu, Z., Jagtap, A. D., Karniadakis, G. E., and Kawaguchi, K · 2022
Later among the works it cites.
Adaptive discrete communication bottlenecks with dynamic vector quantization
Liu, D., Lamb, A., Ji, X., Notsawo, P., Mozer, M., Bengio, Y., and Kawaguchi, K · 2022
Later among the works it cites.
Pac-bayes compression bounds so tight that they can explain generalization
Lotfi, S., Finzi, M., Kapoor, S., Potapczynski, A., Goldblum, M., and Wilson, A. G · 2022
Later among the works it cites.
Graph structure learning with variational information bottleneck
Sun, Q., Li, J., Peng, H., Wu, J., Fu, X., Ji, C., and Philip, S. Y · 2022
Later among the works it cites.
On rademacher complexity-based generalization bounds for deep learning
Truong, L. V · 2022
Later among the works it cites.
Vision transformer with information bottleneck for fine-grained visual classification
Su, T., Song, C., and Cheng, J · 2023
Closest in time.
Discrete key-value bottleneck
Träuble, F., Goyal, A., Rahaman, N., Mozer, M., Kawaguchi, K., Bengio, Y., and Schölkopf, B · 2023
Closest in time.