Fetching the paper…
Reading the bibliography…
Much of the field of Machine Learning exhibits a prominent set of failure modes, including vulnerability to adversarial examples, poor out-of-distribution (OoD) detection, miscalibration, and willingness to memorize random labelings of datasets.
On Variational Bounds of Mutual Information
Poole, B., Ozair, S., van den Oord, A., Alemi, A. A., and Tucker, G · 1905
Earlier work this paper cites.
A Mathematical Theory of Communication
Shannon, C. E · 1948
Earlier work this paper cites.
A new outlook on shannon’s information measures
Yeung, R. W · 1991
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
Predictive information
Bialek, W. and Tishby, N · 1999
Earlier work this paper cites.
The information bottleneck method
Tishby, N., Pereira, F. C., and Bialek, W · 2000
Earlier work this paper cites.
Elements of information theory 2nd edition
Cover, T. M. and Thomas, J. A · 2006
Earlier work this paper cites.
Extracting and composing robust features with denoising autoencoders
Vincent, P., Larochelle, H., Bengio, Y., and Manzagol, P.-A · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A. and Hinton, G · 2009
Earlier work this paper cites.
Learning and generalization with the information bottleneck
Shamir, O., Sabato, S., and Tishby, N · 2010
Earlier work this paper cites.
Anantharam, V., Gohari, A., Kamath, S., and Nair, C · 2013
Earlier work this paper cites.
Auto-encoding variational Bayes
Kingma, D. P. and Welling, M · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R · 2014
Earlier work this paper cites.
Fast and accurate deep network learning by exponential linear units (elus)
Clevert, D.-A., Unterthiner, T., and Hochreiter, S · 2015
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Goodfellow, I. J., Shlens, J., and Szegedy, C · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. and Ba, J · 2015
Earlier work this paper cites.
A Baseline for Detecting Misclassified and Out-of-Distribution Examples in Neural Networks
Hendrycks, D. and Gimpel, K · 2016
Cited alongside, same era.
Adversarial machine learning at scale
Kurakin, A., Goodfellow, I., and Bengio, S · 2016
Cited alongside, same era.
Simple and Scalable Uncertainty Estimation using Deep Ensembles
Lakshminarayanan, B., Pritzel, A., and Blundell, C · 2016
Cited alongside, same era.
Wavenet: A generative model for raw audio
Van Den Oord, A., Dieleman, S., Zen, H., Simonyan, K., Vinyals, O., Graves, A., Kalchbrenner, N., Senior, A. W., and Kavukcuoglu, K · 2016
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O · 2016
How (Not) To Train Your Neural Network Using the Information Bottleneck Principle
Amjad, R. A. and Geiger, B. C · 2018
Later among the works it cites.
Obfuscated gradients give a false sense of security: Circumventing defenses to adversarial examples
Athalye, A., Carlini, N., and Wagner, D · 2018
Later among the works it cites.
Autoaugment: Learning augmentation policies from data
Cubuk, E. D., Zoph, B., Mane, D., Vasudevan, V., and Le, Q. V · 2018
Later among the works it cites.
Learning Confidence for Out-of-Distribution Detection in Neural Networks
DeVries, T. and Taylor, G. W · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deep Variational Information Bottleneck
Alemi, A. A., Fischer, I., Dillon, J. V., and Murphy, K · 2017
Cited alongside, same era.
Learners that use little information
Bassily, R., Moran, S., Nachum, I., Shafer, J., and Yehudayoff, A · 2017
Cited alongside, same era.
Towards evaluating the robustness of neural networks
Carlini, N. and Wagner, D · 2017
Cited alongside, same era.
Accurate, large minibatch sgd: training imagenet in 1 hour
Goyal, P., Dollár, P., Girshick, R., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K · 2017
Cited alongside, same era.
On Calibration of Modern Neural Networks
Guo, C., Pleiss, G., Sun, Y., and Weinberger, K. Q · 2017
Cited alongside, same era.
Training Confidence-calibrated Classifiers for Detecting Out-of-Distribution Samples
Lee, K., Lee, H., Lee, K., and Shin, J · 2017
Cited alongside, same era.
Enhancing The Reliability of Out-of-distribution Image Detection in Neural Networks
Liang, S., Li, Y., and Srikant, R · 2017
Cited alongside, same era.
Figurnov, M., Mohamed, S., and Mnih, A · 2018
Later among the works it cites.
Scan: Learning hierarchical compositional visual concepts
Higgins, I., Sonnerat, N., Matthey, L., Pal, A., Burgess, C. P., Bošnjak, M., Shanahan, M., Botvinick, M., Hassabis, D., and Lerchner, A · 2018
Later among the works it cites.
Learning deep representations by mutual information estimation and maximization
Hjelm, R. D., Fedorov, A., Lavoie-Marchildon, S., Grewal, K., Trischler, A., and Bengio, Y · 2018
Later among the works it cites.
A simple unified framework for detecting out-of-distribution samples and adversarial attacks
Lee, K., Lee, K., Lee, H., and Shin, J · 2018
Later among the works it cites.
Do deep generative models know what they don’t know?
Nalisnick, E., Matsukawa, A., Teh, Y. W., Gorur, D., and Lakshminarayanan, B · 2018
Later among the works it cites.
Representation learning with contrastive predictive coding
Oord, A. v. d., Li, Y., and Vinyals, O · 2018
Later among the works it cites.
Technical report on the cleverhans v2.1.0 adversarial examples library
Papernot, N., Faghri, F., Carlini, N., Goodfellow, I., Feinman, R., Kurakin, A., Xie, C., Sharma, Y., Brown, T., Roy, A., Matyasko, A., Behzadan, V., Hambardzumyan, K., Zhang, Z., Juang, Y.-L., Li, Z., Sheatsley, R., Garg, A., Uesato, J., Gierke, W., Dong, Y., Berthelot, D., Hendricks, P., Rauber, J., and Long, R · 2018
Later among the works it cites.
Do cifar-10 classifiers generalize to cifar-10?
Recht, B., Roelofs, R., Schmidt, L., and Shankar, V · 2018
Later among the works it cites.
Generative models of visually grounded imagination
Vedantam, R., Fischer, I., Huang, J., and Murphy, K · 2018
Later among the works it cites.
Scaling provable adversarial defenses
Wong, E., Schmidt, F., Metzen, J. H., and Kolter, J. Z · 2018
Later among the works it cites.
Certified adversarial robustness via randomized smoothing
Cohen, J. M., Rosenfeld, E., and Kolter, J. Z · 2019
Later among the works it cites.
Adversarial examples are not bugs, they are features
Ilyas, A., Santurkar, S., Tsipras, D., Engstrom, L., Tran, B., and Madry, A · 2019
Later among the works it cites.
Learnability for the Information Bottleneck
Wu, T., Fischer, I., Chuang, I., and Tegmark, M · 2019
Later among the works it cites.