2018

Sufficient Conditions for Idealised Models to Have No Adversarial Examples: a Theoretical and Empirical Study with Bayesian Neural Networks

Gal, Yarin, Smith, Lewis

Understand

We prove, under two sufficient conditions, that idealised models can have no adversarial examples.

  • We discuss which idealised models satisfy our conditions, and show that idealised Bayesian neural networks (BNNs) satisfy these.
  • We continue by studying near-idealised BNNs using HMC inference, demonstrating the theoretical ideas in practice.
  • We experiment with HMC on synthetic data derived from MNIST for which we know the ground-truth image density, showing that near-perfect epistemic uncertainty correlates to density under image manifold, and that adversarial images lie off the manifold in our setting.

Reading the bibliography…