2017

Taming Non-stationary Bandits: A Bayesian Approach

Raj, Vishnu, Kalyani, Sheetal

Understand

We consider the multi armed bandit problem in non-stationary environments.

  • Based on the Bayesian method, we propose a variant of Thompson Sampling which can be used in both rested and restless bandit scenarios.
  • Applying discounting to the parameters of prior distribution, we describe a way to systematically reduce the effect of past observations.
  • Further, we derive the exact expression for the probability of picking sub-optimal arms.

Reading the bibliography…