2017

A Multi-Armed Bandit Approach for Online Expert Selection in Markov Decision Processes

Mazumdar, Eric, Dong, Roy, Royo, Vicenç Rúbies et al.

Understand

We formulate a multi-armed bandit (MAB) approach to choosing expert policies online in Markov decision processes (MDPs).

  • Given a set of expert policies trained on a state and action space, the goal is to maximize the cumulative reward of our agent.
  • The hope is to quickly find the best expert in our set.
  • The MAB formulation allows us to quantify the performance of an algorithm in terms of the regret incurred from not choosing the best expert from the beginning.

Reading the bibliography…