2016

Conservative Contextual Linear Bandits

Kazerouni, Abbas, Ghavamzadeh, Mohammad, Abbasi-Yadkori, Yasin et al.

Understand

Safety is a desirable property that can immensely increase the applicability of learning algorithms in real-world decision-making problems.

  • It is much easier for a company to deploy an algorithm that is safe, i.e., guaranteed to perform at least as well as a baseline.
  • In this paper, we study the issue of safety in contextual linear bandits that have application in many different fields including personalized ad recommendation in online marketing.
  • We formulate a notion of safety for this class of algorithms.

Built on

  • Finite-time analysis of the multiarmed bandit problem

    P. Auer, N. Cesa-Bianchi, and P. Fischer · 2002

    Earlier work this paper cites.

  • Stochastic linear optimization under bandit feedback

    V. Dani, T. Hayes, and S. Kakade · 2008

    Earlier work this paper cites.

  • Linearly parameterized bandits

    P. Rusmevichientong and J. Tsitsiklis · 2010

    Earlier work this paper cites.

  • Improved algorithms for linear stochastic bandits

    Y. Abbasi-Yadkori, D. Pál, and C. Szepesvári · 2011

    Earlier work this paper cites.

  • Contextual bandits with linear payoff functions

    W. Chu, L. Li, L. Reyzin, and R. Schapire · 2011

    Earlier work this paper cites.

Similar

  • Counterfactual reasoning and learning systems: The example of computational advertising

    L. Bottou, J. Peters, J. Quinonero-Candela, D. Charles, D. Chickering, E. Portugaly, D. Ray, P. Simard, and E. Snelson · 2013

    Cited alongside, same era.

  • Learning to optimize via posterior sampling

    D. Russo and B. Van Roy · 2014

    Cited alongside, same era.

  • Batch learning from logged bandit feedback through counterfactual risk minimization

    A. Swaminathan and T. Joachims · 2015

    Cited alongside, same era.

  • Counterfactual risk minimization: Learning from logged bandit feedback

    A. Swaminathan and T. Joachims · 2015

    Cited alongside, same era.

  • Building personal ad recommendation systems for life-time value optimization with guarantees

    G. Theocharous, P. Thomas, and M. Ghavamzadeh · 2015

    Cited alongside, same era.

Then

  • High confidence off-policy evaluation

    P. Thomas, G. Theocharous, and M. Ghavamzadeh · 2015

    Later among the works it cites.

  • High confidence policy improvement

    P. Thomas, G. Theocharous, and M. Ghavamzadeh · 2015

    Later among the works it cites.

  • Doubly robust off-policy value evaluation for reinforcement learning

    N. Jiang and L. Li · 2016

    Closest in time.

  • Safe policy improvement by minimizing robust baseline regret

    M. Petrik, M. Ghavamzadeh, and Y. Chow · 2016

    Closest in time.

  • Conservative bandits

    Y. Wu, R. Shariff, T. Lattimore, and C. Szepesvári · 2016

    Closest in time.

Beyond the bibliography

alphaXiv searches the wider corpus for related work and actual follow-ups.

Open on alphaXiv

alphaXiv is searching for related work…