Understand
Safety is a desirable property that can immensely increase the applicability of learning algorithms in real-world decision-making problems.
- It is much easier for a company to deploy an algorithm that is safe, i.e., guaranteed to perform at least as well as a baseline.
- In this paper, we study the issue of safety in contextual linear bandits that have application in many different fields including personalized ad recommendation in online marketing.
- We formulate a notion of safety for this class of algorithms.
Built on
Finite-time analysis of the multiarmed bandit problem
P. Auer, N. Cesa-Bianchi, and P. Fischer · 2002
Earlier work this paper cites.
Stochastic linear optimization under bandit feedback
V. Dani, T. Hayes, and S. Kakade · 2008
Earlier work this paper cites.
Linearly parameterized bandits
P. Rusmevichientong and J. Tsitsiklis · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Y. Abbasi-Yadkori, D. Pál, and C. Szepesvári · 2011
Earlier work this paper cites.
Contextual bandits with linear payoff functions
W. Chu, L. Li, L. Reyzin, and R. Schapire · 2011
Earlier work this paper cites.
Similar
Counterfactual reasoning and learning systems: The example of computational advertising
L. Bottou, J. Peters, J. Quinonero-Candela, D. Charles, D. Chickering, E. Portugaly, D. Ray, P. Simard, and E. Snelson · 2013
Cited alongside, same era.
Learning to optimize via posterior sampling
D. Russo and B. Van Roy · 2014
Cited alongside, same era.
Batch learning from logged bandit feedback through counterfactual risk minimization
A. Swaminathan and T. Joachims · 2015
Cited alongside, same era.
Counterfactual risk minimization: Learning from logged bandit feedback
A. Swaminathan and T. Joachims · 2015
Cited alongside, same era.
Building personal ad recommendation systems for life-time value optimization with guarantees
G. Theocharous, P. Thomas, and M. Ghavamzadeh · 2015
Cited alongside, same era.
Then
High confidence off-policy evaluation
P. Thomas, G. Theocharous, and M. Ghavamzadeh · 2015
Later among the works it cites.
High confidence policy improvement
P. Thomas, G. Theocharous, and M. Ghavamzadeh · 2015
Later among the works it cites.
Doubly robust off-policy value evaluation for reinforcement learning
N. Jiang and L. Li · 2016
Closest in time.
Safe policy improvement by minimizing robust baseline regret
M. Petrik, M. Ghavamzadeh, and Y. Chow · 2016
Closest in time.
Conservative bandits
Y. Wu, R. Shariff, T. Lattimore, and C. Szepesvári · 2016
Closest in time.
Beyond the bibliography
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…