Fetching the paper…
Reading the bibliography…
We investigate two perturbation approaches to overcome conservatism that optimism based algorithms chronically suffer from in practice.
Approximation to bayes risk in repeated play
James Hannan · 1957
Earlier work this paper cites.
Handbook of mathematical functions. 1965, 1964
Milton Abramowitz and Irene A Stegun · 1964
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
Tze Leung Lai and Herbert Robbins · 1985
Earlier work this paper cites.
Reinforcement learning with immediate rewards and linear hypotheses
Naoki Abe, Alan W Biermann, and Philip M Long · 2003
Earlier work this paper cites.
Efficient algorithms for online decision problems
Adam Kalai and Santosh Vempala · 2005
Earlier work this paper cites.
Stochastic linear optimization under bandit feedback
Varsha Dani, Thomas P Hayes, and Sham M Kakade · 2008
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Yasin Abbasi-Yadkori, Dávid Pál, and Csaba Szepesvári · 2011
Earlier work this paper cites.
An empirical evaluation of thompson sampling
Olivier Chapelle and Lihong Li · 2011
Earlier work this paper cites.
Thompson sampling for contextual bandits with linear payoffs
Shipra Agrawal and Navin Goyal · 2013
Earlier work this paper cites.
Online linear optimization via smoothing
Jacob Abernethy, Chansoo Lee, Abhinav Sinha, and Ambuj Tewari · 2014
Cited alongside, same era.
Stochastic multi-armed-bandit problem with non-stationary rewards
Omar Besbes, Yonatan Gur, and Assaf Zeevi · 2014
Cited alongside, same era.
Fighting bandits with a new kind of smoothness
Jacob D Abernethy, Chansoo Lee, and Ambuj Tewari · 2015
Cited alongside, same era.
Linear thompson sampling revisited
Marc Abeille, Alessandro Lazaric, et al · 2017
Cited alongside, same era.
Attribution modeling increases efficiency of bidding in display advertising
Eustache Diemert, Julien Meynet, Pierre Galland, and Damien Lefortier · 2017
Cited alongside, same era.
Efficient contextual bandits in non-stationary worlds
Haipeng Luo, Chen-Yu Wei, Alekh Agarwal, and John Langford · 2018
Cited alongside, same era.
Hedging the drift: Learning to optimize under non-stationarity
Wang Chi Cheung, David Simchi-Levi, and Ruihao Zhu · 2019
Closest in time.
On the optimality of perturbations in stochastic and adversarial multi-armed bandit problems
Baekjin Kim and Ambuj Tewari · 2019
Closest in time.
Perturbed-history exploration in stochastic linear bandits
Branislav Kveton, Csaba Szepesvari, Mohammad Ghavamzadeh, and Craig Boutilier · 2019
Closest in time.
Weighted linear bandits for non-stationary environments
Yoan Russac, Claire Vernade, and Olivier Cappé · 2019
Closest in time.
Randomized exploration in generalized linear bandits
Branislav Kveton, Manzil Zaheer, Csaba Szepesvari, Lihong Li, Mohammad Ghavamzadeh, and Craig Boutilier · 2020
Closest in time.
Old dog learns new tricks: Randomized ucb for bandit problems
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Recent advances in multiarmed bandits for sequential decision making
Shipra Agrawal · 2019
Cited alongside, same era.
A new algorithm for non-stationary contextual bandits: Efficient, optimal, and parameter-free
Yifang Chen, Chung-Wei Lee, Haipeng Luo, and Chen-Yu Wei · 2019
Cited alongside, same era.
Sharan Vaswani, Abbas Mehrabian, Audrey Durand, and Branislav Kveton · 2020
Closest in time.
A simple approach for non-stationary linear bandits
Peng Zhao, Lijun Zhang, Yuan Jiang, and Zhi-Hua Zhou · 2020
Closest in time.
Non-stationary linear bandits revisited
Peng Zhao and Lijun Zhang · 2021
Closest in time.