Fetching the paper…
Reading the bibliography…
We study bandit model selection in stochastic environments.
“Exploration-Exploitation Tradeoff Using Variance Estimates in Multi-Armed Bandits”
Jean-Yves Audibert, R“’emi Munos and Csaba Szepesv“’ari · 1902
Earlier work this paper cites.
“Finite-Time Analysis of the Multiarmed Bandit Problem”
Peter Auer, Nicolo Cesa-Bianchi and Paul Fischer · 2002
Earlier work this paper cites.
“Stochastic Linear Optimization under Bandit Feedback”
Varsha Dani, Thomas. Hayes and Sham. Kakade · 2008
Earlier work this paper cites.
“An Analysis of Model-Based Interval Estimation for Markov Decision Processes”
Alexander Strehl and Michael Littman · 2008
Earlier work this paper cites.
“Near-Optimal Regret Bounds for Reinforcement Learning”
Thomas Jaksch, Ronald Ortner and Peter Auer · 2010
Earlier work this paper cites.
“A Contextual Bandit Approach to Personalized News Article Recommendation”
Lihong Li, Wei Chu, John Langford and Robert Schapire · 2010
Earlier work this paper cites.
“Improved Algorithms for Linear Stochastic Bandits”
Yasin Abbasi-Yadkori, D“’avid P“’al and Csaba Szepesv“’ari · 2011
Earlier work this paper cites.
“Contextual Bandits with Linear Payoff Functions”
Wei Chu, Lihong Li, Lev Reyzin and Robert Schapire · 2011
Earlier work this paper cites.
“Adaptive Bandits: Towards the Best History-Dependent Strategy”
Maillard Odalric and R“’emi Munos · 2011
Earlier work this paper cites.
“Online-to-Confidence-Set Conversions and Application to Sparse Stochastic Bandits”
Yasin Abbasi-Yadkori, David P“’al and Csaba Szepesv“’ari · 2012
Earlier work this paper cites.
“Regret Analysis of Stochastic and Nonstochastic Multi-Armed Bandit Problems”
S“’ebastien Bubeck and Nicol“‘o Cesa-Bianchi · 2012
Earlier work this paper cites.
“The Best of Both Worlds: Stochastic and Adversarial Bandits”
S“’ebastien Bubeck and Aleksandrs Slivkins · 2012
Cited alongside, same era.
“Bandit Theory Meets Compressed Sensing for High Dimensional Stochastic Linear Bandit”
Alexandra Carpentier and R“’emi Munos · 2012
Cited alongside, same era.
“Evaluation and Analysis of the Performance of the EXP3 Algorithm in Stochastic Environments”
Yevgeny Seldin, Csaba Szepesvari, Peter Auer and Yasin Abbasi-Yadkori · 2013
Cited alongside, same era.
“Corralling a Band of Bandit Algorithms”
Alekh Agarwal, Haipeng Luo, Behnam Neyshabur and Robert Schapire · 2017
Cited alongside, same era.
“Why is Posterior Sampling Better than Optimism for Reinforcement Learning?”
Ian Osband and Benjamin Van · 2017
Cited alongside, same era.
“Nonparametric Stochastic Contextual Bandits”
Melody Guan and Heinrich Jiang · 2018
“Provably Efficient Reinforcement Learning with Linear Function Approximation”
Chi Jin, Zhuoran Yang, Zhaoran Wang and Michael Jordan · 2020
Closest in time.
“Bandit Algorithms”
Tor Lattimore and Csaba Szepesv“’ari · 2020
Closest in time.
“Learning with Good Feature Representations in Bandits and in RL with a Generative Model”
Tor Lattimore, Csaba Szepesvari and Gellert Weisz · 2020
Closest in time.
“Regret Bound Balancing and Elimination for Model Selection in Bandits and RL”
Aldo Pacchiano, Christoph Dann, Claudio Gentile and Peter Bartlett · 2020
Closest in time.
“Learning Near Optimal Policies with Low Inherent Bellman Error”
Andrea Zanette, Alessandro Lazaric, Mykel Kochenderfer and Emma Brunskill · 2020
Closest in time.
“Corralling Stochastic Bandit Algorithms”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Almost Optimal Algorithms for Linear Stochastic Bandits with Heavy-Tailed Payoffs”
Han Shao, Xiaotian Yu, Irwin King and Michael Lyu · 2018
Cited alongside, same era.
“Model Selection for Contextual Bandits”
Dylan Foster, Akshay Krishnamurthy and Haipeng Luo · 2019
Cited alongside, same era.
“Introduction to Multi-Armed Bandits”
Aleksandrs Slivkins · 2019
Cited alongside, same era.
“Regret Balancing for Bandit and RL Model Selection”
Yasin Abbasi-Yadkori, Aldo Pacchiano and My Phan · 2020
Cited alongside, same era.
“Beyond UCB: Optimal and Efficient Contextual Bandits with Regression Oracles”
Dylan Foster and Alexander Rakhlin · 2020
Cited alongside, same era.
Raman Arora, Teodor Marinov and Mehryar Mohri · 2021
Closest in time.
“Dynamic Balancing for Model Selection in Bandits and RL”
Ashok Cutkosky, Christoph Dann, Abhimanyu Das, Claudio Gentile, Aldo Pacchiano and Manish Purohit · 2021
Closest in time.
“Online Model Selection for Reinforcement Learning with Function Approximation”
Jonathan Lee, Aldo Pacchiano, Vidya Muthukumar, Weihao Kong and Emma Brunskill · 2021
Closest in time.
“Best of Both Worlds Model Selection”
Aldo Pacchiano, Christoph Dann and Claudio Gentile · 2022
Closest in time.
“Provably Optimal Algorithms for Generalized Linear Contextual Bandits”
Lihong Li, Yu Lu and Dengyong Zhou · 2080
Closest in time.