Fetching the paper…
Reading the bibliography…
We study the problem of reinforcement learning (RL) with low (policy) switching cost - a problem well-motivated by real-life RL applications in which deployments of new policies are costly and the number of policy updates must be low.
Reinforcement learning: An introduction , volume 1
Richard S Sutton and Andrew G Barto · 1998
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Peter Auer, Nicolo Cesa-Bianchi, and Paul Fischer · 2002
Earlier work this paper cites.
R-max-a general polynomial time algorithm for near-optimal reinforcement learning
Ronen I Brafman and Moshe Tennenholtz · 2002
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Michael Kearns and Satinder Singh · 2002
Earlier work this paper cites.
Empirical bernstein bounds and sample variance penalization
Andreas Maurer and Massimiliano Pontil · 2009
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Earlier work this paper cites.
Nearly minimax optimal reward-free reinforcement learning
Zihan Zhang, Simon S Du, and Xiangyang Ji · 2010
Earlier work this paper cites.
Almost optimal model-free reinforcement learningvia reference-advantage decomposition
Zihan Zhang, Yuan Zhou, and Xiangyang Ji · 2010
Earlier work this paper cites.
Online learning with switching costs and other adaptive adversaries
Nicolo Cesa-Bianchi, Ofer Dekel, and Ohad Shamir · 2013
Earlier work this paper cites.
(more) efficient reinforcement learning via posterior sampling
Ian Osband, Daniel Russo, and Benjamin Van Roy · 2013
Earlier work this paper cites.
Batched bandit problems
Vianney Perchet, Philippe Rigollet, Sylvain Chassang, and Erik Snowberg · 2016
Earlier work this paper cites.
Machine-learning-assisted materials discovery using failed experiments
Paul Raccuglia, Katherine C Elbert, Philip DF Adler, Casey Falk, Malia B Wenny, Aurelio Mollo, Matthias Zeller, Sorelle A Friedler, Joshua Schrier, and Alexander J Norquist · 2016
Earlier work this paper cites.
Posterior sampling for reinforcement learning: worst-case regret bounds
Shipra Agrawal and Randy Jia · 2017
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Cited alongside, same era.
Unifying pac and regret: Uniform pac bounds for episodic reinforcement learning
Christoph Dann, Tor Lattimore, and Emma Brunskill · 2017
Cited alongside, same era.
Is q-learning provably efficient?
Chi Jin, Zeyuan Allen-Zhu, Sebastien Bubeck, and Michael I Jordan · 2018
Cited alongside, same era.
Learning to optimize join queries with deep reinforcement learning
Sanjay Krishnan, Zongheng Yang, Ken Goldberg, Joseph Hellerstein, and Ion Stoica · 2018
Cited alongside, same era.
Experience-driven networking: A deep reinforcement learning based approach
Deployment-efficient reinforcement learning via model-based offline optimization
Tatsuya Matsushima, Hiroki Furuta, Yutaka Matsuo, Ofir Nachum, and Shixiang Gu · 2020
Later among the works it cites.
On reward-free reinforcement learning with linear function approximation
Ruosong Wang, Simon S Du, Lin Yang, and Russ R Salakhutdinov · 2020
Later among the works it cites.
Provably efficient reward-agnostic navigation with linear value iteration
Andrea Zanette, Alessandro Lazaric, Mykel J Kochenderfer, and Emma Brunskill · 2020
Later among the works it cites.
Reinforcement learning based recommender systems: A survey
M Mehdi Afsar, Trafford Crump, and Behrouz Far · 2021
Later among the works it cites.
Regret bounds for batched bandits
Hossein Esfandiari, Amin Karbasi, Abbas Mehrabian, and Vahab Mirrokni · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhiyuan Xu, Jian Tang, Jingsong Meng, Weiyi Zhang, Yanzhi Wang, Chi Harold Liu, and Dejun Yang · 2018
Cited alongside, same era.
Provably efficient q-learning with low switching cost
Yu Bai, Tengyang Xie, Nan Jiang, and Yu-Xiang Wang · 2019
Cited alongside, same era.
Policy certificates: Towards accountable reinforcement learning
Christoph Dann, Lihong Li, Wei Wei, and Emma Brunskill · 2019
Cited alongside, same era.
Batched multi-armed bandits problem
Zijun Gao, Yanjun Han, Zhimei Ren, and Zhengqing Zhou · 2019
Cited alongside, same era.
Phase transitions and cyclic phenomena in bandits with switching constraints
David Simchi-Levi and Yunzong Xu · 2019
Cited alongside, same era.
Tighter problem-dependent regret bounds in reinforcement learning without domain knowledge using value function bounds
Andrea Zanette and Emma Brunskill · 2019
Cited alongside, same era.
Reward-free exploration for reinforcement learning
Chi Jin, Akshay Krishnamurthy, Max Simchowitz, and Tiancheng Yu
Cited in the paper.
Minbo Gao, Tianle Xie, Simon S Du, and Lin F Yang · 2021
Later among the works it cites.
Adaptive reward-free exploration
Emilie Kaufmann, Pierre Ménard, Omar Darwiche Domingues, Anders Jonsson, Edouard Leurent, and Michal Valko · 2021
Later among the works it cites.
Fast active learning for pure exploration in reinforcement learning
Pierre Ménard, Omar Darwiche Domingues, Anders Jonsson, Emilie Kaufmann, Edouard Leurent, and Michal Valko · 2021
Later among the works it cites.
Provably efficient reinforcement learning with linear function approximation under adaptivity constraints
Tianhao Wang, Dongruo Zhou, and Quanquan Gu · 2021
Later among the works it cites.
Reinforcement learning in healthcare: A survey
Chao Yu, Jiming Liu, Shamim Nemati, and Guosheng Yin · 2021
Later among the works it cites.
Towards deployment-efficient reinforcement learning: Lower bound and optimality
Jiawei Huang, Jinglin Chen, Li Zhao, Tao Qin, Nan Jiang, and Tie-Yan Liu · 2022
Closest in time.