Fetching the paper…
Reading the bibliography…
We present an efficient algorithm for model-free episodic reinforcement learning on large (potentially continuous) state-action spaces.
Efficient Model-free Reinforcement Learning in Metric Spaces
Zhao Song and Wen Sun · 1905
Earlier work this paper cites.
Reinforcement Leaning in Feature Space: Matrix Bandit, Kernels, and Regret Bound
Lin F. Yang and Mengdi Wang · 1905
Earlier work this paper cites.
Learning from delayed rewards
Christopher John Cornish Hellaby Watkins · 1989
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Martin L. Puterman · 1994
Earlier work this paper cites.
On choosing and bounding probability metrics
Alison L Gibbs and Francis Edward Su · 2002
Earlier work this paper cites.
Ambulance location and relocation models
Luce Brotcorne, Gilbert Laporte, and Frederic Semet · 2003
Earlier work this paper cites.
Exploration in metric state spaces
Sham Kakade, Michael Kearns, and John Langford · 2003
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Peter Auer, Thomas Jaksch, and Ronald Ortner · 2009
Earlier work this paper cites.
Online optimization in x-armed bandits
Sébastien Bubeck, Gilles Stoltz, Csaba Szepesvári, and Rémi Munos · 2009
Earlier work this paper cites.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
Sébastien Bubeck, Nicolo Cesa-Bianchi, et al · 2012
Earlier work this paper cites.
Collaborative learning in networks
Winter Mason and Duncan J Watts · 2012
Earlier work this paper cites.
Online regret bounds for undiscounted continuous reinforcement learning
Ronald Ortner and Daniil Ryabko · 2012
Cited alongside, same era.
Adaptive aggregation for reinforcement learning in average reward markov decision processes
Ronald Ortner · 2013
Cited alongside, same era.
Model-based reinforcement learning and the eluder dimension
Ian Osband and Benjamin Van Roy · 2014
Cited alongside, same era.
Improved regret bounds for undiscounted continuous reinforcement learning
K. Lakshmanan, Ronald Ortner, and Daniil Ryabko · 2015
Cited alongside, same era.
Contextual Bandits with Similarity Information
Aleksandrs Slivkins · 2015
Cited alongside, same era.
Resource management with deep reinforcement learning
Hongzi Mao, Mohammad Alizadeh, Ishai Menache, and Srikanth Kandula · 2016
Cited alongside, same era.
Online optimization in cloud resource provisioning: Predictions, regrets, and algorithms
Joshua Comden, Sijie Yao, Niangjun Chen, Haipeng Xing, and Zhenhua Liu · 2019
Closest in time.
Q-learning with ucb exploration is sample efficient for infinite-horizon mdp
Kefan Dong, Yuanhao Wang, Xiaoyu Chen, and Liwei Wang · 2019
Closest in time.
Simon S Du, Yuping Luo, Ruosong Wang, and Hanrui Zhang · 2019
Closest in time.
Stochastic lipschitz q-learning, 2019
Xu Zhu David Dunson · 2019
Closest in time.
Bandits and experts in metric spaces
Robert Kleinberg, Aleksandrs Slivkins, and Eli Upfal · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Cited alongside, same era.
Is Q-learning provably efficient?
Jin C, Jordan M.I, Allen-Zhu Z, Bubeck S, and NeurIPS 2018 32nd Conference on Neural Information Processing Systems · 2018
Cited alongside, same era.
Bandit algorithms
Tor Lattimore and Csaba Szepesvári · 2018
Cited alongside, same era.
Q-learning with nearest neighbors
Devavrat Shah and Qiaomin Xie · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Cited alongside, same era.
Max Simchowitz and Kevin Jamieson · 2019
Closest in time.
Introduction to multi-armed bandits, 2019
Aleksandrs Slivkins · 2019
Closest in time.
Towards Practical Lipschitz Stochastic Bandits
Tianyu Wang, Weicheng Ye, Dawei Geng, and Cynthia Rudin · 2019
Closest in time.
Nonparametric contextual bandits in an unknown metric space, 2019
Nirandika Wanigasekara and Christina Lee Yu · 2019
Closest in time.
Sample-optimal parametric q-learning using linearly additive features
Lin Yang and Mengdi Wang · 2019
Closest in time.
Learning to control in metric space with optimal regret
Lin F Yang, Chengzhuo Ni, and Mengdi Wang · 2019
Closest in time.