Fetching the paper…
Reading the bibliography…
In this paper, we consider the problem of online learning of Markov decision processes (MDPs) with very large state spaces.
On general minimax theorems
Maurice Sion et al · 1958
Earlier work this paper cites.
On minimum volume ellipsoids containing part of a given ellipsoid
Michael J Todd · 1982
Earlier work this paper cites.
Rates of convergence in the central limit theorem for empirical processes
Pascal Massart · 1986
Earlier work this paper cites.
Learning from delayed rewards
Christopher John Cornish Hellaby Watkins · 1989
Earlier work this paper cites.
On learning sets and functions
Balas K Natarajan · 1989
Earlier work this paper cites.
Q-learning
Christopher JCH Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Decision theoretic generalizations of the pac model for neural net and other learning applications
David Haussler · 1992
Earlier work this paper cites.
Asynchronous stochastic approximation and q-learning
John N Tsitsiklis · 1994
Earlier work this paper cites.
Dynamic programming and optimal control
Dimitri P Bertsekas, Dimitri P Bertsekas, Dimitri P Bertsekas, and Dimitri P Bertsekas · 1995
Earlier work this paper cites.
Gambling in a rigged casino: The adversarial multi-armed bandit problem
Peter Auer, Nicolo Cesa-Bianchi, Yoav Freund, and Robert E Schapire · 1995
Earlier work this paper cites.
Stable function approximation in dynamic programming
Geoffrey J Gordon · 1995
Earlier work this paper cites.
Sphere packing numbers for subsets of the boolean n-cube with bounded vapnik-chervonenkis dimension
David Haussler · 1995
Earlier work this paper cites.
An analysis of bid-price controls for network revenue management
Kalyan Talluri and Garrett Van Ryzin · 1998
Earlier work this paper cites.
Learning rates for q-learning
Eyal Even-Dar and Yishay Mansour · 2003
Earlier work this paper cites.
Pac model-free reinforcement learning
Alexander L Strehl, Lihong Li, Eric Wiewiora, John Langford, and Michael L Littman · 2006
Earlier work this paper cites.
Approximate Dynamic Programming: Solving the curses of dimensionality
Warren B Powell · 2007
Earlier work this paper cites.
Dynamic bid prices in revenue management
Daniel Adelman · 2007
Cited alongside, same era.
Learning near-optimal policies with bellman-residual minimization based fitted policy iteration and a single sample path
András Antos, Csaba Szepesvári, and Rémi Munos · 2008
Cited alongside, same era.
Finite-time bounds for fitted value iteration
Rémi Munos and Csaba Szepesvári · 2008
Cited alongside, same era.
Efficient optimal learning for contextual bandits
Miroslav Dudik, Daniel Hsu, Satyen Kale, Nikos Karampatziakis, John Langford, Lev Reyzin, and Tong Zhang · 2011
Cited alongside, same era.
Speedy q-learning
Mohammad Gheshlaghi Azar, Remi Munos, Mohammad Ghavamzadeh, and Hilbert Kappen · 2011
Cited alongside, same era.
Finite-sample analysis of least-squares policy iteration
Alessandro Lazaric, Mohammad Ghavamzadeh, and Rémi Munos · 2012
Cited alongside, same era.
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Later among the works it cites.
Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates
Shixiang Gu, Ethan Holly, Timothy Lillicrap, and Sergey Levine · 2017
Later among the works it cites.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Later among the works it cites.
Efficient reinforcement learning in deterministic systems with value function generalization
Zheng Wen and Benjamin Van Roy · 2017
Later among the works it cites.
Is Q-learning provably efficient?
Chi Jin, Zeyuan Allen-Zhu, Sebastien Bubeck, and Michael I Jordan · 2018
Later among the works it cites.
Variance reduced value iteration and faster algorithms for solving markov decision processes
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Convergence of stochastic processes
David Pollard · 2012
Cited alongside, same era.
A probabilistic theory of pattern recognition
Luc Devroye, László Györfi, and Gábor Lugosi · 2013
Cited alongside, same era.
Taming the monster: A fast and simple algorithm for contextual bandits
Alekh Agarwal, Daniel Hsu, Satyen Kale, John Langford, Lihong Li, and Robert Schapire · 2014
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
Unifying count-based exploration and intrinsic motivation
Marc Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Remi Munos · 2016
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Cited alongside, same era.
Aaron Sidford, Mengdi Wang, Xian Wu, and Yinyu Ye · 2018
Later among the works it cites.
On oracle-efficient pac rl with rich observations
Christoph Dann, Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, John Langford, and Robert E Schapire · 2018
Later among the works it cites.
Near-optimal time and sample complexities for solving discounted markov decision process with a generative model
Aaron Sidford, Mengdi Wang, Xian Wu, Lin F Yang, and Yinyu Ye · 2018
Later among the works it cites.
Global convergence of policy gradient methods for linearized control problems
Maryam Fazel, Rong Ge, Sham M Kakade, and Mehran Mesbahi · 2018
Later among the works it cites.
Andrea Zanette and Emma Brunskill · 2019
Closest in time.
Information-theoretic considerations in batch reinforcement learning
Jinglin Chen and Nan Jiang · 2019
Closest in time.
Provably efficient reinforcement learning with linear function approximation
Chi Jin, Zhuoran Yang, Zhaoran Wang, and Michael I Jordan · 2019
Closest in time.
Provably efficient rl with rich observations via latent state decoding
Simon S. Du, Akshay Krishnamurthy, Nan Jiang, Alekh Agarwal, Miroslav Dudik, and John Langford · 2019
Closest in time.
Martin J Wainwright · 2019
Closest in time.
https://arxiv.org/abs/1905.10389
Yang Lin and Wang Mengdi · 2019
Closest in time.