Fetching the paper…
Reading the bibliography…
We introduce the technique of adaptive discretization to design an efficient model-based episodic reinforcement learning algorithm in large (potentially continuous) state-action spaces.
Learning from delayed rewards
Christopher John Cornish Hellaby Watkins · 1989
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Martin L. Puterman · 1994
Earlier work this paper cites.
Reinforcement learning with soft state aggregation
Satinder P Singh, Tommi Jaakkola, and Michael I Jordan · 1995
Earlier work this paper cites.
Decision tree function approximation in reinforcement learning
Larry D Pyeatt, Adele E Howe, et al · 2001
Earlier work this paper cites.
Exploration in metric state spaces
Sham Kakade, Michael J Kearns, and John Langford · 2003
Earlier work this paper cites.
Basis function adaptation in temporal difference reinforcement learning
Ishai Menache, Shie Mannor, and Nahum Shimkin · 2005
Earlier work this paper cites.
Automatic basis function construction for approximate dynamic programming and reinforcement learning
Philipp W Keller, Shie Mannor, and Doina Precup · 2006
Earlier work this paper cites.
Evolutionary function approximation for reinforcement learning
Shimon Whiteson and Peter Stone · 2006
Earlier work this paper cites.
Self-optimizing memory controllers: A reinforcement learning approach
Engin Ipek, Onur Mutlu, José F. Martínez, and Rich Caruana · 2008
Earlier work this paper cites.
Dctcp: Efficient packet transport for the commoditized data center
Mohammad Alizadeh, Albert Greenberg, Dave Maltz, Jitu Padhye, Parveen Patel, Balaji Prabhakar, Sudipta Sengupta, and Murari Sridharan · 2010
Earlier work this paper cites.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
Sébastien Bubeck, Nicolo Cesa-Bianchi, et al · 2012
Earlier work this paper cites.
Elements of information theory
Thomas M Cover and Joy A Thomas · 2012
Earlier work this paper cites.
Collaborative learning in networks
Winter Mason and Duncan J Watts · 2012
Earlier work this paper cites.
pfabric: Minimal near-optimal datacenter transport
Mohammad Alizadeh, Shuang Yang, Milad Sharif, Sachin Katti, Nick McKeown, Balaji Prabhakar, and Scott Shenker · 2013
Earlier work this paper cites.
Minimax PAC bounds on the sample complexity of reinforcement learning with a generative model
Mohammad Gheshlaghi Azar, Rémi Munos, and Hilbert J Kappen · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
Eluder dimension and the sample complexity of optimistic exploration
Daniel Russo and Benjamin Van Roy · 2013
Earlier work this paper cites.
Model-based reinforcement learning and the eluder dimension
Ian Osband and Benjamin Van Roy · 2014
Cited alongside, same era.
Contextual bandits with similarity information
Aleksandrs Slivkins · 2014
Cited alongside, same era.
Sample complexity of episodic fixed-horizon reinforcement learning
Christoph Dann and Emma Brunskill · 2015
Cited alongside, same era.
Improved regret bounds for undiscounted continuous reinforcement learning
Kailasam Lakshmanan, Ronald Ortner, and Daniil Ryabko · 2015
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Provably efficient reinforcement learning with linear function approximation
Chi Jin, Zhuoran Yang, Zhaoran Wang, and Michael I Jordan · 2019
Later among the works it cites.
Bandits and experts in metric spaces
Robert Kleinberg, Aleksandrs Slivkins, and Eli Upfal · 2019
Later among the works it cites.
Non-asymptotic gap-dependent regret bounds for tabular mdps
Max Simchowitz and Kevin G Jamieson · 2019
Later among the works it cites.
Adaptive discretization for episodic reinforcement learning in metric spaces
Sean R. Sinclair, Siddhartha Banerjee, and Christina Lee Yu · 2019
Later among the works it cites.
Introduction to multi-armed bandits
Aleksandrs Slivkins · 2019
Later among the works it cites.
Efficient model-free reinforcement learning in metric spaces
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Cited alongside, same era.
Adaptive state space partitioning of markov decision processes for elastic resource management
K. Lolos, I. Konstantinou, V. Kantere, and N. Koziris · 2017
Cited alongside, same era.
Mastering chess and shogi by self-play with a general reinforcement learning algorithm
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, et al · 2017
Cited alongside, same era.
Cellular network traffic scheduling with deep reinforcement learning
Sandeep Chinchali, Pan Hu, Tianshu Chu, Manu Sharma, Manu Bansal, Rakesh Misra, Marco Pavone, and Sachin Katti · 2018
Cited alongside, same era.
Is Q-Learning provably efficient?
Chi Jin, Zeyuan Allen-Zhu, Sebastien Bubeck, and Michael I Jordan · 2018
Cited alongside, same era.
Competitive caching with machine learned advice
Thodoris Lykouris and Sergei Vassilvitskii · 2018
Cited alongside, same era.
Zhao Song and Wen Sun · 2019
Later among the works it cites.
Optimism in reinforcement learning with generalized linear function approximation
Yining Wang, Ruosong Wang, Simon S Du, and Akshay Krishnamurthy · 2019
Later among the works it cites.
Nonparametric contextual bandits in metric spaces with unknown metric
Nirandika Wanigasekara and Christina Yu · 2019
Later among the works it cites.
Sharp asymptotic and finite-sample rates of convergence of empirical measures in wasserstein distance
Jonathan Weed, Francis Bach, et al · 2019
Later among the works it cites.
Learning to control in metric space with optimal regret
Lin F Yang, Chengzhuo Ni, and Mengdi Wang · 2019
Later among the works it cites.
Andrea Zanette and Emma Brunskill · 2019
Later among the works it cites.
Limiting extrapolation in linear approximate value iteration
Andrea Zanette, Alessandro Lazaric, Mykel J Kochenderfer, and Emma Brunskill · 2019
Later among the works it cites.
Model-based reinforcement learning with value-targeted regression
Alex Ayoub, Zeyu Jia, Csaba Szepesvari, Mengdi Wang, and Lin F Yang · 2020
Closest in time.
Provably adaptive reinforcement learning in metric spaces, 2020
Tongyi Cao and Akshay Krishnamurthy · 2020
Closest in time.
Regret bounds for kernel-based reinforcement learning
Omar Darwiche Domingues, Pierre Ménard, Matteo Pirotta, Emilie Kaufmann, and Michal Valko · 2020
Closest in time.
Provably efficient reinforcement learning with general value function approximation
Ruosong Wang, Ruslan Salakhutdinov, and Lin F Yang · 2020
Closest in time.