Residual algorithms: Reinforcement learning with function approximation
Leemon Baird · 1995
Earlier work this paper cites.
A generalized reinforcement-learning model: Convergence and applications
Michael L. Littman and Csaba Szepesvari · 1996
Earlier work this paper cites.
Value function approximation in zero-sum markov games
Michail G. Lagoudakis and Ronald Parr · 2002
Earlier work this paper cites.
Correlated Q-learning
Amy Greenwald, Keith Hall, and Roberto Serrano · 2003
Earlier work this paper cites.
PAC model-free reinforcement learning
Alexander L Strehl, Lihong Li, Eric Wiewiora, John Langford, and Michael L Littman · 2006
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Earlier work this paper cites.
Contextual bandit learning with predictable rewards
Alekh Agarwal, Miroslav Dudík, Satyen Kale, John Langford, and Robert E. Schapire · 2012
Earlier work this paper cites.
Playing atari with deep reinforcement learningep reinforcement learning
Original
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, DaanWierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
Eluder dimension and the sample complexity of optimistic exploration
Daniel Russo and Benjamin Van Roy · 2013
Earlier work this paper cites.
Model-based reinforcement learning and the eluder dimension
Original
Ian Osband and Benjamin Van Roy · 2014
Earlier work this paper cites.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Earlier work this paper cites.
Approximate dynamic programming for two-player zero-sum markov games
Julien Perolat, Bruno Scherrer, Bilal Piot, and Olivier Pietquin · 2015
Earlier work this paper cites.
Approximate dynamic programming for two-player zero-sum Markov games
Julien Perolat, Bruno Scherrer, Bilal Piot, and Olivier Pietquin · 2015
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, and Marc Lanctot · 2015
Earlier work this paper cites.
Contextual decision processes with low bellman rank are pac-learnable
Original
Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, John Langford, and Robert E. Schapire · 2016
Earlier work this paper cites.
Safe, multi-agent, reinforcement learning for autonomous driving
Original
Shai Shalev-Shwartz, Shaked Shammah, and Amnon Shashua · 2016
Earlier work this paper cites.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Earlier work this paper cites.
Contextual decision processes with low bellman rank are PAC-learnable
Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, John Langford, and Robert E Schapire · 2017
Earlier work this paper cites.
Learning nash equilibrium for general-sum markov games from batch data
Julien Pérolat, Florian Strub, Bilal Piot, and Olivier Pietquin · 2017
Earlier work this paper cites.
Super-human ai for strategic reasoning: Beating top pros in heads-up no-limit texas hold’em
Tuomas Sandholm · 2017
Earlier work this paper cites.