Fetching the paper…
Reading the bibliography…
Reinforcement learning agents have traditionally been evaluated on small toy problems.
Neuronlike adaptive elements that can solve difficult learning control problems
A.G. Barto, R.S. Sutton, and C.W. Anderson · 1983
Earlier work this paper cites.
Learning from delayed rewards
Christopher John Cornish Hellaby Watkins · 1989
Earlier work this paper cites.
A reinforcement learning method for maximizing undiscounted rewards
Anton Schwartz · 1993
Earlier work this paper cites.
On-line Q-learning using connectionist systems
Gavin A Rummery and Mahesan Niranjan · 1994
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
Leemon Baird et al · 1995
Earlier work this paper cites.
Stable function approximation in dynamic programming
Geoffrey J Gordon · 1995
Earlier work this paper cites.
Reinforcement learning with selective perception and hidden state
Andrew Kachites McCallum · 1996
Cited alongside, same era.
Reinforcement learning with replacing eligibility traces
Satinder P Singh and Richard S Sutton · 1996
Cited alongside, same era.
Least-squares policy iteration
Michail G Lagoudakis and Ronald Parr · 2003
Cited alongside, same era.
Accelerating reinforcement learning through implicit imitation
Bob Price and Craig Boutilier · 2003
Cited alongside, same era.
A convergent o(n) algorithm for off-policy temporal-difference learning with linear function approximation
Richard S Sutton, Cs. Szepesvari, and H. R. Maei · 2008
Cited alongside, same era.
A survey of robot learning from demonstration
Brenna D Argall, Sonia Chernova, Manuela Veloso, and Brett Browning · 2009
Cited alongside, same era.
Gq ( λ \lambda ): A general gradient algorithm for temporal-difference prediction learning with eligibility traces
Hamid Reza Maei and Richard S Sutton · 2010
Later among the works it cites.
Game-Independent AI Agents for Playing Atari 2600 Console Games
Yavar Naddaf · 2010
Later among the works it cites.
Imitation learning by coaching
He He, Hal Daume III, and Jason Eisner · 2012
Later among the works it cites.
The arcade learning environment: An evaluation platform for general agents
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2013
Later among the works it cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…