Fetching the paper…
Reading the bibliography…
In typical reinforcement learning (RL), the environment is assumed given and the goal of the learning is to identify an optimal policy for the agent taking actions through its interactions with the environment.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
Richard S Sutton · 1990
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Prioritized sweeping: Reinforcement learning with less data and less time
Andrew W Moore and Christopher G Atkeson · 1993
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
Michael L Littman · 1994
Earlier work this paper cites.
Reinforcement learning: A survey
Leslie Pack Kaelbling, Michael L Littman, and Andrew W Moore · 1996
Earlier work this paper cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 1998
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A McAllester, Satinder P Singh, Yishay Mansour, et al · 1999
Earlier work this paper cites.
Algorithmic mechanism design
Noam Nisan and Amir Ronen · 2001
Earlier work this paper cites.
Reinforcement learning for true adaptive traffic signal control
Baher Abdulhai, Rob Pringle, and Grigoris J Karakoulas · 2003
Earlier work this paper cites.
Nash q-learning for general-sum stochastic games
Junling Hu and Michael P Wellman · 2003
Cited alongside, same era.
Traffic signal timing optimisation based on genetic algorithm approach, including drivers’ routing
Halim Ceylan and Michael GH Bell · 2004
Cited alongside, same era.
A framework for sequential planning in multi-agent settings
Piotr J Gmytrasiewicz and Prashant Doshi · 2005
Cited alongside, same era.
Robust reinforcement learning
Jun Morimoto and Kenji Doya · 2005
Cited alongside, same era.
The complexity of the elementary interface: shopping space
Alan Penn · 2005
Cited alongside, same era.
Automatic design of balanced board games
Vincent Hom and Joe Marks · 2007
Cited alongside, same era.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Later among the works it cites.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Later among the works it cites.
A comprehensive survey on safe reinforcement learning
Javier Garcıa and Fernando Fernández · 2015
Later among the works it cites.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Later among the works it cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Later among the works it cites.
keras-rl
Matthias Plappert · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
An experiment in automatic game design
Julian Togelius and Jurgen Schmidhuber · 2008
Cited alongside, same era.
Reward design via online gradient ascent
Jonathan Sorg, Richard L Lewis, and Satinder P Singh · 2010
Cited alongside, same era.
Search-based procedural content generation: A taxonomy and survey
Julian Togelius, Georgios N Yannakakis, Kenneth O Stanley, and Cameron Browne · 2011
Cited alongside, same era.
A comprehensive survey of multiagent reinforcement learning
Lucian Busoniu and Bart De Schutter
Cited in the paper.
Later among the works it cites.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Later among the works it cites.
Real-time bidding by reinforcement learning in display advertising
Han Cai, Kan Ren, Weinan Zhang, Kleanthis Malialis, Jun Wang, Yong Yu, and Defeng Guo · 2017
Closest in time.
Feudal networks for hierarchical reinforcement learning
Alexander Sasha Vezhnevets, Simon Osindero, Tom Schaul, Nicolas Heess, Max Jaderberg, David Silver, and Koray Kavukcuoglu · 2017
Closest in time.