Fetching the paper…
Reading the bibliography…
Monte Carlo Tree Search (MCTS) methods have proven powerful in planning for sequential decision-making problems such as Go and video games, but their performance can be poor when the planning depth and sampling trajectories are limited or when the rewards are sparse.
An on-line algorithm for dynamic reinforcement learning and planning in reactive environments
Jürgen Schmidhuber · 1990
Earlier work this paper cites.
TD-Gammon, a self-teaching backgammon program, achieves master-level play
Gerald Tesauro · 1994
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Andrew Y Ng, Daishi Harada, and Stuart Russell · 1999
Earlier work this paper cites.
Making rational decisions using adaptive utility elicitation
Urszula Chajewska, Daphne Koller, and Ronald Parr · 2000
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Andrew Y Ng and Stuart J Russell · 2000
Earlier work this paper cites.
The optimal reward baseline for gradient-based reinforcement learning
Lex Weaver and Nigel Tao · 2001
Earlier work this paper cites.
Bandit based monte-carlo planning
Levente Kocsis and Csaba Szepesvári · 2006
Earlier work this paper cites.
Potential-based shaping in model-based reinforcement learning
John Asmuth, Michael L Littman, and Robert Zinkov · 2008
Earlier work this paper cites.
Transpositions and move groups in Monte Carlo tree search
Benjamin E Childs, James H Brodeur, and Levente Kocsis · 2008
Earlier work this paper cites.
Learning deep architectures for AI
Yoshua Bengio · 2009
Earlier work this paper cites.
Convolutional deep belief networks for scalable unsupervised learning of hierarchical representations
Honglak Lee, Roger Grosse, Rajesh Ranganath, and Andrew Y Ng · 2009
Cited alongside, same era.
Intrinsically motivated reinforcement learning: An evolutionary perspective
Satinder Singh, Richard Lewis, Andrew Barto, and Jonathan Sorg · 2010
Cited alongside, same era.
Reward design via online gradient ascent
Jonathan Sorg, Richard L Lewis, and Satinder Singh · 2010
Cited alongside, same era.
Internal rewards mitigate agent boundedness
Jonathan Sorg, Satinder Singh, and Richard L Lewis · 2010
Cited alongside, same era.
Monte Carlo tree search and rapid action value estimation in computer Go
Sylvain Gelly and David Silver · 2011
Cited alongside, same era.
Sketch-based linear value function approximation
Marc Bellemare, Joel Veness, and Michael Bowling · 2012
State aggregation in Monte Carlo tree search
Jesse Hostetler, Alan Fern, and Tom Dietterich · 2014
Later among the works it cites.
Improving UCT planning via approximate homomorphisms
Nan Jiang, Satinder Singh, and Richard Lewis · 2014
Later among the works it cites.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2015
Later among the works it cites.
Classical planning with simulators: results on the atari video games
Nir Lipovetzky, Miquel Ramirez, and Hector Geffner · 2015
Later among the works it cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, and et al · 2015
Later among the works it cites.
Action-conditional video prediction using deep networks in ATARI games
Junhyuk Oh, Xiaoxiao Guo, Honglak Lee, Richard Lewis, and Satinder Singh · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Investigating contingency awareness using ATARI 2600 games
Marc G Bellemare, Joel Veness, and Michael Bowling · 2012
Cited alongside, same era.
Bayesian learning of recursively factored environments
Marc Bellemare, Joel Veness, and Michael Bowling · 2013
Cited alongside, same era.
The Arcade Learning Environment: an evaluation platform for general agents
Marc G Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2013
Cited alongside, same era.
Deep learning for real-time ATARI game play using offline Monte-Carlo tree search planning
Xiaoxiao Guo, Satinder Singh, Honglak Lee, Richard Lewis, and Xiaoshi Wang · 2014
Cited alongside, same era.
Deep learning in neural networks: An overview
Jürgen Schmidhuber · 2015
Later among the works it cites.
Trust region policy optimization
John Schulman, Sergey Levine, Philipp Moritz, Michael I Jordan, and Pieter Abbeel · 2015
Later among the works it cites.
Improving exploration in UCT using local manifolds
Sriram Srinivasan, Erik Talvitie, and Michael Bowling · 2015
Later among the works it cites.
Incentivizing exploration in reinforcement learning with deep predictive models
Bradly C Stadie, Sergey Levine, and Pieter Abbeel · 2015
Later among the works it cites.