Fetching the paper…
Reading the bibliography…
Planning and reinforcement learning are two key approaches to sequential decision making.
Dynamic programming
Richard Bellman · 1966
Earlier work this paper cites.
Some studies in machine learning using the game of checkers. II - Recent progress
Arthur L Samuel · 1967
Earlier work this paper cites.
Heuristic and analytic processes in reasoning
Jonathan St BT Evans · 1984
Earlier work this paper cites.
Dyna, an integrated architecture for learning, planning, and reacting
Richard S Sutton · 1991
Earlier work this paper cites.
Learning to act using real-time dynamic programming
Andrew G Barto, Steven J Bradtke, and Satinder P Singh · 1995
Earlier work this paper cites.
Exploration strategies for model-based learning in multi-agent systems: Exploration strategies
David Carmel and Shaul Markovitch · 1999
Earlier work this paper cites.
Deep blue
Murray Campbell, A Joseph Hoane Jr, and Feng-hsiung Hsu · 2002
Earlier work this paper cites.
Maximizing versus satisficing: Happiness is a matter of choice
Barry Schwartz, Andrew Ward, John Monterosso, Sonja Lyubomirsky, Katherine White, and Darrin R Lehman · 2002
Earlier work this paper cites.
World-championship-caliber Scrabble
Brian Sheppard · 2002
Earlier work this paper cites.
Maps of bounded rationality: Psychology for behavioral economics
Daniel Kahneman · 2003
Earlier work this paper cites.
Uncertainty-based competition between prefrontal and dorsolateral striatal systems for behavioral control
Nathaniel D Daw, Yael Niv, and Peter Dayan · 2005
Cited alongside, same era.
Bootstrapping from game tree search
Joel Veness, David Silver, Alan Blair, and William Uther · 2009
Cited alongside, same era.
Thinking, fast and slow
Daniel Kahneman · 2011
Cited alongside, same era.
A survey of monte carlo tree search methods
Cameron B Browne, Edward Powley, Daniel Whitehouse, Simon M Lucas, Peter I Cowling, Philipp Rohlfshagen, Stephen Tavener, Diego Perez, Spyridon Samothrakis, and Simon Colton · 2012
Cited alongside, same era.
Adam: A Method for Stochastic Optimization
Diederik Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Markov Decision Processes.: Discrete Stochastic Dynamic Programming
Martin L Puterman · 2014
Cited alongside, same era.
Artificial intelligence: a modern approach
Stuart J Russell and Peter Norvig · 2016
Later among the works it cites.
Thinking fast and slow with deep learning and tree search
Thomas Anthony, Zheng Tian, and David Barber · 2017
Later among the works it cites.
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Later among the works it cites.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
Kurtland Chua, Roberto Calandra, Rowan McAllister, and Sergey Levine · 2018
Later among the works it cites.
Beyond the One-Step Greedy Approach in Reinforcement Learning
Yonathan Efroni, Gal Dalal, Bruno Scherrer, and Shie Mannor · 2018
Later among the works it cites.
A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning to Search Better than Your Teacher
Kai-Wei Chang, Akshay Krishnamurthy, Alekh Agarwal, Hal Daume, and John Langford · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Cited alongside, same era.
Pybullet, a python module for physics simulation for games, robotics and machine learning
Erwin Coumans and Yunfei Bai · 2016
Cited alongside, same era.
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, et al · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Later among the works it cites.
Multi-Step Greedy and Approximate Real Time Dynamic Programming
Yonathan Efroni, Mohammad Ghavamzadeh, and Shie Mannor · 2019
Later among the works it cites.
Benchmarking Model-Based Reinforcement Learning
Tingwu Wang, Xuchan Bao, Ignasi Clavera, Jerrick Hoang, Yeming Wen, Eric Langlois, Shunshi Zhang, Guodong Zhang, Pieter Abbeel, and Jimmy Ba · 2019
Later among the works it cites.