Fetching the paper…
Reading the bibliography…
Humans learn to play video games significantly faster than the state-of-the-art reinforcement learning (RL) algorithms.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
William R Thompson · 1933
Earlier work this paper cites.
An introduction to computational learning theory
Michael J Kearns, Umesh Virkumar Vazirani, and Umesh Vazirani · 1994
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
R-max-a general polynomial time algorithm for near-optimal reinforcement learning
Ronen I Brafman and Moshe Tennenholtz · 2002
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Michael Kearns and Satinder Singh · 2002
Earlier work this paper cites.
A sparse sampling algorithm for near-optimal planning in large markov decision processes
Michael Kearns, Yishay Mansour, and Andrew Y Ng · 2002
Earlier work this paper cites.
Generalizing plans to new environments in relational mdps
Carlos Guestrin, Daphne Koller, Chris Gearhart, and Neal Kanodia · 2003
Earlier work this paper cites.
Bandit based monte-carlo planning
Levente Kocsis and Csaba Szepesvári · 2006
Earlier work this paper cites.
Bias and variance approximation in value function estimates
Shie Mannor, Duncan Simester, Peng Sun, and John N Tsitsiklis · 2007
Earlier work this paper cites.
Monte-carlo tree search: A new framework for game ai
Guillaume Chaslot, Sander Bakkes, Istvan Szita, and Pieter Spronck · 2008
Earlier work this paper cites.
An object-oriented representation for efficient reinforcement learning
Carlos Diuk, Andre Cohen, and Michael L Littman · 2008
Earlier work this paper cites.
Knows what it knows: a framework for self-aware learning
Lihong Li, Michael L Littman, and Thomas J Walsh · 2008
Earlier work this paper cites.
An analysis of model-based interval estimation for markov decision processes
Alexander L Strehl and Michael L Littman · 2008
Earlier work this paper cites.
Provably efficient learning with typed parametric models
Emma Brunskill, Bethany R Leffler, Lihong Li, Michael L Littman, and Nicholas Roy · 2009
Earlier work this paper cites.
Reinforcement learning in finite mdps: Pac analysis
Alexander L Strehl, Lihong Li, and Michael L Littman · 2009
Earlier work this paper cites.
Monte-carlo tree search
Guillaume Maurice Jean-Bernard Chaslot Chaslot · 2010
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Earlier work this paper cites.
Nested rollout policy adaptation for monte carlo tree search
Christopher D Rosin · 2011
Cited alongside, same era.
A survey of monte carlo tree search methods
Cameron B Browne, Edward Powley, Daniel Whitehouse, Simon M Lucas, Peter I Cowling, Philipp Rohlfshagen, Stephen Tavener, Diego Perez, Spyridon Samothrakis, and Simon Colton · 2012
Cited alongside, same era.
Efficient bayes-adaptive reinforcement learning using sample-based search
Arthur Guez, David Silver, and Peter Dayan · 2012
Cited alongside, same era.
The arcade learning environment: An evaluation platform for general agents
Marc G Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2013
Cited alongside, same era.
Object focused q-learning for autonomous agents
Luis C Cobo, Charles L Isbell, and Andrea L Thomaz · 2013
Cited alongside, same era.
Bayesian reinforcement learning: A survey
Mohammad Ghavamzadeh, Shie Mannor, Joelle Pineau, Aviv Tamar, et al · 2015
Deep visual foresight for planning robot motion
Chelsea Finn and Sergey Levine · 2017
Later among the works it cites.
Noisy networks for exploration
Meire Fortunato, Mohammad Gheshlaghi Azar, Bilal Piot, Jacob Menick, Ian Osband, Alex Graves, Vlad Mnih, Remi Munos, Demis Hassabis, Olivier Pietquin, et al · 2017
Later among the works it cites.
Mask r-cnn
Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick · 2017
Later among the works it cites.
Rainbow: Combining improvements in deep reinforcement learning
Matteo Hessel, Joseph Modayil, Hado Van Hasselt, Tom Schaul, Georg Ostrovski, Will Dabney, Dan Horgan, Bilal Piot, Mohammad Azar, and David Silver · 2017
Later among the works it cites.
Deep q-learning from demonstrations
Todd Hester, Matej Vecerik, Olivier Pietquin, Marc Lanctot, Tom Schaul, Bilal Piot, Dan Horgan, John Quan, Andrew Sendonaris, Gabriel Dulac-Arnold, et al · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Fast r-cnn
Ross Girshick · 2015
Cited alongside, same era.
The dependence of effective planning horizon on model accuracy
Nan Jiang, Alex Kulesza, Satinder Singh, and Richard Lewis · 2015
Cited alongside, same era.
Deepmpc: Learning deep latent features for model predictive control
Ian Lenz, Ross A Knepper, and Ashutosh Saxena · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
Unifying count-based exploration and intrinsic motivation
Marc Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Remi Munos · 2016
Cited alongside, same era.
Towards deep symbolic reinforcement learning
Marta Garnelo, Kai Arulkumaran, and Murray Shanahan · 2016
Cited alongside, same era.
Later among the works it cites.
Building machines that learn and think like people
Brenden M Lake, Tomer D Ullman, Joshua B Tenenbaum, and Samuel J Gershman · 2017
Later among the works it cites.
Count-based exploration with neural density models
Georg Ostrovski, Marc G Bellemare, Aaron van den Oord, and Rémi Munos · 2017
Later among the works it cites.
Yolo9000: better, faster, stronger
Joseph Redmon and Ali Farhadi · 2017
Later among the works it cites.
Melrose Roderick, Christopher Grimm, and Stefanie Tellex · 2017
Later among the works it cites.
Mastering chess and shogi by self-play with a general reinforcement learning algorithm
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, et al · 2017
Later among the works it cites.
# exploration: A study of count-based exploration for deep reinforcement learning
Haoran Tang, Rein Houthooft, Davis Foote, Adam Stooke, OpenAI Xi Chen, Yan Duan, John Schulman, Filip DeTurck, and Pieter Abbeel · 2017
Later among the works it cites.
Human learning in atari
Pedro A Tsividis, Thomas Pouncy, Jacqueline L Xu, Joshua B Tenenbaum, and Samuel J Gershman · 2017
Later among the works it cites.
Imagination-augmented agents for deep reinforcement learning
Théophane Weber, Sébastien Racanière, David P Reichert, Lars Buesing, Arthur Guez, Danilo Jimenez Rezende, Adria Puigdomenech Badia, Oriol Vinyals, Nicolas Heess, Yujia Li, et al · 2017
Later among the works it cites.
Playing hard exploration games by watching youtube
Yusuf Aytar, Tobias Pfaff, David Budden, Tom Le Paine, Ziyu Wang, and Nando de Freitas · 2018
Closest in time.
Efficient exploration through bayesian deep q-networks
Kamyar Azizzadenesheli, Emma Brunskill, and Animashree Anandkumar · 2018
Closest in time.
Investigating human priors for playing video games
Rachit Dubey, Pulkit Agrawal, Deepak Pathak, Thomas L Griffiths, and Alexei A Efros · 2018
Closest in time.