Fetching the paper…
Reading the bibliography…
A generally intelligent agent must be able to teach itself how to solve problems in complex domains with minimal human supervision.
Dynamic Programming
Richard Bellman · 1957
Earlier work this paper cites.
Thistlethwaite’s 52-move algorithm
Morwen Thistlethwaite · 1981
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Richard S. Sutton · 1988
Earlier work this paper cites.
Q-learning
Christopher JCH Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Superflip requires 20 face turns
Michael Reid · 1995
Earlier work this paper cites.
Temporal difference learning and td-gammon
Gerald Tesauro · 1995
Earlier work this paper cites.
Finding optimal solutions to rubik’s cube using pattern databases
Richard E. Korf · 1997
Earlier work this paper cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 1998
Earlier work this paper cites.
Evolutionary algorithms for reinforcement learning
David E Moriarty, Alan C Schultz, and John J Grefenstette · 1999
Earlier work this paper cites.
Efficient selectivity and backup operators in monte-carlo tree search
Rémi Coulom · 2006
Earlier work this paper cites.
Bandit based monte-carlo planning
Levente Kocsis and Csaba Szepesvári · 2006
Earlier work this paper cites.
Efficient selectivity and backup operators in monte-carlo tree search
Rémi Coulom · 2007
Earlier work this paper cites.
Twenty-six moves suffice for rubik’s cube
Daniel Kunkle and Gene Cooperman · 2007
Cited alongside, same era.
Rubik’s cube can be solved in 34 quarter turns
Silviu Radu · 2007
Cited alongside, same era.
Twenty-two moves suffice for rubik’s cube®
Tomas Rokicki · 2010
Cited alongside, same era.
From machine learning to machine reasoning, 2011
Leon Bottou · 2011
Cited alongside, same era.
The rubik cube and gp temporal sequence learning: an initial study
Peter Lichodzijewski and Malcolm Heywood · 2011
Cited alongside, same era.
On the scalability of parallel uct
Richard B. Segal · 2011
Cited alongside, same era.
Deep learning for real-time atari game play using offline monte-carlo tree search planning
Neural combinatorial optimization with reinforcement learning, 2016
Irwan Bello, Hieu Pham, Quoc V. Le, Mohammad Norouzi, and Samy Bengio · 2016
Later among the works it cites.
Deep learning
Ian Goodfellow, Yoshua Bengio, Aaron Courville, and Yoshua Bengio · 2016
Later among the works it cites.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy P Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Later among the works it cites.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Christopher J. Maddison, Arthur Guez, Laurent Sifre, George van den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, Sander Dieleman, Dominik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy Lillicrap, Madeleine Leach, Koray Kavukcuoglu, Thore Graepel, and Demis Hassabis · 2016
Later among the works it cites.
Discovering rubik’s cube subgroups using coevolutionary gp: A five twist experiment
Robert J Smith, Stephen Kelly, and Malcolm I Heywood · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xiaoxiao Guo, Satinder Singh, Honglak Lee, Richard L Lewis, and Xiaoshi Wang · 2014
Cited alongside, same era.
Overview of mini-batch gradient descent, 2014
Geoffrey Hinton, Nitsh Srivastava, and Kevin Swersky · 2014
Cited alongside, same era.
God’s number is 26 in the quarter-turn metric
Tomas Rokicki · 2014
Cited alongside, same era.
The diameter of the rubik’s cube group is twenty
Tomas Rokicki, Herbert Kociemba, Morley Davidson, and John Dethridge · 2014
Cited alongside, same era.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
Thinking fast and slow with deep learning and tree search, 2017
Thomas Anthony, Zheng Tian, and David Barber · 2017
Later among the works it cites.
Rubik’s cube solver
Andrew Brown · 2017
Later among the works it cites.
Deep heuristic-learning in the rubik’s cube domain: an experimental evaluation
Robert Brunetto and Otakar Trunda · 2017
Later among the works it cites.
Mastering chess and shogi by self-play with a general reinforcement learning algorithm
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, et al · 2017
Later among the works it cites.
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, Yutian Chen, Timothy Lillicrap, Fan Hui, Laurent Sifre, George van den Driessche, Thore Graepel, and Demis Hassabis · 2017
Later among the works it cites.
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Later among the works it cites.
Kociemba
Maxim Tsoy · 2018
Closest in time.