Fetching the paper…
Reading the bibliography…
Reinforcement learning is a general and powerful framework with which to study and implement artificial intelligence.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
William R Thompson · 1933
Earlier work this paper cites.
Theory of Games and Economic Behavior
Oskar Morgenstern and John von Neumann · 1944
Earlier work this paper cites.
A proposal for the Dartmouth summer research project on artificial intelligence
John McCarthy, Marvin Minsky, Nathaniel Rochester, and Claude Shannon · 1955
Earlier work this paper cites.
Mind Children: The Future of Robot and Human Intelligence
Hans Moravec · 1988
Earlier work this paper cites.
Time-inconsistent preferences and consumer self-control
Stephen J Hoch and George Loewenstein · 1991
Earlier work this paper cites.
Simultaneous map building and localization for an autonomous mobile robot
J. J. Leonard and H. F. Durrant-Whyte · 1991
Earlier work this paper cites.
Curious model-building control systems
J. Schmidhuber · 1991
Earlier work this paper cites.
The role of exploration in learning control
S. Thrun · 1992
Earlier work this paper cites.
Q-learning
Christopher JCH Watkins and Peter Dayan · 1992
Earlier work this paper cites.
The coming technological singularity
Vernor Vinge · 1993
Earlier work this paper cites.
Dynamic Programming and Optimal Control
Dimitri P Bertsekas and John Tsitsiklis · 1995
Earlier work this paper cites.
Convolutional networks for images, speech, and time-series
Y. LeCun and Y. Bengio · 1995
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Richard S. Sutton and Andrew G. Barto · 1998
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour · 1999
Earlier work this paper cites.
Monte carlo localization for mobile robots
Sebastian Thrun, Dieter Fox, Wolfram Burgard, and Frank Dellaert · 1999
Earlier work this paper cites.
A theory of universal artificial intelligence based on algorithmic complexity
Marcus Hutter · 2000
Earlier work this paper cites.
A Bayesian framework for reinforcement learning
Malcolm Strens · 2000
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Peter Auer, Nicolò Cesa-Bianchi, and Paul Fischer · 2002
Earlier work this paper cites.
Self-optimizing and Pareto-optimal policies in general environments based on Bayes-mixtures
Marcus Hutter · 2002
Earlier work this paper cites.
Information Theory, Inference & Learning Algorithms
David J. C. MacKay · 2002
Earlier work this paper cites.
A gentle introduction to the universal algorithmic agent AIXI
Marcus Hutter · 2003
Earlier work this paper cites.
Probability Theory: The Logic of Science
Edwin T Jaynes · 2003
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Pieter Abbeel and Andrew Y Ng · 2004
Earlier work this paper cites.
Reinforcement Learning and Its Relationship to Supervised Learning , pages 45–63
Andrew G. Barto and Thomas G. Dietterich · 2004
Earlier work this paper cites.
Universal Artificial Intelligence
Marcus Hutter · 2005
Earlier work this paper cites.
The Singularity is Near: When Humans Transcend Biology
Ray Kurzweil · 2005
Earlier work this paper cites.
Pattern Recognition and Machine Learning
Christopher M Bishop · 2006
Earlier work this paper cites.
Bandit based monte-carlo planning
Levente Kocsis and Csaba Szepesvári · 2006
Earlier work this paper cites.
The Dartmouth College artificial intelligence conference: The next fifty years
James Moor · 2006
Earlier work this paper cites.
Machine Super Intelligence
Shane Legg · 2008
Cited alongside, same era.
An Introduction to Kolmogorov Complexity and Its Applications
Ming Li and Paul M. B. Vitányi · 2008
Cited alongside, same era.
The basic AI drives
Stephen M Omohundro · 2008
Cited alongside, same era.
Near-optimal regret bounds for reinforcement learning
Peter Auer, Thomas Jaksch, and Ronald Ortner · 2009
Cited alongside, same era.
The Elements of Statistical Learning
Trevor Hastie, Robert Tibshirani, and Jerome Friedman · 2009
Cited alongside, same era.
Open problems in universal induction & intelligence
Marcus Hutter · 2009
Cited alongside, same era.
AI Winter
Frederic P. Miller, Agnes F. Vandome, and John McBrewster · 2009
Reinforcejs, 2015
Andrej Karpathy · 2015
Later among the works it cites.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Later among the works it cites.
Bad universal priors and notions of optimality
Jan Leike and Marcus Hutter · 2015
Later among the works it cites.
Machine learning applications in genetics and genomics
Maxwell W. Libbrecht and William Stafford Noble · 2015
Later among the works it cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Later among the works it cites.
Fast and accurate recurrent neural network acoustic models for speech recognition
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Optimality issues of universal greedy agents with static priors
Laurent Orseau · 2010
Cited alongside, same era.
A comprehensive survey of data mining-based fraud detection research
Clifton Phua, Vincent C. S. Lee, Kate Smith-Miles, and Ross W. Gayler · 2010
Cited alongside, same era.
Artificial Intelligence. A Modern Approach
Stuart J Russell and Peter Norvig · 2010
Cited alongside, same era.
Monte-Carlo planning in large POMDPs
David Silver and Joel Veness · 2010
Cited alongside, same era.
Asymptotically optimal agents
Tor Lattimore and Marcus Hutter · 2011
Cited alongside, same era.
Hasim Sak, Andrew W. Senior, Kanishka Rao, and Françoise Beaufays · 2015
Later among the works it cites.
Deep learning in neural networks: An overview
Jürgen Schmidhuber · 2015
Later among the works it cites.
Rationality, optimism and guarantees in general reinforcement learning
Peter Sunehag and Marcus Hutter · 2015
Later among the works it cites.
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jonathon Shlens, and Zbigniew Wojna · 2015
Later among the works it cites.
Compress and control
Joel Veness, Marc G Bellemare, Marcus Hutter, Alvin Chua, and Guillaume Desjardins · 2015
Later among the works it cites.
Rationality: From AI to Zombies
Eliezer Yudkowsky · 2015
Later among the works it cites.
Theano: A Python framework for fast computation of mathematical expressions
Rami Al-Rfou, Guillaume Alain, Amjad Almahairi, Christof Angermueller, Dzmitry Bahdanau, Nicolas Ballas, Frédéric Bastien, Justin Bayer, Anatoly Belikov, Alexander Belopolsky, Yoshua Bengio, Arnaud Bergeron, James Bergstra, Valentin Bisson, Josh Bleecher Snyder, Nicolas Bouchard, Nicolas Boulanger-Lewandowski, Xavier Bouthillier, Alexandre de Brébisson, Olivier Breuleux, Pierre-Luc Carrier, Kyunghyun Cho, Jan Chorowski, Paul Christiano, Tim Cooijmans, Marc-Alexandre Côté, Myriam Côté, Aaron Courville, Yann N. Dauphin, Olivier Delalleau, Julien Demouth, Guillaume Desjardins, Sander Dieleman, Laurent Dinh, Mélanie Ducoffe, Vincent Dumoulin, Samira Ebrahimi Kahou, Dumitru Erhan, Ziye Fan, Orhan Firat, Mathieu Germain, Xavier Glorot, Ian Goodfellow, Matt Graham, Caglar Gulcehre, Philippe Hamel, Iban Harlouchet, Jean-Philippe Heng, Balázs Hidasi, Sina Honari, Arjun Jain, Sébastien Jean, Kai Jia, Mikhail Korobov, Vivek Kulkarni, Alex Lamb, Pascal Lamblin, Eric Larsen, César Laurent, Sean Lee, Simon Lefrancois, Simon Lemieux, Nicholas Léonard, Zhouhan Lin, Jesse A. Livezey, Cory Lorenz, Jeremiah Lowin, Qianli Ma, Pierre-Antoine Manzagol, Olivier Mastropietro, Robert T. McGibbon, Roland Memisevic, Bart van Merriënboer, Vincent Michalski, Mehdi Mirza, Alberto Orlandi, Christopher Pal, Razvan Pascanu, Mohammad Pezeshki, Colin Raffel, Daniel Renshaw, Matthew Rocklin, Adriana Romero, Markus Roth, Peter Sadowski, John Salvatier, François Savard, Jan Schlüter, John Schulman, Gabriel Schwartz, Iulian Vlad Serban, Dmitriy Serdyuk, Samira Shabanian, Étienne Simon, Sigurd Spieckermann, S. Ramana Subramanyam, Jakub Sygnowski, Jérémie Tanguay, Gijs van Tulder, Joseph Turian, Sebastian Urban, Pascal Vincent, Francesco Visin, Harm de Vries, David Warde-Farley, Dustin J. Webb, Matthew Willson, Kelvin Xu, Lijun Xue, Li Yao, Saizheng Zhang, and Ying Zhang · 2016
Later among the works it cites.
Unifying count-based exploration and intrinsic motivation
Marc G. Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Rémi Munos · 2016
Later among the works it cites.
Chord diagram example, 2016
Michael Bostock · 2016
Later among the works it cites.
OpenAI Gym, 2016
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Later among the works it cites.
Keras-js, 2016
Leon Chen · 2016
Later among the works it cites.
Avoiding wireheading with value reinforcement learning
Tom Everitt and Marcus Hutter · 2016
Later among the works it cites.
Self-modification of policy and utility function in rational agents
Tom Everitt, Daniel Filan, Mayank Daswani, and Marcus Hutter · 2016
Later among the works it cites.
What we learned in Seoul with AlphaGo
Google · 2016
Later among the works it cites.
Preparing for the future of artificial intelligence, 2016
John Holdren, Ed Felten, Terah Lyons, and Michael Garris · 2016
Later among the works it cites.
Playing FPS games with deep reinforcement learning
Guillaume Lample and Devandra Singh Chaplot · 2016
Later among the works it cites.
Thompson sampling is asymptotically optimal in general environments
Jan Leike, Tor Lattimore, Laurent Orseau, and Marcus Hutter · 2016
Later among the works it cites.
Death and suicide in universal artificial intelligence
Jarryd Martin, Tom Everitt, and Marcus Hutter · 2016
Later among the works it cites.
Future progress in artificial intelligence: A survey of expert opinion
Vincent C Müller and Nick Bostrom · 2016
Later among the works it cites.
Mastering the game of Go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George van den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, Sander Dieleman, Dominik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy Lillicrap, Madeleine Leach, Koray Kavukcuoglu, Thore Graepel, and Demis Hassabis · 2016
Later among the works it cites.
Tensorflow Playground, 2016
Daniel Smilkov and Shan Carter · 2016
Later among the works it cites.
WaveNet: A Generative Model for Raw Audio
A. van den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu · 2016
Later among the works it cites.
How to use t-sne effectively
Martin Wattenberg, Fernanda Viégas, and Ian Johnson · 2016
Later among the works it cites.