Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) problems are often phrased in terms of Markov decision processes (MDPs).
Über formal unentscheidbare Sätze der Principia Mathematica und verwandter Systeme I
Kurt Gödel · 1931
Earlier work this paper cites.
Über formal unentscheidbare Sätze der Principia Mathematica und verwandter Systeme I
Kurt Gödel · 1931
Earlier work this paper cites.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
William R Thompson · 1933
Earlier work this paper cites.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
William R Thompson · 1933
Earlier work this paper cites.
Studies in the logic of confirmation (I.)
Carl G Hempel · 1945
Earlier work this paper cites.
Studies in the logic of confirmation (I.)
Carl G Hempel · 1945
Earlier work this paper cites.
Programming a computer for playing chess
Claude E Shannon · 1950
Earlier work this paper cites.
Programming a computer for playing chess
Claude E Shannon · 1950
Earlier work this paper cites.
Introduction to Metamathematics
Stephen Cole Kleene · 1952
Earlier work this paper cites.
Introduction to Metamathematics
Stephen Cole Kleene · 1952
Earlier work this paper cites.
Stochastic Processes
Joseph L. Doob · 1953
Earlier work this paper cites.
Stochastic Processes
Joseph L. Doob · 1953
Earlier work this paper cites.
Positive reinforcement produced by electrical stimulation of septal area and other regions of rat brain
James Olds and Peter Milner · 1954
Earlier work this paper cites.
Positive reinforcement produced by electrical stimulation of septal area and other regions of rat brain
James Olds and Peter Milner · 1954
Earlier work this paper cites.
A proposal for the Dartmouth summer research project on artificial intelligence
John McCarthy, Marvin Minsky, Nathaniel Rochester, and Claude Shannon · 1955
Earlier work this paper cites.
A proposal for the Dartmouth summer research project on artificial intelligence
John McCarthy, Marvin Minsky, Nathaniel Rochester, and Claude Shannon · 1955
Earlier work this paper cites.
The paradox of confirmation
Irving John Good · 1960
Earlier work this paper cites.
The paradox of confirmation
Irving John Good · 1960
Earlier work this paper cites.
Le Problème Logique de L’Induction
Jean Nicod · 1961
Earlier work this paper cites.
Le Problème Logique de L’Induction
Jean Nicod · 1961
Earlier work this paper cites.
Merging of opinions with increasing information
David Blackwell and Lester Dubins · 1962
Earlier work this paper cites.
Merging of opinions with increasing information
David Blackwell and Lester Dubins · 1962
Earlier work this paper cites.
The paradox of confirmation
John L Mackie · 1963
Earlier work this paper cites.
The paradox of confirmation
John L Mackie · 1963
Earlier work this paper cites.
A formal theory of inductive inference. Parts 1 and 2
Ray Solomonoff · 1964
Earlier work this paper cites.
A formal theory of inductive inference. Parts 1 and 2
Ray Solomonoff · 1964
Earlier work this paper cites.
Speculations concerning the first ultraintelligent machine
Irving John Good · 1965
Earlier work this paper cites.
Speculations concerning the first ultraintelligent machine
Irving John Good · 1965
Earlier work this paper cites.
The white shoe is a red herring
Irving John Good · 1967
Earlier work this paper cites.
The white shoe: No red herring
Carl G Hempel · 1967
Earlier work this paper cites.
Mathematical Logic
Joseph R Shoenfield · 1967
Earlier work this paper cites.
The white shoe is a red herring
Irving John Good · 1967
Earlier work this paper cites.
The white shoe: No red herring
Carl G Hempel · 1967
Earlier work this paper cites.
Mathematical Logic
Joseph R Shoenfield · 1967
Earlier work this paper cites.
The paradoxes of confirmation: A survey
Richard G Swinburne · 1971
Earlier work this paper cites.
The paradoxes of confirmation: A survey
Richard G Swinburne · 1971
Earlier work this paper cites.
On the notion of a random sequence
Leonid A Levin · 1973
Earlier work this paper cites.
On the notion of a random sequence
Leonid A Levin · 1973
Earlier work this paper cites.
A universal algorithm for sequential data compression
Jacob Ziv and Abraham Lempel · 1977
Earlier work this paper cites.
Complexity-based induction systems: Comparisons and convergence theorems
Ray Solomonoff · 1978
Earlier work this paper cites.
Complexity-based induction systems: Comparisons and convergence theorems
Ray Solomonoff · 1978
Earlier work this paper cites.
Bandit processes and dynamic allocation indices
John Gittins · 1979
Earlier work this paper cites.
Bandit processes and dynamic allocation indices
John Gittins · 1979
Earlier work this paper cites.
Universal upper bound on the entropy-to-energy ratio for bounded systems
Jacob D Bekenstein · 1981
Earlier work this paper cites.
Universal upper bound on the entropy-to-energy ratio for bounded systems
Jacob D Bekenstein · 1981
Earlier work this paper cites.
On the relation between descriptional complexity and algorithmic probability
Péter Gács · 1983
Earlier work this paper cites.
On the relation between descriptional complexity and algorithmic probability
Péter Gács · 1983
Earlier work this paper cites.
The complexity of Markov decision processes
Christos H Papadimitriou and John N Tsitsiklis · 1987
Earlier work this paper cites.
The complexity of Markov decision processes
Christos H Papadimitriou and John N Tsitsiklis · 1987
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Richard Sutton · 1988
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Richard Sutton · 1988
Earlier work this paper cites.
Q-learning
Christopher JCH Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Q-learning
Christopher JCH Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Rational learning leads to Nash equilibrium
Ehud Kalai and Ehud Lehrer · 1993
Earlier work this paper cites.
The coming technological singularity
Vernor Vinge · 1993
Earlier work this paper cites.
Rational learning leads to Nash equilibrium
Ehud Kalai and Ehud Lehrer · 1993
Earlier work this paper cites.
The coming technological singularity
Vernor Vinge · 1993
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Yoshua Bengio, Patrice Simard, and Paolo Frasconi · 1994
Earlier work this paper cites.
Weak and strong merging of opinions
Ehud Kalai and Ehud Lehrer · 1994
Earlier work this paper cites.
Diffusions, Markov Processes, and Martingales: Volume 1, Foundations
Chris Rogers and David Williams · 1994
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Yoshua Bengio, Patrice Simard, and Paolo Frasconi · 1994
Earlier work this paper cites.
Weak and strong merging of opinions
Ehud Kalai and Ehud Lehrer · 1994
Earlier work this paper cites.
Diffusions, Markov Processes, and Martingales: Volume 1, Foundations
Chris Rogers and David Williams · 1994
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
Leemon Baird · 1995
Earlier work this paper cites.
Dynamic Programming and Optimal Control
Dimitri P Bertsekas and John Tsitsiklis · 1995
Earlier work this paper cites.
Temporal difference learning and TD-Gammon
Gerald Tesauro · 1995
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
Leemon Baird · 1995
Earlier work this paper cites.
Dynamic Programming and Optimal Control
Dimitri P Bertsekas and John Tsitsiklis · 1995
Earlier work this paper cites.
Temporal difference learning and TD-Gammon
Gerald Tesauro · 1995
Earlier work this paper cites.
Merging and learning
Ehud Lehrer and Rann Smorodinsky · 1996
Earlier work this paper cites.
Optimality criteria in reinforcement learning
Sridhar Mahadevan · 1996
Earlier work this paper cites.
Merging and learning
Ehud Lehrer and Rann Smorodinsky · 1996
Earlier work this paper cites.
Optimality criteria in reinforcement learning
Sridhar Mahadevan · 1996
Earlier work this paper cites.
Prediction, optimization, and learning in repeated games
John H Nachbar · 1997
Earlier work this paper cites.
An analysis of temporal-difference learning with function approximation
John N Tsitsiklis and Benjamin Van Roy · 1997
Earlier work this paper cites.
Prediction, optimization, and learning in repeated games
John H Nachbar · 1997
Earlier work this paper cites.
An analysis of temporal-difference learning with function approximation
John N Tsitsiklis and Benjamin Van Roy · 1997
Earlier work this paper cites.
Bayesian Q-learning
Richard Dearden, Nir Friedman, and Stuart Russell · 1998
Earlier work this paper cites.
The Theory of Learning in Games
Drew Fudenberg and David K Levine · 1998
Earlier work this paper cites.
The maximum speed of dynamical evolution
Norman Margolus and Lev B Levitin · 1998
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Richard S. Sutton and Andrew G. Barto · 1998
Earlier work this paper cites.
Bayesian Q-learning
Richard Dearden, Nir Friedman, and Stuart Russell · 1998
Earlier work this paper cites.
The Theory of Learning in Games
Drew Fudenberg and David K Levine · 1998
Earlier work this paper cites.
The maximum speed of dynamical evolution
Norman Margolus and Lev B Levitin · 1998
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Richard S. Sutton and Andrew G. Barto · 1998
Earlier work this paper cites.
On the undecidability of probabilistic planning and infinite-horizon partially observable Markov decision problems
Omid Madani, Steve Hanks, and Anne Condon · 1999
Earlier work this paper cites.
Inductive logic and the ravens paradox
Patrick Maher · 1999
Earlier work this paper cites.
The role of absolute continuity in “merging of opinions” and “rational learning”
Ronald I Miller and Chris William Sanchirico · 1999
Earlier work this paper cites.
On the undecidability of probabilistic planning and infinite-horizon partially observable Markov decision problems
Omid Madani, Steve Hanks, and Anne Condon · 1999
Earlier work this paper cites.
Inductive logic and the ravens paradox
Patrick Maher · 1999
Earlier work this paper cites.
The role of absolute continuity in “merging of opinions” and “rational learning”
Ronald I Miller and Chris William Sanchirico · 1999
Earlier work this paper cites.
A theory of universal artificial intelligence based on algorithmic complexity
Marcus Hutter · 2000
Earlier work this paper cites.
Complexity of finite-horizon Markov decision process problems
Martin Mundhenk, Judy Goldsmith, Christopher Lusena, and Eric Allender · 2000
Earlier work this paper cites.
A Bayesian framework for reinforcement learning
Malcolm Strens · 2000
Earlier work this paper cites.
A theory of universal artificial intelligence based on algorithmic complexity
Marcus Hutter · 2000
Earlier work this paper cites.
Complexity of finite-horizon Markov decision process problems
Martin Mundhenk, Judy Goldsmith, Christopher Lusena, and Eric Allender · 2000
Earlier work this paper cites.
A Bayesian framework for reinforcement learning
Malcolm Strens · 2000
Earlier work this paper cites.
Rational and convergent learning in stochastic games
Michael Bowling and Manuela Veloso · 2001
Earlier work this paper cites.
On the impossibility of predicting the behavior of rational agents
Dean P Foster and H Peyton Young · 2001
Earlier work this paper cites.
Reinforcement learning with function approximation converges to a region
Geoffrey J Gordon · 2001
Earlier work this paper cites.
Creating friendly AI 1.0: The analysis and design of benevolent goal architectures
Eliezer Yudkowsky · 2001
Earlier work this paper cites.
Rational and convergent learning in stochastic games
Michael Bowling and Manuela Veloso · 2001
Earlier work this paper cites.
On the impossibility of predicting the behavior of rational agents
Dean P Foster and H Peyton Young · 2001
Earlier work this paper cites.
Reinforcement learning with function approximation converges to a region
Geoffrey J Gordon · 2001
Earlier work this paper cites.
Creating friendly AI 1.0: The analysis and design of benevolent goal architectures
Eliezer Yudkowsky · 2001
Earlier work this paper cites.
Existential risks
Nick Bostrom · 2002
Earlier work this paper cites.
The speed prior: A new simplicity measure yielding near-optimal computable predictions
Jürgen Schmidhuber · 2002
Earlier work this paper cites.
Random walk—1-dimensional
Eric W Weisstein · 2002
Earlier work this paper cites.
Existential risks
Nick Bostrom · 2002
Earlier work this paper cites.
The speed prior: A new simplicity measure yielding near-optimal computable predictions
Jürgen Schmidhuber · 2002
Earlier work this paper cites.
Random walk—1-dimensional
Eric W Weisstein · 2002
Earlier work this paper cites.
Ethical issues in advanced artificial intelligence
Nick Bostrom · 2003
Earlier work this paper cites.
A gentle introduction to the universal algorithmic agent AIXI
Marcus Hutter · 2003
Earlier work this paper cites.
Probability Theory: The Logic of Science
Edwin T Jaynes · 2003
Earlier work this paper cites.
On the Sample Complexity of Reinforcement Learning
Sham Machandranath Kakade · 2003
Earlier work this paper cites.
Autonomous helicopter flight via reinforcement learning
HJ Kim, Michael I Jordan, Shankar Sastry, and Andrew Y Ng · 2003
Earlier work this paper cites.
On the undecidability of probabilistic planning and related stochastic optimization problems
Omid Madani, Steve Hanks, and Anne Condon · 2003
Earlier work this paper cites.
Learning predictive state representations
Satinder Singh, Michael L Littman, Nicholas K Jong, David Pardoe, and Peter Stone · 2003
Earlier work this paper cites.
Ethical issues in advanced artificial intelligence
Nick Bostrom · 2003
Earlier work this paper cites.
A gentle introduction to the universal algorithmic agent AIXI
Marcus Hutter · 2003
Earlier work this paper cites.
Probability Theory: The Logic of Science
Edwin T Jaynes · 2003
Earlier work this paper cites.
On the Sample Complexity of Reinforcement Learning
Sham Machandranath Kakade · 2003
Earlier work this paper cites.
Autonomous helicopter flight via reinforcement learning
HJ Kim, Michael I Jordan, Shankar Sastry, and Andrew Y Ng · 2003
Earlier work this paper cites.
On the undecidability of probabilistic planning and related stochastic optimization problems
Omid Madani, Steve Hanks, and Anne Condon · 2003
Earlier work this paper cites.
Learning predictive state representations
Satinder Singh, Michael L Littman, Nicholas K Jong, David Pardoe, and Peter Stone · 2003
Earlier work this paper cites.
Bayesianism, infinite decisions, and binding
Frank Arntzenius, Adam Elga, and John Hawthorne · 2004
Earlier work this paper cites.
Predictive state representations: A new theory for modeling dynamical systems
Satinder Singh, Michael R James, and Matthew R Rudary · 2004
Earlier work this paper cites.
Hempel’s raven paradox: A lacuna in the standard Bayesian solution
Peter BM Vranas · 2004
Earlier work this paper cites.
All of Statistics
Larry Wasserman · 2004
Earlier work this paper cites.
Bayesianism, infinite decisions, and binding
Frank Arntzenius, Adam Elga, and John Hawthorne · 2004
Earlier work this paper cites.
Predictive state representations: A new theory for modeling dynamical systems
Satinder Singh, Michael R James, and Matthew R Rudary · 2004
Earlier work this paper cites.
Hempel’s raven paradox: A lacuna in the standard Bayesian solution
Peter BM Vranas · 2004
Earlier work this paper cites.
All of Statistics
Larry Wasserman · 2004
Earlier work this paper cites.
Universal Artificial Intelligence
Marcus Hutter · 2005
Earlier work this paper cites.
The Singularity is Near: When Humans Transcend Biology
Ray Kurzweil · 2005
Earlier work this paper cites.
Beliefs in repeated games
John H Nachbar · 2005
Earlier work this paper cites.
Universal Artificial Intelligence
Marcus Hutter · 2005
Earlier work this paper cites.
The Singularity is Near: When Humans Transcend Biology
Ray Kurzweil · 2005
Cited alongside, same era.
Beliefs in repeated games
John H Nachbar · 2005
Cited alongside, same era.
Pattern Recognition and Machine Learning
Christopher M Bishop · 2006
Cited alongside, same era.
Prediction, Learning, and Games
Nicolo Cesa-Bianchi and Gábor Lugosi · 2006
Cited alongside, same era.
Elements of Information Theory
Thomas M Cover and Joy A Thomas · 2006
Cited alongside, same era.
Is there an elegant universal theory of prediction?
Shane Legg · 2006
Cited alongside, same era.
Pattern Recognition and Machine Learning
Christopher M Bishop · 2006
Markov Decision Processes
Martin L Puterman · 2014
Later among the works it cites.
Aligning superintelligence with human interests: A technical research agenda
Nate Soares and Benja Fallenstein · 2014
Later among the works it cites.
Responses to catastrophic AGI risk: A survey
Kaj Sotala and Roman V Yampolskiy · 2014
Later among the works it cites.
Superintelligence: Paths, Dangers, Strategies
Nick Bostrom · 2014
Later among the works it cites.
Ethical guidelines for a superintelligence
Ernest Davis · 2014
Later among the works it cites.
Extreme state aggregation beyond MDPs
Marcus Hutter · 2014
Later among the works it cites.
General time consistent discounting
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Prediction, Learning, and Games
Nicolo Cesa-Bianchi and Gábor Lugosi · 2006
Cited alongside, same era.
Elements of Information Theory
Thomas M Cover and Joy A Thomas · 2006
Cited alongside, same era.
Is there an elegant universal theory of prediction?
Shane Legg · 2006
Cited alongside, same era.
The Minimum Description Length Principle
Peter D. Grünwald · 2007
Cited alongside, same era.
Universal semimeasures: An introduction
Nicholas J Hay · 2007
Cited alongside, same era.
Tor Lattimore and Marcus Hutter · 2014
Later among the works it cites.
Generalized Thompson sampling for sequential decision-making and causal inference
Pedro A Ortega and Daniel A Braun · 2014
Later among the works it cites.
Markov Decision Processes
Martin L Puterman · 2014
Later among the works it cites.
Aligning superintelligence with human interests: A technical research agenda
Nate Soares and Benja Fallenstein · 2014
Later among the works it cites.
Responses to catastrophic AGI risk: A survey
Kaj Sotala and Roman V Yampolskiy · 2014
Later among the works it cites.
Generic Reinforcement Learning Beyond Small MDPs
Mayank Daswani · 2015
Later among the works it cites.
A definition of happiness for reinforcement learning agents
Mayank Daswani and Jan Leike · 2015
Later among the works it cites.
Sequential extensions of causal and evidential decision theory
Tom Everitt, Jan Leike, and Marcus Hutter · 2015
Later among the works it cites.
Agents using speed priors
Daniel Filan · 2015
Later among the works it cites.
Elon musk donates $10m to keep AI beneficial
Future of Life Institute · 2015
Later among the works it cites.
The latest chapter for the self-driving car: mastering city street driving
Google · 2015
Later among the works it cites.
Thompson sampling for learning parameterized Markov decision processes
Aditya Gopalan and Shie Mannor · 2015
Later among the works it cites.
Deep recurrent Q-learning for partially observable MDPs
Matthew Hausknecht and Peter Stone · 2015
Later among the works it cites.
Transcending complacency on superintelligent machines
Stephen Hawking, Max Tegmark, Stuart Russell, and Frank Wilczek · 2015
Later among the works it cites.
Memory-based control with recurrent neural networks
Nicolas Heess, Jonathan J Hunt, Timothy P Lillicrap, and David Silver · 2015
Later among the works it cites.
Ultimate Automizer with array interpolation (competition contribution)
Matthias Heizmann, Daniel Dietsch, Jan Leike, Betim Musa, and Andreas Podelski · 2015
Later among the works it cites.
50’000€ prize for compressing human knowledge
Marcus Hutter · 2015
Later among the works it cites.
Deep Blue
IBM · 2015
Later among the works it cites.
A computer called Watson
IBM · 2015
Later among the works it cites.
On Martin-Löf (non-)convergence of Solomonoff’s universal mixture
Tor Lattimore and Marcus Hutter · 2015
Later among the works it cites.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Later among the works it cites.
Ranking templates for linear loops
Jan Leike and Matthias Heizmann · 2015
Later among the works it cites.
Emphatic temporal-difference learning
A Rupam Mahmood, Huizhen Yu, Martha White, and Richard Sutton · 2015
Later among the works it cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Later among the works it cites.
Massively parallel methods for deep reinforcement learning
Arun Nair, Praveen Srinivasan, Sam Blackwell, Cagdas Alcicek, Rory Fearon, Alessandro De Maria, Vedavyas Panneershelvam, Mustafa Suleyman, Charles Beattie, Stig Petersen, Shane Legg, Volodymyr Mnih, Koray Kavukcuoglu, and David Silver · 2015
Later among the works it cites.
Research priorities for robust and beneficial artificial intelligence
Stuart Russell, Daniel Dewey, and Max Tegmark · 2015
Later among the works it cites.
Deep learning in neural networks: An overview
Jürgen Schmidhuber · 2015
Later among the works it cites.
The Technological Singularity
Murray Shanahan · 2015
Later among the works it cites.
Formalizing two problems of realistic world-models
Nate Soares · 2015
Later among the works it cites.
Rationality, optimism and guarantees in general reinforcement learning
Peter Sunehag and Marcus Hutter · 2015
Later among the works it cites.
Compress and control
Joel Veness, Marc G Bellemare, Marcus Hutter, Alvin Chua, and Guillaume Desjardins · 2015
Later among the works it cites.
On convergence of emphatic temporal-difference learning
Huizhen Yu · 2015
Later among the works it cites.
Generic Reinforcement Learning Beyond Small MDPs
Mayank Daswani · 2015
Later among the works it cites.
A definition of happiness for reinforcement learning agents
Mayank Daswani and Jan Leike · 2015
Later among the works it cites.
Sequential extensions of causal and evidential decision theory
Tom Everitt, Jan Leike, and Marcus Hutter · 2015
Later among the works it cites.
Agents using speed priors
Daniel Filan · 2015
Later among the works it cites.
Elon musk donates $10m to keep AI beneficial
Future of Life Institute · 2015
Later among the works it cites.
The latest chapter for the self-driving car: mastering city street driving
Google · 2015
Later among the works it cites.
Thompson sampling for learning parameterized Markov decision processes
Aditya Gopalan and Shie Mannor · 2015
Later among the works it cites.
Deep recurrent Q-learning for partially observable MDPs
Matthew Hausknecht and Peter Stone · 2015
Later among the works it cites.
Transcending complacency on superintelligent machines
Stephen Hawking, Max Tegmark, Stuart Russell, and Frank Wilczek · 2015
Later among the works it cites.
Memory-based control with recurrent neural networks
Nicolas Heess, Jonathan J Hunt, Timothy P Lillicrap, and David Silver · 2015
Later among the works it cites.
Ultimate Automizer with array interpolation (competition contribution)
Matthias Heizmann, Daniel Dietsch, Jan Leike, Betim Musa, and Andreas Podelski · 2015
Later among the works it cites.
50’000€ prize for compressing human knowledge
Marcus Hutter · 2015
Later among the works it cites.
Deep Blue
IBM · 2015
Later among the works it cites.
A computer called Watson
IBM · 2015
Later among the works it cites.
On Martin-Löf (non-)convergence of Solomonoff’s universal mixture
Tor Lattimore and Marcus Hutter · 2015
Later among the works it cites.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Later among the works it cites.
Ranking templates for linear loops
Jan Leike and Matthias Heizmann · 2015
Later among the works it cites.
Emphatic temporal-difference learning
A Rupam Mahmood, Huizhen Yu, Martha White, and Richard Sutton · 2015
Later among the works it cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Later among the works it cites.
Massively parallel methods for deep reinforcement learning
Arun Nair, Praveen Srinivasan, Sam Blackwell, Cagdas Alcicek, Rory Fearon, Alessandro De Maria, Vedavyas Panneershelvam, Mustafa Suleyman, Charles Beattie, Stig Petersen, Shane Legg, Volodymyr Mnih, Koray Kavukcuoglu, and David Silver · 2015
Later among the works it cites.
Research priorities for robust and beneficial artificial intelligence
Stuart Russell, Daniel Dewey, and Max Tegmark · 2015
Later among the works it cites.
Deep learning in neural networks: An overview
Jürgen Schmidhuber · 2015
Later among the works it cites.
The Technological Singularity
Murray Shanahan · 2015
Later among the works it cites.
Formalizing two problems of realistic world-models
Nate Soares · 2015
Later among the works it cites.
Rationality, optimism and guarantees in general reinforcement learning
Peter Sunehag and Marcus Hutter · 2015
Later among the works it cites.
Compress and control
Joel Veness, Marc G Bellemare, Marcus Hutter, Alvin Chua, and Guillaume Desjardins · 2015
Later among the works it cites.
On convergence of emphatic temporal-difference learning
Huizhen Yu · 2015
Later among the works it cites.
AI researchers on AI risk
Scott Alexander · 2016
Closest in time.
Interruptibility and corrigibility for AIXI and Monte Carlo agents
Stuart Armstrong and Laurent Orseau · 2016
Closest in time.
Increasing the action gap: New operators for reinforcement learning
Marc G Bellemare, Georg Ostrovski, Arthur Guez, Philip S Thomas, and Rémi Munos · 2016
Closest in time.
Strategic implications of openness in AI development
Nick Bostrom · 2016
Closest in time.
Amnon H Eden · 2016
Closest in time.
Avoiding wireheading with value reinforcement learning
Tom Everitt and Marcus Hutter · 2016
Closest in time.
Self-modification of policy and utility function in rational agents
Tom Everitt, Daniel Filan, Mayank Daswani, and Marcus Hutter · 2016
Closest in time.
Loss bounds and time complexity for speed priors
Daniel Filan, Jan Leike, and Marcus Hutter · 2016
Closest in time.
Learning to communicate to solve riddles with deep distributed recurrent Q-networks
Jakob N Foerster, Yannis M Assael, Nando de Freitas, and Shimon Whiteson · 2016
Closest in time.
2015 project grants recommended for funding
Future of Life Institute · 2016
Closest in time.
Research priorities for robust and beneficial artificial intelligence: An open letter
Future of Life Institute · 2016
Closest in time.
Deep learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Closest in time.
What we learned in Seoul with AlphaGo
Google · 2016
Closest in time.
Deep reinforcement learning from self-play in imperfect-information games
Johannes Heinrich and David Silver · 2016
Closest in time.
Ultimate Automizer with two-track proofs (competition contribution)
Matthias Heizmann, Daniel Dietsch, Marius Greitschus, Jan Leike, Betim Musa, Claus Schätzle, and Andreas Podelski · 2016
Closest in time.
Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation
Tejas D Kulkarni, Karthik R Narasimhan, Ardavan Saeedi, and Joshua B Tenenbaum · 2016
Closest in time.
Regret analysis of the finite-horizon Gittins index strategy for multi-armed bandits
Tor Lattimore · 2016
Closest in time.
Future of AI 6. Discussion of ‘Superintelligence: Paths, Dangers, Strategies’
Neil Lawrence · 2016
Closest in time.
Geometric nontermination arguments
Jan Leike and Matthias Heizmann · 2016
Closest in time.
On the computability of Solomonoff induction and AIXI
Jan Leike and Marcus Hutter · 2016
Closest in time.
State of the art control of Atari games using shallow reinforcement learning
Yitao Liang, Marlos C Machado, Erik Talvitie, and Michael Bowling · 2016
Closest in time.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2016
Closest in time.
Death and suicide in universal artificial intelligence
Jarryd Martin, Tom Everitt, and Marcus Hutter · 2016
Closest in time.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adrià Puigdomènech Badia, Mehdi Mirza, Alex Graves, Timothy P Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Closest in time.
Future progress in artificial intelligence: A survey of expert opinion
Vincent C Müller and Nick Bostrom · 2016
Closest in time.
Is A.I. an existential threat to humanity?
Andrew Ng · 2016
Closest in time.
About OpenAI
OpenAI · 2016
Closest in time.
Safely interruptible agents
Laurent Orseau and Stuart Armstrong · 2016
Closest in time.
Prioritized experience replay
Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver · 2016
Closest in time.
Mastering the game of Go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George van den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, Sander Dieleman, Dominik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy Lillicrap, Madeleine Leach, Koray Kavukcuoglu, Thore Graepel, and Demis Hassabis · 2016
Closest in time.
Putnam’s diagonal argument and the impossibility of a universal learning machine
Tom F Sterkenburg · 2016
Closest in time.
Deep reinforcement learning with double Q-learning
Hado van Hasselt, Arthur Guez, and David Silver · 2016
Closest in time.
The singularity may never be near
Toby Walsh · 2016
Closest in time.
Dueling network architectures for deep reinforcement learning
Ziyu Wang, Nando de Freitas, Tom Schaul, Matteo Hessel, Hado van Hasselt, and Marc Lanctot · 2016
Closest in time.
Graying the black box: Understanding DQNs
Tom Zahavy, Nir Ben Zrihem, and Shie Mannor · 2016
Closest in time.
AI researchers on AI risk
Scott Alexander · 2016
Closest in time.
Interruptibility and corrigibility for AIXI and Monte Carlo agents
Stuart Armstrong and Laurent Orseau · 2016
Closest in time.
Increasing the action gap: New operators for reinforcement learning
Marc G Bellemare, Georg Ostrovski, Arthur Guez, Philip S Thomas, and Rémi Munos · 2016
Closest in time.
Strategic implications of openness in AI development
Nick Bostrom · 2016
Closest in time.
Amnon H Eden · 2016
Closest in time.
Avoiding wireheading with value reinforcement learning
Tom Everitt and Marcus Hutter · 2016
Closest in time.
Self-modification of policy and utility function in rational agents
Tom Everitt, Daniel Filan, Mayank Daswani, and Marcus Hutter · 2016
Closest in time.
Loss bounds and time complexity for speed priors
Daniel Filan, Jan Leike, and Marcus Hutter · 2016
Closest in time.
Learning to communicate to solve riddles with deep distributed recurrent Q-networks
Jakob N Foerster, Yannis M Assael, Nando de Freitas, and Shimon Whiteson · 2016
Closest in time.
2015 project grants recommended for funding
Future of Life Institute · 2016
Closest in time.
Research priorities for robust and beneficial artificial intelligence: An open letter
Future of Life Institute · 2016
Closest in time.
Deep learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Closest in time.
What we learned in Seoul with AlphaGo
Google · 2016
Closest in time.
Deep reinforcement learning from self-play in imperfect-information games
Johannes Heinrich and David Silver · 2016
Closest in time.
Ultimate Automizer with two-track proofs (competition contribution)
Matthias Heizmann, Daniel Dietsch, Marius Greitschus, Jan Leike, Betim Musa, Claus Schätzle, and Andreas Podelski · 2016
Closest in time.
Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation
Tejas D Kulkarni, Karthik R Narasimhan, Ardavan Saeedi, and Joshua B Tenenbaum · 2016
Closest in time.
Regret analysis of the finite-horizon Gittins index strategy for multi-armed bandits
Tor Lattimore · 2016
Closest in time.
Future of AI 6. Discussion of ‘Superintelligence: Paths, Dangers, Strategies’
Neil Lawrence · 2016
Closest in time.
Geometric nontermination arguments
Jan Leike and Matthias Heizmann · 2016
Closest in time.
On the computability of Solomonoff induction and AIXI
Jan Leike and Marcus Hutter · 2016
Closest in time.
State of the art control of Atari games using shallow reinforcement learning
Yitao Liang, Marlos C Machado, Erik Talvitie, and Michael Bowling · 2016
Closest in time.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2016
Closest in time.
Death and suicide in universal artificial intelligence
Jarryd Martin, Tom Everitt, and Marcus Hutter · 2016
Closest in time.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adrià Puigdomènech Badia, Mehdi Mirza, Alex Graves, Timothy P Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Closest in time.
Future progress in artificial intelligence: A survey of expert opinion
Vincent C Müller and Nick Bostrom · 2016
Closest in time.
Is A.I. an existential threat to humanity?
Andrew Ng · 2016
Closest in time.
About OpenAI
OpenAI · 2016
Closest in time.
Safely interruptible agents
Laurent Orseau and Stuart Armstrong · 2016
Closest in time.
Prioritized experience replay
Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver · 2016
Closest in time.
Mastering the game of Go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George van den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, Sander Dieleman, Dominik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy Lillicrap, Madeleine Leach, Koray Kavukcuoglu, Thore Graepel, and Demis Hassabis · 2016
Closest in time.
Putnam’s diagonal argument and the impossibility of a universal learning machine
Tom F Sterkenburg · 2016
Closest in time.
Deep reinforcement learning with double Q-learning
Hado van Hasselt, Arthur Guez, and David Silver · 2016
Closest in time.
The singularity may never be near
Toby Walsh · 2016
Closest in time.
Dueling network architectures for deep reinforcement learning
Ziyu Wang, Nando de Freitas, Tom Schaul, Matteo Hessel, Hado van Hasselt, and Marc Lanctot · 2016
Closest in time.
Graying the black box: Understanding DQNs
Tom Zahavy, Nir Ben Zrihem, and Shie Mannor · 2016
Closest in time.
A universal algorithm for sequential data compression
Jacob Ziv and Abraham Lempel · 2026
Closest in time.