Fetching the paper…
Reading the bibliography…
Reinforcement Learning agents are expected to eventually perform well.
Discussion of Dr Gittins’ paper
Peter Whittle · 1979
Earlier work this paper cites.
Universal Artificial Intelligence: Sequential Decisions based on Algorithmic Probability
Marcus Hutter · 2005
Earlier work this paper cites.
Information-theoretic approach to interactive learning
Susanne Still · 2009
Earlier work this paper cites.
Monte-carlo planning in large pomdps
David Silver and Joel Veness · 2010
Earlier work this paper cites.
Asymptotically optimal agents
Tor Lattimore and Marcus Hutter · 2011
Earlier work this paper cites.
Intrinsically motivated model learning for a developing curious agent
Todd Hester and Peter Stone · 2012
Earlier work this paper cites.
Universal knowledge-seeking agents for stochastic environments
Laurent Orseau, Tor Lattimore, and Marcus Hutter · 2013
Cited alongside, same era.
Bayesian reinforcement learning with exploration
Tor Lattimore and Marcus Hutter · 2014
Cited alongside, same era.
General time consistent discounting
Tor Lattimore and Marcus Hutter · 2014
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
Curiosity-driven exploration in deep reinforcement learning via bayesian neural networks
Rein Houthooft, Xi Chen, Yan Duan, John Schulman, Filip De Turck, and Pieter Abbeel · 2016
Cited alongside, same era.
Thompson sampling is asymptotically optimal in general environments
Jan Leike, Tor Lattimore, Laurent Orseau, and Marcus Hutter · 2016
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Later among the works it cites.
Universal reinforcement learning algorithms: survey and experiments
John Aslanides, Jan Leike, and Marcus Hutter · 2017
Later among the works it cites.
AIXIjs: A software demo for general reinforcement learning
John Aslanides · 2017
Later among the works it cites.
Mastering chess and shogi by self-play with a general reinforcement learning algorithm
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, et al · 2017
Later among the works it cites.
Rainbow: Combining improvements in deep reinforcement learning
Matteo Hessel, Joseph Modayil, Hado van Hasselt, Tom Schaul, Georg Ostrovski, Will Dabney, Dan Horgan, Bilal Piot, Mohammad Azar, and David Silver · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Later among the works it cites.