Fetching the paper…
Reading the bibliography…
The reinforcement learning (RL) framework formalizes the notion of learning with interactions.
Feature Reinforcement Learning: Part I. Unstructured MDPs
Marcus Hutter · 1946
Earlier work this paper cites.
Q-learning
Christopher J. C. H. Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Dynamic programming: An overview , volume 35
Dimitri Bertsekas and John Tsitsiklis · 1995
Earlier work this paper cites.
Instance-Based Utile Distinctions for Reinforcement Learning with Hidden State
Andrew McCallum · 1995
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
Leslie Pack Kaelbling, Michael L. Littman, and Anthony R. Cassandra · 1998
Earlier work this paper cites.
An analysis of direct reinforcement learning in non-Markovian domains
Mark D Pendrith and Michael J McGarity · 1998
Earlier work this paper cites.
Towards a unified theory of state abstraction for MDPs
Lihong Li, Thomas J. Walsh, and Michael L. Littman · 2006
Earlier work this paper cites.
Reinforcement learning in finite MDPs: PAC analysis
Alexander L. Strehl, Hong Li, and Michael L. Littman · 2009
Cited alongside, same era.
Approximate Dynamic Programming: Solving the Curses of Dimensionality: Second Edition , volume 136
Warren B. Powell · 2011
Cited alongside, same era.
Universal Artificial Intellegence
Marcus Hutter · 2012
Cited alongside, same era.
General time consistent discounting
Tor Lattimore and Marcus Hutter · 2013
Cited alongside, same era.
Near-optimal PAC bounds for discounted MDPs
Tor Lattimore and Marcus Hutter · 2014
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Extreme state aggregation beyond Markov decision processes
Marcus Hutter · 2016
Later among the works it cites.
Mastering the game of Go with deep neural networks and tree search
David Silver, Aja Huang, Chris J. Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, Sander Dieleman, Dominik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy Lillicrap, Madeleine Leach, Koray Kavukcuoglu, Thore Graepel, and Demis Hassabis · 2016
Later among the works it cites.
Reinforcement Learning: An Introduction
Richard S. Sutton and Andrew G. Barto · 2018
Later among the works it cites.
A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, Timothy Lillicrap, Karen Simonyan, and Demis Hassabis · 2018
Later among the works it cites.
On Q-learning convergence for non-Markov decision processes
Sultan Javed Majeed and Marcus Hutter · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Near optimal behavior via approximate state abstraction
David Abel, D. Ellis Hershkowitz, and Michael L. Littman · 2016
Cited alongside, same era.
Performance Guarantees for Homomorphisms beyond Markov Decision Processes
Sultan Javed Majeed and Marcus Hutter · 2019
Later among the works it cites.