Fetching the paper…
Reading the bibliography…
A key question in reinforcement learning is how an intelligent agent can generalize knowledge across different inputs.
Adaptive Control Processes: A Guided Tour , volume 2045
Richard E Bellman · 1961
Earlier work this paper cites.
Q Q -learning
Christopher J.C.H. Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Improving generalization for temporal difference learning: The successor representation
Peter Dayan · 1993
Earlier work this paper cites.
Generalization in reinforcement learning: Safely approximating the value function
Justin A Boyan and Andrew W Moore · 1995
Earlier work this paper cites.
Generalization in reinforcement learning: Successful examples using sparse coarse coding
Richard S Sutton · 1996
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Richard S. Sutton and Andrew G. Barto · 1998
Earlier work this paper cites.
Approximate equivalence of Markov decision processes
Eyal Even-Dar and Yishay Mansour · 2003
Earlier work this paper cites.
Equivalence notions and model minimization in Markov decision processes
Robert Givan, Thomas Dean, and Matthew Greig · 2003
Earlier work this paper cites.
Metrics for finite markov decision processes
Norm Ferns, Prakash Panangaden, and Doina Precup · 2004
Earlier work this paper cites.
Neural fitted Q iteration–first experiences with a data efficient neural reinforcement learning method
Martin Riedmiller · 2005
Earlier work this paper cites.
Towards a unified theory of state abstraction for MDPs
Lihong Li, Thomas J Walsh, and Michael L Littman · 2006
Earlier work this paper cites.
An analysis of linear models, linear value-function approximation, and feature selection for reinforcement learning
Ronald Parr, Lihong Li, Gavin Taylor, Christopher Painter-Wakefield, and Michael L Littman · 2008
Earlier work this paper cites.
Dyna-style planning with linear function approximation and prioritized sweeping
Richard S. Sutton, Csaba Szepesvári, Alborz Geramifard, and Michael Bowling · 2008
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Earlier work this paper cites.
Dynamic programming and optimal control 3rd edition, volume ii
Dimitri P Bertsekas · 2011
Earlier work this paper cites.
Bisimulation metrics for continuous Markov decision processes
Norm Ferns, Prakash Panangaden, and Doina Precup · 2011
Earlier work this paper cites.
Value function approximation in reinforcement learning using the Fourier basis
George Konidaris, Sarah Osentoski, and Philip Thomas · 2011
Cited alongside, same era.
Approximate policy iteration with linear action models
Hengshuai Yao and Csaba Szepesvári · 2012
Cited alongside, same era.
(More) efficient reinforcement learning via posterior sampling
Ian Osband, Daniel Russo, and Benjamin Van Roy · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2014
Cited alongside, same era.
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015
Martín Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S. Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Ian Goodfellow, Andrew Harp, Geoffrey Irving, Michael Isard, Yangqing Jia, Rafal Jozefowicz, Lukasz Kaiser, Manjunath Kudlur, Josh Levenberg, Dandelion Mané, Rajat Monga, Sherry Moore, Derek Murray, Chris Olah, Mike Schuster, Jonathon Shlens, Benoit Steiner, Ilya Sutskever, Kunal Talwar, Paul Tucker, Vincent Vanhoucke, Vijay Vasudevan, Fernanda Viégas, Oriol Vinyals, Pete Warden, Martin Wattenberg, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng · 2015
The hippocampus as a predictive map
Kimberly L Stachenfeld, Matthew M Botvinick, and Samuel J Gershman · 2017
Later among the works it cites.
Imagination-augmented agents for deep reinforcement learning
Théophane Weber, Sébastien Racanière, David P Reichert, Lars Buesing, Arthur Guez, Danilo Jimenez Rezende, Adria Puigdomènech Badia, Oriol Vinyals, Nicolas Heess, Yujia Li, et al · 2017
Later among the works it cites.
Deep reinforcement learning with successor features for navigation across similar environments
Jingwei Zhang, Jost Tobias Springenberg, Joschka Boedecker, and Wolfram Burgard · 2017
Later among the works it cites.
State abstractions for lifelong reinforcement learning
David Abel, Dilip Arumugam, Lucas Lehnert, and Michael Littman · 2018
Later among the works it cites.
Lipschitz continuity in model-based reinforcement learning
Kavosh Asadi, Dipendra Misra, and Michael Littman · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Representation discovery for mdps using bisimulation metrics
Sherry Shanshan Ruan, Gheorghe Comanici, Prakash Panangaden, and Doina Precup · 2015
Cited alongside, same era.
Near optimal behavior via approximate state abstraction
David Abel, David Hershkowitz, and Michael Littman · 2016
Cited alongside, same era.
Successor features for transfer in reinforcement learning
André Barreto, Rémi Munos, Tom Schaul, and David Silver · 2016
Cited alongside, same era.
Deep successor reinforcement learning
Tejas D Kulkarni, Ardavan Saeedi, Simanta Gautam, and Samuel J Gershman · 2016
Cited alongside, same era.
Linear feature encoding for reinforcement learning
Zhao Song, Ronald E Parr, Xuejun Liao, and Lawrence Carin · 2016
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Cited alongside, same era.
Advantages and limitations of using successor features for transfer in reinforcement learning
Lucas Lehnert, Stefanie Tellex, and Michael L Littman · 2017
Cited alongside, same era.
Transfer in deep reinforcement learning using successor features and generalised policy improvement
Andre Barreto, Diana Borsa, John Quan, Tom Schaul, David Silver, Matteo Hessel, Daniel Mankowitz, Augustin Zidek, and Remi Munos · 2018
Later among the works it cites.
Learning robust rewards with adverserial inverse reinforcement learning
Justin Fu, Katie Luo, and Sergey Levine · 2018
Later among the works it cites.
Is Q-learning provably efficient?
Chi Jin, Zeyuan Allen-Zhu, Sebastien Bubeck, and Michael I Jordan · 2018
Later among the works it cites.
Foundations of Machine Learning
Mehryar Mohri, Afshin Rostamizadeh, and Ameet Talwalkar · 2018
Later among the works it cites.
Reinforcement Learning: An Introduction
Richard S Sutton and Andrew G Barto · 2018
Later among the works it cites.
Learning the reward function for a misspecified model
Erik Talvitie · 2018
Later among the works it cites.
State abstraction as compression in apprenticeship learning
David Abel, Dilip Arumugam, Kavosh Asadi, Yuu Jinnai, Michael L. Littman, and Lawson L.S. Wong · 2019
Closest in time.
Combined reinforcement learning via abstract representations
Vincent François-Lavet, Yoshua Bengio, Doina Precup, and Joelle Pineau · 2019
Closest in time.
DeepMDP: Learning continuous latent space models for representation learning
Carles Gelada, Saurabh Kumar, Jacob Buckman, Ofir Nachum, and Marc G Bellemare · 2019
Closest in time.
Reward predictive representations generalize across tasks in reinforcement learning
Lucas Lehnert, Michael J Frank, and Michael L Littman · 2019
Closest in time.
Mastering atari, go, chess and shogi by planning with a learned model
Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert, Karen Simonyan, Laurent Sifre, Simon Schmitt, Arthur Guez, Edward Lockhart, Demis Hassabis, Thore Graepel, et al · 2019
Closest in time.