Fetching the paper…
Reading the bibliography…
Can simple algorithms with a good representation solve challenging reinforcement learning problems? In this work, we answer this question in the affirmative, where we take "simple learning algorithm" to be tabular Q-Learning, the "good representations" to be a learned state abstraction, and "challenging problems" to be continuous control tasks.
A proposal for the Dartmouth summer research project on artificial intelligence, August 31, 1955
McCarthy, J., Minsky, M. L., Rochester, N., and Shannon, C. E · 1955
Earlier work this paper cites.
Pattern-recognizing control systems
Widrow, B. and Smith, F. W · 1964
Earlier work this paper cites.
An algorithm for finding best matches in logarithmic time
Friedman, J. H., Bentley, J. L., and Finkel, R. A · 1976
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Sutton, R. S · 1988
Earlier work this paper cites.
Intelligence without representation
Brooks, R. A · 1991
Earlier work this paper cites.
Input generalization in delayed reinforcement learning: An algorithm and performance comparisons
Chapman, D. and Kaelbling, L. P · 1991
Earlier work this paper cites.
Adaptive state space quantisation for reinforcement learning of collision-free navigation
Krose, B. J. and Van Dam, J. W · 1992
Earlier work this paper cites.
Q-learning
Watkins, C. J. and Dayan, P · 1992
Earlier work this paper cites.
Feudal reinforcement learning
Dayan, P. and Hinton, G. E · 1993
Earlier work this paper cites.
The parti-game algorithm for variable resolution reinforcement learning in multidimensional state-spaces
Moore, A. W · 1994
Earlier work this paper cites.
Generalization in reinforcement learning: Safely approximating the value function
Boyan, J. A. and Moore, A. W · 1995
Earlier work this paper cites.
Reinforcement learning with soft state aggregation
Singh, S. P., Jaakkola, T., and Jordan, M. I · 1995
Earlier work this paper cites.
Temporal difference learning and TD-gammon
Tesauro, G · 1995
Earlier work this paper cites.
Learning to use selective attention and short-term memory in sequential tasks
McCallum, A. K · 1996
Earlier work this paper cites.
Reinforcement learning with hierarchies of machines
Parr, R. and Russell, S. J · 1998
Earlier work this paper cites.
Learning to learn
Thrun, S. and Pratt, L · 1998
Earlier work this paper cites.
Tree based discretization for continuous state space reinforcement learning
Uther, W. T. and Veloso, M. M · 1998
Earlier work this paper cites.
Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning
Sutton, R. S., Precup, D., and Singh, S · 1999
Earlier work this paper cites.
Hierarchical reinforcement learning with the MAXQ value function decomposition
Dietterich, T. G · 2000
Earlier work this paper cites.
Rademacher and Gaussian complexities: Risk bounds and structural results
Bartlett, P. L. and Mendelson, S · 2002
Earlier work this paper cites.
On the Sample Complexity of Reinforcement Learning
Kakade, S. M · 2003
Earlier work this paper cites.
Dynamic programming for structured continuous markov decision problems
Feng, Z., Dearden, R., Meuleau, N., and Washington, R · 2004
Earlier work this paper cites.
Adaptive state space partitioning for reinforcement learning
Lee, I. S. and Lau, H. Y · 2004
Cited alongside, same era.
Behavior transfer for value-function-based reinforcement learning
Taylor, M. E. and Stone, P · 2005
Cited alongside, same era.
Autonomous shaping: Knowledge transfer in reinforcement learning
Konidaris, G. and Barto, A · 2006
Cited alongside, same era.
Towards a unified theory of state abstraction for MDPs
Li, L., Walsh, T. J., and Littman, M. L · 2006
Cited alongside, same era.
Transferring state abstractions between MDPs
Walsh, T. J., Li, L., and Littman, M. L · 2006
Cited alongside, same era.
Building portable options: Skill transfer in reinforcement learning
Konidaris, G. and Barto, A. G · 2007
Cited alongside, same era.
Near optimal behavior via approximate state abstraction
Abel, D., Hershkowitz, D. E., and Littman, M. L · 2016
Later among the works it cites.
OpenAI Gym, 2016
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Later among the works it cites.
State of the art control of Atari games using shallow reinforcement learning
Liang, Y., Machado, M. C., Talvitie, E., and Bowling, M · 2016
Later among the works it cites.
A vector-contraction inequality for Rademacher complexities
Maurer, A · 2016
Later among the works it cites.
Learning to reinforcement learn
Wang, J. X., Kurth-Nelson, Z., Tirumala, D., Soyer, H., Leibo, J. Z., Munos, R., Blundell, C., Kumaran, D., and Botvinick, M · 2016
Later among the works it cites.
Successor features for transfer in reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Adaptive tile coding for value function approximation, 2007
Whiteson, S., Taylor, M. E., and Stone, P · 2007
Cited alongside, same era.
Multi-task reinforcement learning: A hierarchical bayesian approach
Wilson, A., Fern, A., Ray, S., and Tadepalli, P · 2007
Cited alongside, same era.
Transfer in variable-reward hierarchical reinforcement learning
Mehta, N., Natarajan, S., Tadepalli, P., and Fern, A · 2008
Cited alongside, same era.
Transfer learning for reinforcement learning domains: A survey
Taylor, M. E. and Stone, P · 2009
Cited alongside, same era.
Automatic state abstraction from demonstration
Cobo, L. C., Zang, P., Isbell Jr, C. L., and Thomaz, A. L · 2011
Cited alongside, same era.
Automatic task decomposition and state abstraction from demonstration
Cobo, L. C., Isbell Jr, C. L., and Thomaz, A. L · 2012
Cited alongside, same era.
Barreto, A., Dabney, W., Munos, R., Hunt, J. J., Schaul, T., van Hasselt, H. P., and Silver, D · 2017
Later among the works it cites.
Model-agnostic meta-learning for fast adaptation of deep networks
Finn, C., Abbeel, P., and Levine, S · 2017
Later among the works it cites.
DARLA: Improving zero-shot transfer in reinforcement learning
Higgins, I., Pal, A., Rusu, A. A., Matthey, L., Burgess, C. P., Pritzel, A., Botvinick, M., Blundell, C., and Lerchner, A · 2017
Later among the works it cites.
Deep variational Bayes filters: Unsupervised learning of state space models from raw data
Karl, M., Soelch, M., Bayer, J., and van der Smagt, P · 2017
Later among the works it cites.
Value prediction network
Oh, J., Singh, S., and Lee, H · 2017
Later among the works it cites.
Distral: Robust multitask reinforcement learning
Teh, Y., Bapst, V., Czarnecki, W. M., Quan, J., Kirkpatrick, J., Hadsell, R., Heess, N., and Pascanu, R · 2017
Later among the works it cites.
A deep hierarchical approach to lifelong learning in minecraft
Tessler, C., Givony, S., Zahavy, T., Mankowitz, D. J., and Mannor, S · 2017
Later among the works it cites.
State abstractions for lifelong reinforcement learning
Abel, D., Arumugam, D., Lehnert, L., and Littman, M. L · 2018
Later among the works it cites.
State representation learning for control: An overview
Lesort, T., Díaz-Rodríguez, N., Goudou, J.-F., and Filliat, D · 2018
Later among the works it cites.
On Q-learning convergence for non-Markov decision processes
Majeed, S. J. and Hutter, M · 2018
Later among the works it cites.
State abstraction synthesis for discrete models of continuous domains
Menashe, J. and Stone, P · 2018
Later among the works it cites.
Foundations of machine learning
Mohri, M., Rostamizadeh, A., and Talwalkar, A · 2018
Later among the works it cites.
Meta reinforcement learning with latent variable gaussian processes
Sæmundsson, S., Hofmann, K., and Deisenroth, M. P · 2018
Later among the works it cites.
Reinforcement Learning: An Introduction
Sutton, R. S. and Barto, A. G · 2018
Later among the works it cites.
State abstraction as compression in apprenticeship learning
Abel, D., Arumugam, D., Asadi, K., Jinnai, Y., Littman, M. L., and Wong, L. L · 2019
Later among the works it cites.
On the necessity of abstraction
Konidaris, G · 2019
Later among the works it cites.