Fetching the paper…
Reading the bibliography…
Current reinforcement learning (RL) methods can successfully learn single tasks but often generalize poorly to modest perturbations in task domain or training procedure.
Learning to predict by the methods of temporal differences
Sutton, Richard S · 1988
Earlier work this paper cites.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
Sutton, Richard · 1990
Earlier work this paper cites.
Function optimization using connectionist reinforcement learning algorithms
J. Williams, Ronald and Peng, Jing · 1991
Earlier work this paper cites.
Improving generalization for temporal difference learning: The successor representation
Dayan, Peter · 1993
Earlier work this paper cites.
Reinforcement learning: A survey
Kaelbling, Leslie Pack, Littman, Michael L, and Moore, Andrew W · 1996
Earlier work this paper cites.
Long short-term memory
Hochreiter, Sepp and Schmidhuber, Jürgen · 1997
Earlier work this paper cites.
Reinforcement learning: An introduction , volume 1
Sutton, Richard S and Barto, Andrew G · 1998
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Sutton, Richard S, Precup, Doina, and Singh, Satinder · 1999
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, Richard S, McAllester, David A, Singh, Satinder P, and Mansour, Yishay · 2000
Earlier work this paper cites.
Mosaic model for sensorimotor learning and control
Haruno, Masahiko, Wolpert, Daniel M., and Kawato, Mitsuo · 2001
Earlier work this paper cites.
Predictive representations of state
Littman, Michael L and Sutton, Richard S · 2002
Earlier work this paper cites.
Efficient selectivity and backup operators in monte-carlo tree search
Coulom, Remi · 2006
Earlier work this paper cites.
Transferring state abstractions between mdps
Walsh, Thomas J., Li, Lihong, and Littman, Michael L · 2006
Earlier work this paper cites.
Multi-task feature learning
Argyriou, Andreas, Evgeniou, Theodoros, and Pontil, Massimiliano · 2007
Earlier work this paper cites.
General game learning using knowledge transfer
Banerjee, Bikramjit and Stone, Peter · 2007
Earlier work this paper cites.
Transfer learning for reinforcement learning domains: A survey
Taylor, Matthew E. and Stone, Peter · 2009
Cited alongside, same era.
Deep sparse rectifier neural networks
Glorot, Xavier, Bordes, Antoine, and Bengio, Yoshua · 2011
Cited alongside, same era.
Informing sequential clinical decision-making through reinforcement learning: an empirical study
Shortreed, S. M., Laber, E., Lizotte, D. J., Stroup, S., Pineau, J., and Murphy, S · 2011
Cited alongside, same era.
Mujoco: A physics engine for model-based control
Todorov, Emanuel, Erez, Tom, and Tassa, Yuval · 2012
Cited alongside, same era.
Reinforcement learning in robotics: A survey
Kober, Jens, Bagnell, J. Andrew, and Peters, Jan · 2013
Cited alongside, same era.
Fast and accurate deep network learning by exponential linear units (elus)
Benchmarking deep reinforcement learning for continuous control
Duan, Yan, Chen, Xi, Houthooft, Rein, Schulman, John, and Abbeel, Pieter · 2016
Later among the works it cites.
Deep Visual Foresight for Planning Robot Motion
Finn, C. and Levine, S · 2016
Later among the works it cites.
Generalizing skills with semi-supervised reinforcement learning
Finn, Chelsea, Yu, Tianhe, Fu, Justin, Abbeel, Pieter, and Levine, Sergey · 2016
Later among the works it cites.
Reinforcement learning with unsupervised auxiliary tasks
Jaderberg, Max, Mnih, Volodymyr, Czarnecki, Wojciech Marian, Schaul, Tom, Leibo, Joel Z, Silver, David, and Kavukcuoglu, Koray · 2016
Later among the works it cites.
Deep Successor Reinforcement Learning
Kulkarni, T. D., Saeedi, A., Gautam, S., and Gershman, S. J · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Clevert, Djork-Arné, Unterthiner, Thomas, and Hochreiter, Sepp · 2015
Cited alongside, same era.
Predictive state representations with state space partitioning
Liu, Yunlong, Tang, Yun, and Zeng, Yifeng · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, Volodymyr, Kavukcuoglu, Koray, Silver, David, Rusu, Andrei A, Veness, Joel, Bellemare, Marc G, Graves, Alex, Riedmiller, Martin, Fidjeland, Andreas K, Ostrovski, Georg, et al · 2015
Cited alongside, same era.
Trust region policy optimization
Schulman, John, Levine, Sergey, Moritz, Philipp, Jordan, Michael I., and Abbeel, Pieter · 2015
Cited alongside, same era.
Mazebase: A sandbox for learning from games
Sukhbaatar, Sainbayar, Szlam, Arthur, Synnaeve, Gabriel, Chintala, Soumith, and Fergus, Rob · 2015
Cited alongside, same era.
Embed to Control: A Locally Linear Latent Dynamics Model for Control from Raw Images
Watter, M., Springenberg, J. T., Boedecker, J., and Riedmiller, M · 2015
Cited alongside, same era.
Model-based Reinforcement Learning with Parametrized Physical Models and Optimism-Driven Exploration
Xie, C., Patil, S., Moldovan, T., Levine, S., and Abbeel, P · 2015
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Mnih, Volodymyr, Badia, Adria Puigdomenech, Mirza, Mehdi, Graves, Alex, Lillicrap, Timothy, Harley, Tim, Silver, David, and Kavukcuoglu, Koray · 2016
Later among the works it cites.
Learning to reinforcement learn
Wang, J. X, Kurth-Nelson, Z., Tirumala, D., Soyer, H., Leibo, J. Z, Munos, R., Blundell, C., Kumaran, D., and Botvinick, M · 2016
Later among the works it cites.
Practical learning of predictive state representations
Downey, Carlton, Hefny, Ahmed, and Gordon, Geoffrey · 2017
Later among the works it cites.
Treeqn and atreec: Differentiable tree planning for deep reinforcement learning
Farquhar, Gregory, Rocktäschel, Tim, Igl, Maximilian, and Whiteson, Shimon · 2017
Later among the works it cites.
Eigenoption Discovery through the Deep Successor Representation
Machado, Marlos C., Rosenbaum, Clemens, Guo, Xiaoxiao, Liu, Miao, Tesauro, Gerald, and Campbell, Murray · 2017
Later among the works it cites.
Value Prediction Network
Oh, J., Singh, S., and Lee, H · 2017
Later among the works it cites.
Proximal policy optimization algorithms
Schulman, John, Wolski, Filip, Dhariwal, Prafulla, Radford, Alec, and Klimov, Oleg · 2017
Later among the works it cites.
Elf: An extensive, lightweight and flexible research platform for real-time strategy games
Tian, Yuandong, Gong, Qucheng, Shang, Wenling, Wu, Yuxin, and Zitnick, C. Lawrence · 2017
Later among the works it cites.
Imagination-augmented agents for deep reinforcement learning
Weber, Theophane, Racanière, Sébastien, Reichert, David P., Buesing, Lars, Guez, Arthur, Rezende, Danilo Jimenez, Badia, Adrià Puigdomènech, Vinyals, Oriol, Heess, Nicolas, Li, Yujia, Pascanu, Razvan, Battaglia, Peter, Silver, David, and Wierstra, Daan · 2017
Later among the works it cites.
Deep reinforcement learning that matters
Henderson, Peter, Riashat Islam, Philip Bachman, Pineau, Joelle, Precup, Doina, and Meger, David · 2018
Closest in time.