Fetching the paper…
Reading the bibliography…
Recently, researchers have made significant progress combining the advances in deep learning for learning feature representations with reinforcement learning.
On induced stability
Stephenson, A · 1908
Earlier work this paper cites.
Dynamic Programming
Bellman, R · 1957
Earlier work this paper cites.
Error decorrelation: a technique for matching a class of functions
Donaldson, P. E. K · 1960
Earlier work this paper cites.
Pattern recognition and adaptive control
Widrow, B · 1964
Earlier work this paper cites.
BOXES: An experiment in adaptive control
Michie, D. and Chambers, R. A · 1968
Earlier work this paper cites.
Life at low Reynolds number
Purcell, E. M · 1977
Earlier work this paper cites.
Computer control of a double inverted pendulum
Furuta, K., Okutani, T., and Sone, H · 1978
Earlier work this paper cites.
3D balance in legged locomotion: modeling and simulation for the one-legged case
Murthy, S. S. and Raibert, M. H · 1984
Earlier work this paper cites.
Efficient memory-based learning for robot control
Moore, A · 1990
Earlier work this paper cites.
A case study in approximate linearization: The Acrobot example
Murray, R. M. and Hauser, J · 1991
Earlier work this paper cites.
Animation of dynamic legged locomotion
Raibert, M. H. and Hodgins, J. K · 1991
Earlier work this paper cites.
SWITCHBOARD: Telephone speech corpus for research and development
Godfrey, J. J., Holliman, E. C., and McDaniel, J · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
DARPA TIMIT acoustic-phonetic continuous speech corpus CD-ROM. NIST speech disc 1-1.1
Garofolo, J. S., Lamel, L. F., Fisher, W. M., Fiscus, J. G., and Pallett, D. S · 1993
Earlier work this paper cites.
Swinging up the Acrobot: An example of intelligent control
DeJong, G. and Spong, M. W · 1994
Earlier work this paper cites.
Neuro-dynamic programming: an overview
Bertsekas, Dimitri P and Tsitsiklis, John N · 1995
Earlier work this paper cites.
Temporal difference learning and TD-Gammon
Tesauro, G · 1995
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
2-d pole balancing with recurrent evolutionary networks
Gomez, F. and Miikkulainen, R · 1998
Earlier work this paper cites.
The MNIST database of handwritten digits, 1998
LeCun, Y., Cortes, C., and Burges, C · 1998
Earlier work this paper cites.
Reinforcement learning with hierarchies of machines
Parr, Ronald and Russell, Stuart · 1998
Earlier work this paper cites.
Stochastic real-valued reinforcement learning to solve a nonlinear control problem
Kimura, H. and Kobayashi, S · 1999
Earlier work this paper cites.
The cross-entropy method for combinatorial and continuous optimization
Rubinstein, R · 1999
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Sutton, Richard S, Precup, Doina, and Singh, Satinder · 1999
Earlier work this paper cites.
Hierarchical reinforcement learning with the MAXQ value function decomposition
Dietterich, T. G · 2000
Earlier work this paper cites.
Reinforcement learning in continuous time and space
Doya, K · 2000
Earlier work this paper cites.
The Aurora experimental framework for the performance evaluation of speech recognition systems under noisy conditions
Hirsch, H.-G. and Pearce, D · 2000
Cited alongside, same era.
Reinforcement learning with long short-term memory
Bakker, B · 2001
Cited alongside, same era.
Completely derandomized self-adaptation in evolution strategies
Hansen, N. and Ostermeier, A · 2001
Cited alongside, same era.
A database of human segmented natural images and its application to evaluating segmentation algorithms and measuring ecological statistics
Martin, D., C. Fowlkes, D. Tal, and Malik, J · 2001
Cited alongside, same era.
Reinforcement learning using neural networks, with applications to motor control
Coulom, Rémi · 2002
Cited alongside, same era.
A natural policy gradient
Kakade, S. M · 2002
ApproxRL: A Matlab toolbox for approximate RL and DP
Busoniu, L · 2010
Later among the works it cites.
The pascal visual object classes (VOC) challenge
Everingham, M., Van Gool, L., Williams, C. K. I., Winn, J., and Zisserman, A · 2010
Later among the works it cites.
Relative entropy policy search
Peters, J., Mülling, K., and Altün, Y · 2010
Later among the works it cites.
SkyAI: Highly modularized reinforcement learning library
Yamaguchi, A. and Ogasawara, T · 2010
Later among the works it cites.
Box2D: A 2D physics engine for games, 2011
Catto, E · 2011
Later among the works it cites.
Infinite horizon model predictive control for nonlinear periodic tasks
Erez, Tom, Tassa, Yuval, and Todorov, Emanuel · 2011
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Policy Gradient Toolbox
Peters, J · 2002
Cited alongside, same era.
Covariant policy search
Bagnell, J. A. and Schneider, J · 2003
Cited alongside, same era.
Policy gradient methods for robot control
Peters, J., Vijaykumar, S., and Schaal, S · 2003
Cited alongside, same era.
ε \varepsilon -MDPs: Learning in varying environments
Szita, I., Takács, B., and Lörincz, A · 2003
Cited alongside, same era.
Reinforcement learning benchmarks and bake-offs ii
Dutech, Alain, Edmunds, Timothy, Kok, Jelle, Lagoudakis, Michail, Littman, Michael, Riedmiller, Martin, Russell, Bryan, Scherrer, Bruno, Sutton, Richard, Timmer, Stephan, et al · 2005
Cited alongside, same era.
Solving partially observable reinforcement learning problems with recurrent neural networks
Schäfer, A. M. and Udluft, S · 2005
Cited alongside, same era.
Maja machine learning framework
Metzen, J. M. and Edgington, M · 2011
Later among the works it cites.
Deep neural networks for acoustic modeling in speech recognition
Hinton, G., Deng, L., Yu, D., Mohamed, A.-R., Jaitly, N., Senior, A., Vanhoucke, V., Nguyen, P., Dahl, T. S. G., and Kingsbury, B · 2012
Later among the works it cites.
ImageNet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G · 2012
Later among the works it cites.
CLS2: Closed loop simulation system
Riedmiller, M., Blum, M., and Lampe, T · 2012
Later among the works it cites.
Synthesis and stabilization of complex behaviors through online trajectory optimization
Tassa, Yuval, Erez, Tom, and Todorov, Emanuel · 2012
Later among the works it cites.
MuJoCo: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Later among the works it cites.
RLLib: Lightweight standard and on/off policy reinforcement learning library (C++)
Abeyruwan, S · 2013
Later among the works it cites.
The Arcade Learning Environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Later among the works it cites.
A survey on policy search for robotics, foundations and trends in robotics
Deisenroth, M. P., Neumann, G., and Peters, J · 2013
Later among the works it cites.
The open-source TEXPLORE code release for reinforcement learning on robots
Hester, T. and Stone, P · 2013
Later among the works it cites.
Guided policy search
Levine, S. and Koltun, V · 2013
Later among the works it cites.
dotrl: A platform for rapid reinforcement learning methods development and validation
Papis, B. and Wawrzyński, P · 2013
Later among the works it cites.
Policy evaluation with temporal differences: A survey and comparison
Dann, C., Neumann, G., and Peters, J · 2014
Later among the works it cites.
The reinforcement learning competition 2014
Dimitrakakis, Christos, Li, Guangliang, and Tziortziotis, Nikoalos · 2014
Later among the works it cites.
Deep learning for real-time Atari game play using offline monte-carlo tree search planning
Guo, X., Singh, S., Lee, H., Lewis, R. L., and Wang, X · 2014
Later among the works it cites.
End-to-end training of deep visuomotor policies
Levine, S., Finn, C., Darrell, T., and Abbeel, P · 2015
Later among the works it cites.
Continuous control with deep reinforcement learning
Lillicrap, T., Hunt, J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2015
Later among the works it cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., and Hassabis, D · 2015
Later among the works it cites.
Learning of non-parametric control policies with high-dimensional state features
van Hoof, H., Peters, J., and Neumann, G · 2015
Later among the works it cites.
Embed to control: A locally linear latent dynamics model for control from raw images
Watter, M., Springenberg, J., Boedecker, J., and Riedmiller, M · 2015
Later among the works it cites.