Fetching the paper…
Reading the bibliography…
Learning by interaction is the key to skill acquisition for most living organisms, which is formally called Reinforcement Learning (RL).
Dormand, J. R. and P. J. Prince, A family of embedded Runge-Kutta formulae. Journal of Computational and Applied Mathematics
1980
Earlier work this paper cites.
Ken Perlin. An image synthesizer. ACM SIGGRAPH Computer Graphics,
1985
Earlier work this paper cites.
Richard Sutton. Integrated architectures for learning, planning, and reacting based on approximating dynamic programming. International Conference on Machine Learning, ICML,
1990
Earlier work this paper cites.
Leslie Kaelbli, Michael Littman, and Andrew Moore. Reinforcement learning: a survey. Journal of Articial Intelligence Research, JAIR,
1996
Earlier work this paper cites.
Sepp Hochreiter and Jürgen Schmidhuber. Long short-term memory. Neural Computation
1997
Earlier work this paper cites.
Weiwei Li and Emanuel Todorov. Iterative linear quadratic regulator design for nonlinear biological movement systems. International Conference on Informatics in Control, Automation, and Robotics, ICINCO,
2004
Earlier work this paper cites.
2004
Earlier work this paper cites.
Pascal Vincent, Hugo Larochelle, Isabelle Lajoie, Yoshua Bengio, and Pierre-Antoine Manzagol. Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion. The Journal of Machine Learning Research, JMLR,
2010
Earlier work this paper cites.
Mark Deisenroth and Carl Rasmussen. PILCO: A model-based and data-efficient approach to policy search. International Conference on Machine Learning, ICML,
2011
Earlier work this paper cites.
Cameron Browne, Edward Powley, Daniel Whitehouse, Simon Lucas, Peter Cowling, et al. A survey of monte carlo tree search methods. IEEE Transactions on Computational Intelligence and AI in Games,
2012
Earlier work this paper cites.
Yuval Tassa, Tom Erez, and Emanuel Todorov. Synthesis and stabilization of complex behaviors through online trajectory optimization. International Conference on Intelligent Robots and Systems, IROS,
2012
Earlier work this paper cites.
Stephane Ross and J. Andrew Bagnell. Agnostic system identification for model-based reinforcement learning. International Conference on Machine Learning, ICML,
2012
Earlier work this paper cites.
Sergey Levine and Vtruladlen Koltun. Guided policy search. International Conference on Machine Learning, ICML,
2013
Earlier work this paper cites.
Zdravko Botev1, Dirk Kroese, Reuven Rubinstein, and Pierre L’Ecuyer. The cross-entropy method for optimization. Handbook of Statistics, volume 31, chapter 3,
2013
Earlier work this paper cites.
Joschika Boedecker, Jost Springenberg, Jan Wulfing, and Martin Riedmiller. Approximate real-time optimal control based on sparse gaussian process models. IEEE Symposium on Adaptive Dynamic Programming and Reinforcement Learning, ADPRL
2014
Cited alongside, same era.
Diederik P. Kingma and Max Welling. Auto-encoding variational Bayes. International Conference on Learning Representations, ICLR,
2014
Cited alongside, same era.
Sergey Levine and Pieter Abbeel. Learning neural network policies with guided policy search under unknown dynamics. Advances in Neural Information Processing Systems, NeurIPS,
2014
Cited alongside, same era.
Junyoung Chung, Kyle Kastner, Laurent Dinh, Kratarth Goel, Aaron Courville, Yoshua Bengio. A recurrent latent variable model for sequential data. Advances in neural information processing systems,
2015
Cited alongside, same era.
David Ha and Jürgen Schmidhuber. World models. arXiv preprint arXiv:1803.10122
2018
Later among the works it cites.
Thanard Kurutach, Ignasi Clavera, Yan Duan, Aviv Tamar, and Pieter Abbeel. Model-ensemble trust-region policy optimization. International Conference on Learning Representations, ICLR,
2018
Later among the works it cites.
Tian Qi Chen, Yulia Rubanova, Jesse Bettencourt, and David Duvenaud. Neural ordinary differential equations. Advances in Neural Information Processing Systems, NeurIPS,
2018
Later among the works it cites.
Brandon Amos, Ivan Dario Jimenez Rodriguez, Jacob Sacks, Byron Boots, and J. Zico Kolter. Differentiable MPC for end-to-end planning and control. Advances in Neural Information Processing Systems, NeurIPS,
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei Rusu, et al. Human-level control through deep reinforcement learning. Nature,
2015
Cited alongside, same era.
Manuel Watter, Jost Tobias Springenberg, Joschka Boedecker, and Martin Riedmiller . Embed to control: a locally linear latent dynamics model for control from raw images. Advances in Neural Information Processing Systems, NeurIPS,
2015
Cited alongside, same era.
2015
Cited alongside, same era.
Shixiang Gu, Timothy Lillicrap, Ilya Sutskever, and Sergey Levine. Continuous deep Q-learning with model-based acceleration. International Conference on Machine Learning, ICML,
2016
Cited alongside, same era.
2016
Cited alongside, same era.
Charles Qi, Hao Su, Kaichun Mo, and Leonidas Guibas. Pointnet: Deep learning on point sets for 3d classification and segmentation. Proc. IEEE Computer Vision and Pattern Recognition, CVPR,
2017
Cited alongside, same era.
2017
Cited alongside, same era.
Thomas Anthony, Zheng Tian, and David Barber. Thinking fast and slow with deep learning and tree search. Advances in Neural Information Processing Systems, NeurIPS,
2017
Cited alongside, same era.
2019
Later among the works it cites.
Danijar Hafner, Timothy Lillicrap, Ian Fischer, Ruben Villegas, David Ha, Honglak Lee, and James Davidson. Learning latent dynamics for planning from pixels. International Conference on Machine Learning, ICML,
2019
Later among the works it cites.
2019
Later among the works it cites.
OpenAI. Dota 2 with Large Scale Deep Reinforcement Learning. arXiv preprint arXiv:1912.06680
2019
Later among the works it cites.
OpenAI. Solving Rubik’s Cube with a Robot Hand. arXiv preprint arXiv:1910.07113
2019
Later among the works it cites.
2019
Later among the works it cites.
Oier Mees, Markus Merklinger, Gabriel Kalweit, and Wolfram Burgard. Adversarial Skill Networks: Unsupervised Robot Skill Learning from Videos International Conference on Robotics and Automation, ICRA,
2020
Later among the works it cites.
2020
Later among the works it cites.
Kanishka Rao, Chris Harris, Alex Irpan, Sergey Levine, Julian Ibarz, and Mohi Khansari. RL-CycleGAN: Reinforcement Learning Aware Simulation-To-Real. Conference on Computer Vision and Pattern Recognition, CVPR,
2020
Later among the works it cites.