Fetching the paper…
Reading the bibliography…
Model-free deep reinforcement learning has been shown to exhibit good performance in domains ranging from video games to simulated robotic manipulation and locomotion.
A. J. Ijspeert, J. Nakanishi, and S. Schaal, “Learning attractor landscapes for learning motor primitives,” in Advances in neural information processing systems , 2003, pp. 1547–1554
2003
Earlier work this paper cites.
H. J. Kappen, “Path integrals and symmetry breaking for optimal control theory,” Journal of Statistical Mechanics: Theory And Experiment , vol. 2005, no. 11, p. P11011, 2005
2005
Earlier work this paper cites.
J. Peters and S. Schaal, “Policy gradient methods for robotics,” in Intelligent Robots and Systems, 2006 IEEE/RSJ International Conference on . IEEE, 2006, pp. 2219–2225
2006
Earlier work this paper cites.
F. Guenter, M. Hersch, S. Calinon, and A. Billard, “Reinforcement learning for imitating constrained reaching movements,” Advanced Robotics , vol. 21, no. 13, pp. 1521–1544, 2007
2007
Earlier work this paper cites.
E. Todorov, “Linearly-solvable Markov decision problems,” in Advances in Neural Information Processing Systems . MIT Press, 2007, pp. 1369–1376
2007
Earlier work this paper cites.
B. D. Ziebart, A. L. Maas, J. A. Bagnell, and A. K. Dey, “Maximum entropy inverse reinforcement learning,” in AAAI Conference on Artificial Intelligence , 2008, pp. 1433–1438
2008
Earlier work this paper cites.
J. Peters and S. Schall, “Reinforcement learning of motor skills with policy gradients,” Neural networks , vol. 21, no. 4, pp. 682–692, 2008
2008
Earlier work this paper cites.
E. Todorov, “General duality between optimal control and estimation,” in IEEE Conf. on Decision and Control . IEEE, 2008, pp. 4286–4292
2008
Earlier work this paper cites.
P. Pastor, H. Hoffmann, T. Asfour, and S. Schaal, “Learning and generalization of motor skills by learning from demonstration,” in Robotics and Automation, 2009. ICRA’09. IEEE International Conference on . IEEE, 2009, pp. 763–768
2009
Earlier work this paper cites.
M. Toussaint, “Robot trajectory optimization using approximate inference,” in Int. Conf. on Machine Learning . ACM, 2009, pp. 1049–1056
2009
Earlier work this paper cites.
E. Todorov, “Compositionality of optimal control laws,” in Advances in Neural Information Processing Systems , 2009, pp. 1856–1864
2009
Earlier work this paper cites.
E. Theodorou, J. Buchli, and S. Schaal, “Reinforcement learning of motor skills in high dimensions: A path integral approach,” in Robotics and Automation (ICRA), 2010 IEEE International Conference on . IEEE, 2010, pp. 2397–2403
2010
Earlier work this paper cites.
K. Rawlik, M. Toussaint, and S. Vijayakumar, “On stochastic optimal control and reinforcement learning by approximate inference,” Proceedings of Robotics: Science and Systems VIII , 2012
2012
Earlier work this paper cites.
E. Todorov, T. Erez, and Y. Tassa, “MuJoCo: A physics engine for model-based control,” in Intelligent Robots and Systems (IROS), 2012 IEEE/RSJ International Conference on . IEEE, 2012, pp. 5026–5033
2012
Cited alongside, same era.
A. J. Ijspeert, J. Nakanishi, H. Hoffmann, P. Pastor, and S. Schaal, “Dynamical movement primitives: learning attractor models for motor behaviors,” Neural computation , vol. 25, no. 2, pp. 328–373, 2013
2013
Cited alongside, same era.
2015
Cited alongside, same era.
2015
Cited alongside, same era.
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
A. Ghadirzadeh, A. Maki, D. Kragic, and M. Björkman, “Deep predictive policy training using reinforcement learning,” in Intelligent Robots and Systems (IROS), 2017 IEEE/RSJ International Conference on . IEEE, 2017, pp. 2351–2358
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
2016
Cited alongside, same era.
S. Levine, C. Finn, T. Darrell, and P. Abbeel, “End-to-end training of deep visuomotor policies,” Journal of Machine Learning Research , vol. 17, no. 39, pp. 1–40, 2016
2016
Cited alongside, same era.
2016
Cited alongside, same era.
S. Gu, T. Lillicrap, I. Sutskever, and S. Levine, “Continuous deep Q-learning with model-based acceleration,” in Int. Conf. on Machine Learning , 2016, pp. 2829–2838
2016
Cited alongside, same era.
R. Fox, A. Pakman, and N. Tishby, “Taming the noise in reinforcement learning via soft updates,” in Conf. on Uncertainty in Artificial Intelligence , 2016
2016
Cited alongside, same era.
Q. Liu and D. Wang, “Stein variational gradient descent: A general purpose bayesian inference algorithm,” in Advances In Neural Information Processing Systems , 2016, pp. 2370–2378
2016
Cited alongside, same era.
2016
Cited alongside, same era.
J. Tobin, R. Fong, A. Ray, J. Schneider, W. Zaremba, and P. Abbeel, “Domain randomization for transferring deep neural networks from simulation to the real world,” in Intelligent Robots and Systems (IROS), 2017 IEEE/RSJ International Conference on . IEEE, 2017, pp. 23–30
2017
Later among the works it cites.
M. Andrychowicz, D. Crow, A. Ray, J. Schneider, R. Fong, P. Welinder, B. McGrew, J. Tobin, P. Abbeel, and W. Zaremba, “Hindsight experience replay,” in Advances in Neural Information Processing Systems , 2017, pp. 5055–5065
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
O. Nachum, M. Norouzi, K. Xu, and D. Schuurmans, “Bridging the gap between value and policy based reinforcement learning,” in Advances in Neural Information Processing Systems , 2017, pp. 2772–2782
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.