Fetching the paper…
Reading the bibliography…
Deep reinforcement learning has recently made significant progress in solving computer games and robotic control tasks.
R. Neuneier and O. Mihatsch, “Risk sensitive reinforcement learning,” in Neural Information Processing Systems (NIPS) , 1998
1998
Earlier work this paper cites.
B. Wymann, E. Espié, C. Guionneau, C. Dimitrakakis, R. Coulom, and A. Sumner, “Torcs, the open racing car simulator,” Software available at http://torcs. sourceforge. net , 2000
2000
Earlier work this paper cites.
T. M. Moldovan and P. Abbeel, “Safe exploration in markov decision processes,” in International Conference on Machine Learning (ICML) , 2012
2012
Earlier work this paper cites.
A. Aswani, P. Bouffard, and C. Tomlin, “Extensions of learning-based model predictive control for real-time application to a quadrotor helicopter,” in American Control Conference (ACC), 2012 . IEEE, 2012, pp. 4661–4666
2012
Earlier work this paper cites.
A. Aswani, H. Gonzalez, S. S. Sastry, and C. Tomlin, “Provably safe and robust learning-based model predictive control,” Automatica , vol. 49, no. 5, pp. 1216–1226, 2013
2013
Earlier work this paper cites.
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al. , “Human-level control through deep reinforcement learning,” Nature , vol. 518, no. 7540, pp. 529–533, 2015
2015
Earlier work this paper cites.
J. Schulman, S. Levine, P. Moritz, M. I. Jordan, and P. Abbeel, “Trust region policy optimization,” in International Conference on Machine Learning (ICML) , 2015
2015
Earlier work this paper cites.
I. J. Goodfellow, J. Shlens, and C. Szegedy, “Explaining and harnessing adversarial examples,” in International Conference on Learning Representations (ICLR) , 2015
2015
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in International Conference on Learning Representations (ICLR) , 2015
2015
Earlier work this paper cites.
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, et al. , “Mastering the game of go with deep neural networks and tree search,” Nature , vol. 529, no. 7587, pp. 484–489, 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra, “Continuous control with deep reinforcement learning,” in International Conference on Learning Representations (ICLR) , 2016
2016
Earlier work this paper cites.
Y. Duan, X. Chen, R. Houthooft, J. Schulman, and P. Abbeel, “Benchmarking deep reinforcement learning for continuous control,” in International Conference on Machine Learning (ICML) , 2016
2016
Earlier work this paper cites.
A. Tamar, D. Di Castro, and S. Mannor, “Learning the variance of the reward-to-go,” Journal of Machine Learning Research , vol. 17, no. 13, pp. 1–36, 2016
2016
Earlier work this paper cites.
I. Osband, C. Blundell, A. Pritzel, and B. Van Roy, “Deep exploration via bootstrapped dqn,” in Advances in Neural Information Processing Systems , 2016, pp. 4026–4034
2016
Cited alongside, same era.
S. Carpin, Y.-L. Chow, and M. Pavone, “Risk aversion in finite markov decision processes using total cost criteria and average value at risk,” in IEEE International Conference on Robotics and Automation (ICRA) , 2016
2016
Cited alongside, same era.
M. Bellemare, S. Srinivasan, G. Ostrovski, T. Schaul, D. Saxton, and R. Munos, “Unifying count-based exploration and intrinsic motivation,” in Advances in Neural Information Processing Systems , 2016, pp. 1471–1479
2016
Cited alongside, same era.
Y. You, X. Pan, Z. Wang, and C. Lu, “Virtual to real reinforcement learning for autonomous driving,” British Machine Vision Conference , 2017
2017
Cited alongside, same era.
J. Tobin, R. Fong, A. Ray, J. Schneider, W. Zaremba, and P. Abbeel, “Domain randomization for transferring deep neural networks from simulation to the real world,” in IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , 2017
2017
Later among the works it cites.
G.-H. Liu, A. Siravuru, S. Prabhakar, M. Veloso, and G. Kantor, “Learning end-to-end multimodal sensor policies for autonomous navigation,” in Conference on Robot Learning (CoRL) , 2017
2017
Later among the works it cites.
S. Ebrahimi, A. Rohrbach, and T. Darrell, “Gradient-free policy architecture search and adaptation,” in Conference on Robot Learning (CoRL) , 2017
2017
Later among the works it cites.
Y. Zhu, Z. Wang, J. Merel, A. Rusu, T. Erez, S. Cabi, S. Tunyasuvunakool, J. Kramar, R. Hadsell, N. de Freitas, and N. Heess, “Reinforcement and imitation learning for diverse visuomotor skills,” in Robotics: Science and Systems (RSS) , 2018
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
F. Sadeghi and S. Levine, “CAD2RL: real single-image flight without a single real image,” in Robotics: Science and Systems , 2017
2017
Cited alongside, same era.
L. Pinto, J. Davidson, R. Sukthankar, and A. Gupta, “Robust adversarial reinforcement learning,” International Conference on Machine Learning (ICML) , 2017
2017
Cited alongside, same era.
A. Mandlekar, Y. Zhu, A. Garg, L. Fei-Fei, and S. Savarese, “Adversarially robust policy learning: Active construction of physically-plausible perturbations,” in IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , 2017
2017
Cited alongside, same era.
L. Pinto, J. Davidson, and A. Gupta, “Supervision via competition: Robot adversaries for learning tasks,” in IEEE International Conference on Robotics and Automation (ICRA) , 2017
2017
Cited alongside, same era.
2017
Cited alongside, same era.
J. Achiam, D. Held, A. Tamar, and P. Abbeel, “Constrained policy optimization,” in International Conference on Machine Learning (ICML) , 2017
2017
Cited alongside, same era.
D. Held, Z. McCarthy, M. Zhang, F. Shentu, and P. Abbeel, “Probabilistically safe policy transfer,” in IEEE International Conference on Robotics and Automation (ICRA) , 2017
2017
Cited alongside, same era.
A. Rajeswaran, S. Ghotra, B. Ravindran, and S. Levine, “Epopt: Learning robust neural network policies using model ensembles,” in International Conference on Learning Representations (ICLR) , 2017
2017
Cited alongside, same era.
A. Nagabandi, G. Kahn, R. S. Fearing, and S. Levine, “Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuning,” in IEEE International Conference on Robotics and Automation (ICRA) , 2018
2018
Later among the works it cites.
T. Bansal, J. Pachocki, S. Sidor, I. Sutskever, and I. Mordatch, “Emergent complexity via multi-agent competition,” in International Conference on Learning Representations (ICLR) , 2018
2018
Later among the works it cites.
S. Paul, K. Chatzilygeroudis, K. Ciosek, J.-B. Mouret, M. A. Osborne, and S. Whiteson, “Alternating optimisation and quadrature for robust control,” in AAAI 2018-The Thirty-Second AAAI Conference on Artificial Intelligence , 2018
2018
Later among the works it cites.
Y. Chow, M. Ghavamzadeh, L. Janson, and M. Pavone, “Risk-constrained reinforcement learning with percentile risk criteria,” Journal of Machine Learning Research , 2018
2018
Later among the works it cites.
M. Plappert, R. Houthooft, P. Dhariwal, S. Sidor, R. Y. Chen, X. Chen, T. Asfour, P. Abbeel, and M. Andrychowicz, “Parameter space noise for exploration,” in International Conference on Learning Representations (ICLR) , 2018
2018
Later among the works it cites.
M. Fortunato, M. G. Azar, B. Piot, J. Menick, I. Osband, A. Graves, V. Mnih, R. Munos, D. Hassabis, O. Pietquin, C. Blundell, and S. Legg, “Noisy networks for exploration,” in International Conference on Learning Representations (ICLR) , 2018
2018
Later among the works it cites.
B. Eysenbach, S. Gu, J. Ibarz, and S. Levine, “Leave no trace: Learning to reset for safe and autonomous reinforcement learning,” in International Conference on Learning Representations (ICLR) , 2018
2018
Later among the works it cites.
X. B. Peng, M. Andrychowicz, W. Zaremba, and P. Abbeel, “Sim-to-real transfer of robotic control with dynamics randomization,” in IEEE International Conference on Robotics and Automation (ICRA) , 2018
2018
Later among the works it cites.
A. Amini, L. Paull, T. Balch, S. Karaman, and D. Rus, “Learning steering bounds for parallel autonomous systems,” in IEEE International Conference on Robotics and Automation (ICRA) , 2018
2018
Later among the works it cites.