Fetching the paper…
Reading the bibliography…
Many real-world physical control systems are required to satisfy constraints upon deployment.
Constrained Markov decision processes , volume 7
Altman, E · 1999
Earlier work this paper cites.
Robust dynamic programming
Iyengar, G. N · 2005
Earlier work this paper cites.
Policy gradients with variance related risk criteria
Di Castro, D., Tamar, A., and Mannor, S · 2012
Earlier work this paper cites.
Scaling up robust mdps using function approximation
Tamar, A., Mannor, S., and Xu, H · 2014
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., and Hassabis, D · 2015
Earlier work this paper cites.
Transfer from simulation to real world through learning deep inverse dynamics model
Christiano, P. F., Shah, Z., Mordatch, I., Schneider, J., Blackwell, T., Tobin, J., Abbeel, P., and Zaremba, W · 2016
Earlier work this paper cites.
Constrained policy optimization
Achiam, J., Held, D., Tamar, A., and Abbeel, P · 2017
Earlier work this paper cites.
A distributional perspective on reinforcement learning
Bellemare, M. G., Dabney, W., and Munos, R · 2017
Earlier work this paper cites.
Mastering the game of Go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., Chen, Y., Lillicrap, T., Hui, F., Sifre, L., van den Driessche, G., Graepel, T., and Hassabis, D · 2017
Earlier work this paper cites.
A deep hierarchical approach to lifelong learning in minecraft
Tessler, C., Givony, S., Zahavy, T., Mankowitz, D. J., and Mannor, S · 2017
Cited alongside, same era.
Mutual alignment transfer learning
Wulfmeier, M., Posner, I., and Abbeel, P · 2017
Cited alongside, same era.
Learning dexterous in-hand manipulation
Andrychowicz, M., Baker, B., Chociej, M., Jozefowicz, R., McGrew, B., Pachocki, J., Petron, A., Plappert, M., Powell, G., Ray, A., et al · 2018
Cited alongside, same era.
Distributed distributional deterministic policy gradients
Barth-Maron, G., Hoffman, M. W., Budden, D., Dabney, W., Horgan, D., Tb, D., Muldal, A., Heess, N., and Lillicrap, T · 2018
Cited alongside, same era.
Value constrained model-free continuous control
Bohez, S., Abdolmaleki, A., Neunert, M., Buchli, J., Heess, N., and Hadsell, R · 2019
Later among the works it cites.
A bayesian approach to robust reinforcement learning
Derman, E., Mankowitz, D. J., Mann, T. A., and Mannor, S · 2019
Later among the works it cites.
Challenges of real-world reinforcement learning
Dulac-Arnold, G., Mankowitz, D. J., and Hester, T · 2019
Later among the works it cites.
Robust reinforcement learning for continuous control with model misspecification
Mankowitz, D. J., Levine, N., Jeong, R., Abdolmaleki, A., Springenberg, J. T., Mann, T. A., Hester, T., and Riedmiller, M. A · 2019
Later among the works it cites.
A distributional view on multi-objective policy optimization
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Derman, E., Mankowitz, D. J., Mann, T. A., and Mannor, S · 2018
Cited alongside, same era.
Learning robust options
Mankowitz, D. J., Mann, T. A., Bacon, P.-L., Precup, D., and Mannor, S · 2018
Cited alongside, same era.
Sim-to-real transfer of robotic control with dynamics randomization
Peng, X. B., Andrychowicz, M., Zaremba, W., and Abbeel, P · 2018
Cited alongside, same era.
Sample-efficient reinforcement learning via difference models
Rastogi, D., Koryakovskiy, I., and Kober, J · 2018
Cited alongside, same era.
Tassa, Y., Doron, Y., Muldal, A., Erez, T., Li, Y., de Las Casas, D., Budden, D., Abdolmaleki, A., Merel, J., Lefrancq, A., Lillicrap, T. P., and Riedmiller, M. A · 2018
Cited alongside, same era.
Reward constrained policy optimization
Tessler, C., Mankowitz, D. J., and Mannor, S · 2018
Cited alongside, same era.
Relative entropy regularized policy iteration
Abdolmaleki, A., Springenberg, J. T., Degrave, J., Bohez, S., Tassa, Y., Belov, D., Heess, N., and Riedmiller, M. A
Cited in the paper.
Maximum a posteriori policy optimisation
Abdolmaleki, A., Springenberg, J. T., Tassa, Y., Munos, R., Heess, N., and Riedmiller, M
Cited in the paper.
Abdolmaleki, A., Huang, S. H., Hasenclever, L., Neunert, M., Song, H. F., Zambelli, M., Martins, M. F., Heess, N., Hadsell, R., and Riedmiller, M · 2020
Closest in time.
Balancing constraints and rewards with meta-gradient d4pg, 2020
Calian, D. A., Mankowitz, D. J., Zahavy, T., Xu, Z., Oh, J., Levine, N., and Mann, T · 2020
Closest in time.
Scalable neural learning for verifiable consistency with temporal specifications, 2020
Dathathri, S., Welbl, J., Dvijotham, K. D., Kumar, R., Kanade, A., Uesato, J., Gowal, S., Huang, P.-S., and Kohli, P · 2020
Closest in time.
An empirical investigation of the challenges of real-world reinforcement learning
Dulac-Arnold, G., Levine, N., Mankowitz, D. J., Li, J., Paduraru, C., Gowal, S., and Hester, T · 2020
Closest in time.
Exploration-exploitation in constrained mdps, 2020
Efroni, Y., Mannor, S., and Pirotta, M · 2020
Closest in time.
Certified adversarial robustness for deep reinforcement learning, 2020
Everett, M., Lutjens, B., and How, J. P · 2020
Closest in time.