Soft-robust actor-critic policy-gradient
Original
Derman, E., Mankowitz, D. J., Mann, T. A., and Mannor, S · 2018
Later among the works it cites.
IMPALA: Scalable distributed deep-RL with importance weighted actor-learner architectures
Espeholt, L., Soyer, H., Munos, R., Simonyan, K., Mnih, V., Ward, T., Doron, Y., Firoiu, V., Harley, T., Dunning, I., Legg, S., and Kavukcuoglu, K · 2018
Later among the works it cites.
More robust doubly robust off-policy evaluation
Original
Farajtabar, M., Chow, Y., and Ghavamzadeh, M · 2018
Later among the works it cites.
Learning latent dynamics for planning from pixels
Original
Hafner, D., Lillicrap, T., Fischer, I., Villegas, R., Ha, D., Lee, H., and Davidson, J · 2018
Later among the works it cites.
Deep reinforcement learning that matters
Henderson, P., Islam, R., Bachman, P., Pineau, J., Precup, D., and Meger, D · 2018
Later among the works it cites.
Deep q-learning from demonstrations
Hester, T., Vecerik, M., Pietquin, O., Lanctot, M., Schaul, T., Piot, B., Horgan, D., Quan, J., Sendonaris, A., Osband, I., Dulac-Arnold, G., Agapiou, J., Leibo, J. Z., and Gruslys, A · 2018
Later among the works it cites.
Distributed prioritized experience replay
Original
Horgan, D., Quan, J., Budden, D., Barth-Maron, G., Hessel, M., van Hasselt, H., and Silver, D · 2018
Later among the works it cites.
Optimizing agent behavior over long time scales by transporting value
Original
Hung, C.-C., Lillicrap, T., Abramson, J., Wu, Y., Mirza, M., Carnevale, F., Ahuja, A., and Wayne, G · 2018
Later among the works it cites.
Deep reinforcement learning doesn’t work yet
Irpan, A · 2018
Later among the works it cites.
Learning robust options
Mankowitz, D. J., Mann, T. A., Bacon, P.-L., Precup, D., and Mannor, S · 2018
Later among the works it cites.
Learning from delayed outcomes with intermediate observations
Original
Mann, T. A., Gowal, S., Jiang, R., Hu, H., Lakshminarayanan, B., and György, A · 2018
Later among the works it cites.
Methods for interpreting and understanding deep neural networks
Montavon, G., Samek, W., and Müller, K.-R · 2018
Later among the works it cites.
Deep online learning via meta-learning: Continual adaptation for model-based RL
Original
Nagabandi, A., Finn, C., and Levine, S · 2018
Later among the works it cites.
Sim-to-real transfer of robotic control with dynamics randomization
Peng, X. B., Andrychowicz, M., Zaremba, W., and Abbeel, P · 2018
Later among the works it cites.
Manipulating and measuring model interpretability
Original
Poursabzi-Sangdeh, F., Goldstein, D. G., Hofman, J. M., Vaughan, J. W., and Wallach, H. M · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Later among the works it cites.
Deepmind control suite
Original
Tassa, Y., Doron, Y., Muldal, A., Erez, T., Li, Y., Casas, D. d. L., Budden, D., Abdolmaleki, A., Merel, J., Lefrancq, A., et al · 2018
Later among the works it cites.
Reward constrained policy optimization
Original
Tessler, C., Mankowitz, D. J., and Mannor, S · 2018
Later among the works it cites.
Programmatically interpretable reinforcement learning
Original
Verma, A., Murali, V., Singh, R., Kohli, P., and Chaudhuri, S · 2018
Later among the works it cites.
Safe exploration and optimization of constrained mdps using gaussian processes
Wachi, A., Sui, Y., Yue, Y., and Ono, M · 2018
Later among the works it cites.
Learn what not to learn: Action elimination with deep reinforcement learning
Zahavy, T., Haroush, M., Merlis, N., Mankowitz, D. J., and Mannor, S · 2018
Later among the works it cites.
Value constrained model-free continuous control
Original
Bohez, S., Abdolmaleki, A., Neunert, M., Buchli, J., Heess, N., and Hadsell, R · 2019
Closest in time.