Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T. P., Harley, T., Silver, D., and Kavukcuoglu, K · 2016
Cited alongside, same era.
Learning state representation for deep actor-critic control
Munk, J., Kober, J., and Babuska, R · 2016
Cited alongside, same era.
Loss is its own reward: Self-supervision for reinforcement learning
Shelhamer, E., Mahmoudieh, P., Argus, M., and Darrell, T · 2016
Cited alongside, same era.
beta-vae: Learning basic visual concepts with a constrained variational framework
Higgins, I., Matthey, L., Pal, A., Burgess, C., Glorot, X., Botvinick, M., Mohamed, S., and Lerchner, A · 2017
Cited alongside, same era.
Reinforcement learning with unsupervised auxiliary tasks
Jaderberg, M., Mnih, V., Czarnecki, W., Schaul, T., Leibo, J. Z., Silver, D., and Kavukcuoglu, K · 2017
Cited alongside, same era.
Visual interaction networks
Watters, N., Tacchetti, A., Weber, T., Pascanu, R., Battaglia, P. W., and Zoran, D · 2017
Cited alongside, same era.
Distributional policy gradients
Barth-Maron, G., Hoffman, M. W., Budden, D., Dabney, W., Horgan, D., TB, D., Muldal, A., Heess, N., and Lillicrap, T · 2018
Cited alongside, same era.
Learning actionable representations from visual observations
Dwibedi, D., Tompson, J., Lynch, C., and Sermanet, P · 2018
Cited alongside, same era.
Recurrent world models facilitate policy evolution
Ha, D. and Schmidhuber, J · 2018
Cited alongside, same era.
Soft actor-critic algorithms and applications
Original
Haarnoja, T., Zhou, A., Hartikainen, K., Tucker, G., Ha, S., Tan, J., Kumar, V., Zhu, H., Gupta, A., Abbeel, P., et al · 2018
Cited alongside, same era.
Darla: Improving zero-shot transfer in reinforcement learning
Higgins, I., Pal, A., Rusu, A. A., Matthey, L., Burgess, C. P., Pritzel, A., Botvinick, M., Blundell, C., and Lerchner, A
Cited in the paper.
Decoupling dynamics and reward for transfer learning
Zhang, A., Satija, H., and Pineau, J
Cited in the paper.