Wasserstein gan, 2017
Arjovsky, M., Chintala, S., and Bottou, L · 2017
Cited alongside, same era.
Reinforcement learning with deep energy-based policies, 2017
Haarnoja, T., Tang, H., Abbeel, P., and Levine, S · 2017
Cited alongside, same era.
Sequence tutor: Conservative fine-tuning of sequence generation models with kl-control
Jaques, N., Gu, S., Bahdanau, D., Hernández-Lobato, J. M., Turner, R. E., and Eck, D · 2017
Cited alongside, same era.
Bridging the gap between value and policy based reinforcement learning
Original
Nachum, O., Norouzi, M., Xu, K., and Schuurmans, D · 2017
Cited alongside, same era.
Behavior regularized offline reinforcement learning
Original
Wu, Y., Tucker, G., and Nachum, O · 2017
Cited alongside, same era.
Image-to-image translation for cross-domain disentanglement, 2018
Gonzalez-Garcia, A., van de Weijer, J., and Bengio, Y · 2018
Cited alongside, same era.
Discriminator-actor-critic: Addressing sample inefficiency and reward bias in adversarial imitation learning, 2018
Kostrikov, I., Agrawal, K. K., Dwibedi, D., Levine, S., and Tompson, J · 2018
Cited alongside, same era.
Spectral normalization for generative adversarial networks
Miyato, T., Kataoka, T., Koyama, M., and Yoshida, Y · 2018
Cited alongside, same era.
Off-policy deep reinforcement learning without exploration
Fujimoto, S., Meger, D., and Precup, D · 2019
Cited alongside, same era.
Soft actor-critic algorithms and applications, 2019
Haarnoja, T., Zhou, A., Hartikainen, K., Tucker, G., Ha, S., Tan, J., Kumar, V., Zhu, H., Gupta, A., Abbeel, P., and Levine, S · 2019
Cited alongside, same era.
Way off-policy batch deep reinforcement learning of implicit human preferences in dialog
Original
Jaques, N., Ghandeharioun, A., Shen, J. H., Ferguson, C., Lapedriza, A., Jones, N., Gu, S., and Picard, R · 2019
Cited alongside, same era.
Stabilizing off-policy q-learning via bootstrapping error reduction
Kumar, A., Fu, J., Soh, M., Tucker, G., and Levine, S · 2019
Cited alongside, same era.