Continuous deep q-learning with model-based acceleration
Gu, S., Lillicrap, T., Sutskever, I., and Levine, S · 2016
Later among the works it cites.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K · 2016
Later among the works it cites.
Pgq: Combining policy gradient and q-learning
Original
O’Donoghue, B., Munos, R., Kavukcuoglu, K., and Mnih, V · 2016
Later among the works it cites.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al · 2016
Later among the works it cites.
Data-efficient off-policy policy evaluation for reinforcement learning
Thomas, P. and Brunskill, E · 2016
Later among the works it cites.
The reactor: A sample-efficient actor-critic architecture
Original
Gruslys, A., Azar, M. G., Bellemare, M. G., and Munos, R · 2017
Later among the works it cites.
Time limits in reinforcement learning
Original
Pardo, F., Tavakoli, A., Levdik, V., and Kormushev, P · 2017
Later among the works it cites.
Proximal policy optimization algorithms
Original
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Later among the works it cites.
Policy gradient methods for reinforcement learning with function approximation and action-dependent baselines
Original
Thomas, P. S. and Brunskill, E · 2017
Later among the works it cites.
Backpropagation through the void: Optimizing control variates for black-box gradient estimation
Grathwohl, W., Choi, D., Wu, Y., Roeder, G., and Duvenaud, D · 2018
Closest in time.
Action-dependent control variates for policy optimization via stein identity
Liu, H., Feng, Y., Mao, Y., Zhou, D., Peng, J., and Liu, Q · 2018
Closest in time.
Variance reduction for policy gradient with action-dependent factorized baselines
Wu, C., Rajeswaran, A., Duan, Y., Kumar, V., Bayen, A. M., Kakade, S., Mordatch, I., and Abbeel, P · 2018
Closest in time.