Gradient diversity empowers distributed learning
Original
D. Yin, A. Pananjady, M. Lam, D. S. Papailiopoulos, K. Ramchandran, and P. Bartlett · 2017
Later among the works it cites.
RUDDER: Return decomposition for delayed rewards
Original
J. A. Arjona-Medina, M. Gillhofer, M. Widrich, T. Unterthiner, J. Brandstetter, and S. Hochreiter · 2018
Later among the works it cites.
Adapting auxiliary losses using gradient similarity
Original
Y. Du, W. M. Czarnecki, S. M. Jayakumar, R. Pascanu, and B. Lakshminarayanan · 2018
Later among the works it cites.
IMPALA: Scalable distributed Deep-RL with importance weighted actor-learner architectures
Original
L. Espeholt, H. Soyer, R. Munos, K. Simonyan, V. Mnih, T. Ward, Y. Doron, V. Firoiu, T. Harley, I. Dunning, et al · 2018
Later among the works it cites.
Dynamic task prioritization for multitask learning
M. Guo, A. Haque, D.-A. Huang, S. Yeung, and L. Fei-Fei · 2018
Later among the works it cites.
Rainbow: Combining improvements in deep reinforcement learning
M. Hessel, J. Modayil, H. Van Hasselt, T. Schaul, G. Ostrovski, W. Dabney, D. Horgan, B. Piot, M. Azar, and D. Silver · 2018
Later among the works it cites.
Multi-task deep reinforcement learning with popart
Original
M. Hessel, H. Soyer, L. Espeholt, W. Czarnecki, S. Schmitt, and H. van Hasselt · 2018
Later among the works it cites.
Openai five
OpenAI · 2018
Later among the works it cites.
Learning dexterous in-hand manipulation
Original
OpenAI, M. Andrychowicz, B. Baker, M. Chociej, R. Józefowicz, B. McGrew, J. W. Pachocki, J. Pachocki, A. Petron, M. Plappert, G. Powell, A. Ray, J. Schneider, S. Sidor, J. Tobin, P. Welinder, L. Weng, and W. Zaremba · 2018
Later among the works it cites.
The Barbados 2018 list of open issues in continual learning
Original
T. Schaul, H. van Hasselt, J. Modayil, M. White, A. White, P. Bacon, J. Harb, S. Mourad, M. G. Bellemare, and D. Precup · 2018
Later among the works it cites.
A general reinforcement learning algorithm that masters chess, shogi, and go through self-play
D. Silver, T. Hubert, J. Schrittwieser, I. Antonoglou, M. Lai, A. Guez, M. Lanctot, L. Sifre, D. Kumaran, T. Graepel, et al · 2018
Later among the works it cites.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2018
Later among the works it cites.
The value function polytope in reinforcement learning
Original
R. Dadashi, A. A. Taïga, N. L. Roux, D. Schuurmans, and M. G. Bellemare · 2019
Closest in time.
Learning to learn without forgetting by maximizing transfer and minimizing interference
M. Riemer, I. Cases, R. Ajemian, M. Liu, I. Rish, Y. Tu, , and G. Tesauro · 2019
Closest in time.
AlphaStar: Mastering the Real-Time Strategy Game StarCraft II
O. Vinyals, I. Babuschkin, J. Chung, M. Mathieu, M. Jaderberg, W. M. Czarnecki, A. Dudzik, A. Huang, P. Georgiev, R. Powell, T. Ewalds, D. Horgan, M. Kroiss, I. Danihelka, J. Agapiou, J. Oh, V. Dalibard, D. Choi, L. Sifre, Y. Sulsky, S. Vezhnevets, J. Molloy, T. Cai, D. Budden, T. Paine, C. Gulcehre, Z. Wang, T. Pfaff, T. Pohlen, Y. Wu, D. Yogatama, J. Cohen, K. McKinney, O. Smith, T. Schaul, T. Lillicrap, C. Apps, K. Kavukcuoglu, D. Hassabis, and D. Silver · 2019
Closest in time.