Fetching the paper…
Reading the bibliography…
Conventional Reinforcement Learning (RL) algorithms usually have one single agent learning to solve the task independently.
B. L. Smith and J. T. MacGregor, “What is collaborative learning,” 1992
1992
Earlier work this paper cites.
E. G. Cohen, “Restructuring the classroom: Conditions for productive small groups,”
1994
Earlier work this paper cites.
R. S. Sutton and A. G. Barto, “Reinforcement learning: An introduction,” 1998
1998
Earlier work this paper cites.
H.-G. Beyer and H.-P. Schwefel, “Evolution strategies–a comprehensive introduction,”
2002
Earlier work this paper cites.
T. Rückstiess, F. Sehnke, T. Schaul, D. Wierstra, Y. Sun, and J. Schmidhuber, “Exploring parameter space in reinforcement learning,”
2010
Earlier work this paper cites.
J. Lehman and K. O. Stanley, “Abandoning objectives: Evolution through the search for novelty alone,”
2011
Earlier work this paper cites.
E. Todorov, T. Erez, and Y. Tassa, “Mujoco: A physics engine for model-based control,” in
2012
Earlier work this paper cites.
S. Doncieux and J.-B. Mouret, “Behavioral diversity with multiple behavioral distances,” in
2013
Earlier work this paper cites.
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski,
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu, “Asynchronous methods for deep reinforcement learning,” in
2016
Earlier work this paper cites.
I. Osband, C. Blundell, A. Pritzel, and B. Van Roy, “Deep exploration via bootstrapped dqn,” in
2016
Earlier work this paper cites.
R. Houthooft, X. Chen, Y. Duan, J. Schulman, F. De Turck, and P. Abbeel, “Vime: Variational information maximizing exploration,” in
2016
Earlier work this paper cites.
M. Bellemare, S. Srinivasan, G. Ostrovski, T. Schaul, D. Saxton, and R. Munos, “Unifying count-based exploration and intrinsic motivation,” in
2016
Earlier work this paper cites.
J. K. Pugh, L. B. Soros, and K. O. Stanley, “Quality diversity: A new frontier for evolutionary computation,”
2016
Cited alongside, same era.
2017
Cited alongside, same era.
T. Haarnoja, H. Tang, P. Abbeel, and S. Levine, “Reinforcement learning with deep energy-based policies,” in
2017
Cited alongside, same era.
2017
Cited alongside, same era.
2017
2018
Later among the works it cites.
I. Adamski, R. Adamski, T. Grel, A. Jedrych, K. Kaczmarek, and H. Michalewski, “Distributed deep reinforcement learning: Learn how to play atari games in 21 minutes,” in
2018
Later among the works it cites.
2018
Later among the works it cites.
E. Conti, V. Madhavan, F. P. Such, J. Lehman, K. Stanley, and J. Clune, “Improving exploration in evolution strategies for deep reinforcement learning via a population of novelty-seeking agents,” in
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
J. Achiam and S. Sastry, “Surprise-based intrinsic motivation for deep reinforcement learning,”
2017
Cited alongside, same era.
I. Osband, D. Russo, Z. Wen, and B. Van Roy, “Deep exploration via randomized value functions,”
2017
Cited alongside, same era.
H. Tang, R. Houthooft, D. Foote, A. Stooke, O. X. Chen, Y. Duan, J. Schulman, F. DeTurck, and P. Abbeel, “# exploration: A study of count-based exploration for deep reinforcement learning,” in
2017
Cited alongside, same era.
G. Ostrovski, M. G. Bellemare, A. van den Oord, and R. Munos, “Count-based exploration with neural density models,” in
2017
Cited alongside, same era.
D. Pathak, P. Agrawal, A. A. Efros, and T. Darrell, “Curiosity-driven exploration by self-supervised prediction,” in
2017
Cited alongside, same era.
R. Lowe, Y. I. Wu, A. Tamar, J. Harb, O. P. Abbeel, and I. Mordatch, “Multi-agent actor-critic for mixed cooperative-competitive environments,” in
2017
Cited alongside, same era.
L. Espeholt, H. Soyer, R. Munos, K. Simonyan, V. Mnih, T. Ward, Y. Doron, V. Firoiu, T. Harley, I. Dunning,
2018
Cited alongside, same era.
2018
Later among the works it cites.
S. Fujimoto, D. Meger, and D. Precup, “Off-policy deep reinforcement learning without exploration,”
2018
Later among the works it cites.
E. Liang, R. Liaw, R. Nishihara, P. Moritz, R. Fox, K. Goldberg, J. Gonzalez, M. Jordan, and I. Stoica, “Rllib: Abstractions for distributed reinforcement learning,” in
2018
Later among the works it cites.
M. Chevalier-Boisvert, L. Willems, and S. Pal, “Minimalistic gridworld environment for openai gym.”
2018
Later among the works it cites.
C. Tessler, G. Tennenholtz, and S. Mannor, “Distributional policy optimization: An alternative approach for continuous control,” in
2019
Later among the works it cites.
K. Ciosek, Q. Vuong, R. Loftin, and K. Hofmann, “Better exploration with optimistic actor critic,” in
2019
Later among the works it cites.
M. Masood and F. Doshi-Velez, “Diversity-inducing policy gradient: using maximum mean discrepancy to find a set of diverse policies,” in
2019
Later among the works it cites.
Y. Zhang, W. Yu, and G. Turk, “Learning novel policies for tasks,” in
2019
Later among the works it cites.
2019
Later among the works it cites.
S. Schmitt, M. Hessel, and K. Simonyan, “Off-policy actor-critic with shared experience replay,”
2019
Later among the works it cites.
A. Cohen, X. Qiao, L. Yu, E. Way, and X. Tong, “Diverse exploration via conjugate policies for policy gradient methods,” in
2019
Later among the works it cites.
D. Ye, Z. Liu, M. Sun, B. Shi, P. Zhao, H. Wu, H. Yu, S. Yang, X. Wu, Q. Guo,
2019
Later among the works it cites.