Fetching the paper…
Reading the bibliography…
In cooperative stochastic games multiple agents work towards learning joint optimal actions in an unknown environment to achieve a common goal.
R. B. Diddigi, D. S. K. Reddy, P. KJ, and S. Bhatnagar, “Actor-critic algorithms for constrained multi-agent reinforcement learning,” in Proceedings of the 18th International Conference on Autonomous Agents and MultiAgent Systems , 2019, pp. 1931–1933
1933
Earlier work this paper cites.
M. L. Littman, “Markov games as a framework for multi-agent reinforcement learning,” in Machine Learning Proceedings 1994 . Elsevier, 1994, pp. 157–163
1994
Earlier work this paper cites.
V. S. Borkar, “Stochastic approximation with two time scales,” Systems & Control Letters , vol. 29, no. 5, pp. 291–294, 1997
1997
Earlier work this paper cites.
R. S. Sutton and A. G. Barto, Introduction to reinforcement learning . MIT press Cambridge, 1998, vol. 135
1998
Earlier work this paper cites.
E. Altman, Constrained Markov decision processes . CRC Press, 1999, vol. 7
1999
Earlier work this paper cites.
M. Lauer and M. Riedmiller, “An algorithm for distributed reinforcement learning in cooperative multi-agent systems,” in In Proceedings of the Seventeenth International Conference on Machine Learning . Citeseer, 2000
2000
Earlier work this paper cites.
M. L. Littman, “Value-function reinforcement learning in markov games,” Cognitive Systems Research , vol. 2, no. 1, pp. 55–66, 2001
2001
Earlier work this paper cites.
V. S. Borkar, “An actor-critic algorithm for constrained markov decision processes,” Systems & control letters , vol. 54, no. 3, pp. 207–213, 2005
2005
Earlier work this paper cites.
L. Busoniu, R. Babuska, and B. De Schutter, “A comprehensive survey of multiagent reinforcement learning,” IEEE Transactions on Systems, Man, And Cybernetics-Part C: Applications and Reviews, 38 (2), 2008 , 2008
2008
Cited alongside, same era.
S. Bhatnagar, “An actor–critic algorithm with function approximation for discounted cost constrained markov decision processes,” Systems & Control Letters , vol. 59, no. 12, pp. 760–766, 2010
2010
Cited alongside, same era.
K. Lakshmanan and S. Bhatnagar, “A novel q-learning algorithm with function approximation for constrained markov decision processes,” in Communication, Control, and Computing (Allerton), 2012 50th Annual Allerton Conference on . IEEE, 2012, pp. 400–405
2012
Cited alongside, same era.
S. Bhatnagar and K. Lakshmanan, “An online actor–critic algorithm with function approximation for constrained markov decision processes,” Journal of Optimization Theory and Applications , vol. 153, no. 3, pp. 688–708, 2012
2012
R. Lowe, Y. Wu, A. Tamar, J. Harb, P. Abbeel, and I. Mordatch, “Multi-agent actor-critic for mixed cooperative-competitive environments,” in Advances in Neural Information Processing Systems , 2017, pp. 6379–6390
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski et al. , “Human-level control through deep reinforcement learning,” Nature , vol. 518, no. 7540, p. 529, 2015
2015
Cited alongside, same era.
S. Levine, C. Finn, T. Darrell, and P. Abbeel, “End-to-end training of deep visuomotor policies,” The Journal of Machine Learning Research , vol. 17, no. 1, pp. 1334–1373, 2016
2016
Cited alongside, same era.
J. Foerster, I. A. Assael, N. de Freitas, and S. Whiteson, “Learning to communicate with deep multi-agent reinforcement learning,” in Advances in Neural Information Processing Systems , 2016, pp. 2137–2145
2016
Cited alongside, same era.
2017
Cited alongside, same era.
J. K. Gupta, M. Egorov, and M. Kochenderfer, “Cooperative multi-agent control using deep reinforcement learning,” in International Conference on Autonomous Agents and Multiagent Systems . Springer, 2017, pp. 66–83
2017
Later among the works it cites.
2017
Later among the works it cites.
A. Tampuu, T. Matiisen, D. Kodelja, I. Kuzovkin, K. Korjus, J. Aru, J. Aru, and R. Vicente, “Multiagent cooperation and competition with deep reinforcement learning,” PloS one , vol. 12, no. 4, p. e0172395, 2017
2017
Later among the works it cites.
J. Foerster, R. Y. Chen, M. Al-Shedivat, S. Whiteson, P. Abbeel, and I. Mordatch, “Learning with opponent-learning awareness,” in Proceedings of the 17th International Conference on Autonomous Agents and MultiAgent Systems . International Foundation for Autonomous Agents and Multiagent Systems, 2018, pp. 122–130
2018
Later among the works it cites.