Fetching the paper…
Reading the bibliography…
We study decentralized policy learning in Markov games where we control a single agent to play with nonstationary and possibly adversarial opponents.
Dota 2 with large scale deep reinforcement learning
Berner, C., Brockman, G., Chan, B., Cheung, V., Dębiak, P., Dennison, C., Farhi, D., Fischer, Q., Hashme, S., Hesse, C., et al. (2019) · 1912
Earlier work this paper cites.
Optimism in reinforcement learning with generalized linear function approximation
Wang, Y., Wang, R., Du, S. S., and Krishnamurthy, A. (2019) · 1912
Earlier work this paper cites.
Game theory
Fudenberg, D. and Tirole, J. (1991) · 1991
Earlier work this paper cites.
Markov decision processes: Discrete stochastic dynamic programming
Puterman, M. L. (1994) · 1994
Earlier work this paper cites.
A decision-theoretic generalization of on-line learning and an application to boosting
Freund, Y. and Schapire, R. E. (1997) · 1997
Earlier work this paper cites.
Actor-critic algorithms
Konda, V. and Tsitsiklis, J. (1999) · 1999
Earlier work this paper cites.
R-max-a general polynomial time algorithm for near-optimal reinforcement learning
Brafman, R. I. and Tennenholtz, M. (2002) · 2002
Earlier work this paper cites.
Exploration-exploitation in constrained MDPs
Efroni, Y., Mannor, S., and Pirotta, M. (2020) · 2003
Earlier work this paper cites.
Learning, regret minimization, and equilibria
Blum, A. and Monsour, Y. (2007) · 2007
Earlier work this paper cites.
The theory and practice of online learning
Anderson, T. (2008) · 2008
Earlier work this paper cites.
Gaussian process optimization in the bandit setting: No regret and experimental design
Srinivas, N., Krause, A., Kakade, S. M., and Seeger, M. (2009) · 2009
Earlier work this paper cites.
Near-optimal no-regret algorithms for zero-sum games
Daskalakis, C., Deckelbaum, A., and Kim, A. (2011) · 2011
Earlier work this paper cites.
Swarm robotics: a review from the swarm engineering perspective
Brambilla, M., Ferrante, E., Birattari, M., and Dorigo, M. (2013) · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M. (2013) · 2013
Earlier work this paper cites.
Eluder dimension and the sample complexity of optimistic exploration
Russo, D. and Van Roy, B. (2013) · 2013
Earlier work this paper cites.
Deep learning
LeCun, Y., Bengio, Y., and Hinton, G. (2015) · 2015
Earlier work this paper cites.
Introduction to online convex optimization
Hazan, E. et al. (2016) · 2016
Earlier work this paper cites.
Safe, multi-agent, reinforcement learning for autonomous driving
Shalev-Shwartz, S., Shammah, S., and Shashua, A. (2016) · 2016
Cited alongside, same era.
Mastering the game of Go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al. (2016) · 2016
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Azar, M. G., Osband, I., and Munos, R. (2017) · 2017
Cited alongside, same era.
Cooperative multi-agent control using deep reinforcement learning
Gupta, J. K., Egorov, M., and Kochenderfer, M. (2017) · 2017
Cited alongside, same era.
Contextual decision processes with low bellman rank are pac-learnable
Jiang, N., Krishnamurthy, A., Agarwal, A., Langford, J., and Schapire, R. E. (2017) · 2017
Cited alongside, same era.
Bandit algorithms
Lattimore, T. and Szepesvári, C. (2020) · 2020
Later among the works it cites.
Learning zero-sum simultaneous-move markov games using function approximation and correlated equilibrium
Xie, Q., Chen, Y., Wang, Z., and Yang, Z. (2020) · 2020
Later among the works it cites.
Multi-agent reinforcement learning: A review of challenges and applications
Canese, L., Cardarilli, G. C., Di Nunzio, L., Fazzolari, R., Giardino, D., Re, M., and Spanò, S. (2021) · 2021
Later among the works it cites.
Provably efficient safe exploration via primal-dual policy optimization
Ding, D., Wei, X., Yang, Z., Wang, Z., and Jovanovic, M. (2021) · 2021
Later among the works it cites.
Bilinear classes: A structural framework for provable generalization in rl
Du, S. S., Kakade, S. M., Lee, J. D., Lovett, S., Mahajan, G., Sun, W., and Wang, R. (2021) · 2021
Later among the works it cites.
Towards general function approximation in zero-sum markov games
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Online reinforcement learning in stochastic games
Wei, C.-Y., Hong, Y.-T., and Lu, C.-J. (2017) · 2017
Cited alongside, same era.
Is multiagent deep reinforcement learning the answer or the question? a brief survey
Hernandez-Leal, P., Kartal, B., and Taylor, M. E. (2018) · 2018
Cited alongside, same era.
Qmix: Monotonic value function factorisation for deep multi-agent reinforcement learning
Rashid, T., Samvelyan, M., Schroeder, C., Farquhar, G., Foerster, J., and Whiteson, S. (2018) · 2018
Cited alongside, same era.
A survey and critique of multiagent deep reinforcement learning
Hernandez-Leal, P., Kartal, B., and Taylor, M. E. (2019) · 2019
Cited alongside, same era.
Reinforcement learning with convex constraints
Miryoosefi, S., Brantley, K., Daume III, H., Dudik, M., and Schapire, R. E. (2019) · 2019
Cited alongside, same era.
Constrained reinforcement learning has zero duality gap
Paternain, S., Chamon, L., Calvo-Fullana, M., and Ribeiro, A. (2019) · 2019
Cited alongside, same era.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., Choi, D. H., Powell, R., Ewalds, T., Georgiev, P., et al. (2019) · 2019
Cited alongside, same era.
Huang, B., Lee, J. D., Wang, Z., and Yang, Z. (2021) · 2021
Later among the works it cites.
Global convergence of multi-agent policy gradient in markov potential games
Leonardos, S., Overman, W., Panageas, I., and Piliouras, G. (2021) · 2021
Later among the works it cites.
Learning policies with zero or bounded constraint violation for constrained mdps
Liu, T., Zhou, R., Kalathil, D., Kumar, P., and Tian, C. (2021) · 2021
Later among the works it cites.
On improving model-free algorithms for decentralized multi-agent reinforcement learning
Mao, W., Yang, L. F., Zhang, K., and Başar, T. (2021) · 2021
Later among the works it cites.
A simple reward-free approach to constrained reinforcement learning
Miryoosefi, S. and Jin, C. (2021) · 2021
Later among the works it cites.
Decentralized q-learning in zero-sum markov games
Sayin, M., Zhang, K., Leslie, D., Basar, T., and Ozdaglar, A. (2021) · 2021
Later among the works it cites.
Online learning in unknown markov games
Tian, Y., Wang, Y., Yu, T., and Sra, S. (2021) · 2021
Later among the works it cites.
Bellman-consistent pessimism for offline reinforcement learning
Xie, T., Cheng, C.-A., Jiang, N., Mineiro, P., and Agarwal, A. (2021) · 2021
Later among the works it cites.
Provably efficient algorithms for multi-objective competitive rl
Yu, T., Tian, Y., Zhang, J., and Sra, S. (2021) · 2021
Later among the works it cites.
Ding, D., Wei, C.-Y., Zhang, K., and Jovanović, M. R. (2022) · 2022
Closest in time.
Multi-agent deep reinforcement learning: a survey
Gronauer, S. and Diepold, K. (2022) · 2022
Closest in time.
Learning markov games with adversarial opponents: Efficient algorithms and fundamental limits
Liu, Q., Wang, Y., and Jin, C. (2022) · 2022
Closest in time.