Fetching the paper…
Reading the bibliography…
This paper makes progress towards learning Nash equilibria in two-player zero-sum Markov games from offline data.
Feature-based q-learning for two-player stochastic games
Jia, Z., Yang, L. F., and Wang, M. (2019) · 1906
Earlier work this paper cites.
Dota 2 with large scale deep reinforcement learning
Berner, C., Brockman, G., Chan, B., Cheung, V., Debiak, P., Dennison, C., Farhi, D., Fischer, Q., Hashme, S., Hesse, C., et al. (2019) · 1912
Earlier work this paper cites.
Non-cooperative games
Nash, J. (1951) · 1951
Earlier work this paper cites.
Stochastic games
Shapley, L. S. (1953) · 1953
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
Littman, M. L. (1994) · 1994
Earlier work this paper cites.
Zero-sum two-person games
Raghavan, T. (1994) · 1994
Earlier work this paper cites.
Adaptive game playing using multiplicative weights
Freund, Y. and Schapire, R. E. (1999) · 1999
Earlier work this paper cites.
Friend-or-foe q-learning in general-sum games
Littman, M. L. et al. (2001) · 2001
Earlier work this paper cites.
Value function approximation in zero-sum markov games
Lagoudakis, M. G. and Parr, R. (2002) · 2002
Earlier work this paper cites.
Nash q-learning for general-sum stochastic games
Hu, J. and Wellman, M. P. (2003) · 2003
Earlier work this paper cites.
Error bounds for approximate policy iteration
Munos, R. (2003) · 2003
Earlier work this paper cites.
Mathematical Statistics
Shao, J. (2003) · 2003
Earlier work this paper cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Levine, S., Kumar, A., Tucker, G., and Fu, J. (2020) · 2005
Earlier work this paper cites.
MOPO: Model-based offline policy optimization
Yu, T., Thomas, G., Yu, L., Ermon, S., Zou, J., Levine, S., Finn, C., and Ma, T. (2020) · 2005
Earlier work this paper cites.
Off-policy exploitability-evaluation in two-player zero-sum markov games
Abe, K. and Kaneko, Y. (2020) · 2007
Earlier work this paper cites.
On stefan banach and some of his results
Ciesielski, K. (2007) · 2007
Earlier work this paper cites.
Provably good batch reinforcement learning without great exploration
Liu, Y., Swaminathan, A., Agarwal, A., and Brunskill, E. (2020) · 2007
Earlier work this paper cites.
Performance bounds in ℓ p \ell_{p} -norm for approximate value iteration
Munos, R. (2007) · 2007
Earlier work this paper cites.
The complexity of computing a nash equilibrium
Daskalakis, C., Goldberg, P. W., and Papadimitriou, C. H. (2009) · 2009
Earlier work this paper cites.
Introduction to nonparametric estimation
Tsybakov, A. B. and Zaiats, V. (2009) · 2009
Earlier work this paper cites.
An overview of multi-agent reinforcement learning from game theoretical perspective
Yang, Y. and Wang, J. (2020) · 2011
Earlier work this paper cites.
Minimax PAC bounds on the sample complexity of reinforcement learning with a generative model
Azar, M. G., Munos, R., and Kappen, H. J. (2013) · 2013
Earlier work this paper cites.
On the complexity of approximating a nash equilibrium
Daskalakis, C. (2013) · 2013
Earlier work this paper cites.
Strategy iteration is strongly polynomial for 2-player turn-based stochastic games with a constant discount factor
Hansen, T. D., Miltersen, P. B., and Zwick, U. (2013) · 2013
Earlier work this paper cites.
Optimization, learning, and games with predictable sequences
Rakhlin, S. and Sridharan, K. (2013) · 2013
Earlier work this paper cites.
Approximate dynamic programming for two-player zero-sum markov games
Perolat, J., Scherrer, B., Piot, B., and Pietquin, O. (2015) · 2015
Cited alongside, same era.
Twenty lectures on algorithmic game theory
Roughgarden, T. (2016) · 2016
Cited alongside, same era.
Safe, multi-agent, reinforcement learning for autonomous driving
Shalev-Shwartz, S., Shammah, S., and Shashua, A. (2016) · 2016
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Azar, M. G., Osband, I., and Munos, R. (2017) · 2017
Cited alongside, same era.
Dynamic programming and optimal control (4th edition)
Bertsekas, D. P. (2017) · 2017
Cited alongside, same era.
Gaussian estimation: Sequence and wavelet models
Johnstone, I. M. (2017) · 2017
Towards tight bounds on the sample complexity of average-reward MDPs
Jin, Y. and Sidford, A. (2021) · 2021
Later among the works it cites.
Breaking the sample complexity barrier to regret-optimal model-free reinforcement learning
Li, G., Shi, L., Chen, Y., Gu, Y., and Chi, Y. (2021) · 2021
Later among the works it cites.
A sharp analysis of model-based reinforcement learning with self-play
Liu, Q., Yu, T., Bai, Y., and Jin, C. (2021) · 2021
Later among the works it cites.
Bridging offline reinforcement learning and imitation learning: A tale of pessimism
Rashidinejad, P., Zhu, B., Ma, C., Jiao, J., and Russell, S. (2021) · 2021
Later among the works it cites.
When can we learn general-sum markov games with a large number of players sample-efficiently?
Song, Z., Mei, S., and Bai, Y. (2021) · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Online reinforcement learning in stochastic games
Wei, C.-Y., Hong, Y.-T., and Lu, C.-J. (2017) · 2017
Cited alongside, same era.
High-dimensional probability: An introduction with applications in data science
Vershynin, R. (2018) · 2018
Cited alongside, same era.
Emergent tool use from multi-agent autocurricula
Baker, B., Kanitscheider, I., Markov, T., Wu, Y., Powell, G., McGrew, B., and Mordatch, I. (2019) · 2019
Cited alongside, same era.
Superhuman AI for multiplayer poker
Brown, N. and Sandholm, T. (2019) · 2019
Cited alongside, same era.
Human-level performance in 3D multiplayer games with population-based reinforcement learning
Jaderberg, M., Czarnecki, W. M., Dunning, I., Marris, L., Lever, G., Castaneda, A. G., Beattie, C., Rabinowitz, N. C., Morcos, A. S., and Ruderman, A. (2019) · 2019
Cited alongside, same era.
Grandmaster level in Starcraft II using multi-agent reinforcement learning
Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., Choi, D. H., Powell, R., Ewalds, T., and Georgiev, P. (2019) · 2019
Cited alongside, same era.
Online learning in unknown markov games
Tian, Y., Wang, Y., Yu, T., and Sra, S. (2021) · 2021
Later among the works it cites.
Pessimistic model-based offline reinforcement learning under partial coverage
Uehara, M. and Sun, W. (2021) · 2021
Later among the works it cites.
Sample-efficient reinforcement learning for linearly-parameterized mdps with a generative model
Wang, B., Yan, Y., and Fan, J. (2021) · 2021
Later among the works it cites.
Last-iterate convergence of decentralized optimistic gradient descent/ascent in infinite-horizon competitive markov games
Wei, C.-Y., Lee, C.-W., Zhang, M., and Luo, H. (2021) · 2021
Later among the works it cites.
Policy finetuning: Bridging sample-efficient offline and online reinforcement learning
Xie, T., Jiang, N., Wang, H., Xiong, C., and Bai, Y. (2021) · 2021
Later among the works it cites.
Towards instance-optimal offline reinforcement learning with pessimism
Yin, M. and Wang, Y.-X. (2021) · 2021
Later among the works it cites.
Provable benefits of actor-critic methods for offline reinforcement learning
Zanette, A., Wainwright, M. J., and Brunskill, E. (2021) · 2021
Later among the works it cites.
Provably efficient policy gradient methods for two-player zero-sum Markov games
Zhao, Y., Tian, Y., Lee, J. D., and Du, S. S. (2021) · 2021
Later among the works it cites.
Almost optimal algorithms for two-player zero-sum linear mixture markov games
Chen, Z., Zhou, D., and Gu, Q. (2022) · 2022
Closest in time.
The complexity of infinite-horizon general-sum stochastic games
Jin, Y., Muthukumar, V., and Sidford, A. (2022) · 2022
Closest in time.
Lu, M., Min, Y., Wang, Z., and Yang, Z. (2022) · 2022
Closest in time.
Provably efficient reinforcement learning in decentralized general-sum Markov games
Mao, W. and Başar, T. (2022) · 2022
Closest in time.
Optimal variance-reduced stochastic approximation in banach spaces
Mou, W., Khamaru, K., Wainwright, M. J., Bartlett, P. L., and Jordan, M. I. (2022) · 2022
Closest in time.
Pessimistic q-learning for offline reinforcement learning: Towards optimal sample complexity
Shi, L., Li, G., Wei, Y., Chen, Y., and Chi, Y. (2022) · 2022
Closest in time.
On gap-dependent bounds for offline reinforcement learning
Wang, X., Cui, Q., and Du, S. S. (2022) · 2022
Closest in time.
Xiong, W., Zhong, H., Shi, C., Shen, C., Wang, L., and Zhang, T. (2022) · 2022
Closest in time.
Provably efficient offline reinforcement learning with trajectory-wise reward
Xu, T. and Liang, Y. (2022) · 2022
Closest in time.
Pessimistic minimax value iteration: Provably efficient equilibrium learning from offline datasets
Zhong, H., Xiong, W., Tan, J., Wang, L., Zhang, T., Wang, Z., and Yang, Z. (2022) · 2022
Closest in time.
The complexity of markov equilibrium in stochastic games
Daskalakis, C., Golowich, N., and Zhang, K. (2023) · 2023
Closest in time.
The efficacy of pessimism in asynchronous q-learning
Yan, Y., Li, G., Chen, Y., and Fan, J. (2023) · 2023
Closest in time.