Fetching the paper…
Reading the bibliography…
Multi-agent reinforcement learning (MARL) lies at the heart of a plethora of applications involving the interaction of a group of agents in a shared unknown environment.
Non-cooperative games
J. F. Nash · 1950
Earlier work this paper cites.
Stochastic games
L. S. Shapley · 1953
Earlier work this paper cites.
A new family of optimal adaptive controllers for markov chains
P. Kumar and A. Becker · 1982
Earlier work this paper cites.
Correlated equilibrium as an expression of bayesian rationality
R. J. Aumann · 1987
Earlier work this paper cites.
Adaptive treatment allocation and the multi-armed bandit problem
T. L. Lai · 1987
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
M. L. Littman · 1994
Earlier work this paper cites.
Quantal response equilibria for normal form games
R. D. McKelvey and T. R. Palfrey · 1995
Earlier work this paper cites.
Notes on equilibria in symmetric games
S.-F. Cheng, D. M. Reeves, Y. Vorobeychik, and M. P. Wellman · 2004
Earlier work this paper cites.
Policy optimization for Markov games: Unified framework and faster convergence
R. Zhang, Q. Liu, H. Wang, C. Xiong, N. Li, and Y. Bai · 2004
Earlier work this paper cites.
A comprehensive survey of multiagent reinforcement learning
L. Busoniu, R. Babuska, and B. De Schutter · 2008
Earlier work this paper cites.
Stochastic linear optimization under bandit feedback
V. Dani, T. P. Hayes, and S. M. Kakade · 2008
Earlier work this paper cites.
The complexity of computing a Nash equilibrium
C. Daskalakis, P. W. Goldberg, and C. H. Papadimitriou · 2009
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Y. Abbasi-Yadkori, D. Pál, and C. Szepesvári · 2011
Earlier work this paper cites.
Finite-time analysis of the multi-armed bandit problem with known trend
D. Bouneffouf · 2016
Earlier work this paper cites.
Last-iterate convergence: Zero-sum games and constrained min-max optimization
C. Daskalakis and I. Panageas · 2018
Earlier work this paper cites.
Is q-learning provably efficient?
C. Jin, Z. Allen-Zhu, S. Bubeck, and M. I. Jordan · 2018
Earlier work this paper cites.
Cycles in adversarial regularized learning
P. Mertikopoulos, C. Papadimitriou, and G. Piliouras · 2018
Earlier work this paper cites.
A tutorial on thompson sampling
D. J. Russo, B. Van Roy, A. Kazerouni, I. Osband, Z. Wen, et al · 2018
Earlier work this paper cites.
Optimism in reinforcement learning with generalized linear function approximation
Y. Wang, R. Wang, S. S. Du, and A. Krishnamurthy · 2019
Earlier work this paper cites.
Sample-optimal parametric q-learning using linearly additive features
L. Yang and M. Wang · 2019
Earlier work this paper cites.
Model-based reinforcement learning with value-targeted regression
A. Ayoub, Z. Jia, C. Szepesvari, M. Wang, and L. Yang · 2020
Cited alongside, same era.
Provably efficient exploration in policy optimization
Q. Cai, Z. Yang, C. Jin, and Z. Wang · 2020
Cited alongside, same era.
Model-based reinforcement learning with value-targeted regression
Z. Jia, L. Yang, C. Szepesvari, and M. Wang · 2020
Cited alongside, same era.
Provably efficient reinforcement learning with linear function approximation
C. Jin, Z. Yang, Z. Wang, and M. I. Jordan · 2020
Cited alongside, same era.
Exploration through reward biasing: Reward-biased maximum likelihood estimation for stochastic multi-armed bandits
X. Liu, P.-C. Hsieh, Y. H. Hung, A. Bhattacharya, and P. Kumar · 2020
Cited alongside, same era.
On the global convergence rates of softmax policy gradient methods
Towards general function approximation in zero-sum Markov games
B. Huang, J. D. Lee, Z. Wang, and Z. Yang · 2022
Later among the works it cites.
Minimax-optimal multi-agent rl in markov games with a generative model
G. Li, Y. Chi, Y. Wei, and Y. Chen · 2022
Later among the works it cites.
Representation learning for general-sum low-rank markov games
C. Ni, Y. Song, X. Zhang, C. Jin, and M. Wang · 2022
Later among the works it cites.
Efficient model-based multi-agent reinforcement learning via optimistic equilibrium computation
P. G. Sessa, M. Kamgarpour, and A. Krause · 2022
Later among the works it cites.
S. Sokota, R. D’Orazio, J. Z. Kolter, N. Loizou, M. Lanctot, I. Mitliagkas, N. Brown, and C. Kroer · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Mei, C. Xiao, C. Szepesvari, and D. Schuurmans · 2020
Cited alongside, same era.
Sample complexity of reinforcement learning using linearly combined model ensembles
A. Modi, N. Jiang, A. Tewari, and S. Singh · 2020
Cited alongside, same era.
Linear last-iterate convergence in constrained saddle-point optimization
C.-Y. Wei, C.-W. Lee, M. Zhang, and H. Luo · 2020
Cited alongside, same era.
Learning zero-sum simultaneous-move markov games using function approximation and correlated equilibrium
Q. Xie, Y. Chen, Z. Wang, and Z. Yang · 2020
Cited alongside, same era.
Sample-efficient learning of stackelberg equilibria in general-sum games
Y. Bai, C. Jin, H. Wang, and C. Xiong · 2021
Cited alongside, same era.
Fast policy extragradient methods for competitive games with entropy regularization
S. Cen, Y. Wei, and Y. Chi · 2021
Cited alongside, same era.
Bilinear classes: A structural framework for provable generalization in rl
S. Du, S. Kakade, J. Lee, S. Lovett, G. Mahajan, W. Sun, and R. Wang · 2021
Cited alongside, same era.
Vo q q l: Towards optimal regret in model-free RL with nonlinear function approximation
A. Agarwal, Y. Jin, and T. Zhang · 2023
Later among the works it cites.
Faster last-iterate convergence of policy optimization in zero-sum markov games
S. Cen, Y. Chi, S. S. Du, and L. Xiao · 2023
Later among the works it cites.
Breaking the curse of multiagents in a large state space: Rl in markov games with independent linear function approximation
Q. Cui, K. Zhang, and S. Du · 2023
Later among the works it cites.
Regret minimization and convergence to equilibria in general-sum Markov games
L. Erez, T. Lancewicki, U. Sherman, T. Koren, and Y. Mansour · 2023
Later among the works it cites.
A survey of uncertainty in deep neural networks
J. Gawlikowski, C. R. N. Tassi, M. Ali, J. Lee, M. Humt, J. Feng, A. Kruspe, R. Triebel, P. Jung, R. Roscher, et al · 2023
Later among the works it cites.
Provably efficient reinforcement learning in decentralized general-sum Markov games
W. Mao and T. Başar · 2023
Later among the works it cites.
Nash learning from human feedback
R. Munos, M. Valko, D. Calandriello, M. G. Azar, M. Rowland, Z. D. Guo, Y. Tang, M. Geist, T. Mesnard, A. Michi, et al · 2023
Later among the works it cites.
Breaking the curse of multiagency: Provably efficient decentralized multi-agent rl with function approximation
Y. Wang, Q. Liu, Y. Bai, and C. Jin · 2023
Later among the works it cites.
Linear convergence of natural policy gradient methods with log-linear policies
R. Yuan, S. S. Du, R. M. Gower, A. Lazaric, and L. Xiao · 2023
Later among the works it cites.
Policy mirror descent for regularized reinforcement learning: A generalized framework with linear convergence
W. Zhan, S. Cen, B. Huang, Y. Chen, J. D. Lee, and Y. Chi · 2023
Later among the works it cites.
Near-optimal policy optimization for correlated equilibrium in general-sum Markov games
Y. Cai, H. Luo, C.-Y. Wei, and W. Zheng · 2024
Later among the works it cites.
Value-incentivized preference optimization: A unified approach to online and offline rlhf
S. Cen, J. Mei, K. Goshvadi, H. Dai, T. Yang, S. Yang, D. Schuurmans, Y. Chi, and B. Dai · 2024
Later among the works it cites.
Refined sample complexity for markov games with independent linear function approximation
Y. Dai, Q. Cui, and S. S. Du · 2024
Later among the works it cites.
Maximize to explore: One objective function fusing estimation, planning, and exploration
Z. Liu, M. Lu, W. Xiong, H. Zhong, H. Hu, S. Zhang, S. Zheng, Z. Yang, and Z. Wang · 2024
Later among the works it cites.
A minimaximalist approach to reinforcement learning from human feedback
G. Swamy, C. Dann, R. Kidambi, Z. S. Wu, and A. Agarwal · 2024
Later among the works it cites.