Fetching the paper…
Reading the bibliography…
We present the first study on provably efficient randomized exploration in cooperative multi-agent reinforcement learning (MARL).
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
W. R. Thompson · 1933
Earlier work this paper cites.
Handbook of mathematical functions with formulas, graphs, and mathematical tables , volume 55
M. Abramowitz and I. A. Stegun · 1968
Earlier work this paper cites.
Planning, learning and coordination in multiagent decision processes
C. Boutilier · 1996
Earlier work this paper cites.
Exponential convergence of langevin distributions and their discrete approximations
G. O. Roberts and R. L. Tweedie · 1996
Earlier work this paper cites.
A bayesian framework for reinforcement learning
M. Strens · 2000
Earlier work this paper cites.
The complexity of decentralized control of markov decision processes
D. S. Bernstein, R. Givan, N. Immerman, and S. Zilberstein · 2002
Earlier work this paper cites.
Parallel reinforcement learning
R. M. Kretchmar · 2002
Earlier work this paper cites.
Coordination in multiagent reinforcement learning: A bayesian approach
G. Chalkiadakis and C. Boutilier · 2003
Earlier work this paper cites.
Opportunities for multiagent systems and multiagent reinforcement learning in traffic control
A. L. Bazzan · 2009
Earlier work this paper cites.
Transfer learning for reinforcement learning domains: A survey
M. E. Taylor and P. Stone · 2009
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
L. Li, W. Chu, J. Langford, and R. E. Schapire · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Y. Abbasi-Yadkori, D. Pál, and C. Szepesvári · 2011
Earlier work this paper cites.
An empirical evaluation of thompson sampling
O. Chapelle and L. Li · 2011
Earlier work this paper cites.
Contextual bandits with linear payoff functions
W. Chu, L. Li, L. Reyzin, and R. Schapire · 2011
Earlier work this paper cites.
Evolving intelligent mario controller by reinforcement learning
J.-J. Tsay, C.-C. Chen, and J.-J. Hsu · 2011
Earlier work this paper cites.
Matrix analysis
R. A. Horn and C. R. Johnson · 2012
Earlier work this paper cites.
Distributed exploration in multi-armed bandits
E. Hillel, Z. S. Karnin, T. Koren, R. Lempel, and O. Somekh · 2013
Earlier work this paper cites.
(more) efficient reinforcement learning via posterior sampling
I. Osband, D. Russo, and B. Van Roy · 2013
Earlier work this paper cites.
Analysis and geometry of Markov diffusion operators , volume 103
D. Bakry, I. Gentil, M. Ledoux, et al · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, et al · 2015
Earlier work this paper cites.
Deep reinforcement learning with double qlearning
H. V. Hasselt, A. Guez, and D. Silver · 2016
Earlier work this paper cites.
On distributed cooperative decision-making in multiarmed bandits
P. Landgren, V. Srivastava, and N. E. Leonard · 2016
Earlier work this paper cites.
Linear thompson sampling revisited
M. Abeille and A. Lazaric · 2017
Earlier work this paper cites.
Optimistic posterior sampling for reinforcement learning: worst-case regret bounds
S. Agrawal and R. Jia · 2017
Earlier work this paper cites.
Theoretical guarantees for approximate sampling from smooth and log-concave densities
A. S. Dalalyan · 2017
Earlier work this paper cites.
Provably optimal algorithms for generalized linear contextual bandits
L. Li, Y. Lu, and D. Zhou · 2017
Earlier work this paper cites.
Why is posterior sampling better than optimism for reinforcement learning?
I. Osband and B. Van Roy · 2017
Earlier work this paper cites.
Coordinated exploration in concurrent reinforcement learning
M. Dimakopoulou and B. V. Roy · 2018
Earlier work this paper cites.
Scalable coordinated exploration in concurrent reinforcement learning
M. Dimakopoulou, I. Osband, and B. V. Roy · 2018
Earlier work this paper cites.
Noisy networks for exploration
M. Fortunato, M. G. Azar, B. Piot, et al · 2018
Cited alongside, same era.
Randomized prior functions for deep reinforcement learning
I. Osband, J. Aslanides, and A. Cassirer · 2018
Cited alongside, same era.
Deep bayesian bandits showdown: An empirical comparison of bayesian deep networks for thompson sampling
C. Riquelme, G. Tucker, and J. Snoek · 2018
Cited alongside, same era.
High-dimensional probability: An introduction with applications in data science , volume 47
R. Vershynin · 2018
Cited alongside, same era.
Global convergence of langevin dynamics based algorithms for nonconvex optimization
P. Xu, J. Chen, D. Zou, and Q. Gu · 2018
Cited alongside, same era.
Fully decentralized multi-agent reinforcement learning with networked agents
K. Zhang, Z. Yang, H. Liu, T. Zhang, and T. Başar · 2018
Faster convergence of stochastic gradient langevin dynamics for non-log-concave sampling
D. Zou, P. Xu, and Q. Gu · 2021
Later among the works it cites.
Society of agents: Regret bounds of concurrent thompson sampling
Y. Chen, P. Dong, Q. Bai, M. Dimakopoulou, W. Xu, and Z. Zhou · 2022
Later among the works it cites.
Cascaded gaps: Towards logarithmic regret for risk-sensitive reinforcement learning
Y. Fei and R. Xu · 2022
Later among the works it cites.
Hyperdqn: A randomized exploration method for deep reinforcement learning
Z. Li, Y. Li, Y. Zhang, T. Zhang, and Z.-Q. Luo · 2022
Later among the works it cites.
Provably efficient multi-agent reinforcement learning with fully decentralized communication
J. Lidard, U. Madhushani, and N. E. Leonard · 2022
Later among the works it cites.
A self-play posterior sampling algorithm for zero-sum markov games
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Perturbed-history exploration in stochastic multi-armed bandits, 2019
B. Kveton, C. Szepesvari, M. Ghavamzadeh, and C. Boutilier · 2019
Cited alongside, same era.
Lifelong federated reinforcement learning: a learning architecture for navigation in cloud robotic systems
B. Liu, L. Wang, and M. Liu · 2019
Cited alongside, same era.
Worst-case regret bounds for exploration via randomized value functions
D. Russo · 2019
Cited alongside, same era.
On multi-agent learning in team sports games
Y. Zhao, I. Borovikov, J. Rupert, C. Somers, and A. Beirami · 2019
Cited alongside, same era.
Posterior sampling for multi-agent reinforcement learning: solving extensive games with imperfect information
Y. Zhou, J. Li, and J. Zhu · 2019
Cited alongside, same era.
Provably efficient exploration in policy optimization
Q. Cai, Z. Yang, C. Jin, and Z. Wang · 2020
Cited alongside, same era.
W. Xiong, H. Zhong, C. Shi, C. Shen, and T. Zhang · 2022
Later among the works it cites.
Langevin monte carlo for contextual bandits
P. Xu, H. Zheng, E. V. Mazumdar, K. Azizzadenesheli, and A. Anandkumar · 2022
Later among the works it cites.
The surprising effectiveness of PPO in cooperative multi-agent games
C. Yu, A. Velu, E. Vinitsky, J. Gao, Y. Wang, A. Bayen, and Y. Wu · 2022
Later among the works it cites.
Guts: Generalized uncertainty-aware thompson sampling for multi-agent active search
N. A. Bakshi, T. Gupta, R. Ghods, and J. Schneider · 2023
Later among the works it cites.
Exploration in deep reinforcement learning: From single-agent to multiagent domain
J. Hao, T. Yang, H. Tang, C. Bai, J. Liu, Z. Meng, P. Liu, and Z. Wang · 2023
Later among the works it cites.
Tight regret and complexity bounds for thompson sampling via langevin monte carlo
T. Huix, M. Zhang, and A. Durmus · 2023
Later among the works it cites.
Thompson sampling with less exploration is fast and optimal
T. Jin, X. Yang, X. Xiao, and P. Xu · 2023
Later among the works it cites.
Langevin thompson sampling with logarithmic communication: Bandits and reinforcement learning
A. Karbasi, N. L. Kuang, Y. Ma, and S. Mitra · 2023
Later among the works it cites.
Posterior sampling with delayed feedback for reinforcement learning with linear function approximation
N. Kuang, M. Yin, et al · 2023
Later among the works it cites.
Cooperative multi-agent reinforcement learning: asynchronous communication and linear function approximation
Y. Min, J. He, T. Wang, and Q. Gu · 2023
Later among the works it cites.
Towards a complete analysis of langevin monte carlo: Beyond poincaré inequality, 2023
A. Mousavi-Hosseini, T. Farghly, Y. He, K. Balasubramanian, and M. A. Erdogdu · 2023
Later among the works it cites.
On sample-efficient offline reinforcement learning: Data diversity, posterior sampling and beyond
T. Nguyen-Tang and R. Arora · 2023
Later among the works it cites.
Posterior sampling for competitive rl: Function approximation and partial observation
S. Qiu, Z. Dai, H. Zhong, Z. Wang, Z. Yang, and T. Zhang · 2023
Later among the works it cites.
Why one strategy does not fit all: a systematic review on exploration–exploitation in different organizational archetypes
C. Rojas-Córdova, A. J. Williamson, J. A. Pertuze, and G. Calvo · 2023
Later among the works it cites.
Sustaingym: A benchmark suite of reinforcement learning for sustainability applications
C. Yeh, V. Li, R. Datta, J. Arroyo, N. Christianson, C. Zhang, Y. Chen, M. Hosseini, A. Golmohammadi, Y. Shi, Y. Yue, and A. Wierman · 2023
Later among the works it cites.
Asynchronous multi-agent reinforcement learning for efficient real-time multi-robot cooperative exploration
C. Yu, X. Yang, J. Gao, J. Chen, Y. Li, J. Liu, Y. Xiang, R. Huang, H. Yang, Y. Wu, and Y. Wang · 2023
Later among the works it cites.
Global convergence of localized policy iteration in networked multi-agent reinforcement learning
Y. Zhang, G. Qu, P. Xu, Y. Lin, Z. Chen, and A. Wierman · 2023
Later among the works it cites.
A theoretical analysis of optimistic proximal policy optimization in linear markov decision processes
H. Zhong and T. Zhang · 2023
Later among the works it cites.
A bayesian learning algorithm for unknown zero-sum stochastic games with an arbitrary opponent
M. Jafarnia-Jahromi, R. Jain, and A. Nayyar · 2024
Closest in time.
Finite-time frequentist regret bounds of multi-agent thompson sampling on sparse hypergraphs
T. Jin, H.-L. Hsu, W. Chang, and P. Xu · 2024
Closest in time.
Cell-free xl-mimo meets multi-agent reinforcement learning: Architectures, challenges, and future directions
Z. Liu, J. Zhang, Z. Liu, H. Du, Z. Wang, D. Niyato, M. Guizani, and B. Ai · 2024
Closest in time.
Randomized exploration in generalized linear bandits
B. Kveton, M. Zaheer, C. Szepesvari, L. Li, M. Ghavamzadeh, and C. Boutilier · 2076
Closest in time.
Randomized exploration in generalized linear bandits
B. Kveton, M. Zaheer, C. Szepesvari, L. Li, M. Ghavamzadeh, and C. Boutilier · 2076
Closest in time.