Fetching the paper…
Reading the bibliography…
Tackling overestimation in $Q$-learning is an important problem that has been extensively studied in single-agent reinforcement learning, but has received comparatively little attention in the multi-agent setting.
Q-learning
C. J. Watkins and P. Dayan · 1992
Earlier work this paper cites.
Issues in using function approximation for reinforcement learning
S. Thrun and A. Schwartz · 1993
Earlier work this paper cites.
An analysis of temporal-difference learning with function approximation
J. N. Tsitsiklis and B. Van Roy · 1997
Earlier work this paper cites.
Optimal and approximate q-value functions for decentralized pomdps
F. A. Oliehoek, M. T. Spaan, and N. Vlassis · 2008
Earlier work this paper cites.
Double q-learning
H. Hasselt · 2010
Earlier work this paper cites.
An overview of recent progress in the study of distributed multi-agent coordination
Y. Cao, W. Yu, W. Ren, and G. Chen · 2012
Earlier work this paper cites.
Coordinated multi-robot exploration under communication constraints using decentralized markov decision processes
L. Matignon, L. Jeanpierre, and A.-I. Mouaddib · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2013
Earlier work this paper cites.
H. Van Hasselt · 2013
Earlier work this paper cites.
Continuous control with deep reinforcement learning
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2015
Earlier work this paper cites.
Unifying count-based exploration and intrinsic motivation
M. Bellemare, S. Srinivasan, G. Ostrovski, T. Schaul, D. Saxton, and R. Munos · 2016
Earlier work this paper cites.
Learning to play in a day: Faster deep reinforcement learning by optimality tightening
F. S. He, Y. Liu, A. G. Schwing, and J. Peng · 2016
Earlier work this paper cites.
Multi-agent reinforcement learning as a rehearsal for decentralized planning
L. Kraemer and B. Banerjee · 2016
Earlier work this paper cites.
A concise introduction to decentralized POMDPs
F. A. Oliehoek, C. Amato, et al · 2016
Earlier work this paper cites.
Deep reinforcement learning with double q-learning
H. Van Hasselt, A. Guez, and D. Silver · 2016
Earlier work this paper cites.
Averaged-dqn: Variance reduction and stabilization for deep reinforcement learning
O. Anschel, N. Baram, and N. Shimkin · 2017
Cited alongside, same era.
Hypernetworks
D. Ha, A. Dai, and Q. V. Le · 2017
Cited alongside, same era.
Multi-agent actor-critic for mixed cooperative-competitive environments
R. Lowe, Y. I. Wu, A. Tamar, J. Harb, O. P. Abbeel, and I. Mordatch · 2017
Cited alongside, same era.
Count-based exploration with neural density models
G. Ostrovski, M. G. Bellemare, A. Oord, and R. Munos · 2017
Cited alongside, same era.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Cited alongside, same era.
Counterfactual multi-agent policy gradients
J. Foerster, G. Farquhar, T. Afouras, N. Nardelli, and S. Whiteson · 2018
The starcraft multi-agent challenge
M. Samvelyan, T. Rashid, C. S. de Witt, G. Farquhar, N. Nardelli, T. G. Rudner, C.-M. Hung, P. H. Torr, J. Foerster, and S. Whiteson · 2019
Later among the works it cites.
Qtran: Learning to factorize with transformation for cooperative multi-agent reinforcement learning
K. Son, D. Kim, W. J. Kang, D. E. Hostallero, and Y. Yi · 2019
Later among the works it cites.
Revisiting the softmax bellman operator: New benefits and new perspective
Z. Song, R. Parr, and L. Carin · 2019
Later among the works it cites.
Deep multi-agent reinforcement learning for decentralized continuous cooperative control
C. S. de Witt, B. Peng, P.-A. Kamienny, P. Torr, W. Böhmer, and S. Whiteson · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Addressing function approximation error in actor-critic methods
S. Fujimoto, H. Hoof, and D. Meger · 2018
Cited alongside, same era.
Revisiting the arcade learning environment: Evaluation protocols and open problems for general agents
M. C. Machado, M. G. Bellemare, E. Talvitie, J. Veness, M. Hausknecht, and M. Bowling · 2018
Cited alongside, same era.
Self-imitation learning
J. Oh, Y. Guo, S. Singh, and H. Lee · 2018
Cited alongside, same era.
Qmix: Monotonic value function factorisation for deep multi-agent reinforcement learning
T. Rashid, M. Samvelyan, C. Schroeder, G. Farquhar, J. Foerster, and S. Whiteson · 2018
Cited alongside, same era.
Value-decomposition networks for cooperative multi-agent learning based on team reward
P. Sunehag, G. Lever, A. Gruslys, W. M. Czarnecki, V. F. Zambaldi, M. Jaderberg, M. Lanctot, N. Sonnerat, J. Z. Leibo, K. Tuyls, et al · 2018
Cited alongside, same era.
Reducing overestimation bias in multi-agent domains using double centralized critics
J. Ackermann, V. Gabler, T. Osa, and M. Sugiyama · 2019
Cited alongside, same era.
Y. Gan, Z. Zhang, and X. Tan · 2020
Later among the works it cites.
Uneven: Universal value exploration for multi-agent reinforcement learning
T. Gupta, A. Mahajan, B. Peng, W. Böhmer, and S. Whiteson · 2020
Later among the works it cites.
Ai-qmix: Attention and imagination for dynamic multi-agent reinforcement learning
S. Iqbal, C. A. S. de Witt, B. Peng, W. Böhmer, S. Whiteson, and F. Sha · 2020
Later among the works it cites.
Maxmin q-learning: Controlling the estimation bias of q-learning
Q. Lan, Y. Pan, A. Fyshe, and M. White · 2020
Later among the works it cites.
Softmax deep double deterministic policy gradients
L. Pan, Q. Cai, and L. Huang · 2020
Later among the works it cites.
Reinforcement learning with dynamic boltzmann softmax updates
L. Pan, Q. Cai, Q. Meng, W. Chen, and L. Huang · 2020
Later among the works it cites.
Weighted qmix: Expanding monotonic value function factorisation for deep multi-agent reinforcement learning
T. Rashid, G. Farquhar, B. Peng, and S. Whiteson · 2020
Later among the works it cites.
Monotonic value function factorisation for deep multi-agent reinforcement learning
T. Rashid, M. Samvelyan, C. S. De Witt, G. Farquhar, J. Foerster, and S. Whiteson · 2020
Later among the works it cites.
Qatten: A general framework for cooperative multiagent reinforcement learning
Y. Yang, J. Hao, B. Liao, K. Shao, G. Chen, W. Liu, and H. Tang · 2020
Later among the works it cites.
Evolving reinforcement learning algorithms
J. D. Co-Reyes, Y. Miao, D. Peng, Q. V. Le, S. Levine, H. Lee, and A. Faust · 2021
Closest in time.
Qplex: Duplex dueling multi-agent q-learning
J. Wang, Z. Ren, T. Liu, Y. Yu, and C. Zhang · 2021
Closest in time.