Fetching the paper…
Reading the bibliography…
We propose a reinforcement learning algorithm for stationary mean-field games, where the goal is to learn a pair of mean-field state and stationary policy that constitutes the Nash equilibrium.
Learning to collaborate in markov decision processes
Radanovic, G., Devidze, R., Parkes, D. C., and Singla, A. (2019) · 1901
Earlier work this paper cites.
Global optimality guarantees for policy gradient methods
Bhandari, J. and Russo, D. (2019) · 1906
Earlier work this paper cites.
On the convergence of model free learning in mean field games
Elie, R., Pérolat, J., Laurière, M., Geist, M., and Pietquin, O. (2019) · 1907
Earlier work this paper cites.
Value iteration algorithm for mean-field games
Anahtarci, B., Kariksiz, C. D., and Saldi, N. (2019b) · 1909
Earlier work this paper cites.
Adaptive trust region policy optimization: Global convergence and faster rates for regularized mdps
Shani, L., Efroni, Y., and Mannor, S. (2019) · 1909
Earlier work this paper cites.
Actor-critic provably finds nash equilibria of linear-quadratic mean-field games
Fu, Z., Yang, Z., Chen, Y., and Wang, Z. (2019) · 1910
Earlier work this paper cites.
Mastering atari, go, chess and shogi by planning with a learned model
Schrittwieser, J., Antonoglou, I., Hubert, T., Simonyan, K., Sifre, L., Schmitt, S., Guez, A., Lockhart, E., Hassabis, D., Graepel, T., et al. (2019) · 1911
Earlier work this paper cites.
Multi-agent reinforcement learning: A selective overview of theories and algorithms
Zhang, K., Yang, Z., and Başar, T. (2019) · 1911
Earlier work this paper cites.
Fitted Q-learning in mean-field games
Anahtarci, B., Kariksiz, C. D., and Saldi, N. (2019a) · 1912
Earlier work this paper cites.
Dota 2 with large scale deep reinforcement learning
Berner, C., Brockman, G., Chan, B., Cheung, V., Dębiak, P., Dennison, C., Farhi, D., Fischer, Q., Hashme, S., Hesse, C., et al. (2019) · 1912
Earlier work this paper cites.
Equilibrium points in n-person games
Nash, J. F. (1950) · 1950
Earlier work this paper cites.
Iterative solution of games by fictitious play
Brown, G. W. (1951) · 1951
Earlier work this paper cites.
Q-learning
Watkins, C. J. and Dayan, P. (1992) · 1992
Earlier work this paper cites.
Dynamic noncooperative game theory
Başar, T. and Olsder, G. J. (1998) · 1998
Earlier work this paper cites.
Actor-critic algorithms
Konda, V. R. and Tsitsiklis, J. N. (2000) · 2000
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Kakade, S. and Langford, J. (2002) · 2002
Earlier work this paper cites.
Social diversity and social preferences in mixed-motive reinforcement learning
McKee, K. R., Gemp, I., McWilliams, B., Duéñez-Guzmán, E. A., Hughes, E., and Leibo, J. Z. (2020) · 2002
Earlier work this paper cites.
Q-learning in regularized mean-field games
Anahtarci, B., Kariksiz, C. D., and Saldi, N. (2020) · 2003
Earlier work this paper cites.
A general framework for learning mean-field games
Guo, X., Hu, A., Xu, R., and Zhang, J. (2020) · 2003
Earlier work this paper cites.
Individual and mass behaviour in large population stochastic wireless power control problems: centralized and nash equilibrium solutions
Huang, M., Caines, P. E., and Malhamé, R. P. (2003) · 2003
Earlier work this paper cites.
Multiagent reinforcement learning for multi-robot systems: A survey
Yang, E. and Gu, D. (2004) · 2004
Earlier work this paper cites.
Decentralized reinforcement learning control of a robotic manipulator
Busoniu, L., De Schutter, B., and Babuska, R. (2006) · 2006
Earlier work this paper cites.
Optimal rates for the regularized least-squares algorithm
Caponnetto, A. and De Vito, E. (2007) · 2007
Earlier work this paper cites.
Fast global convergence of natural policy gradient methods with entropy regularization
Cen, S., Cheng, C., Chen, Y., Wei, Y., and Chi, Y. (2020) · 2007
Earlier work this paper cites.
Dynamic regret of policy optimization in non-stationary environments
Fei, Y., Yang, Z., Wang, Z., and Xie, Q. (2020) · 2007
Earlier work this paper cites.
Large-population cost-coupled LQG problems with nonuniform agents: individual-mass behavior and decentralized ε \varepsilon -Nash equilibria
Huang, M., Caines, P. E., and Malhamé, R. P. (2007) · 2007
Cited alongside, same era.
Mean field games
Lasry, J.-M. and Lions, P.-L. (2007) · 2007
Cited alongside, same era.
Fictitious play for mean field games: Continuous time analysis and applications
Perrin, S., Pérolat, J., Laurière, M., Geist, M., Elie, R., and Pietquin, O. (2020) · 2007
Cited alongside, same era.
If multi-agent learning is the answer, what is the question?
Shoham, Y., Powers, R., and Grenager, T. (2007) · 2007
Cited alongside, same era.
A hilbert space embedding for distributions
Smola, A., Gretton, A., Song, L., and Schölkopf, B. (2007) · 2007
Cited alongside, same era.
A comprehensive survey of multiagent reinforcement learning
A distributional perspective on reinforcement learning
Bellemare, M. G., Dabney, W., and Munos, R. (2017) · 2017
Later among the works it cites.
A survey of learning in multiagent environments: Dealing with non-stationarity
Hernandez-Leal, P., Kaisers, M., Baarslag, T., and de Cote, E. M. (2017) · 2017
Later among the works it cites.
Multi-agent reinforcement learning in sequential social dilemmas
Leibo, J. Z., Zambaldi, V., Lanctot, M., Marecki, J., and Graepel, T. (2017) · 2017
Later among the works it cites.
Distributed learning with regularized least squares
Lin, S.-B., Guo, X., and Zhou, D.-X. (2017) · 2017
Later among the works it cites.
Bridging the gap between value and policy based reinforcement learning
Nachum, O., Norouzi, M., Xu, K., and Schuurmans, D. (2017) · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Busoniu, L., Babuska, R., and De Schutter, B. (2008) · 2008
Cited alongside, same era.
A kernel method for the two-sample problem
Gretton, A., Borgwardt, K. M., Rasch, M. J., Schölkopf, B., and Smola, A. (2008) · 2008
Cited alongside, same era.
Multiagent reinforcement learning for urban traffic control using coordination graphs
Kuyer, L., Whiteson, S., Bakker, B., and Vlassis, N. (2008) · 2008
Cited alongside, same era.
Finite-time bounds for fitted value iteration
Munos, R. and Szepesvári, C. (2008) · 2008
Cited alongside, same era.
An introduction to multiagent systems
Wooldridge, M. (2009) · 2009
Cited alongside, same era.
Discrete time, finite state space mean field games
Gomes, D. A., Mohr, J., and Souza, R. R. (2010) · 2010
Cited alongside, same era.
Mean field games and applications
Guéant, O., Lasry, J.-M., and Lions, P.-L. (2011) · 2010
Cited alongside, same era.
Later among the works it cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. (2017) · 2017
Later among the works it cites.
Mastering the game of Go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., and Bolton, A. (2017) · 2017
Later among the works it cites.
Decision-theoretic planning under anonymity in agent populations
Sonu, E., Chen, Y., and Doshi, P. (2017) · 2017
Later among the works it cites.
A finite time analysis of temporal difference learning with linear function approximation
Bhandari, J., Russo, D., and Singal, R. (2018) · 2018
Later among the works it cites.
Emergent communication through negotiation
Cao, K., Lazaridou, A., Lanctot, M., Leibo, J. Z., Tuyls, K., and Clark, S. (2018) · 2018
Later among the works it cites.
Probabilistic Theory of Mean Field Games with Applications I-II
Carmona, R. and Delarue, F. (2018) · 2018
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S. (2018) · 2018
Later among the works it cites.
Is multiagent deep reinforcement learning the answer or the question? a brief survey
Hernandez-Leal, P., Kartal, B., and Taylor, M. E. (2018) · 2018
Later among the works it cites.
Decentralized reinforcement learning of robot behaviors
Leottau, D. L., Ruiz-del Solar, J., and Babuška, R. (2018) · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G. (2018) · 2018
Later among the works it cites.
A theory of regularized markov decision processes
Geist, M., Scherrer, B., and Pietquin, O. (2019) · 2019
Later among the works it cites.
Learning mean-field games
Guo, X., Hu, A., Xu, R., and Zhang, J. (2019) · 2019
Later among the works it cites.
Social influence as intrinsic motivation for multi-agent deep reinforcement learning
Jaques, N., Lazaridou, A., Hughes, E., Gulcehre, C., Ortega, P., Strouse, D., Leibo, J. Z., and De Freitas, N. (2019) · 2019
Later among the works it cites.
Approximate nash equilibria in partially observed stochastic games with mean-field interactions
Saldi, N., Başar, T., and Raginsky, M. (2019) · 2019
Later among the works it cites.
Reinforcement learning in stationary mean-field games
Subramanian, J. and Mahajan, A. (2019) · 2019
Later among the works it cites.
Grandmaster level in StarCraft II using multi-agent reinforcement learning
Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., Choi, D. H., Powell, R., Ewalds, T., Georgiev, P., et al. (2019) · 2019
Later among the works it cites.
Optimality and approximation with policy gradient methods in markov decision processes
Agarwal, A., Kakade, S. M., Lee, J. D., and Mahajan, G. (2020) · 2020
Closest in time.
Approximate equilibrium computation for discrete-time linear-quadratic mean-field games
uz Zaman, M. A., Zhang, K., Miehling, E., and Başar, T. (2020) · 2020
Closest in time.
Discrete-time ergodic mean-field games with average reward on compact spaces
Więcek, P. (2020) · 2020
Closest in time.