In: Advances in neural information processing systems, pp. 689–699 (2017)
Brown, N., Sandholm, T.: Safe and nested subgame solving for imperfect-information games · 2017
Later among the works it cites.
arXiv preprint arXiv:1708.04782 (2017)
Original
Vinyals, O., Ewalds, T., Bartunov, S., Georgiev, P., Vezhnevets, A.S., Yeo, M., Makhzani, A., Küttler, H., Agapiou, J., Schrittwieser, J., et al.: Starcraft II: A new challenge for reinforcement learning · 2017
Later among the works it cites.
arXiv preprint arXiv:1707.01068 (2017)
Original
Lerer, A., Peysakhovich, A.: Maintaining cooperation in complex social dilemmas using deep reinforcement learning · 2017
Later among the works it cites.
Proceedings of the ACM on Measurement and Analysis of Computing Systems
Chen, Y., Su, L., Xu, J.: Distributed statistical machine learning in adversarial settings: Byzantine gradient descent · 2017
Later among the works it cites.
https://blog.openai.com/openai-five/ (2018)
OpenAI: Openai five · 2018
Later among the works it cites.
arXiv preprint arXiv:1810.05587 (2018)
Original
Hernandez-Leal, P., Kartal, B., Taylor, M.E.: A survey and critique of multiagent deep reinforcement learning · 2018
Later among the works it cites.
In: International Conference on Machine Learning, pp. 5867–5876 (2018)
Zhang, K., Yang, Z., Liu, H., Zhang, T., Başar, T.: Fully decentralized multi-agent reinforcement learning with networked agents · 2018
Later among the works it cites.
arXiv preprint arXiv:1804.05464 (2018)
Original
Mazumdar, E., Ratliff, L.J.: On the convergence of gradient-based learning in continuous games · 2018
Later among the works it cites.
arXiv preprint arXiv:1812.11794 (2018)
Original
Nguyen, T.T., Nguyen, N.D., Nahavandi, S.: Deep reinforcement learning for multi-agent systems: A review of challenges, solutions and applications · 2018
Later among the works it cites.
In: IEEE Conference on Decision and Control, pp. 2771–2776 (2018)
Zhang, K., Yang, Z., Başar, T.: Networked multi-agent reinforcement learning in continuous spaces · 2018
Later among the works it cites.
arXiv preprint arXiv:1812.02783 (2018)
Original
Zhang, K., Yang, Z., Liu, H., Zhang, T., Başar, T.: Finite-sample analyses for fully decentralized multi-agent reinforcement learning · 2018
Later among the works it cites.
In: International Conference on Machine Learning, pp. 2284–2293 (2018)
Jiang, D., Ekwedike, E., Liu, H.: Feedback-based tree search for reinforcement learning · 2018
Later among the works it cites.
MIT Press (2018)
Sutton, R.S., Barto, A.G.: Reinforcement Learning: An Introduction · 2018
Later among the works it cites.
arXiv preprint arXiv:1801.01290 (2018)
Original
Haarnoja, T., Zhou, A., Abbeel, P., Levine, S.: Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor · 2018
Later among the works it cites.
In: IEEE Conference on Decision and Control, pp. 2759–2764 (2018)
Yang, Z., Zhang, K., Hong, M., Başar, T.: A finite sample analysis of the actor-critic algorithm · 2018
Later among the works it cites.
In: International Conference on Learning Representations (2018)
Valcarcel Macua, S., Zazo, J., Zazo, S.: Learning parametric closed-loop policies for Markov potential games · 2018
Later among the works it cites.
In: Advances in Neural Information Processing Systems, pp. 9649–9660 (2018)
Wai, H.T., Yang, Z., Wang, Z., Hong, M.: Multi-agent reinforcement learning via double averaging primal-dual optimization · 2018
Later among the works it cites.
In: Advances in Neural Information Processing Systems, pp. 3422–3435 (2018)
Srinivasan, S., Lanctot, M., Zambaldi, V., Pérolat, J., Tuyls, K., Munos, R., Bowling, M.: Actor-critic policy optimization in partially observable multiagent environments · 2018
Later among the works it cites.
arXiv preprint arXiv:1812.03239 (2018)
Original
Chen, T., Zhang, K., Giannakis, G.B., Başar, T.: Communication-efficient distributed reinforcement learning · 2018
Later among the works it cites.
In: International Conference on Machine Learning, pp. 1802–1811 (2018)
Grover, A., Al-Shedivat, M., Gupta, J., Burda, Y., Edwards, H.: Learning policy representations in multiagent systems · 2018
Later among the works it cites.
In: Workshop at International Conference on Learning Representations (2018)
Gao, C., Mueller, M., Hayward, R.: Adversarial policy gradient for alternating Markov games · 2018
Later among the works it cites.
In: International Conference on Autonomous Agents and Multi-Agent Systems, pp. 2085–2087 (2018)
Sunehag, P., Lever, G., Gruslys, A., Czarnecki, W.M., Zambaldi, V., Jaderberg, M., Lanctot, M., Sonnerat, N., Leibo, J.Z., Tuyls, K., et al.: Value-decomposition networks for cooperative multi-agent learning based on team reward · 2018
Later among the works it cites.
In: International Conference on Machine learning, pp. 681–689 (2018)
Rashid, T., Samvelyan, M., De Witt, C.S., Farquhar, G., Foerster, J., Whiteson, S.: QMIX: Monotonic value function factorisation for deep multi-agent reinforcement learning · 2018
Later among the works it cites.
In: AAAI Conference on Artificial Intelligence (2018)
Foerster, J.N., Farquhar, G., Afouras, T., Nardelli, N., Whiteson, S.: Counterfactual multi-agent policy gradients · 2018
Later among the works it cites.
In: International Conference on Machine Learning, pp. 1233–1242 (2018)
Dibangoye, J., Buffet, O.: Learning to act in decentralized partially observable MDPs · 2018
Later among the works it cites.
In: IEEE Conference on Decision and Control, pp. 1967–1972 (2018)
Lee, D., Yoon, H., Hovakimyan, N.: Primal-dual algorithm for distributed reinforcement learning: Distributed GTD · 2018
Later among the works it cites.
In: International Conference on Artificial Intelligence and Statistics (2018)
Perolat, J., Piot, B., Pietquin, O.: Actor-critic fictitious play in simultaneous move multistage games · 2018
Later among the works it cites.
Springer (2018)
Başar, T., Zaccour, G.: Handbook of Dynamic Game Theory · 2018
Later among the works it cites.
In: International Conference on Machine Learning, pp. 5571–5580 (2018)
Yang, Y., Luo, R., Li, M., Zhou, M., Zhang, W., Wang, J.: Mean field multi-agent reinforcement learning · 2018
Later among the works it cites.
In: Workshop on Adaptive and Learning Agents at International Conference on Autonomous Agents and Multi-Agent Systems (2018)
Subramanian, J., Seraj, R., Mahajan, A.: Reinforcement learning for mean-field teams · 2018
Later among the works it cites.
IEEE Journal of Selected Topics in Signal Processing
Zhang, K., Shi, W., Zhu, H., Dall’Anese, E., Başar, T.: Dynamic power distribution system management with a locally connected communication network · 2018
Later among the works it cites.
Transportation Research Part C: Emerging Technologies
Zhang, K., Lu, L., Lei, C., Zhu, H., Ouyang, Y.: Dynamic operations and pricing of electric unmanned aerial vehicle systems and power networks · 2018
Later among the works it cites.
arXiv preprint arXiv:1801.03326 (2018)
Original
Ciosek, K., Whiteson, S.: Expected policy gradients for reinforcement learning · 2018
Later among the works it cites.
arXiv preprint arXiv:1810.07792 (2018)
Original
Cassano, L., Yuan, K., Sayed, A.H.: Multi-agent fully decentralized off-policy learning with linear convergence rates · 2018
Later among the works it cites.
In: IEEE International Conference on Acoustics, Speech and Signal Processing, pp. 2286–2290 (2018)
Ying, B., Yuan, K., Sayed, A.H.: Convergence of variance-reduced learning under random reshuffling · 2018
Later among the works it cites.
In: Conference On Learning Theory, pp. 1691–1692 (2018)
Bhandari, J., Russo, D., Singal, R.: A finite time analysis of temporal difference learning with linear function approximation · 2018
Later among the works it cites.
In: International Joint Conference on Neural Networks, pp. 1–6 (2018)
Zhang, Q., Zhao, D., Lewis, F.L.: Model-free reinforcement learning for fully cooperative multi-agent graphical games · 2018
Later among the works it cites.
International Journal of Robotics Research pp. 1–22 (2018)
Best, G., Cliff, O.M., Patten, T., Mettu, R.R., Fitch, R.: Dec-MCTS: Decentralized planning for multi-robot active perception · 2018
Later among the works it cites.
In: Advances in Neural Information Processing Systems, pp. 5186–5196 (2018)
Sidford, A., Wang, M., Wu, X., Yang, L., Ye, Y.: Near-optimal time and sample complexities for solving Markov decision processes with a generative model · 2018
Later among the works it cites.
Games and Economic Behavior
Li, Z., Tewari, A.: Sampled fictitious play is hannan consistent · 2018
Later among the works it cites.
arXiv preprint arXiv:1810.04433 (2018)
Original
Zhou, Y., Ren, T., Li, J., Yan, D., Zhu, J.: Lazy-CFR: A fast regret minimization algorithm for extensive games with imperfect information · 2018
Later among the works it cites.
arXiv preprint arXiv:1804.09045 (2018)
Original
Kovařík, V., Lisỳ, V.: Analysis of hannan consistent selection for Monte Carlo tree search in simultaneous move games · 2018
Later among the works it cites.
arXiv preprint arXiv:1805.05751 (2018)
Original
Adolphs, L., Daneshmand, H., Lucchi, A., Hofmann, T.: Local saddle point optimization: A curvature exploitation approach · 2018
Later among the works it cites.
In: Advances in Neural Information Processing Systems, pp. 9236–9246 (2018)
Daskalakis, C., Panageas, I.: The limit points of (optimistic) gradient descent in min-max optimization · 2018
Later among the works it cites.
In: International Conference on Machine Learning, pp. 363–372 (2018)
Balduzzi, D., Racaniere, S., Martens, J., Foerster, J., Tuyls, K., Graepel, T.: The mechanics of n-player differentiable games · 2018
Later among the works it cites.
arXiv preprint arXiv:1812.02878 (2018)
Original
Sanjabi, M., Razaviyayn, M., Lee, J.D.: Solving non-convex non-concave min-max games under Polyak- · 2018
Later among the works it cites.
SIAM Journal on Control and Optimization
Saldi, N., Başar, T., Raginsky, M.: Markov–Nash equilibria in mean-field games with discounted cost · 2018
Later among the works it cites.
arXiv preprint arXiv:1808.03929 (2018)
Original
Saldi, N., Başar, T., Raginsky, M.: Discrete-time risk-sensitive mean-field games · 2018
Later among the works it cites.
In: International Joint Conference on Artificial Intelligence, pp. 562–568 (2018)
Yang, B., Liu, M.: Keeping in touch with collaborative UAVs: A deep reinforcement learning approach · 2018
Later among the works it cites.
arXiv preprint arXiv:1803.07250 (2018)
Original
Pham, H.X., La, H.M., Feil-Seifer, D., Nefian, A.: Cooperative and distributed reinforcement learning of drones for field coverage · 2018
Later among the works it cites.
In: SAI Intelligent Systems Conference, pp. 1169–1177 (2018)
Tožička, J., Szulyovszky, B., de Chambrier, G., Sarwal, V., Wani, U., Gribulis, M.: Application of deep reinforcement learning to UAV fleet control · 2018
Later among the works it cites.
In: AAAI Conference on Artificial Intelligence (2018)
Mordatch, I., Abbeel, P.: Emergence of grounded compositional language in multi-agent populations · 2018
Later among the works it cites.
In: Advances in Neural Information Processing Systems, pp. 7254–7264 (2018)
Jiang, J., Lu, Z.: Learning attentional communication for multi-agent cooperation · 2018
Later among the works it cites.
arXiv preprint arXiv:1810.09202
Original
Jiang, J., Dun, C., Lu, Z.: Graph convolutional reinforcement learning for multi-agent cooperation · 2018
Later among the works it cites.
arXiv preprint arXiv:1803.10357 (2018)
Original
Celikyilmaz, A., Bosselut, A., He, X., Choi, Y.: Deep communicating agents for abstractive summarization · 2018
Later among the works it cites.
arXiv preprint arXiv:1810.11187 (2018)
Original
Das, A., Gervet, T., Romoff, J., Batra, D., Parikh, D., Rabbat, M., Pineau, J.: TarMAC: Targeted multi-agent communication · 2018
Later among the works it cites.
arXiv preprint arXiv:1804.03984 (2018)
Original
Lazaridou, A., Hermann, K.M., Tuyls, K., Clark, S.: Emergence of linguistic communication from referential games with symbolic and pixel input · 2018
Later among the works it cites.
In: Advances in neural information processing systems, pp. 3326–3336 (2018)
Hughes, E., Leibo, J.Z., Phillips, M., Tuyls, K., Dueñez-Guzman, E., Castañeda, A.G., Dunning, I., Zhu, T., McKee, K., Koster, R., et al.: Inequity aversion improves cooperation in intertemporal social dilemmas · 2018
Later among the works it cites.
arXiv preprint arXiv:1802.06509 (2018)
Original
Arora, S., Cohen, N., Hazan, E.: On the optimization of deep networks: Implicit acceleration by overparameterization · 2018
Later among the works it cites.
In: Advances in Neural Information Processing Systems, pp. 8157–8166 (2018)
Li, Y., Liang, Y.: Learning overparameterized neural networks via stochastic gradient descent on structured data · 2018
Later among the works it cites.
arXiv preprint arXiv:1812.03565 (2018)
Original
Tu, S., Recht, B.: The gap between model-based and model-free methods on the linear quadratic regulator: An asymptotic viewpoint · 2018
Later among the works it cites.
arXiv preprint arXiv:1810.10207 (2018)
Original
Lin, Q., Liu, M., Rafique, H., Yang, T.: Solving weakly-convex-weakly-concave saddle-point problems as weakly-monotone variational inequality · 2018
Later among the works it cites.
arXiv preprint arXiv:1803.01498 (2018)
Original
Yin, D., Chen, Y., Ramchandran, K., Bartlett, P.: Byzantine-robust distributed learning: Towards optimal statistical rates · 2018
Later among the works it cites.
https://deepmind.com/blog/alphastar-mastering-real-time-strategy-game-starcraft-ii/ (2019)
Vinyals, O., Babuschkin, I., Chung, J., Mathieu, M., Jaderberg, M., Czarnecki, W.M., Dudzik, A., Huang, A., Georgiev, P., Powell, R., Ewalds, T., Horgan, D., Kroiss, M., Danihelka, I., Agapiou, J., Oh, J., Dalibard, V., Choi, D., Sifre, L., Sulsky, Y., Vezhnevets, S., Molloy, J., Cai, T., Budden, D., Paine, T., Gulcehre, C., Wang, Z., Pfaff, T., Pohlen, T., Wu, Y., Yogatama, D., Cohen, J., McKinney, K., Smith, O., Schaul, T., Lillicrap, T., Apps, C., Kavukcuoglu, K., Hassabis, D., Silver, D.: AlphaStar: Mastering the Real-Time Strategy Game StarCraft II · 2019
Closest in time.
In: International Conference on Autonomous Agents and Multi-Agent Systems, pp. 251–259 (2019)
Subramanian, J., Mahajan, A.: Reinforcement learning in stationary mean-field games · 2019
Closest in time.
arXiv preprint arXiv:1903.09569 (2019)
Original
Zhang, L., Wang, W., Li, S., Pan, G.: Monte Carlo neural fictitious self-play: Approach to approximate Nash equilibrium of imperfect-information games · 2019
Closest in time.
arXiv preprint arXiv:1902.00618 (2019)
Original
Jin, C., Netrapalli, P., Jordan, M.I.: Minmax optimization: Stable limit points of gradient descent ascent are locally optimal · 2019
Closest in time.
In: Advances in Neural Information Processing Systems (2019)
Zhang, K., Yang, Z., Başar, T.: Policy optimization provably converges to Nash equilibria in zero-sum linear quadratic games · 2019
Closest in time.
arXiv preprint arXiv:1908.11071 (2019)
Original
Sidford, A., Wang, M., Yang, L.F., Ye, Y.: Solving discounted stochastic two-player games with near-optimal time and sample complexity · 2019
Closest in time.
arXiv preprint arXiv:1903.05812 (2019)
Original
Yongacoglu, B., Arslan, G., Yüksel, S.: Learning team-optimality for decentralized stochastic control and dynamic games · 2019
Closest in time.
In: IEEE American Control Conference, pp. 167–172 (2019)
Zhang, K., Miehling, E., Başar, T.: Online planning for decentralized stochastic control with partial history sharing · 2019
Closest in time.
arXiv preprint arXiv:1908.03963 (2019)
Original
Oroojlooy Jadid, A., Hajinezhad, D.: A review of cooperative multi-agent deep reinforcement learning · 2019
Closest in time.
arXiv preprint arXiv:1902.05213 (2019)
Original
Shah, D., Xie, Q., Xu, Z.: On reinforcement learning using Monte-Carlo tree search with supervised learning: Non-asymptotic analysis · 2019
Closest in time.
arXiv preprint arXiv:1906.08383 (2019)
Original
Zhang, K., Koppel, A., Zhu, H., Başar, T.: Global convergence of policy gradient methods to (almost) locally optimal policies · 2019
Closest in time.
arXiv preprint arXiv:1908.00261 (2019)
Original
Agarwal, A., Kakade, S.M., Lee, J.D., Mahajan, G.: Optimality and approximation with policy gradient methods in Markov decision processes · 2019
Closest in time.
arXiv preprint arXiv:1906.10306 (2019)
Original
Liu, B., Cai, Q., Yang, Z., Wang, Z.: Neural proximal/trust region policy optimization attains globally optimal policy · 2019
Closest in time.
arXiv preprint arXiv:1909.01150 (2019)
Original
Wang, L., Cai, Q., Yang, Z., Wang, Z.: Neural policy gradient methods: Global optimality and rates of convergence · 2019
Closest in time.
In: International Conference on Machine Learning, pp. 1626–1635 (2019)
Doan, T., Maguluri, S., Romberg, J.: Finite-time analysis of distributed TD (0) with linear function approximation on multi-agent reinforcement learning · 2019
Closest in time.
arXiv preprint arXiv:1906.00190 (2019)
Original
Omidshafiei, S., Hennes, D., Morrill, D., Munos, R., Perolat, J., Lanctot, M., Gruslys, A., Lespiau, J.B., Tuyls, K.: Neural replicator dynamics · 2019
Closest in time.
arXiv preprint arXiv:1908.09453 (2019)
Original
Lanctot, M., Lockhart, E., Lespiau, J.B., Zambaldi, V., Upadhyay, S., Pérolat, J., Srinivasan, S., Timbers, F., Tuyls, K., Omidshafiei, S., et al.: Openspiel: A framework for reinforcement learning in games · 2019
Closest in time.
In: International Conference on Learning Representations (2019)
Kim, D., Moon, S., Hostallero, D., Kang, W.J., Lee, T., Son, K., Yi, Y.: Learning to schedule communication in multi-agent reinforcement learning · 2019
Closest in time.
In: IEEE Conference on Decision and Control (2019)
Lin, Y., Zhang, K., Yang, Z., Wang, Z., Başar, T., Sandhu, R., Liu, J.: A communication-efficient multi-agent actor-critic algorithm for distributed reinforcement learning · 2019
Closest in time.
In: Real-world Sequential Decision Making Workshop at International Conference on Machine Learning (2019)
Ren, J., Haupt, J.: A communication efficient hierarchical distributed optimization algorithm for multi-agent reinforcement learning · 2019
Closest in time.
In: AAAI Conference on Artificial Intelligence (2019)
Kim, W., Cho, M., Sung, Y.: Message-dropout: An efficient training method for multi-agent deep reinforcement learning · 2019
Closest in time.
In: AAAI Conference on Artificial Intelligence (2019)
Li, S., Wu, Y., Cui, X., Dong, H., Fang, F., Russell, S.: Robust multi-agent reinforcement learning via minimax deep deterministic policy gradient · 2019
Closest in time.
In: Advances in Neural Information Processing Systems, pp. 9482–9493 (2019)
Zhang, X., Zhang, K., Miehling, E., Basar, T.: Non-cooperative inverse reinforcement learning · 2019
Closest in time.
In: IEEE Conference on Decision and Control (2019)
Qu, G., Li, N.: Exploiting fast decaying and locality in multi-agent MDP with tree dependence structure · 2019
Closest in time.
arXiv preprint arXiv:1907.12530 (2019)
Original
Doan, T.T., Maguluri, S.T., Romberg, J.: Finite-time performance of distributed temporal difference learning with linear function approximation · 2019
Closest in time.
arXiv preprint arXiv:1903.06372 (2019)
Original
Suttle, W., Yang, Z., Zhang, K., Wang, Z., Başar, T., Liu, J.: A multi-agent off-policy actor-critic algorithm for distributed reinforcement learning · 2019
Closest in time.
In: International Conference on Machine Learning, pp. 5887–5896 (2019)
Son, K., Kim, D., Kang, W.J., Hostallero, D.E., Yi, Y.: QTRAN: Learning to factorize with transformation for cooperative multi-agent reinforcement learning · 2019
Closest in time.
arXiv preprint arXiv:1910.04295 (2019)
Original
Carmona, R., Laurière, M., Tan, Z.: Linear-quadratic mean-field reinforcement learning: Convergence of policy gradient methods · 2019
Closest in time.
arXiv preprint arXiv:1910.12802 (2019)
Original
Carmona, R., Laurière, M., Tan, Z.: Model-free mean-field reinforcement learning: Mean-field MDP and mean-field Q-learning · 2019
Closest in time.
Automatica (2019)
Zhang, K., Liu, Y., Liu, J., Liu, M., Başar, T.: Distributed learning of average belief over networks using sequential observations · 2019
Closest in time.
arXiv preprint arXiv:1903.09255 (2019)
Original
Zhang, Y., Zavlanos, M.M.: Distributed off-policy actor-critic reinforcement learning with policy consensus · 2019
Closest in time.
In: Conference on Learning Theory, pp. 2803–2830 (2019)
Srikant, R., Ying, L.: Finite-time error bounds for linear stochastic approximation and TD learning · 2019
Closest in time.
arXiv preprint arXiv:1901.00137 (2019)
Original
Yang, Z., Xie, Y., Wang, Z.: A theoretical analysis of deep Q-learning · 2019
Closest in time.
arXiv preprint arXiv:1906.00423 (2019)
Original
Jia, Z., Yang, L.F., Wang, M.: Feature-based Q-learning for two-player stochastic games · 2019
Closest in time.
In: AAAI Conference on Artificial Intelligence, vol. 33, pp. 2157–2164 (2019)
Schmid, M., Burch, N., Lanctot, M., Moravcik, M., Kadlec, R., Bowling, M.: Variance reduction in Monte Carlo counterfactual regret minimization (VR-MCCFR) for extensive form games using baselines · 2019
Closest in time.
In: International Conference on Machine Learning, pp. 793–802 (2019)
Brown, N., Lerer, A., Gross, S., Sandholm, T.: Deep counterfactual regret minimization · 2019
Closest in time.
Journal of Artificial Intelligence Research
Burch, N., Moravcik, M., Schmid, M.: Revisiting CFR+ and alternating updates · 2019
Closest in time.
arXiv preprint arXiv:1903.05614 (2019)
Original
Lockhart, E., Lanctot, M., Pérolat, J., Lespiau, J.B., Morrill, D., Timbers, F., Tuyls, K.: Computing approximate equilibria in sequential adversarial games by exploitability descent · 2019
Closest in time.
arXiv preprint arXiv:1901.00838 (2019)
Original
Mazumdar, E.V., Jordan, M.I., Sastry, S.S.: On finding local Nash equilibria (and only local Nash equilibria) in zero-sum games · 2019
Closest in time.
arXiv preprint arXiv:1911.04672 (2019)
Original
Bu, J., Ratliff, L.J., Mesbahi, M.: Global convergence of policy gradient for sequential zero-sum linear quadratic dynamic games · 2019
Closest in time.
In: International Conference on Learning Representations (2019)
Mertikopoulos, P., Zenati, H., Lecouat, B., Foo, C.S., Chandrasekhar, V., Piliouras, G.: Optimistic mirror descent in saddle-point problems: Going the extra (gradient) mile · 2019
Closest in time.
arXiv preprint arXiv:1906.01217 (2019)
Original
Fiez, T., Chasnov, B., Ratliff, L.J.: Convergence of learning dynamics in Stackelberg games · 2019
Closest in time.
arXiv preprint arXiv:1902.08297 (2019)
Original
Nouiehed, M., Sanjabi, M., Lee, J.D., Razaviyayn, M.: Solving a class of non-convex min-max games using iterative first order methods · 2019
Closest in time.
arXiv preprint arXiv:1907.03712 (2019)
Original
Mazumdar, E., Ratliff, L.J., Jordan, M.I., Sastry, S.S.: Policy-gradient algorithms have no guarantees of convergence in continuous action and state multi-agent settings · 2019
Closest in time.
Journal of Machine Learning Research
Letcher, A., Balduzzi, D., Racanière, S., Martens, J., Foerster, J.N., Tuyls, K., Graepel, T.: Differentiable game mechanics · 2019
Closest in time.
arXiv preprint arXiv:1906.00731 (2019)
Original
Chasnov, B., Ratliff, L.J., Mazumdar, E., Burden, S.A.: Convergence analysis of gradient-based learning with non-uniform learning rates in non-cooperative multi-agent settings · 2019
Closest in time.
Mathematics of Operations Research (2019)
Saldi, N., Başar, T., Raginsky, M.: Approximate Nash equilibria in partially observed stochastic games with mean-field interactions · 2019
Closest in time.
arXiv preprint arXiv:1908.08793 (2019)
Original
Saldi, N.: Discrete-time average-cost mean-field games on Polish spaces · 2019
Closest in time.
arXiv preprint arXiv:1901.09585 (2019)
Original
Guo, X., Hu, A., Xu, R., Zhang, J.: Learning mean-field games · 2019
Closest in time.
arXiv preprint arXiv:1910.07498 (2019)
Original
Fu, Z., Yang, Z., Chen, Y., Wang, Z.: Actor-critic provably finds Nash equilibria of linear-quadratic mean-field games · 2019
Closest in time.
Journal de Mathématiques Pures et Appliquées (2019)
Hadikhanloo, S., Silva, F.J.: Finite mean field games: Fictitious play and convergence to a first order continuous mean field game · 2019
Closest in time.
arXiv preprint arXiv:1907.02633 (2019)
Original
Elie, R., Pérolat, J., Laurière, M., Geist, M., Pietquin, O.: Approximate fictitious play for mean field games · 2019
Closest in time.
arXiv preprint arXiv:1909.01758 (2019)
Original
Anahtarci, B., Kariksiz, C.D., Saldi, N.: Value iteration algorithm for mean-field games · 2019
Closest in time.
In: IEEE Annual Consumer Communications & Networking Conference, pp. 1–6 (2019)
Shamsoshoara, A., Khaledi, M., Afghah, F., Razi, A., Ashdown, J.: Distributed cooperative spectrum sharing in UAV networks using multi-agent reinforcement learning · 2019
Closest in time.
In: IEEE International Conference on Communications Workshops, pp. 1–6 (2019)
Cui, J., Liu, Y., Nallanathan, A.: The application of multi-agent reinforcement learning in UAV networks · 2019
Closest in time.
IEEE Access (2019)
Qie, H., Shi, D., Shen, T., Xu, X., Li, Y., Wang, L.: Joint optimization of multi-UAV target assignment and path planning based on multi-agent reinforcement learning · 2019
Closest in time.
arXiv preprint arXiv:1904.09067 (2019)
Original
Cogswell, M., Lu, J., Lee, S., Parikh, D., Batra, D.: Emergence of compositional language with deep generational transmission · 2019
Closest in time.
Nature pp. 1–5 (2019)
Vinyals, O., Babuschkin, I., Czarnecki, W.M., Mathieu, M., Dudzik, A., Chung, J., Choi, D.H., Powell, R., Ewalds, T., Georgiev, P., et al.: Grandmaster level in starcraft ii using multi-agent reinforcement learning · 2019
Closest in time.
arXiv preprint arXiv:1905.10027 (2019)
Original
Cai, Q., Yang, Z., Lee, J.D., Wang, Z.: Neural temporal-difference learning converges to global optima · 2019
Closest in time.
In: Conference on Learning Theory, pp. 2898–2933 (2019)
Sun, W., Jiang, N., Krishnamurthy, A., Agarwal, A., Langford, J.: Model-based RL in contextual decision processes: PAC bounds and exponential improvements over model-free approaches · 2019
Closest in time.
Submitted to IEEE American Control Conference (2020)
Zaman, M.A.u., Zhang, K., Miehling, E., Başar, T.: Approximate equilibrium computation for discrete-time linear-quadratic mean-field games · 2020
Closest in time.