Fetching the paper…
Reading the bibliography…
The analysis and control of large-population systems is of great interest to diverse areas of research and engineering, ranging from epidemiology over robotic swarms to economics and finance.
1904
Earlier work this paper cites.
1910
Earlier work this paper cites.
1910
Earlier work this paper cites.
1911
Earlier work this paper cites.
1912
Earlier work this paper cites.
1912
Earlier work this paper cites.
K. Cui and H. Koeppl, “Approximately solving mean field games via entropy-regularized deep reinforcement learning,” in Proc. AISTATS , 2021, pp. 1909–1917
1917
Earlier work this paper cites.
J. Kennedy and R. Eberhart, “Particle swarm optimization,” in Proc. Int. Conf. Neural Netw. , vol. 4, 1995, pp. 1942–1948
1948
Earlier work this paper cites.
R. J. Glauber, “Time-dependent statistics of the Ising model,” J. Math. Phys. , vol. 4, no. 2, pp. 294–307, 1963
1963
Earlier work this paper cites.
R. Bellman, “Dynamic programming,” Science , vol. 153, no. 3731, pp. 34–37, 1966
1966
Earlier work this paper cites.
A. Okubo, “Dynamical aspects of animal grouping: swarms, schools, flocks, and herds,” Adv. Biophys. , vol. 22, pp. 1–94, 1986
1986
Earlier work this paper cites.
D. Fudenberg and J. Tirole, Game theory . MIT press, 1991
1991
Earlier work this paper cites.
J.-A. Meyer and S. W. Wilson, “Task differentiation in polistes wasp colonies: a model for self-organizing groups of robots,” in From Animals to Animats: Proceedings of the First International Conference on Simulation of Adaptive Behavior . MIT Press, 1991, pp. 346–355
1991
Earlier work this paper cites.
C. J. Watkins and P. Dayan, “Q-learning,” Mach. Learn. , vol. 8, no. 3, pp. 279–292, 1992
1992
Earlier work this paper cites.
R. J. Williams, “Simple statistical gradient-following algorithms for connectionist reinforcement learning,” Machine Learning , vol. 8, no. 3, pp. 229–256, 1992
1992
Earlier work this paper cites.
M. Tan, “Multi-agent reinforcement learning: Independent vs. cooperative agents,” in Proc. ICML , 1993, pp. 330–337
1993
Earlier work this paper cites.
J. A. Boyan and M. L. Littman, “Packet routing in dynamically changing networks: a reinforcement learning approach,” in Proc. NIPS , 1993, pp. 671–678
1993
Earlier work this paper cites.
A. Schaerf, Y. Shoham, and M. Tennenholtz, “Adaptive load balancing: A study in multi-agent learning,” J. Artif. Intell. Res. , vol. 2, pp. 475–500, 1994
1994
Earlier work this paper cites.
L. Baird, “Residual algorithms: Reinforcement learning with function approximation,” in Proc. ICML . Elsevier, 1995, pp. 30–37
1995
Earlier work this paper cites.
C. Boutilier, R. Dearden, and M. Goldszmidt, “Exploiting structure in policy construction,” in Proc. IJCAI , 1995, pp. 1104–1111
1995
Earlier work this paper cites.
S. P. Choi and D.-Y. Yeung, “Predictive Q-routing: a memory-based reinforcement learning approach to adaptive traffic control,” in Proc. NIPS , 1995, pp. 945–951
1995
Earlier work this paper cites.
D. Subramanian, P. Druschel, and J. Chen, “Ants and reinforcement learning: a case study in routing in dynamic networks,” in Proc. IJCAI , 1997, pp. 832–838
1997
Earlier work this paper cites.
R. Schoonderwoerd, O. E. Holland, J. L. Bruten, and L. J. Rothkrantz, “Ant-based load balancing in telecommunications networks,” Adaptive Behav. , vol. 5, no. 2, pp. 169–207, 1997
1997
Earlier work this paper cites.
S. KUMAR, “Dual reinforcement Q-routing: An on-line adaptive routing algorithm,” in Proc. Artif. Neural Netw. Eng. Conf. , 1998, pp. 231–238
1998
Earlier work this paper cites.
G. Di Caro and M. Dorigo, “Mobile agents for adaptive routing,” in Proc. Hawaii Int. Conf. Syst. Sci. , vol. 31, 1998, pp. 74–85
1998
Earlier work this paper cites.
R. S. Sutton, D. McAllester, S. Singh, and Y. Mansour, “Policy gradient methods for reinforcement learning with function approximation,” in Proc. NIPS , 1999, pp. 1057–1063
1999
Earlier work this paper cites.
D. H. Wolpert and K. Tumer, “An introduction to collective intelligence,” arXiv:cs/9908014 , 1999. [Online]. Available: https://arxiv.org/abs/cs/9908014
1999
Earlier work this paper cites.
D. H. Wolpert, K. R. Wheeler, and K. Tumer, “Collective intelligence for control of distributed dynamical systems,” EPL , vol. 49, no. 6, p. 708, 2000
2000
Earlier work this paper cites.
C. Guestrin, D. Koller, and R. Parr, “Multiagent planning with factored MDPs,” in Proc. NIPS , vol. 14. MIT Press, 2001, pp. 1523–1530
2001
Earlier work this paper cites.
Y. T. Valdivia, M. M. Vellasco, and M. A. Pacheco, “An adaptive network routing strategy with temporal differences,” Inteligencia Artificial. Revista Iberoamericana de Inteligencia Artificial , vol. 5, no. 12, pp. 85–91, 2001
2001
Earlier work this paper cites.
H. Matsuo and K. Mori, “Accelerated ants routing in dynamic networks,” in Proc. Int. Conf. Softw. Eng. Artif. Intell. Netw. Parallel/Distrib. Comput. , 2001, pp. 333–339
2001
Earlier work this paper cites.
S. Camazine, N. R. Franks, J. Sneyd, E. Bonabeau, J.-L. Deneubourg, and G. Theraula, Self-Organization in Biological Systems . Princeton University Press, 2001
2001
Earlier work this paper cites.
P. Golle, K. Leyton-Brown, I. Mironov, and M. Lillibridge, “Incentives for sharing in peer-to-peer networks,” in Proc. Int. Workshop Electron. Commerce , 2001, pp. 75–87
2001
Earlier work this paper cites.
D. S. Bernstein, R. Givan, N. Immerman, and S. Zilberstein, “The complexity of decentralized control of Markov decision processes,” Math. Oper. Res. , vol. 27, no. 4, pp. 819–840, 2002
2002
Earlier work this paper cites.
I. D. Couzin, J. Krause, R. James, G. D. Ruxton, and N. R. Franks, “Collective memory and spatial sorting in animal groups,” J. Theor. Biol. , vol. 218, no. 1, pp. 1–11, 2002
2002
Earlier work this paper cites.
C. Guestrin, M. Lagoudakis, and R. Parr, “Coordinated reinforcement learning,” in Proc. ICML , vol. 2, 2002, pp. 227–234
2002
Earlier work this paper cites.
R. Albert and A.-L. Barabási, “Statistical mechanics of complex networks,” Rev. Mod. Phys. , vol. 74, no. 1, p. 47, 2002
2002
Earlier work this paper cites.
J. K. Parrish, S. V. Viscido, and D. Grunbaum, “Self-organized fish schools: an examination of emergent properties,” Biol. Bull. , vol. 202, no. 3, pp. 296–305, 2002
2002
Earlier work this paper cites.
A. Castelletti, G. Corani, A. Rizzolli, R. Soncinie-Sessa, and E. Weber, “Reinforcement learning in the operational management of a water system,” in Proc. IFAC Workshop on Model. Contr. in Environmental Issues , 2002, pp. 325–330
2002
Earlier work this paper cites.
P. D. Christofides and J. Chow, “Nonlinear and robust control of PDE systems: Methods and applications to transport-reaction processes,” Appl. Mech. Rev. , vol. 55, no. 2, pp. B29–B30, 2002
2002
Earlier work this paper cites.
C. Guestrin, D. Koller, R. Parr, and S. Venkataraman, “Efficient solution algorithms for factored MDPs,” J. Artif. Intell. Res. , vol. 19, pp. 399–468, 2003
2003
Earlier work this paper cites.
2003
Earlier work this paper cites.
M. Penrose, Random Geometric Graphs . OUP Oxford, 2003, vol. 5
2003
Earlier work this paper cites.
Z. Lin, M. Broucke, and B. Francis, “Local control strategies for groups of mobile autonomous agents,” IEEE Trans. Automat. Contr. , vol. 49, no. 4, pp. 622–629, 2004
2004
Earlier work this paper cites.
C. V. Goldman and S. Zilberstein, “Decentralized control of cooperative systems: Categorization and complexity analysis,” J. Artif. Intell. Res. , vol. 22, pp. 143–174, 2004
2004
Earlier work this paper cites.
E. A. Hansen, D. S. Bernstein, and S. Zilberstein, “Dynamic programming for partially observable stochastic games,” in Proc. AAAI , vol. 4, 2004, pp. 709–715
2004
Earlier work this paper cites.
J. R. Kok and N. Vlassis, “Sparse cooperative Q-learning,” in Proc. ICML , 2004, p. 61
2004
Earlier work this paper cites.
2004
Earlier work this paper cites.
D. Szer, F. Charpillet, and S. Zilberstein, “MAA*: a heuristic search algorithm for solving decentralized POMDPs,” in Proc. UAI , 2005, pp. 576–583
2005
Earlier work this paper cites.
R. Nair, P. Varakantham, M. Tambe, and M. Yokoo, “Networked distributed POMDPs: A synthesis of distributed constraint optimization and POMDPs,” in Proc. AAAI , 2005, p. 133–139
2005
Earlier work this paper cites.
J. Dowling, E. Curran, R. Cunningham, and V. Cahill, “Using feedback in collaborative reinforcement learning to adaptively optimize MANET routing,” IEEE Trans. Syst. Man Cybern. A: Syst. and Humans , vol. 35, no. 3, pp. 360–372, 2005
2005
Earlier work this paper cites.
J. R. Kok and N. Vlassis, “Collaborative multiagent reinforcement learning by payoff propagation,” J. Mach. Learn. Res. , vol. 7, no. 65, pp. 1789–1828, 2006
2006
Earlier work this paper cites.
Y. Kim, R. Nair, P. Varakantham, M. Tambe, and M. Yokoo, “Exploiting locality of interaction in networked distributed POMDPs.” in Proc. AAAI Spring Symp.: Distrib. Plan Sched. Manage. , 2006, pp. 41–48
2006
Earlier work this paper cites.
D. Zhou, J. Huang, and B. Schölkopf, “Learning with hypergraphs: clustering, classification, and embedding,” in Proc. NIPS , 2006, pp. 1601–1608
2006
Earlier work this paper cites.
M. Huang, R. P. Malhamé, and P. E. Caines, “Large population stochastic dynamic games: closed-loop McKean-Vlasov systems and the Nash certainty equivalence principle,” Commun. Inf. Syst. , vol. 6, no. 3, pp. 221–252, 2006
2006
Earlier work this paper cites.
D. Milutinović and P. Lima, “Modeling and optimal centralized control of a large-size robotic population,” IEEE Trans. Robot. , vol. 22, no. 6, pp. 1280–1285, 2006
2006
Earlier work this paper cites.
2006
Earlier work this paper cites.
M. Huang, P. E. Caines, R. P. Malhamé et al. , “Distributed multi-agent decision-making with partial observations: asymptotic Nash equilibria,” in Proc. 17th Internat. Symp. MTNS , 2006, pp. 2725–2730
2006
Earlier work this paper cites.
M. Dorigo, M. Birattari, and T. Stutzle, “Ant colony optimization,” IEEE Comput. Intell. Mag. , vol. 1, no. 4, pp. 28–39, 2006
2006
Earlier work this paper cites.
J.-M. Lasry and P.-L. Lions, “Mean field games,” Japanese J. Math. , vol. 2, no. 1, pp. 229–260, 2007
2007
Earlier work this paper cites.
B. Shucker and J. K. Bennett, “Scalable control of distributed robotic macrosensors,” in Distributed Autonomous Robotic Systems 6 , 2007, pp. 379–388
2007
Earlier work this paper cites.
F. Oliehoek, M. Spaan, S. Whiteson, and N. Vlassis, “Exploiting locality of interaction in factored Dec-POMDPs,” in Proc. AAMAS , May 2008, pp. 517–524
2008
Earlier work this paper cites.
A. Barrat, M. Barthlemy, and A. Vespignani, Dynamical Processes on Complex Networks . USA: Cambridge University Press, 2008
2008
Earlier work this paper cites.
M. Gribaudo, D. Cerotti, and A. Bobbio, “Analysis of on-off policies in sensor networks using interacting Markovian agents,” in Proc. Ann. IEEE Int. Conf. PerCom , 2008, pp. 300–305
2008
Earlier work this paper cites.
C. Daskalakis, P. W. Goldberg, and C. H. Papadimitriou, “The complexity of computing a Nash equilibrium,” SIAM J. Comput. , vol. 39, no. 1, pp. 195–259, 2009
2009
Earlier work this paper cites.
A. Nedic and A. Ozdaglar, “Distributed subgradient methods for multi-agent optimization,” IEEE Trans. Automat. Contr. , vol. 54, no. 1, pp. 48–61, 2009
2009
Earlier work this paper cites.
P. Varshavskaya, L. P. Kaelbling, and D. Rus, “Efficient distributed reinforcement learning through agreement,” in Distributed Autonomous Robotic Systems 8 . Springer, 2009, pp. 367–378
2009
Earlier work this paper cites.
A. Gosavi, “Reinforcement learning: A tutorial survey and recent advances,” INFORMS J. Comput. , vol. 21, no. 2, pp. 178–192, 2009
2009
Earlier work this paper cites.
P. Van Mieghem, J. Omic, and R. Kooij, “Virus spread in networks,” IEEE/ACM Trans. Netw. , vol. 17, no. 1, pp. 1–14, 2009
2009
Earlier work this paper cites.
V. M. Preciado and A. Jadbabaie, “Spectral analysis of virus spreading in random geometric networks,” in Proc. IEEE CDC , 2009, pp. 4802–4807
2009
Earlier work this paper cites.
A. Diaz-Guilera, J. Gómez-Gardenes, Y. Moreno, and M. Nekovee, “Synchronization in random geometric graphs,” Int. J. Bifurcation Chaos , vol. 19, no. 02, pp. 687–693, 2009
2009
Earlier work this paper cites.
P. Pennesi and I. C. Paschalidis, “A distributed actor-critic algorithm and applications to mobile sensor network coordination problems,” IEEE Trans. Automat. Contr. , vol. 55, no. 2, pp. 492–497, 2010
2010
Earlier work this paper cites.
D. Oužecki and D. Jevtić, “Reinforcement learning as adaptive network routing of mobile agents,” in Proc. Int. Conv. MIPRO , 2010, pp. 479–484
2010
Earlier work this paper cites.
Y. Achdou and I. Capuzzo-Dolcetta, “Mean field games: numerical methods,” SIAM J. Numer. Anal. , vol. 48, no. 3, pp. 1136–1162, 2010
2010
Earlier work this paper cites.
C.-H. Yu, J. Werfel, and R. Nagpal, “Collective decision-making in multi-agent systems by implicit leadership,” in Proc. AAMAS , vol. 3, 2010, p. 1189–1196
2010
Earlier work this paper cites.
C. Wu, K. Kumekawa, and T. Kato, “Distributed reinforcement learning approach for vehicular ad hoc networks,” IEICE Trans. Commun. , vol. 93, no. 6, pp. 1431–1442, 2010
2010
Earlier work this paper cites.
2010
Earlier work this paper cites.
2011
Earlier work this paper cites.
A. Kumar, S. Zilberstein, and M. Toussaint, “Scalable multiagent planning using probabilistic inference,” in Proc. IJCAI , 2011, p. 2140–2146
2011
Earlier work this paper cites.
C. Zhang and V. Lesser, “Coordinated multi-agent reinforcement learning in networked distributed POMDPs,” in Proc. AAAI , 2011, pp. 764–770
2011
Earlier work this paper cites.
J. K. Pajarinen and J. T. Peltonen, “Efficient planning for factored infinite-horizon Dec-POMDPs,” in Proc. IJCAI , 2011, p. 325–331
2011
Earlier work this paper cites.
A. Agarwal and J. C. Duchi, “Distributed delayed stochastic optimization,” in Proc. NIPS , 2011, pp. 873–881
2011
Earlier work this paper cites.
D. Jakovetic, J. Xavier, and J. M. Moura, “Cooperative convex optimization in networked systems: Augmented Lagrangian algorithms with directed gossip communication,” IEEE Trans. Sig. Proc. , vol. 59, no. 8, pp. 3889–3902, 2011
2011
Earlier work this paper cites.
L. Rose, S. Lasaulce, S. M. Perlaza, and M. Debbah, “Learning equilibria with partial information in decentralized wireless networks,” IEEE Commun. Mag. , vol. 49, no. 8, pp. 136–142, 2011
2011
Earlier work this paper cites.
A. M. Lopez and D. R. Heisterkamp, “Simulated annealing based hierarchical Q-routing: A dynamic routing protocol,” in Proc. Int. Conf. Inf. Technol. New Gener. , 2011, pp. 791–796
2011
Earlier work this paper cites.
D. Andersson and B. Djehiche, “A maximum principle for sdes of mean-field type,” Appl. Math. Optim. , vol. 63, no. 3, pp. 341–356, 2011
2011
Earlier work this paper cites.
N. Correll and A. Martinoli, “Modeling and designing self-organized aggregation in a swarm of miniature robots,” Int. J. Robot. Res. , vol. 30, no. 5, pp. 615–626, 2011
2011
Earlier work this paper cites.
V. Sperati, V. Trianni, and S. Nolfi, “Self-organised path formation in a swarm of robots,” Swarm Intell. , vol. 5, no. 2, pp. 97–119, 2011
2011
Earlier work this paper cites.
M. A. Montes de Oca, E. Ferrante, A. Scheidler, C. Pinciroli, M. Birattari, and M. Dorigo, “Majority-rule opinion dynamics with differential latency: a mechanism for self-organized collective decision-making,” Swarm Intell. , vol. 5, no. 3, pp. 305–327, 2011
2011
Earlier work this paper cites.
Z. Ma, D. S. Callaway, and I. A. Hiskens, “Decentralized charging control of large populations of plug-in electric vehicles,” IEEE Trans. Contr. Syst. Technol. , vol. 21, no. 1, pp. 67–78, 2011
2011
Earlier work this paper cites.
J. P. Gleeson, “High-accuracy approximation of binary-state dynamics on networks,” Phys. Rev. Lett. , vol. 107, p. 068701, 08 2011
2011
Earlier work this paper cites.
M. Walters, “Random geometric graphs,” Surv. Combinatorics , vol. 392, pp. 365–402, 2011
2011
Earlier work this paper cites.
H. H. Wensink, J. Dunkel, S. Heidenreich, K. Drescher, R. E. Goldstein, H. Löwen, and J. M. Yeomans, “Meso-scale turbulence in living fluids,” PNAS , vol. 109, no. 36, pp. 14 308–14 313, 2012
2012
Earlier work this paper cites.
A. Mahajan, N. C. Martins, M. C. Rotkowitz, and S. Yüksel, “Information structures in optimal decentralized control,” in Proc. IEEE CDC , 2012, pp. 1291–1306
2012
Earlier work this paper cites.
O. Hernández-Lerma and J. B. Lasserre, Discrete-time Markov control processes: basic optimality criteria . Springer Science & Business Media, 2012, vol. 30
2012
Earlier work this paper cites.
S.-Y. Tu and A. H. Sayed, “Diffusion strategies outperform consensus strategies for distributed estimation over adaptive networks,” IEEE Trans. Sig. Proc. , vol. 60, no. 12, pp. 6217–6234, 2012
2012
Earlier work this paper cites.
F. Oliehoek, S. Whiteson, and M. Spaan, “Exploiting structure in cooperative Bayesian games,” in Proc. UAI , August 2012, pp. 654–664
2012
Earlier work this paper cites.
R. A. Haraty and B. Traboulsi, “MANET with the Q-routing protocol,” in Proc. Int. Conf. Netw. , 2012, pp. 187–192
2012
Earlier work this paper cites.
L. Lovász, Large networks and graph limits . Am. Math. Soc., 2012, vol. 60
2012
Earlier work this paper cites.
2012
Earlier work this paper cites.
D. Bruneo, M. Scarpa, A. Bobbio, D. Cerotti, and M. Gribaudo, “Markovian agent modeling swarm intelligence algorithms in wireless sensor networks,” Perform. Eval. , vol. 69, no. 3–4, p. 135–149, Mar. 2012
2012
Earlier work this paper cites.
J. Kober, J. A. Bagnell, and J. Peters, “Reinforcement learning in robotics: A survey,” Int. J. Robot. Res. , vol. 32, no. 11, pp. 1238–1274, 2013
2013
Earlier work this paper cites.
M. Brambilla, E. Ferrante, M. Birattari, and M. Dorigo, “Swarm robotics: A review from the swarm engineering perspective,” Swarm Intell. , vol. 7, no. 1, pp. 1–41, 2013
2013
Earlier work this paper cites.
C. Amato, G. Chowdhary, A. Geramifard, N. K. Üre, and M. J. Kochenderfer, “Decentralized control of partially observable Markov decision processes,” in Proc. IEEE CDC , 2013, pp. 2398–2405
2013
Earlier work this paper cites.
F. A. Oliehoek, S. Whiteson, M. T. Spaan et al. , “Approximate solutions for factored Dec-POMDPs with many agents.” in Proc. AAMAS , 2013, pp. 563–570
2013
Earlier work this paper cites.
F. Wu, S. Zilberstein, and N. R. Jennings, “Monte-carlo expectation maximization for decentralized POMDPs,” in Proc. IJCAI , 2013, pp. 397–403
2013
Earlier work this paper cites.
D. Gamarnik, “Correlation decay method for decision, optimization, and inference in large-scale networks,” in Theory Driven by Influential Applications . INFORMS, 2013, pp. 108–121
2013
Earlier work this paper cites.
H. Sayama, I. Pestov, J. Schmidt, B. J. Bush, C. Wong, J. Yamanoi, and T. Gross, “Modeling complex systems with adaptive networks,” Comput. Math. Appl. , vol. 65, no. 10, pp. 1645–1664, 2013
2013
Cited alongside, same era.
J. P. Gleeson, “Binary-state dynamics on complex networks: Pair approximation and beyond,” Phys. Rev. X , vol. 3, p. 021004, 04 2013
2013
Cited alongside, same era.
A. Bensoussan, J. Frehse, P. Yam et al. , Mean field games and mean field type control theory . Springer, 2013, vol. 101
2013
Cited alongside, same era.
M. Nourian and P. E. Caines, “ ϵ \epsilon -Nash mean field game theory for nonlinear stochastic dynamical systems with major and minor agents,” SIAM J. Contr. Optim. , vol. 51, no. 4, pp. 3302–3331, 2013
2013
Cited alongside, same era.
A. Cavagna and I. Giardina, “Bird flocks as condensed matter,” Annu. Rev. Condens. Matter Phys. , vol. 5, no. 1, pp. 183–207, 2014
Z. Zhang, D. Zhang, and R. C. Qiu, “Deep reinforcement learning for power system applications: An overview,” CSEE J. Power Energy Syst. , vol. 6, no. 1, pp. 213–225, 2019
2019
Later among the works it cites.
C. Yu, X. Wang, X. Xu, M. Zhang, H. Ge, J. Ren, L. Sun, B. Chen, and G. Tan, “Distributed multiagent coordinated learning for autonomous driving in highways based on dynamic coordination graphs,” IEEE Trans. Transp. Syst. , vol. 21, no. 2, pp. 735–748, 2019
2019
Later among the works it cites.
T. Chu, J. Wang, L. Codecà, and Z. Li, “Multi-agent deep reinforcement learning for large-scale traffic signal control,” IEEE Trans. Transp. Syst. , vol. 21, no. 3, pp. 1086–1095, 2019
2019
Later among the works it cites.
J. S. Juul and M. A. Porter, “Hipsters on networks: How a minority group of individuals can lead to an antiestablishment majority,” Phys. Rev. E , vol. 99, p. 022313, 02 2019
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2014
Cited alongside, same era.
M. L. Puterman, Markov decision processes: discrete stochastic dynamic programming . John Wiley & Sons, 2014
2014
Cited alongside, same era.
A. H. Sayed, “Adaptive networks,” Proc. IEEE , vol. 102, no. 4, pp. 460–497, 2014
2014
Cited alongside, same era.
2014
Cited alongside, same era.
N. Şen and P. E. Caines, “Mean field games with partially observed major player and stochastic mean field,” in Proc. IEEE CDC , 2014, pp. 2709–2715
2014
Cited alongside, same era.
M. Gauci, J. Chen, W. Li, T. J. Dodd, and R. Groß, “Self-organized aggregation without computation,” Int. J. Robot. Res. , vol. 33, no. 8, pp. 1145–1161, 2014
2014
Cited alongside, same era.
R. Fujisawa, S. Dobata, K. Sugawara, and F. Matsuno, “Designing pheromone communication in swarm robotics: Group foraging behavior mediated by chemical substance,” Swarm Intell. , vol. 8, no. 3, pp. 227–246, 2014
2014
Cited alongside, same era.
D. Câmara, “Cavalry to the rescue: Drones fleet to help rescuers operations over disasters scenarios,” in Proc. IEEE CAMA , 2014, pp. 1–4
2014
Cited alongside, same era.
I. Iacopini, G. Petri, A. Barrat, and V. Latora, “Simplicial models of social contagion,” Nature Commun. , vol. 10, no. 1, pp. 1–9, 2019
2019
Later among the works it cites.
J. Schrittwieser, I. Antonoglou, T. Hubert, K. Simonyan, L. Sifre, S. Schmitt, A. Guez, E. Lockhart, D. Hassabis, T. Graepel et al. , “Mastering atari, go, chess and shogi by planning with a learned model,” Nature , vol. 588, no. 7839, pp. 604–609, 2020
2020
Later among the works it cites.
D. Wang, T. Fan, T. Han, and J. Pan, “A two-stage reinforcement learning approach for multi-UAV collision avoidance under imperfect sensing,” IEEE Robot. Automat. Lett. , vol. 5, no. 2, pp. 3098–3105, 2020
2020
Later among the works it cites.
P. Muller, S. Omidshafiei, M. Rowland, K. Tuyls, J. Perolat, S. Liu, D. Hennes, L. Marris, M. Lanctot, E. Hughes et al. , “A generalized training approach for multiagent learning,” in Proc. ICLR , 2020, pp. 1–35
2020
Later among the works it cites.
L. Wang, Z. Yang, and Z. Wang, “Breaking the curse of many agents: Provable mean embedding Q-iteration for mean-field reinforcement learning,” in Proc. ICML , 2020, pp. 10 092–10 103
2020
Later among the works it cites.
T. Rashid, G. Farquhar, B. Peng, and S. Whiteson, “Weighted QMIX: Expanding monotonic value function factorisation for deep multi-agent reinforcement learning,” in Proc. NeurIPS , vol. 33, 2020, pp. 10 199–10 210
2020
Later among the works it cites.
W. Böhmer, V. Kurin, and S. Whiteson, “Deep coordination graphs,” in Proc. ICML , 2020, pp. 980–991
2020
Later among the works it cites.
S. Perrin, J. Pérolat, M. Laurière, M. Geist, R. Elie, and O. Pietquin, “Fictitious play for mean field games: Continuous time analysis and applications,” in Proc. NeurIPS , vol. 33, 2020, pp. 13 199–13 213
2020
Later among the works it cites.
D. Linzner and H. Koeppl, “A variational perturbative approach to planning in graph-based Markov decision processes,” in Proc. AAAI , vol. 34, no. 05, 2020, pp. 7203–7210
2020
Later among the works it cites.
G. Qu, A. Wierman, and N. Li, “Scalable reinforcement learning of localized policies for multi-agent networked systems,” in Proc. Learn. Dyn. Contr. , 2020, pp. 256–266
2020
Later among the works it cites.
N. W. Landry and J. G. Restrepo, “The effect of heterogeneity on hypergraph contagion models,” Chaos , vol. 30, no. 10, p. 103117, 2020
2020
Later among the works it cites.
M. A. Porter, “Nonlinearity+ networks: A 2020 vision,” in Emerging frontiers in nonlinear science . Springer, 2020, pp. 131–159
2020
Later among the works it cites.
F. Battiston, G. Cencetti, I. Iacopini, V. Latora, M. Lucas, A. Patania, J.-G. Young, and G. Petri, “Networks beyond pairwise interactions: structure and dynamics,” Phys. Rep. , vol. 874, pp. 1–92, 2020
2020
Later among the works it cites.
A. Bitaillou, B. Parrein, and G. Andrieux, “New results on Q-routing protocol for wireless networks,” in Proc. Int. Conf. Ad Hoc Netw. , 2020, pp. 29–43
2020
Later among the works it cites.
S. Ganapathi Subramanian, P. Poupart, M. E. Taylor, and N. Hegde, “Multi type mean field reinforcement learning,” in Proc. AAMAS , vol. 19, 2020, pp. 411–419
2020
Later among the works it cites.
A. Iglesias, A. Gálvez, and P. Suárez, “Swarm robotics–a case study: bat robotics,” in Nature-Inspired Computation and Swarm Intelligence . Elsevier, 2020, pp. 273–302
2020
Later among the works it cites.
S. Camazine, J.-L. Deneubourg, N. R. Franks, J. Sneyd, G. Theraula, and E. Bonabeau, “Self-organization in biological systems,” in Self-Organization in Biological Systems . Princeton university press, 2020
2020
Later among the works it cites.
W. Lee, N. Vaughan, and D. Kim, “Task allocation into a foraging task with a series of subtasks in swarm robotic system,” IEEE Access , vol. 8, pp. 107 549–107 561, 2020
2020
Later among the works it cites.
S. Kar, R. Rehrmann, A. Mukhopadhyay, B. Alt, F. Ciucu, H. Koeppl, C. Binnig, and A. Rizk, “On the throughput optimization in large-scale batch-processing systems,” Perform. Eval. , vol. 144, p. 102142, 2020
2020
Later among the works it cites.
W. R. KhudaBukhsh, S. Kar, B. Alt, A. Rizk, and H. Koeppl, “Generalized cost-based job scheduling in very large heterogeneous cluster systems,” IEEE Trans. Parallel Distrib. Syst. , vol. 31, no. 11, pp. 2594–2604, 2020
2020
Later among the works it cites.
M. Andrews, S. Borst, J. Lee, E. Martin-Lopez, and K. Palyutina, “Tracking the state of large dynamic networks via reinforcement learning,” in Proc. IEEE INFOCOM , 2020, pp. 416–425
2020
Later among the works it cites.
F. Wei, G. Feng, Y. Sun, Y. Wang, and Y.-C. Liang, “Dynamic network slice reconfiguration by exploiting deep reinforcement learning,” in Proc. ICC , 2020, pp. 1–6
2020
Later among the works it cites.
F. Wei, G. Feng, Y. Sun, Y. Wang, S. Qin, and Y.-C. Liang, “Network slice reconfiguration by exploiting deep reinforcement learning with large action space,” IEEE Trans. Netw. Service Manage. , vol. 17, no. 4, pp. 2197–2211, 2020
2020
Later among the works it cites.
S. Brandi, M. S. Piscitelli, M. Martellacci, and A. Capozzoli, “Deep reinforcement learning to optimise indoor temperature control and heating energy consumption in buildings,” Energy Build. , vol. 224, p. 110225, 2020
2020
Later among the works it cites.
A. Tampuu, T. Matiisen, M. Semikin, D. Fishman, and N. Muhammad, “A survey of end-to-end driving: Architectures and training methods,” IEEE Trans. Neural Netw. Learn. Syst. , vol. 33, no. 4, pp. 1364–1384, 2020
2020
Later among the works it cites.
T. Tanaka, E. Nekouei, A. R. Pedram, and K. H. Johansson, “Linearly solvable mean-field traffic routing games,” IEEE Trans. Automat. Contr. , vol. 66, no. 2, pp. 880–887, 2020
2020
Later among the works it cites.
Y. Xie, Z. Wang, J. Lu, and Y. Li, “Stability analysis and control strategies for a new SIS epidemic model in heterogeneous networks,” Appl. Math. Comput. , vol. 383, p. 125381, 2020
2020
Later among the works it cites.
Z. Abbasi, I. Zamani, A. H. A. Mehra, M. Shafieirad, and A. Ibeas, “Optimal control design of impulsive sqeiar epidemic models with application to COVID-19,” Chaos, Solitons & Fractals , vol. 139, p. 110054, 2020
2020
Later among the works it cites.
A. Q. Ohi, M. Mridha, M. M. Monowar, M. Hamid et al. , “Exploring optimal control of epidemic spread using reinforcement learning,” Sci. Rep. , vol. 10, no. 1, pp. 1–19, 2020
2020
Later among the works it cites.
Z. Wu, S. Pan, F. Chen, G. Long, C. Zhang, and S. Y. Philip, “A comprehensive survey on graph neural networks,” IEEE Trans. Neural Netw. Learn. Syst. , vol. 32, no. 1, pp. 4–24, 2020
2020
Later among the works it cites.
L. Ruiz, L. Chamon, and A. Ribeiro, “Graphon neural networks and the transferability of graph neural networks,” in Proc. NeurIPS , vol. 33, 2020, pp. 1702–1712
2020
Later among the works it cites.
M. Everett, Y. F. Chen, and J. P. How, “Collision avoidance in pedestrian-rich environments with deep reinforcement learning,” IEEE Access , vol. 9, pp. 10 357–10 377, 2021
2021
Later among the works it cites.
A. Charpentier, R. Elie, and C. Remlinger, “Reinforcement learning in economics and finance,” Comput. Econ. , pp. 1–38, 2021
2021
Later among the works it cites.
K. Zhang, Z. Yang, and T. Başar, “Multi-agent reinforcement learning: A selective overview of theories and algorithms,” in Handbook of Reinforcement Learning and Control , K. G. Vamvoudakis, Y. Wan, F. L. Lewis, and D. Cansever, Eds. Cham: Springer International Publishing, 2021, pp. 321–384
2021
Later among the works it cites.
G. Papoudakis, F. Christianos, L. Schäfer, and S. V. Albrecht, “Benchmarking multi-agent deep reinforcement learning algorithms in cooperative tasks,” in Proc. NeurIPS Track Datasets Benchmarks , 2021
2021
Later among the works it cites.
S. Perrin, M. Laurière, J. Pérolat, M. Geist, R. Élie, and O. Pietquin, “Mean field games flock! The reinforcement learning way,” in Proc. IJCAI , 2021, pp. 356–362
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
D. Vasal, R. Mishra, and S. Vishwanath, “Sequential decomposition of graphon mean field games,” in Proc. IEEE ACC , 2021, pp. 730–736
2021
Later among the works it cites.
——, “Mean-field controls with Q-learning for cooperative MARL: convergence and complexity analysis,” SIAM J. Math. Data Sci. , vol. 3, no. 4, pp. 1168–1196, 2021
2021
Later among the works it cites.
Y. Lin, G. Qu, L. Huang, and A. Wierman, “Multi-agent reinforcement learning in stochastic networked systems,” in Proc. NeurIPS , 2021, pp. 7825–7837
2021
Later among the works it cites.
J. Noonan and R. Lambiotte, “Dynamics of majority rule on hypergraphs,” Phys. Rev. E , vol. 104, no. 2, p. 024316, 2021
2021
Later among the works it cites.
F. Battiston, E. Amico, A. Barrat, G. Bianconi, G. Ferraz de Arruda, B. Franceschiello, I. Iacopini, S. Kéfi, V. Latora, Y. Moreno et al. , “The physics of higher-order interactions in complex systems,” Nature Phys. , vol. 17, no. 10, pp. 1093–1098, 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
P. Cong, Y. Zhang, Z. Liu, T. Baker, H. Tawfik, W. Wang, K. Xu, R. Li, and F. Li, “A deep reinforcement learning-based multi-optimality routing scheme for dynamic IoT networks,” Comput. Netw. , vol. 192, p. 108057, 2021
2021
Later among the works it cites.
S. Ganapathi Subramanian, M. E. Taylor, M. Crowley, and P. Poupart, “Partially observable mean field reinforcement learning,” in Proc. AAMAS , vol. 20, 2021, pp. 537–545
2021
Later among the works it cites.
2021
Later among the works it cites.
P. E. Caines, “Mean field games,” in Encyclopedia of systems and control . Springer, 2021, pp. 1197–1202
2021
Later among the works it cites.
2021
Later among the works it cites.
B. Anahtarci, C. D. Kariksiz, and N. Saldi, “Learning in discrete-time average-cost mean-field games,” in Proc. IEEE CDC , 2021, pp. 3048–3053
2021
Later among the works it cites.
2021
Later among the works it cites.
P. Muller, M. Rowland, R. Elie, G. Piliouras, J. Perolat, M. Lauriere, R. Marinier, O. Pietquin, and K. Tuyls, “Learning equilibria in mean-field games: Introducing mean-field PSRO,” in Proc. AAMAS , vol. 20, 2021, p. 926–934
2021
Later among the works it cites.
2021
Later among the works it cites.
K. Cui, A. Tahir, M. Sinzger, and H. Koeppl, “Discrete-time mean field control with environment states,” in Proc. IEEE CDC , 2021, pp. 5239–5246
2021
Later among the works it cites.
T. Zheng, Q. Han, and H. Lin, “Transporting robotic swarms via mean-field feedback control,” IEEE Trans. Automat. Contr. , vol. 67, no. 8, pp. 4170–4177, 2021
2021
Later among the works it cites.
R. Carmona, D. Cooney, C. Graves, and M. Lauriere, “Stochastic graphon games: I. the static case,” Math. Oper. Res. , vol. 47, no. 1, pp. 750–778, 2021
2021
Later among the works it cites.
D. Lacker and A. Soret, “A case study on stochastic games on large graphs in mean field and sparse regimes,” Math. Oper. Res. , vol. 47, no. 2, pp. 1530–1565, 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
M. Schranz, G. A. Di Caro, T. Schmickl, W. Elmenreich, F. Arvin, A. Şekercioğlu, and M. Sende, “Swarm intelligence and cyber-physical systems: Concepts, challenges and future trends,” Swarm Evol. Comput. , vol. 60, p. 100762, 2021
2021
Later among the works it cites.
D. Troullinos, G. Chalkiadakis, I. Papamichail, and M. Papageorgiou, “Collaborative multiagent decision making for lane-free autonomous driving,” in Proc. AAMAS , vol. 20, 2021, pp. 1335–1343
2021
Later among the works it cites.
M. Wang, L. Wu, J. Li, and L. He, “Traffic signal control with reinforcement learning based on region-aware cooperative strategy,” IEEE Trans. Transp. Syst. , 2021
2021
Later among the works it cites.
K. Huang, X. Chen, X. Di, and Q. Du, “Dynamic driving and routing games for autonomous vehicles on networks: A mean field game approach,” Transp. Res. C Emerg. Technol. , vol. 128, p. 103189, 2021
2021
Later among the works it cites.
K. Sugishita, M. A. Porter, M. Beguerisse-Díaz, and N. Masuda, “Opinion dynamics on tie-decay networks,” Phys. Rev. Res. , vol. 3, p. 023249, 06 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
K. Tran and G. Yin, “Optimal control and numerical methods for hybrid stochastic SIS models,” Nonlinear Analysis: Hybrid Systems , vol. 41, p. 101051, 2021
2021
Later among the works it cites.
R. Capobianco, V. Kompella, J. Ault, G. Sharon, S. Jong, S. Fox, L. Meyers, P. R. Wurman, and P. Stone, “Agent-based Markov modeling for improved COVID-19 mitigation policies,” J. Artif. Intell. Res. , vol. 71, pp. 953–992, 2021
2021
Later among the works it cites.
B. R. Kiran, I. Sobh, V. Talpaert, P. Mannion, A. A. Al Sallab, S. Yogamani, and P. Pérez, “Deep reinforcement learning for autonomous driving: A survey,” IEEE Trans. Transp. Syst. , pp. 4909–4926, 2022
2022
Closest in time.
M. M. Afsar, T. Crump, and B. Far, “Reinforcement learning based recommender systems: A survey,” ACM Comput. Surv. , pp. 1–37, 2022
2022
Closest in time.
C. Jin, Q. Liu, Y. Wang, and T. Yu, “V-learning–a simple, efficient, decentralized algorithm for multiagent RL,” in Proc. ICLR 2022 Workshop Gamification Multiagent Solut. , 2022
2022
Closest in time.
W. Fu, C. Yu, Z. Xu, J. Yang, and Y. Wu, “Revisiting some common practices in cooperative multi-agent reinforcement learning,” in Proc. ICML , 2022, pp. 6863–6877
2022
Closest in time.
J. Suarez, Y. Du, I. Mordach, and P. Isola, “Neural MMO v1. 3: A massively multiagent game environment for training and evaluating neural networks,” in Proc. AAMAS , vol. 19, 2020, pp. 2020–2022
2022
Closest in time.
B. Anahtarci, C. D. Kariksiz, and N. Saldi, “Q-learning in regularized mean-field games,” Dyn. Games and Appl. , pp. 1–29, 2022
2022
Closest in time.
J. Pérolat, S. Perrin, R. Elie, M. Laurière, G. Piliouras, M. Geist, K. Tuyls, and O. Pietquin, “Scaling mean field games by online mirror descent,” in Proc. AAMAS , vol. 21, 2022, pp. 1028–1037
2022
Closest in time.
M. Lauriere, S. Perrin, S. Girgin, P. Muller, A. Jain, T. Cabannes, G. Piliouras, J. Perolat, R. Elie, O. Pietquin, and M. Geist, “Scalable deep reinforcement learning algorithms for mean field games,” in Proc. ICML , 2022, pp. 12 078–12 095
2022
Closest in time.
K. Cui and H. Koeppl, “Learning graphon mean field games and approximate Nash equilibria,” in Proc. ICLR , 2022, pp. 1–31
2022
Closest in time.
X.-J. Xu, S. He, and L.-J. Zhang, “Dynamics of the threshold model on hypergraphs,” Chaos , vol. 32, no. 2, p. 023125, 2022
2022
Closest in time.
2022
Closest in time.
N. Nezamoddini and A. Gholami, “A survey of adaptive multi-agent networks and their applications in smart cities,” Smart Cities , vol. 5, no. 1, pp. 318–347, 2022
2022
Closest in time.
S. G. Subramanian, M. E. Taylor, M. Crowley, and P. Poupart, “Decentralized mean field games,” in Proc. AAAI , vol. 36, no. 9, 2022, pp. 9439–9447
2022
Closest in time.
2022
Closest in time.
Y. Chen, L. Zhang, J. Liu, and S. Hu, “Individual-level inverse reinforcement learning for mean field games,” in Proc. AAMAS , 2022, p. 253–262
2022
Closest in time.
X. Guo, A. Hu, R. Xu, and J. Zhang, “A general framework for learning mean-field games,” Math. Oper. Res. , 2022
2022
Closest in time.
X. Guo, R. Xu, and T. Zariphopoulou, “Entropy regularization for mean field games with learning,” Math. Oper. Res. , 2022
2022
Closest in time.
2022
Closest in time.
L. Campi and M. Fischer, “Correlated equilibria and mean field games: a simple model,” Math. Oper. Res. , 2022
2022
Closest in time.
M. Motte and H. Pham, “Mean-field Markov decision processes with common noise and open-loop controls,” Ann. Appl. Probab. , vol. 32, no. 2, pp. 1421–1458, 2022
2022
Closest in time.
W. U. Mondal, M. Agarwal, V. Aggarwal, and S. V. Ukkusuri, “On the approximation of cooperative heterogeneous multi-agent reinforcement learning (MARL) using mean field control (MFC),” J. Mach. Learn. Res. , vol. 23, no. 129, pp. 1–46, 2022
2022
Closest in time.
2022
Closest in time.
K. Cui, W. R. KhudaBukhsh, and H. Koeppl, “Motif-based mean-field approximation of interacting particles on clustered networks,” Phys. Rev. E , vol. 105, p. L042301, 2022
2022
Closest in time.
2022
Closest in time.
M. A. Gkogkas and C. Kuehn, “Graphop mean-field limits for Kuramoto-type models,” SIAM J. Appl. Dyn. Syst. , vol. 21, no. 1, pp. 248–283, 2022
2022
Closest in time.
T. Nie and K. Yan, “Extended mean-field control problem with partial observation,” ESAIM Control Optim. Calc. Var. , vol. 28, p. 17, 2022
2022
Closest in time.
J. Su, S. Yu, B. Li, and Y. Ye, “Distributed and collective intelligence for computation offloading in aerial edge networks,” IEEE Trans. Transp. Syst. , 2022
2022
Closest in time.
R. Ourari, K. Cui, A. Elshamanhory, and H. Koeppl, “Nearest-neighbor-based collision avoidance for quadrotors via reinforcement learning,” in Proc. IEEE ICRA , 2022, pp. 293–300
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
T. Cabannes, M. Laurière, J. Perolat, R. Marinier, S. Girgin, S. Perrin, O. Pietquin, A. M. Bayen, E. Goubault, and R. Elie, “Solving n-player dynamic routing games with congestion: A mean-field approach,” in Proc. AAMAS , vol. 21, 2022, pp. 1557–1559
2022
Closest in time.
A. Aurell, R. Carmona, G. Dayanıklı, and M. Laurière, “Finite state graphon games with applications to epidemics,” Dyn. Games and Appl. , vol. 12, no. 1, pp. 49–81, 2022
2022
Closest in time.
S. Perrin, M. Laurière, J. Pérolat, R. Élie, M. Geist, and O. Pietquin, “Generalization in mean field games by learning master policies,” in Proc. AAAI , vol. 36, no. 9, 2022, pp. 9413–9421
2022
Closest in time.
2022
Closest in time.
G. Qu, Y. Lin, A. Wierman, and N. Li, “Scalable multi-agent reinforcement learning for networked systems with average reward,” in Proc. NeurIPS , vol. 33, 2020, pp. 2074–2086
2086
Closest in time.
P. Sunehag, G. Lever, A. Gruslys, W. M. Czarnecki, V. Zambaldi, M. Jaderberg, M. Lanctot, N. Sonnerat, J. Z. Leibo, K. Tuyls et al. , “Value-decomposition networks for cooperative multi-agent learning based on team reward,” in Proc. AAMAS , vol. 17, 2018, pp. 2085–2087
2087
Closest in time.