Fetching the paper…
Reading the bibliography…
The paper considers a class of multi-agent Markov decision processes (MDPs), in which the network agents respond differently (as manifested by the instantaneous one-stage random costs) to a global controlled state and the control actions of a remote controller.
D. Bertsekas, Dynamic Programming and Stochastic Control . New York, NY: Academic Press, Inc., 1976
1976
Earlier work this paper cites.
J. N. Tsitsiklis, D. P. Bertsekas, and M. Athans, “Distributed asynchronous deterministic and stochastic gradient optimization algorithms,” IEEE Transactions on Automatic Control , vol. 31, no. 9, pp. 803–812, September 1986
1986
Earlier work this paper cites.
J. Jacod and A. Shiryaev, Limit Theorems for Stochastic Processes . Berlin Heidelberg: Springer-Verlag, 1987
1987
Earlier work this paper cites.
B. Mohar, “The Laplacian spectrum of graphs,” in Graph Theory, Combinatorics, and Applications , Y. Alavi, G. Chartrand, O. R. Oellermann, and A. J. Schwenk, Eds. New York: J. Wiley & Sons, 1991, vol. 2, pp. 871–898
1991
Earlier work this paper cites.
S. Yuta and S. Premvuti, “Coordinating autonomous and centralized decision making to achieve cooperative behaviors between multiple mobile robots,” in lEEE/RSJ International Conference on Intelligent Robots and Systems , 7-10 July 1992, pp. 1566–1574
1992
Earlier work this paper cites.
C. Watkins and P. Dayan, “ Q {Q} -learning,” Machine Learning , vol. 8, pp. 279–292, 1992
1992
Earlier work this paper cites.
T. Jaakkola, M. Jordan, and S. Singh, “On the convergence of stochastic iterative dynamic programming algorithms,” Neural Computation , vol. 6, no. 6, pp. 1185–1201, 1994 1992
1992
Earlier work this paper cites.
R. Sutton, A. Barto, and R. Williams, “Reinforcement learning is direct adaptive control,” IEEE Control Systems Magazine , pp. 19 – 22, April 1992
1992
Earlier work this paper cites.
J. Tsitsiklis, “Asynchronous stochastic approximation and Q {Q} -learning,” Machine Learning , vol. 16, pp. 185–202, 1994
1994
Earlier work this paper cites.
M. Littman, “Markov games as a framework for multi-agent reinforcement learning,” in The 11th International Conference on Machine Learning , 1994, pp. 157–163
1994
Earlier work this paper cites.
A. Barto, S. Bradtke, and S. Singh, “Real-time learning and control using asynchronous dynamic programming,” Artificial Intelligence , 1995
1995
Earlier work this paper cites.
M. Littman and C. Szepesvari, “A generalized reinforcement learning model: convergence and applications,” in The 13th International Conference on Machine Learning , 1996, pp. 310–318
1996
Earlier work this paper cites.
M. Veloso, P. Stone, K. Han, and S. Achim, “CMUnited: A team of robotic soccer agents collaborating in an adversarial environment,” in H. Kitano, editor, RoboCup-97: The First Robot World Cup Soccer Games and Conferences . Springer Verlag, 1997, pp. 242–256
1997
Earlier work this paper cites.
F. R. K. Chung, Spectral Graph Theory . Providence, RI : American Mathematical Society, 1997
1997
Earlier work this paper cites.
J. Hu and P. Wellman, “Multiagent reinforcement learning: theoretical framework and an algorithm,” in The 15th International Conference on Machine Learning , 1998, pp. 242–250
1998
Earlier work this paper cites.
C. Claus and C. Boutilier, “The dynamics of reinforcement learning in cooperative multiagent systems,” in The 15th International Conference on Artificial Intelligence , 1998, pp. 746–752
1998
Earlier work this paper cites.
C. Szepesvari, “The asymptotic convergence-rate of Q {Q} -learning,” in Advances in Neural Information Processing Systems , M. Jordan, M. Kearns, and S. Solla, Eds., 1998, vol. 10, p. 1064 1070
1998
Earlier work this paper cites.
M. Lauer and M. Riedmiller, “An algorithm for distributed reinforcement learning in cooperative multi-agent systems,” in The 17th International Conference on Machine Learning , Stanford Univ., Stanford, CA, Jun. 29 - Jul. 2 2000, pp. 535–542
2000
Cited alongside, same era.
M. Bowling and M. Veloso, “Rational and convergent learning in stochastic games,” in The 17th International Joint Conference on Artificial Intelligence , 2001, pp. 1021–1026
2001
Cited alongside, same era.
M. Littman, “Value function reinforcement learning in Markov games,” J. Cogn. Syst. Res. , vol. 2, no. 1, pp. 55–66, 2001
2001
Cited alongside, same era.
C. Guestrin, M. Lagoudakis, and R. Parr, “Coordinated reinforcement learning,” in The 19th International Conference on Machine Learning , Sydney, Australia, Jul. 8-12 2002, pp. 227–234
2002
Cited alongside, same era.
K. Li and J. Baillieul, “Robust and efficient quantization and coding for control of multidimensional linear systems under data rate constraints,” International Journal of Robust and Nonlinear Control Special Issue: Communicating-Agent Networks , vol. 17, no. 10-11, pp. 898–920, July 2007
2007
Later among the works it cites.
2007
Later among the works it cites.
L. Busoniu, R. Babuska, and B. Schutter, “A comprehensive survey of multiagent reinforcement learning,” IEEE Transactions on Systems, Man, and Cybernetics - Part C: Applications and Reviews , vol. 38, no. 2, pp. 156–172, March 2008
2008
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Pynadath and M. Tambe, “The communicative multiagent team decision problem: analyzing teamwork theories and models,” J. Artif. Intell. Res. , vol. 16, pp. 389–423, 2002
2002
Cited alongside, same era.
A. Kapetanakis and D. Kudenko, “Reinforcement learning of coordination in cooperative multi-agent systems,” in The 18th Nat. Conf. Artif. Intell and 14th Conf. Innov. Appl. Artif. Intell. , Menlo Park, CA, Jul. 28 - Aug. 1 2002, pp. 326–331
2002
Cited alongside, same era.
Y. Shoham, R. Powers, and T. Grenager, “Multi-agent reinforcement learning: a critical survey,” May 2003, computer Science Dept., Stanford University, Stanford, CA. [Online]: http://ece.ut.ac.ir/classpages/F85/ControlOfStochasticSystems/res/Multi_Agent_Reinforcement_Learning.pdf
2003
Cited alongside, same era.
A. Jadbabaie, J. Lin, and A. S. Morse, “Coordination of groups of mobile autonomous agents using nearest neighbor rules,” IEEE Transactions on Automatic Control , vol. 48, no. 6, pp. 988–1001, Jun. 2003
2003
Cited alongside, same era.
E. Even-Dar and Y. Mansour, “Learning rates for Q {Q} -learning,” Journal of Machine Learning Research , vol. 5, pp. 1–25, 2003
2003
Cited alongside, same era.
R. Olfati-Saber and R. M. Murray, “Consensus problems in networks of agents with switching topology and time-delays,” IEEE Trans. Automat. Contr. , vol. 49, no. 9, pp. 1520–1533, Sept. 2004
2004
Cited alongside, same era.
S. Tatikonda and S. Mitter, “Control under communication constraints,” IEEE Transactions on Automatic Control , vol. 49, no. 7, pp. 1056 – 1068, July 2004
2004
Cited alongside, same era.
J. Kok, M. Spaan, and N. Vlassis, “Non-communicative multi-robot coordination in dynamic environment,” Robot. Auton. Syst. , vol. 50, no. 2-3, pp. 99–114, 2005
2005
Cited alongside, same era.
2008
Later among the works it cites.
C. G. Lopes and A. H. Sayed, “Diffusion least-mean squares over adaptive networks: Formulation and performance analysis,” IEEE Transactions on Signal Processing , vol. 56, no. 7, pp. 3122–3136, July 2008
2008
Later among the works it cites.
S. Kar and J. M. F. Moura, “Distributed consensus algorithms in sensor networks with imperfect communication: Link failures and channel noise,” IEEE Transactions on Signal Processing , vol. 57, no. 1, pp. 355–369, January 2009
2009
Later among the works it cites.
A. Nedic, A. Olshevsky, A. Ozdaglar, and J. N. Tsitsiklis, “On distributed averaging algorithms and quantization effects,” IEEE Transactions on Automatic Control , no. 11, pp. 2506–2517, Nov. 2009
2009
Later among the works it cites.
A. Nedic and A. Ozdaglar, “Distributed subgradient methods for multi-agent optimization,” IEEE Transactions on Automatic Control , vol. 54, no. 1, p. 48 61, Jan. 2009
2009
Later among the works it cites.
S. S. Ram, A. Nedic, and V. V. Veeravalli, “Incremental stochastic subgradient algorithms for convex optimization,” SIAM Journal on Optimization , vol. 20, no. 2, pp. 691–717, June 2009
2009
Later among the works it cites.
A. G. Dimakis, S. Kar, J. M. F. Moura, M. G. Rabbat, and A. Scaglione, “Gossip algorithms for distributed signal processing,” Proceedings of the IEEE , vol. 98, no. 11, pp. 1847–1864, Nov 2010
2010
Later among the works it cites.
G. Mateos, J. Bazerque, and G. Giannakis, “Distributed sparse linear regression,” IEEE Transactions on Signal Processing , vol. 58, no. 11, pp. 5262–5276, Nov. 2010
2010
Later among the works it cites.
D. Callaway and I. Hiskens, “Achieving controllability of electric loads,” Proceedings of the IEEE , vol. 99, no. 1, pp. 184 – 199, Jan. 2011
2011
Later among the works it cites.
F. Melo and M. Veloso, “Decentralized MDPs with sparse interactions,” Artificial Intelligence , vol. 175, no. 11, pp. 1757–1789, July 2011
2011
Later among the works it cites.
S. Kar and J. M. F. Moura, “Convergence rate analysis of distributed gossip (linear parameter) estimation: Fundamental limits and tradeoffs,” IEEE Journal of Selected Topics in Signal Processing: Signal Processing in Gossiping Algorithms Design and Applications , vol. 5, no. 4, pp. 674–690, August 2011
2011
Later among the works it cites.
2011
Later among the works it cites.
D. Bajovic, D. Jakovetic, J. Moura, J. Xavier, and B. Sinopoli, “Large deviations analysis of consensus+innovations detection in random networks,” in The 49th Annual Allerton Conference on Control, Communication, and Computing , Monticello, IL, Sept. 28 - 30 2011, pp. 151–155
2011
Later among the works it cites.
D. Jakovetic, J. Xavier, and J. Moura, “Cooperative convex optimization in networked systems: Augmented Lagrangian algorithms with directed gossip communication,” IEEE Transactions on Signal Processing , vol. 59, no. 8, pp. 3889–3902, August 2011
2011
Later among the works it cites.