Fetching the paper…
Reading the bibliography…
Communication is a critical factor for the big multi-agent world to stay organized and productive.
Mnih V, Badia A P, Mirza M, et al. Asynchronous methods for deep reinforcement learning[C]//International Conference on Machine Learning. 2016: 1928-1937
1937
Earlier work this paper cites.
Williams, Ronald J. Simple statistical gradient-following algorithms for connectionist reinforcement learning. Machine learning 8.3-4 (1992): 229-256
1992
Earlier work this paper cites.
Tan M. Multi-agent reinforcement learning: Independent vs. cooperative agents[C]//Proceedings of the tenth international conference on machine learning. 1993: 330-337
1993
Earlier work this paper cites.
Vicisano L, Crowcroft J, Rizzo L. TCP-like congestion control for layered multicast data transfer[C]//INFOCOM’98. Seventeenth Annual Joint Conference of the IEEE Computer and Communications Societies. Proceedings. IEEE. IEEE, 1998, 3: 996-1003
1998
Earlier work this paper cites.
Sutton R S, Barto A G. Introduction to reinforcement learning[M]. Cambridge: MIT Press, 1998
1998
Earlier work this paper cites.
Konda, Vijay R., and John N. Tsitsiklis. Actor-critic algorithms. Advances in neural information processing systems. 2000
2000
Earlier work this paper cites.
Elwalid A, Jin C, Low S, et al. MATE: MPLS adaptive traffic engineering[C]//INFOCOM 2001. Twentieth Annual Joint Conference of the IEEE Computer and Communications Societies. Proceedings. IEEE. IEEE, 2001, 3: 1300-1309
2001
Earlier work this paper cites.
Giles C L, Jim K C. Learning communication for multi-agent systems[C]//Workshop on Radical Agent Concepts. Springer Berlin Heidelberg, 2002: 377-390
2002
Earlier work this paper cites.
Wan C Y, Eisenman S B, Campbell A T. CODA: congestion detection and avoidance in sensor networks[C]//Proceedings of the 1st international conference on Embedded networked sensor systems. ACM, 2003: 266-279
2003
Earlier work this paper cites.
Goldman C V, Zilberstein S. Optimizing information exchange in cooperative multi-agent systems[C]// 2003:137-144
2003
Earlier work this paper cites.
Konda V R, Tsitsiklis J N. Onactor-critic algorithms[J]. SIAM journal on Control and Optimization, 2003, 42(4): 1143-1166
2003
Earlier work this paper cites.
Goldman C V, Zilberstein S. Decentralized control of cooperative systems: categorization and complexity analysis[M]. AI Access Foundation, 2004
2004
Earlier work this paper cites.
Kandula S, Katabi D, Davie B, et al. Walking the tightrope: Responsive yet stable traffic engineering[C]//ACM SIGCOMM Computer Communication Review. ACM, 2005, 35(4): 253-264
2005
Earlier work this paper cites.
Lee, Honglak, Chaitanya Ekanadham, and Andrew Y. Ng. Sparse deep belief net model for visual area V2. Advances in neural information processing systems. 2008
2008
Earlier work this paper cites.
Agogino A K, Tumer K. A multiagent approach to managing air traffic flow[J]. Autonomous Agents and Multi-Agent Systems, 2012, 24(1): 1-25
2012
Cited alongside, same era.
Oliehoek F A. Decentralized POMDPs[M]// Reinforcement Learning. Springer Berlin Heidelberg, 2012:471-503
2012
Cited alongside, same era.
Grondman I, Busoniu L, Lopes G A D, et al. A survey of actor-critic reinforcement learning: Standard and natural policy gradients[J]. IEEE Transactions on Systems, Man, and Cybernetics, Part C (Applications and Reviews), 2012, 42(6): 1291-1307
2012
Cited alongside, same era.
Grondman I, Busoniu L, Lopes G A D, et al. A survey of actor-critic reinforcement learning: Standard and natural policy gradients[J]. IEEE Transactions on Systems, Man, and Cybernetics, Part C (Applications and Reviews), 2012, 42(6): 1291-1307
2012
Cited alongside, same era.
Oliehoek F A, Amato C. A concise introduction to decentralized POMDPs[M]. Springer International Publishing, 2016
2016
Later among the works it cites.
2016
Later among the works it cites.
Sukhbaatar S, Fergus R. Learning multiagent communication with backpropagation[C]//Advances in Neural Information Processing Systems. 2016: 2244-2252
2016
Later among the works it cites.
2016
Later among the works it cites.
Tampuu A, Matiisen T, Kodelja D, et al. Multiagent cooperation and competition with deep reinforcement learning[J]. PloS one, 2017, 12(4): e0172395
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2013
Cited alongside, same era.
Lever G. Deterministic policy gradient algorithms[J]. 2014
2014
Cited alongside, same era.
2015
Cited alongside, same era.
Mnih V, Kavukcuoglu K, Silver D, et al. Human-level control through deep reinforcement learning[J]. Nature, 2015, 518(7540): 529-533
2015
Cited alongside, same era.
2015
Cited alongside, same era.
2015
Cited alongside, same era.
2015
Cited alongside, same era.
Makhzani, Alireza, and Brendan J. Frey. Winner-take-all autoencoders. Advances in Neural Information Processing Systems. 2015
2015
Cited alongside, same era.
2017
Closest in time.
2017
Closest in time.
2017
Closest in time.
2017
Closest in time.
2017
Closest in time.
2017
Closest in time.
2017
Closest in time.
Omidshafiei, Shayegan, et al. Deep decentralized multi-task multi-agent reinforcement learning under partial observability. International Conference on Machine Learning. 2017
2017
Closest in time.
2017
Closest in time.