Fetching the paper…
Reading the bibliography…
This paper deals with distributed policy optimization in reinforcement learning, which involves a central controller and a group of learners.
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu, “Asynchronous methods for deep reinforcement learning,” in Proc. Intl. Conf. Machine Learn. , New York City, NY, Jun. 2016, pp. 1928–1937
1937
Earlier work this paper cites.
R. E. Schapire, “The strength of weak learnability,” Machine learning , vol. 5, no. 2, pp. 197–227, 1990
1990
Earlier work this paper cites.
C. J. Watkins and P. Dayan, “Q-learning,” Machine Learn. , vol. 8, no. 3-4, pp. 279–292, May 1992
1992
Earlier work this paper cites.
R. J. Williams, “Simple statistical gradient-following algorithms for connectionist reinforcement learning,” Machine Learning , vol. 8, no. 3-4, pp. 229–256, May 1992
1992
Earlier work this paper cites.
J. A. Boyan and M. L. Littman, “Packet routing in dynamically changing networks: A reinforcement learning approach,” in Proc. Advances in Neural Info. Process. Syst. , Denver, CO, Nov. 1994, pp. 671–678
1994
Earlier work this paper cites.
I. Pinelis, “Optimum bounds for the distributions of martingales in banach spaces,” The Annals of Probability , vol. 22, no. 4, pp. 1679–1706, Oct. 1994
1994
Earlier work this paper cites.
C. Claus and C. Boutilier, “The dynamics of reinforcement learning in cooperative multiagent systems,” in Proc. of the Assoc. for the Advanc. of Artificial Intell. , Orlando, FL, Oct. 1998, pp. 746–752
1998
Earlier work this paper cites.
D. H. Wolpert, K. R. Wheeler, and K. Tumer, “General principles of learning-based multi-agent systems,” in Proc. of the Annual Conf. on Autonomous Agents , Seattle, WA, May 1999, pp. 77–83
1999
Earlier work this paper cites.
J. Schneider, W.-K. Wong, A. Moore, and M. Riedmiller, “Distributed value functions,” in Proc. Intl. Conf. Machine Learn. , Bled, Slovenia, Jun. 1999, pp. 371–378
1999
Earlier work this paper cites.
R. S. Sutton, D. A. McAllester, S. P. Singh, and Y. Mansour, “Policy gradient methods for reinforcement learning with function approximation,” in Proc. Advances in Neural Info. Process. Syst. , Denver, CO, Dec. 2000, pp. 1057–1063
2000
Earlier work this paper cites.
V. R. Konda and J. N. Tsitsiklis, “Actor-critic algorithms,” in Proc. Advances in Neural Info. Process. Syst. , Denver, CO, Dec. 2000, pp. 1008–1014
2000
Earlier work this paper cites.
P. Stone and M. Veloso, “Multiagent systems: A survey from a machine learning perspective,” Autonomous Robots , vol. 8, no. 3, pp. 345–383, Jun. 2000
2000
Earlier work this paper cites.
M. Lauer and M. Riedmiller, “An algorithm for distributed reinforcement learning in cooperative multi-agent systems,” in Proc. Intl. Conf. Machine Learn. , Stanford, CA, Jun. 2000
2000
Earlier work this paper cites.
J. Baxter and P. L. Bartlett, “Infinite-horizon policy-gradient estimation,” J. Artificial Intelligence Res. , vol. 15, pp. 319–350, 2001
2001
Earlier work this paper cites.
S. M. Kakade, “A natural policy gradient,” in Proc. Advances in Neural Info. Process. Syst. , Vancouver, Canada, Dec. 2002, pp. 1531–1538
2002
Earlier work this paper cites.
D. S. Bernstein, R. Givan, N. Immerman, and S. Zilberstein, “The complexity of decentralized control of markov decision processes,” Mathematics of Operations Research , vol. 27, no. 4, pp. 819–840, Nov. 2002
2002
Earlier work this paper cites.
M. P. Deisenroth, Efficient Reinforcement Learning Using Gaussian Processes . Karlsruhe, Germany: KIT Scientific Publishing, 2010, vol. 9
2010
Earlier work this paper cites.
S. Shalev-Shwartz, O. Shamir, N. Srebro, and K. Sridharan, “Learnability, stability and uniform convergence,” J. Machine Learning Res. , vol. 11, pp. 2635–2670, 2010
2010
Earlier work this paper cites.
B. Recht, C. Re, S. Wright, and F. Niu, “Hogwild: A lock-free approach to parallelizing stochastic gradient descent,” in Proc. Advances in Neural Info. Process. Syst. , Granada, Spain, Dec. 2011, pp. 693–701
2011
Earlier work this paper cites.
Y. Li and D. Schuurmans, “Mapreduce for parallel reinforcement learning,” in European Workshop on Reinforcement Learning . Springer, 2011, pp. 309–320
2011
Cited alongside, same era.
D. V. Dimarogonas, E. Frazzoli, and K. H. Johansson, “Distributed event-triggered control for multi-agent systems,” IEEE Trans. Automatic Control , vol. 57, no. 5, pp. 1291–1297, Nov. 2011
2011
Cited alongside, same era.
S. Kar, J. M. Moura, and H. V. Poor, “QD-learning: A collaborative distributed strategy for multi-agent reinforcement learning through Consensus + Innovations,” IEEE Trans. Sig. Proc. , vol. 61, no. 7, pp. 1848–1862, Jul. 2013
2013
Cited alongside, same era.
A. Nayyar, A. Mahajan, and D. Teneketzis, “Decentralized stochastic control with partial history sharing: A common information approach,” IEEE Trans. Automatic Control , vol. 58, no. 7, pp. 1644–1658, Jan. 2013
2013
Cited alongside, same era.
D. Alistarh, D. Grubic, J. Li, R. Tomioka, and M. Vojnovic, “Qsgd: Communication-efficient sgd via gradient quantization and encoding,” in Proc. Advances in Neural Info. Process. Syst. , Long Beach, CA, Dec. 2017, pp. 1709–1720
2017
Later among the works it cites.
M. Papini, M. Pirotta, and M. Restelli, “Adaptive batch size for safe policy gradients,” in Proc. Advances in Neural Info. Process. Syst. , Long beach, CA, Dec. 2017, pp. 3591–3600
2017
Later among the works it cites.
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, “Proximal policy optimization algorithms,” arXiv preprint:1707.06347 , Jul. 2017
2017
Later among the works it cites.
J. K. Gupta, M. Egorov, and M. Kochenderfer, “Cooperative multi-agent control using deep reinforcement learning,” in Intl. Conf. Auto. Agents and Multi-agent Systems , 2017, pp. 66–83
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
V. Vapnik, The Nature of Statistical Learning Theory . Berlin, Germany: Springer, 2013
2013
Cited alongside, same era.
M. Li, D. G. Andersen, J. W. Park, A. J. Smola, A. Ahmed, V. Josifovski, J. Long, E. J. Shekita, and B.-Y. Su, “Scaling distributed machine learning with the parameter server,” in Proc. USENIX Symp. Operating Syst. Design and Implement. , vol. 14, Broomfield, CO, Oct. 2014, pp. 583–598
2014
Cited alongside, same era.
D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. Riedmiller, “Deterministic policy gradient algorithms,” in Proc. Intl. Conf. Machine Learn. , Beijing, China, Jun. 2014
2014
Cited alongside, same era.
A. Tampuu, T. Matiisen, D. Kodelja, I. Kuzovkin, K. Korjus, J. Aru, J. Aru, and R. Vicente, “Multiagent cooperation and competition with deep reinforcement learning,” arXiv preprint:1507.04296 , Jul. 2015
2015
Cited alongside, same era.
A. Nair, P. Srinivasan, S. Blackwell, C. Alcicek, R. Fearon, A. De Maria, V. Panneershelvam, M. Suleyman, C. Beattie, S. Petersen et al. , “Massively parallel methods for deep reinforcement learning,” arXiv preprint:1507.04296 , 2015
2015
Cited alongside, same era.
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz, “Trust region policy optimization,” in Proc. Intl. Conf. Machine Learn. , Lille, France, Jul. 2015, pp. 1889–1897
2015
Cited alongside, same era.
Y. Zhang and X. Lin, “DiSCO: Distributed optimization for self-concordant empirical loss,” in Proc. Intl. Conf. Machine Learn. , Lille, France, Jun. 2015, pp. 362–370
2015
Cited alongside, same era.
S. S. Ponda, L. B. Johnson, A. Geramifard, and J. P. How, “Cooperative mission planning for multi-uav teams,” in Handbook of Unmanned Aerial Vehicles . Springer, 2015, pp. 1447–1490
2015
Cited alongside, same era.
R. Lowe, Y. Wu, A. Tamar, J. Harb, P. Abbeel, and I. Mordatch, “Multi-agent actor-critic for mixed cooperative-competitive environments,” in Proc. Advances in Neural Info. Process. Syst. , Long beach, CA, Dec. 2017
2017
Later among the works it cites.
S. Omidshafiei, J. Pazis, C. Amato, J. P. How, and J. Vian, “Deep decentralized multi-task multi-agent reinforcement learning under partial observability,” in Proc. Intl. Conf. Machine Learn. , Sydney, Australia, Jun. 2017, pp. 2681–2690
2017
Later among the works it cites.
Y. Liu, C. Nowzari, Z. Tian, and Q. Ling, “Asynchronous periodic event-triggered coordination of multi-agent systems,” in Proc. IEEE Conf. Decision and Control , Melbourne, Australia, Dec. 2017, pp. 6696–6701
2017
Later among the works it cites.
R. S. Sutton and A. G. Barto, Reinforcement Learning: An Introduction . Cambridge, MA: MIT Press, 2018
2018
Closest in time.
H.-T. Wai, Z. Yang, Z. Wang, and M. Hong, “Multi-agent reinforcement learning via double averaging primal-dual optimization,” in Proc. Advances in Neural Info. Process. Syst. , Montreal, Canada, Dec. 2018
2018
Closest in time.
K. Zhang, Z. Yang, H. Liu, T. Zhang, and T. Başar, “Fully decentralized multi-agent reinforcement learning with networked agents,” in Proc. Intl. Conf. Machine Learn. , Stockholm, Sweden, Jul. 2018, pp. 5872–5881
2018
Closest in time.
M. I. Jordan, J. D. Lee, and Y. Yang, “Communication-efficient distributed statistical inference,” J. American Statistical Association , vol. to appear, 2018
2018
Closest in time.
2018
Closest in time.
M. Papini, D. Binaghi, G. Canonaco, M. Pirotta, and M. Restelli, “Stochastic variance-reduced policy gradient,” in Proc. Intl. Conf. Machine Learn. , Stockholm, Sweden, Jul. 2018, pp. 4026–4035
2018
Closest in time.
S. U. Stich, “Local SGD converges fast and communicates little,” arXiv preprint:1805.09767 , May 2018
2018
Closest in time.
T. Chen and G. B. Giannakis, “Bandit convex optimization for scalable and dynamic IoT management,” IEEE Internet Things J. , vol. 6, no. 1, pp. 1276–1286, Feb. 2019
2019
Closest in time.
C. Nowzari, E. Garcia, and J. Cortés, “Event-triggered communication and control of networked systems for multi-agent consensus,” Automatica , vol. 105, pp. 1–27, Jul. 2019
2019
Closest in time.
S. Paternain, J. Bazerque, A. Small, and A. Ribeiro, “Stochastic policy gradient ascent in reproducing kernel hilbert spaces,” IEEE Trans. Automatic Control , Oct. 2020
2020
Closest in time.
K. Zhang, A. Koppel, H. Zhu, and T. Basar, “Global convergence of policy gradient methods: A nonconvex optimization perspective,” SIAM Journal on control and Optimization , 2020
2020
Closest in time.