Fetching the paper…
Reading the bibliography…
In this paper, we study the problem of learning to satisfy temporal logic specifications with a group of agents in an unknown environment, which may exhibit probabilistic behaviour.
The Temporal Logic of Programs. In Proceedings of the 18th Annual Symposium on Foundations of Computer Science (FOCS ’77) . IEEE Computer Society, USA, 46–57
Amir Pnueli. 1977 · 1977
Earlier work this paper cites.
Learning to Predict by the Methods of Temporal Differences
Richard S. Sutton. 1988 · 1988
Earlier work this paper cites.
Game Theory
Drew Fudenberg and Jean Tirole. 1991 · 1991
Earlier work this paper cites.
Markov Games As a Framework for Multi-agent Reinforcement Learning. In Proceedings of the Eleventh International Conference on International Conference on Machine Learning (New Brunswick, NJ, USA) (ICML’94) . Morgan Kaufmann Publishers Inc., San Francisco, CA, USA, 157–163
Michael L. Littman. 1994 · 1994
Earlier work this paper cites.
Neuro-dynamic Programming
Dimitri P. Bertsekas and John N. Tsitsiklis. 1996 · 1996
Earlier work this paper cites.
A Theory of Lexicographic Multi-criteria Optimization. In Proceedings of ICECCS 1996: 2nd IEEE International Conference on Engineering of Complex Computer Systems . IEEE Comput. Soc. Press
Mark J. Rentmeesters, Wei K. Tsai, and Kwei-Jay Lin. 1996 · 1996
Earlier work this paper cites.
An Analysis of Temporal-difference Learning with Function Approximation
John N. Tsitsiklis and Benjamin Van Roy. 1997 · 1997
Earlier work this paper cites.
Natural Gradient Works Efficiently in Learning
Shun’ichi Amari. 1998 · 1998
Earlier work this paper cites.
Policy Gradient Methods for Reinforcement Learning with Function Approximation. In Proceedings of the 12th International Conference on Neural Information Processing Systems (Denver, CO) (NIPS’99) . MIT Press, Cambridge, MA, USA, 1057–1063
Richard S. Sutton, David McAllester, Satinder Singh, and Yishay Mansour. 1999 · 1999
Earlier work this paper cites.
Actor-critic Algorithms
Vijay R. Konda and John N. Tsitsiklis. 2000 · 2000
Earlier work this paper cites.
Rational and Convergent Learning in Stochastic Games. In Proceedings of the 17th International Joint Conference on Artificial Intelligence - Volume 2 (Seattle, WA, USA) (IJCAI’01) . Morgan Kaufmann Publishers Inc., San Francisco, CA, USA, 1021–1026
Michael Bowling and Manuela Veloso. 2001 · 2001
Earlier work this paper cites.
A Natural Policy Gradient. In Proceedings of the 14th International Conference on Neural Information Processing Systems: Natural and Synthetic (Vancouver, British Columbia, Canada) (NIPS’01) . MIT Press, Cambridge, MA, USA, 1531–1538
Sham Kakade. 2001 · 2001
Earlier work this paper cites.
Markov Perfect Equilibrium
Eric Maskin and Jean Tirole. 2001 · 2001
Earlier work this paper cites.
Approximately Optimal Approximate Reinforcement Learning. In Proceedings of the Nineteenth International Conference on Machine Learning (ICML ’02) . Morgan Kaufmann Publishers Inc., San Francisco, CA, USA, 267–274
Sham Kakade and John Langford. 2002 · 2002
Earlier work this paper cites.
Reinforcement Learning to Play an Optimal Nash Equilibrium in Team Markov Games. In Proceedings of the 15th International Conference on Neural Information Processing Systems (NIPS’02) . MIT Press, Cambridge, MA, USA, 1603–1610
Xiaofeng Wang and Tuomas Sandholm. 2002 · 2002
Earlier work this paper cites.
Awesome: A General Multiagent Learning Algorithm That Converges in Self-play and Learns a Best Response against Stationary Opponents. In Proceedings of the Twentieth International Conference on International Conference on Machine Learning (ICML’03) . AAAI Press, Washington, DC, USA, 83–90
Vincent Conitzer and Tuomas Sandholm. 2003 · 2003
Earlier work this paper cites.
Martin Zinkevich, Amy Greenwald, and Michael L. Littman. 2005 · 2005
Earlier work this paper cites.
Multi-objective Model Checking of Markov Decision Processes
Kousha Etessami, Marta Kwiatkowska, Moshe Y. Vardi, and Mihalis Yannakakis. 2007 · 2007
Earlier work this paper cites.
Stochastic Approximation
Vivek S. Borkar. 2008 · 2008
Earlier work this paper cites.
A Comprehensive Survey of Multiagent Reinforcement Learning
Lucian Busoniu, Robert Babuska, and Bart De Schutter. 2008 · 2008
Earlier work this paper cites.
Natural Actor-critic
Jan Peters and Stefan Schaal. 2008 · 2008
Cited alongside, same era.
Natural Actor–critic Algorithms
Shalabh Bhatnagar, Richard S. Sutton, Mohammad Ghavamzadeh, and Mark Lee. 2009 · 2009
Cited alongside, same era.
Online Markov Decision Processes
Eyal Even-Dar, Sham. M. Kakade, and Yishay Mansour. 2009 · 2009
Cited alongside, same era.
Rational Synthesis
Dana Fisman, Orna Kupferman, and Yoad Lustig. 2010 · 2010
Cited alongside, same era.
Prism 4.0: Verification of Probabilistic Real-time Systems. In Proc. 23rd International Conference on Computer Aided Verification (CAV’11) (LNCS, Vol. 6806) , G. Gopalakrishnan and S. Qadeer (Eds.). Springer, 585–591
Marta Kwiatkowska, Gethin Norman, and David Parker. 2011 · 2011
Cited alongside, same era.
Game Theory and Multi-agent Reinforcement Learning
Ann Nowé, Peter Vrancx, and Yann-Michaël De Hauwere. 2012 · 2012
Automata Guided Reinforcement Learning with Demonstrations
Xiao Li, Yao Ma, and Calin Belta. 2018 · 2018
Later among the works it cites.
Actor-critic Fictitious Play in Simultaneous Move Multistage Games. In Proceedings of the Twenty-First International Conference on Artificial Intelligence and Statistics (Proceedings of Machine Learning Research, Vol. 84) , Amos Storkey and Fernando Perez-Cruz (Eds.). PMLR, Playa Blanca, Lanzarote, Canary Islands, 919–928
Julien Perolat, Bilal Piot, and Olivier Pietquin. 2018 · 2018
Later among the works it cites.
Reinforcement Learning
Richard S. Sutton and Andrew G. Barto. 2018 · 2018
Later among the works it cites.
Fully Decentralized Multi-agent Reinforcement Learning with Networked Agents. In Proceedings of the 35th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 80) , Jennifer Dy and Andreas Krause (Eds.). PMLR, Stockholmsmässan, Stockholm Sweden, 5872–5881
Kaiqing Zhang, Zhuoran Yang, Han Liu, Tong Zhang, and Tamer Basar. 2018 · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Verification of Markov Decision Processes Using Learning Algorithms
Tomáš Brázdil, Krishnendu Chatterjee, Martin Chmelík, Vojtěch Forejt, Jan Křetiínský, Marta Kwiatkowska, David Parker, and Mateusz Ujma. 2014 · 2014
Cited alongside, same era.
Probably Approximately Correct Mdp Learning and Control with Temporal Logic Constraints. In Robotics: Science and Systems X, University of California, Berkeley, USA, July 12-16, 2014 , Dieter Fox, Lydia E. Kavraki, and Hanna Kurniawati (Eds.)
Jie Fu and Ufuk Topcu. 2014 · 2014
Cited alongside, same era.
A Learning Based Approach to Control Synthesis of Markov Decision Processes for Linear Temporal Logic Specifications. In 53rd IEEE Conference on Decision and Control, CDC 2014, Los Angeles, CA, USA, December 15-17, 2014 . 1091–1096
Dorsa Sadigh, Eric S. Kim, Samuel Coogan, S. Shankar Sastry, and Sanjit A. Seshia. 2014 · 2014
Cited alongside, same era.
Bias in Natural Actor-critic Algorithms. In Proceedings of the 31st International Conference on International Conference on Machine Learning - Volume 32 (ICML’14) . JMLR.org, Beijing, China, I–441–I–448
Philip S. Thomas. 2014 · 2014
Cited alongside, same era.
In Proceedings of the 2015 International Conference on Autonomous Agents and Multiagent Systems (Istanbul, Turkey) (AAMAS ’15) . International Foundation for Autonomous Agents and Multiagent Systems, Richland, SC, 1371–1379
H.L. Prasad, Prashanth L.A., and Shalabh Bhatnagar. 2015 · 2015
Cited alongside, same era.
Limit-deterministic Büchi Automata for Linear Temporal Logic
Salomon Sickert, Javier Esparza, Stefan Jaax, and Jan Křetínský. 2016 · 2016
Cited alongside, same era.
Later among the works it cites.
Pac Statistical Model Checking for Markov Decision Processes and Stochastic Games
Pranav Ashok, Jan Křetínský, and Maximilian Weininger. 2019 · 2019
Later among the works it cites.
Control Synthesis from Linear Temporal Logic Specifications Using Model-free Reinforcement Learning
Alper Kamil Bozkurt, Yu Wang, Michael M. Zavlanos, and Miroslav Pajic. 2019 · 2019
Later among the works it cites.
Omega-regular Objectives in Model-free Reinforcement Learning
Ernst M. Hahn, Mateo Perez, Sven Schewe, Fabio Somenzi, Ashutosh Trivedi, and Dominik Wojtczak. 2019 · 2019
Later among the works it cites.
Certified Reinforcement Learning with Logic Guidance
Mohammadhosein Hasanbeig, Alessandro Abate, and Daniel Kroening. 2019 · 2019
Later among the works it cites.
A Composable Specification Language for Reinforcement Learning Tasks. In Advances in Neural Information Processing Systems 32 , H. Wallach, H. Larochelle, A. Beygelzimer, F. d Alché-Buc, E. Fox, and R. Garnett (Eds.). Curran Associates, Inc., 13041–13051
Kishor Jothimurugan, Rajeev Alur, and Osbert Bastani. 2019 · 2019
Later among the works it cites.
Equilibria-based Probabilistic Model Checking for Concurrent Stochastic Games. In Lecture Notes in Computer Science . Springer International Publishing, 298–315
Marta Kwiatkowska, Gethin Norman, David Parker, and Gabriel Santos. 2019 · 2019
Later among the works it cites.
Multi-agent Reinforcement Learning: A Selective Overview of Theories and Algorithms
Kaiqing Zhang, Zhuoran Yang, and Tamer Başar. 2019 · 2019
Later among the works it cites.
Optimality and Approximation with Policy Gradient Methods in Markov Decision Processes. In Proceedings of Thirty Third Conference on Learning Theory (Proceedings of Machine Learning Research, Vol. 125) , Jacob Abernethy and Shivani Agarwal (Eds.). PMLR, 64–66
Alekh Agarwal, Sham M Kakade, Jason D Lee, and Gaurav Mahajan. 2020 · 2020
Later among the works it cites.
Reward Shaping for Reinforcement Learning with Omega-regular Objectives
Ernst M. Hahn, Mateo Perez, Sven Schewe, Fabio Somenzi, Ashutosh Trivedi, and Dominik Wojtczak. 2020 · 2020
Later among the works it cites.
Deep Reinforcement Learning with Temporal Logics
Mohammadhosein Hasanbeig, Daniel Kroening, and Alessandro Abate. 2020 · 2020
Later among the works it cites.
Extended Markov Games to Learn Multiple Tasks in Multi-Agent Reinforcement Learning. In ECAI 2020 - 24th European Conference on Artificial Intelligence, 29 August-8 September 2020, Santiago de Compostela, Spain, August 29 - September 8, 2020 - Including 10th Conference on Prestigious Applications of Artificial Intelligence (PAIS 2020) (Frontiers in Artificial Intelligence and Applications, Vol. 325) , Giuseppe De Giacomo, Alejandro Catalá, Bistra Dilkina, Michela Milano, Senén Barro, Alberto Bugarín, and Jérôme Lang (Eds.). IOS Press, 139–146
Borja G. León and Francesco Belardinelli. 2020 · 2020
Later among the works it cites.
Is the Policy Gradient a Gradient?. In Proceedings of the 19th International Conference on Autonomous Agents and Multi-Agent Systems (Auckland, New Zealand) (AAMAS ’20) . International Foundation for Autonomous Agents and Multiagent Systems, Richland, SC, 939–947
Chris Nota and Philip S. Thomas. 2020 · 2020
Later among the works it cites.
Ryohei Oura, Ami Sakakibara, and Toshimitsu Ushio. 2020 · 2020
Later among the works it cites.
Scalable Reinforcement Learning of Localized Policies for Multi-agent Networked Systems
Guannan Qu, Adam Wierman, and Na Li. 2020 · 2020
Later among the works it cites.
Lexicographic Multi-objective Reinforcement Learning
Joar Skalse, Lewis Hammond, and Alessandro Abate. 2021 · 2021
Closest in time.