Fetching the paper…
Reading the bibliography…
We provide the first solution for model-free reinforcement learning of {\omega}-regular objectives for Markov decision processes (MDPs).
Old Possum’s Book of Practical Cats
T. S. Eliot · 1939
Earlier work this paper cites.
Automatic verification of probabilistic concurrent finite state programs
M. Y. Vardi · 1985
Earlier work this paper cites.
The Temporal Logic of Reactive and Concurrent Systems *Specification*
Z. Manna and A. Pnueli · 1991
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
M. L. Puterman · 1994
Earlier work this paper cites.
The complexity of probabilistic verification
C. Courcoubetis and M. Yannakakis · 1995
Earlier work this paper cites.
Formal Verification of Probabilistic Systems
L. de Alfaro · 1998
Earlier work this paper cites.
Handbook of Markov Decision Processes: Methods and Applications
A. Hordijk and A. A. Yushkevich · 2002
Earlier work this paper cites.
Principles of Model Checking
C. Baier and J.-P. Katoen · 2008
Cited alongside, same era.
PRISM 4.0: Verification of probabilistic real-time systems
M. Kwiatkowska, G. Norman, and D. Parker · 2011
Cited alongside, same era.
Verification of markov decision processes using learning algorithms
T. Brázdil, K. Chatterjee, M. Chmelík, V. Forejt, J. Křetínský, M. Kwiatkowska, D. Parker, and M. Ujma · 2014
Cited alongside, same era.
Probably approximately correct MDP learning and control with temporal logic constraints
J. Fu and U. Topcu · 2014
Cited alongside, same era.
A learning based approach to control synthesis of Markov decision processes for linear temporal logic specifications
D. Sadigh, E. Kim, S. Coogan, S. S. Sastry, and S. A. Seshia · 2014
Cited alongside, same era.
The Hanoi omega-automata format
T. Babiak, F. Blahoudek, A. Duret-Lutz, J. Klein, J. Křetínský, D. Müller, D. Parker, and J. Strejček · 2015
Learning an optimal control policy for a Markov decision process under linear temporal logic specifications
M. Hiromoto and T. Ushio · 2015
Later among the works it cites.
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Later among the works it cites.
Limit-deterministic Büchi automata for linear temporal logic
S. Sickert, J. Esparza, S. Jaax, and J. Křetínský · 2016
Later among the works it cites.
Reinforcement learning with temporal logic rewards
X. Li, C. I. Vasile, and C. Belta · 2017
Later among the works it cites.
https://automata.tools/hoa/cpphoafparser
cpphoafparser · 2018
Closest in time.
https://gym.openai.com
OpenAI Gym · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Lazy probabilistic model checking without determinisation
E. M. Hahn, G. Li, S. Schewe, A. Turrini, and L. Zhang · 2015
Cited alongside, same era.
Reinforcement Learnging: An Introduction
R. S. Sutton and A. G. Barto · 2018
Closest in time.