Fetching the paper…
Reading the bibliography…
We present the first model-free Reinforcement Learning (RL) algorithm to synthesise policies for an unknown Markov Decision Process (MDP), such that a linear time property is satisfied.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A.P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., Kavukcuoglu, K.: · 1937
Earlier work this paper cites.
The temporal logic of programs
Pnueli, A.: · 1977
Earlier work this paper cites.
On the complexity of omega-automata
Safra, S.: · 1988
Earlier work this paper cites.
Q-learning
Watkins, C.J., Dayan, P.: · 1992
Earlier work this paper cites.
Neuro-dynamic Programming. Volume 1
Bertsekas, D.P., Tsitsiklis, J.N.: · 1996
Earlier work this paper cites.
Reinforcement learning: An introduction. Volume 1
Sutton, R.S., Barto, A.G.: · 1998
Earlier work this paper cites.
Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning
Sutton, R.S., Precup, D., Singh, S.: · 1999
Earlier work this paper cites.
Essentials of stochastic processes. Volume 1
Durrett, R.: · 1999
Earlier work this paper cites.
Model checking of safety properties
Kupferman, O., Vardi, M.Y.: · 2001
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Kearns, M., Singh, S.: · 2002
Earlier work this paper cites.
Deterministic generators and games for LTL fragments
Alur, R., La Torre, S.: · 2004
Earlier work this paper cites.
From nondeterministic Büchi and Streett automata to deterministic parity automata
Piterman, N.: · 2006
Earlier work this paper cites.
An application of reinforcement learning to aerobatic helicopter flight
Abbeel, P., Coates, A., Quigley, M., Ng, A.Y.: · 2007
Earlier work this paper cites.
Principles of Model Checking
Baier, C., Katoen, J.P., Larsen, K.G.: · 2008
Earlier work this paper cites.
Teachable robots: Understanding human teaching behavior to build more effective robot learners
Thomaz, A.L., Breazeal, C.: · 2008
Cited alongside, same era.
An inequality for variances of the discounted rewards
Feinberg, E.A., Fei, J.: · 2009
Cited alongside, same era.
Motion planning and control from temporal logic specifications with probabilistic satisfaction guarantees
Lahijanian, M., Wasniewski, J., Andersson, S.B., Belta, C.: · 2010
Cited alongside, same era.
Optimal path planning for surveillance with temporal-logic constraints
Smith, S.L., Tumová, J., Belta, C., Rus, D.: · 2011
Cited alongside, same era.
PRISM 4.0: Verification of probabilistic real-time systems
Kwiatkowska, M., Norman, G., Parker, D.: · 2011
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A.A., Veness, J., Bellemare, M.G., Graves, A., Riedmiller, M., Fidjeland, A.K., Ostrovski, G., et al.: · 2015
Later among the works it cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A.A., Veness, J., Bellemare, M.G., Graves, A., Riedmiller, M., Fidjeland, A.K., Ostrovski, G., et al.: · 2015
Later among the works it cites.
Correct-by-synthesis reinforcement learning with temporal logic constraints
Wen, M., Ehlers, R., Topcu, U.: · 2015
Later among the works it cites.
Model-based reinforcement learning in continuous environments using real-time constrained optimization
Andersson, O., Heintz, F., Doherty, P.: · 2015
Later among the works it cites.
Multi-agent learning in coverage control games
Hasanbeig, M.: · 2016
Later among the works it cites.
Mastering the game of Go with deep neural networks and tree search
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
van Otterlo, M., Wiering, M.: · 2012
Cited alongside, same era.
Robust control of uncertain Markov decision processes with temporal logic specifications
Wolff, E.M., Topcu, U., Murray, R.M.: · 2012
Cited alongside, same era.
Pareto curves for probabilistic model checking
Forejt, V., Kwiatkowska, M., Parker, D.: · 2012
Cited alongside, same era.
Optimal control of MDPs with temporal logic constraints
Svorenova, M., Cerna, I., Belta, C.: · 2013
Cited alongside, same era.
Markov decision processes: Discrete stochastic dynamic programming
Puterman, M.L.: · 2014
Cited alongside, same era.
Game theory control solution for sensor coverage problem in unknown environment
Rahili, S., Ren, W.: · 2014
Cited alongside, same era.
A learning based approach to control synthesis of Markov decision processes for linear temporal logic specifications
Sadigh, D., Kim, E.S., Coogan, S., Sastry, S.S., Seshia, S.A.: · 2014
Cited alongside, same era.
Silver, D., Huang, A., Maddison, C.J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al.: · 2016
Later among the works it cites.
Limit-deterministic Büchi automata for linear temporal logic
Sickert, S., Esparza, J., Jaax, S., Křetínskỳ, J.: · 2016
Later among the works it cites.
Safety-constrained reinforcement learning for MDPs
Junges, S., Jansen, N., Dehnert, C., Topcu, U., Katoen, J.P.: · 2016
Later among the works it cites.
Reinforcement learning with temporal logic rewards
Li, X., Vasile, C.I., Belta, C.: · 2016
Later among the works it cites.
On synchronous binary log-linear learning and second order Q-learning
Hasanbeig, M., Pavel, L.: · 2017
Later among the works it cites.
Quantitative model-checking of controlled discrete-time Markov processes
Tkachev, I., Mereacre, A., Katoen, J.P., Abate, A.: · 2017
Later among the works it cites.
Verification and repair of control policies for safe reinforcement learning
Pathak, S., Pulina, L., Tacchella, A.: · 2017
Later among the works it cites.
Safe reinforcement learning via shielding
Alshiekh, M., Bloem, R., Ehlers, R., Könighofer, B., Niekum, S., Topcu, U.: · 2017
Later among the works it cites.