Fetching the paper…
Reading the bibliography…
Linear temporal logic (LTL) and, more generally, $\omega$-regular objectives are alternatives to the traditional discount sum and average reward objectives in reinforcement learning (RL), offering the advantage of greater comprehensibility and hence explainability.
Discrete Dynamic Programming
David Blackwell · 1962
Earlier work this paper cites.
A reinforcement learning method for maximizing undiscounted rewards
Anton Schwartz · 1993
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Martin L. Puterman · 1994
Earlier work this paper cites.
Formal Verification of Probabilistic Systems
Luca Alfaro · 1998
Earlier work this paper cites.
Computing minimum and maximum reachability times in probabilistic systems
Luca de Alfaro · 1999
Earlier work this paper cites.
Automata Theory and its Applications
Bakhadyr Khoussainov and Anil Nerode · 2001
Earlier work this paper cites.
Blackwell Optimality
Arie Hordijk and Alexander A. Yushkevich · 2002
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Michael Kearns and Satinder Singh · 2002
Earlier work this paper cites.
R-max - A general polynomial time algorithm for near-optimal reinforcement learning
Ronen I. Brafman and Moshe Tennenholtz · 2003
Earlier work this paper cites.
Reinforcement learning for long-run average cost
Abhijit Gosavi · 2004
Earlier work this paper cites.
Theory of Computation
Dexter Kozen · 2006
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Peter Auer, Thomas Jaksch, and Ronald Ortner · 2008
Earlier work this paper cites.
Principles of Model Checking
Christel Baier and Joost-Pieter Katoen · 2008
Earlier work this paper cites.
Reinforcement learning in finite mdps: PAC analysis
Alexander L. Strehl, Lihong Li, and Michael L. Littman · 2009
Earlier work this paper cites.
Faster and dynamic algorithms for maximal end-component decomposition and related graph problems in probabilistic verification
Krishnendu Chatterjee and Monika Henzinger · 2011
Cited alongside, same era.
Robust control of uncertain Markov Decision Processes with Temporal Logic specifications
Eric M. Wolff, Ufuk Topcu, and Richard M. Murray · 2012
Cited alongside, same era.
Verification of Markov decision processes using learning algorithms
Tomáš Brázdil, Krishnendu Chatterjee, Martin Chmelik, Vojtěch Forejt, Jan Křetínskỳ, Marta Kwiatkowska, David Parker, and Mateusz Ujma · 2014
Cited alongside, same era.
Optimal control of Markov Decision Processes with Linear Temporal Logic constraints
Xuchu Ding, Stephen L. Smith, Calin Belta, and Daniela Rus · 2014
Cited alongside, same era.
Probably approximately correct MDP learning and control with Temporal Logic constraints
Jie Fu and Ufuk Topcu · 2014
Cited alongside, same era.
Deep reinforcement learning with Temporal Logics
Mohammadhosein Hasanbeig, Daniel Kroening, and Alessandro Abate · 2020
Later among the works it cites.
On the expressivity of markov reward
David Abel, Will Dabney, Anna Harutyunyan, Mark K. Ho, Michael L. Littman, Doina Precup, and Satinder Singh · 2021
Later among the works it cites.
Average-reward model-free reinforcement learning: a systematic review and literature mapping, 2021
Vektor Dewanto, George Dunn, Ali Eshragh, Marcus Gallagher, and Fred Roosta · 2021
Later among the works it cites.
Learning and planning in average-reward markov decision processes
Yi Wan, Abhishek Naik, and Richard S. Sutton · 2021
Later among the works it cites.
A framework for transforming specifications in reinforcement learning
Rajeev Alur, Suguman Bansal, Osbert Bastani, and Kishor Jothimurugan · 2022
Later among the works it cites.
Reward Machines
Rodrigo Toro Icarte · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Probability Theory: A Comprehensive Course
Achim Klenke · 2014
Cited alongside, same era.
A learning based approach to control synthesis of Markov Decision Processes for Linear Temporal Logic specifications
Dorsa Sadigh, Eric S. Kim, Samuel Coogan, S. Shankar Sastry, and Sanjit A. Seshia · 2014
Cited alongside, same era.
Limit-deterministic Büchi automata for Linear Temporal Logic
Salomon Sickert, Javier Esparza, Stefan Jaax, and Jan Křetínský · 2016
Cited alongside, same era.
Efficient average reward reinforcement learning using constant shifting values
Shangdong Yang, Yang Gao, Bo An, Hao Wang, and Xingguo Chen · 2016
Cited alongside, same era.
Using reward machines for high-level task specification and decomposition in reinforcement learning
Rodrigo Toro Icarte, Toryn Klassen, Richard Valenzano, and Sheila McIlraith · 2018
Cited alongside, same era.
Reinforcement Learning: An Introduction
Richard S. Sutton and Andrew G. Barto · 2018
Cited alongside, same era.
Control synthesis from Linear Temporal Logic specifications using model-free reinforcement learning
Alper Kamil Bozkurt, Yu Wang, Michael M. Zavlanos, and Miroslav Pajic · 2019
Cited alongside, same era.
Later among the works it cites.
Policy optimization with Linear Temporal Logic constraints
Cameron Voloshin, Hoang Le, Swarat Chaudhuri, and Yisong Yue · 2022
Later among the works it cites.
On the (In)Tractability of Reinforcement Learning for LTL Objectives
Cambridge Yang, Michael L. Littman, and Michael Carbin · 2022
Later among the works it cites.
Optimal probabilistic motion planning with potential infeasible LTL constraints
Mingyu Cai, Shaoping Xiao, Zhijun Li, and Zhen Kan · 2023
Later among the works it cites.
Reducing blackwell and average optimality to discounted mdps via the blackwell discount factor
Julien Grand-Clément and Marek Petrik · 2023
Later among the works it cites.
Certified reinforcement learning with logic guidance
Hosein Hasanbeig, Daniel Kroening, and Alessandro Abate · 2023
Later among the works it cites.
Sample Efficient Model-free Reinforcement Learning from LTL Specifications with Optimality Guarantees
Daqian Shao and Marta Kwiatkowska · 2023
Later among the works it cites.
Eventual discounting temporal logic counterfactual experience replay
Cameron Voloshin, Abhinav Verma, and Yisong Yue · 2023
Later among the works it cites.
A PAC Learning Algorithm for LTL and Omega-Regular Objectives in MDPs
Mateo Perez, Fabio Somenzi, and Ashutosh Trivedi · 2024
Closest in time.