Fetching the paper…
Reading the bibliography…
There are situations in which an agent should receive rewards only after having accomplished a series of previous tasks, that is, rewards are non-Markovian.
D. Angluin, ‘Learning regular sets from queries and counterexamples’, Information and Computation
1987
Earlier work this paper cites.
M. Puterman, Markov Decision Processes: Discrete Dynamic Programming
1994
Earlier work this paper cites.
F. Bacchus, C. Boutilier, and A. Grove, ‘Rewarding behaviors’, in Proceedings of the Thirteenth Natl. Conf. on Artif. Intell
1996
Earlier work this paper cites.
D. Lee and M. Yannakakis, ‘Principles and methods of testing finite state machines - a survey’, Proceedings of the IEEE
1996
Earlier work this paper cites.
S. Thiébaux, C. Gretton, J. Slaney, D. Price, and F. Kabanza, ‘Decision-theoretic planning with non-Markovian rewards’, Artif. Intell. Research
2006
Earlier work this paper cites.
C. Baier and J.-P. Katoen, Principles of Model Checking
2008
Earlier work this paper cites.
M. Shahbaz and R. Groz, ‘Inferring Mealy machines’, in Proceedings of the International Symposium on Formal Methods (FM-09)
2009
Earlier work this paper cites.
C. Amato, B. Bonet, and S. Zilberstein, ‘Finite-state controllers based on Mealy machines for centralized and decentralized POMDPs’, in Proceedings of the Twenty-Fourth AAAI Conf. on Artif. Intell. (AAAI-10)
2010
Earlier work this paper cites.
H. Hasselt, ‘Double q-learning’, in Advances in neural information processing systems 23
2010
Earlier work this paper cites.
D. Kingma and J. Ba, ‘Adam: A method for stochastic optimization’, arXiv preprint arXiv:1412.6980
2014
Earlier work this paper cites.
V. Mnih, K. Kavukcuoglu, D. Silver, A. Rusu, J. Veness, M. Bellemare, A. Graves, M. Riedmiller, A. Fidjeland, G. Ostrovski, et al., ‘Human-level control through deep reinforcement learning’, nature
2015
Cited alongside, same era.
M. Abadi, P. Barham, J. Chen, Z. Chen, A. Davis, J. Dean, M. Devin, S. Ghemawat, G. Irving, M. Isard, et al., ‘Tensorflow: A system for large-scale machine learning’, in Twelfth USENIX symposium on operating systems design and implementation (OSDI 16)
2016
Cited alongside, same era.
M. McTear, Z. Callejas, and D. Griol, The conversational interface
2016
Cited alongside, same era.
keras-rl
M. Plappert · 2016
Cited alongside, same era.
M. Turchetta, F. Berkenkamp, and A. Krause, ‘Safe exploration in finite markov decision processes with gaussian processes’, in Proceedings of the Thirtieth Conference on Neural Information Processing Systems
R. Toro Icarte, T. Klassen, R. Valenzano, and S. McIlraith, ‘Teaching multiple tasks to an RL agent using LTL’, in Proceedings of the Seventeenth Intl. Conf. on Autonomous Agents and Multiagent Systems
2018
Later among the works it cites.
R. Toro Icarte, T. Klassen, R. Valenzano, and S. McIlraith, ‘Using reward machines for high-level task specification and decomposition in reinforcement learning’, in Proceedings of the Thirty-Fifth Intl. Conf. on Machine Learning
2018
Later among the works it cites.
J. Křetínský, G. Pérez, and J.-F. Raskin, ‘Learning-based mean-payoff optimization in an unknown MDP under omega-regular constraints’, in Proceedings of the Twenty-Ninth Intl. Conf. on Concurrency Theory (CONCUR-18)
2018
Later among the works it cites.
A. Camacho, R. Toro Icarte, T. Klassen, R. Valenzano, and S. McIlraith, ‘LTL and beyond: Formal languages for reward function specification in reinforcement learning’, in Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
F. Vaandrager, ‘Model Learning’, Communications of the ACM
2017
Cited alongside, same era.
M. Alshiekh, R. Bloem, R. Ehlers, B. Könighofer, S. Niekum, and U. Topcu, ‘Safe reinforcement learning via shielding’, in Proceedings of the Thirty-Second AAAI Conf. on Artif. Intell. (AAAI-18)
2018
Cited alongside, same era.
R. Brafman, G. De Giacomo, and F. Patrizi, ‘LTL f
2018
Cited alongside, same era.
A. Camacho, O. Chen, S. Sanner, and S. McIlraith, ‘Non-Markovian rewards expressed in LTL: Guiding search via reward shaping (extended version)’, in Proceedings of the First Workshop on Goal Specifications for Reinforcement Learning
2018
Cited alongside, same era.
2019
Later among the works it cites.
R. Cheng, G. Orosz, R. Murray, and J. Burdick, ‘End-to-end safe reinforcement learning through barrier functions for safety-critical continuous control tasks’, in The Thirty-third AAAI Conference on Artificial Intelligence
2019
Later among the works it cites.
G. De Giacomo, M. Favorito, L. Iocchi, and F. Patrizi, ‘Foundations for restraining bolts: Reinforcement learning with LTL f
2019
Later among the works it cites.
M. Hasanbeig, A. Abate, and D. Kroening, ‘Logically-constrained neural fitted q-iteration’, in Proceedings of the Eighteenth Intl. Conf. on Autonomous Agents and Multiagent Systems
2019
Later among the works it cites.
M. Hasanbeig, D. Kroening, and A. Abate, ‘Towards verifiable and safe model-free reinforcement learning’, in Proceedings of the First Workshop on Artificial Intelligence and Formal Verification, Logics, Automata and Synthesis (OVERLAY)
2019
Later among the works it cites.
R. Toro Icarte, E. Waldie, T. Klassen, R. Valenzano, M. Castro, and S. McIlraith, ‘Learning reward machines for partially observable reinforcement learning’, in Proceedings of the Thirty-third Conference on Neural Information Processing Systems
2019
Later among the works it cites.