Fetching the paper…
Reading the bibliography…
The success of reinforcement learning in typical settings is predicated on Markovian assumptions on the reward signal by which an agent learns optimal policies.
A study of grammatical inference
J. J. Horning · 1969
Earlier work this paper cites.
Learning regular sets from queries and counterexamples
D. Angluin · 1987
Earlier work this paper cites.
Reinforcement learning - an introduction
R. S. Sutton and A. G. Barto · 1998
Earlier work this paper cites.
An integrated approach to testing complex systems
O. Niese · 2003
Earlier work this paper cites.
Learning I/O automata
F. Aarts and F. W. Vaandrager · 2010
Earlier work this paper cites.
Grammatical Inference: Learning Automata and Grammars
C. de la Higuera · 2010
Earlier work this paper cites.
Using reward machines for high-level task specification and decomposition in reinforcement learning
R. T. Icarte, T. Q. Klassen, R. A. Valenzano, and S. A. McIlraith · 2018
Earlier work this paper cites.
Extracting automata from recurrent neural networks using queries and counterexamples
G. Weiss, Y. Goldberg, and E. Yahav · 2018
Earlier work this paper cites.
LTL and beyond: Formal languages for reward function specification in reinforcement learning
A. Camacho, R. T. Icarte, T. Q. Klassen, R. A. Valenzano, and S. A. McIlraith · 2019
Cited alongside, same era.
Omega-regular objectives in model-free reinforcement learning
E. M. Hahn, M. Perez, S. Schewe, F. Somenzi, A. Trivedi, and D. Wojtczak · 2019
Cited alongside, same era.
Learning reward machines for partially observable reinforcement learning
R. T. Icarte, E. Waldie, T. Q. Klassen, R. A. Valenzano, M. P. Castro, and S. A. McIlraith · 2019
Cited alongside, same era.
L * {}^{\mbox{*}} -based learning of markov decision processes
M. Tappler, B. K. Aichernig, G. Bacci, M. Eichlseder, and K. G. Larsen · 2019
Cited alongside, same era.
Learning deterministic weighted automata with queries and counterexamples
G. Weiss, Y. Goldberg, and E. Yahav · 2019
Cited alongside, same era.
Learning and solving regular decision processes
Verification-guided tree search
A. Velasquez and D. Melcer · 2020
Later among the works it cites.
Joint inference of reward machines and policies for reinforcement learning
Z. Xu, I. Gavran, Y. Ahmad, R. Majumdar, D. Neider, U. Topcu, and B. Wu · 2020
Later among the works it cites.
Advice-guided reinforcement learning in a non-markovian environment
D. Neider, J. Gaglione, I. Gavran, U. Topcu, B. Wu, and Z. Xu · 2021
Closest in time.
Online learning of non-markovian reward models
G. Rens, J. Raskin, R. Reynouard, and G. Marra · 2021
Closest in time.
Dynamic automaton-guided reward shaping for monte carlo tree search
A. Velasquez, B. Bissey, L. Barak, A. Beckus, I. Alkhouri, D. Melcer, and G. K. Atia · 2021
Closest in time.
Active finite reward automaton inference and reinforcement learning using queries and counterexamples
Z. Xu, B. Wu, A. Ojha, D. Neider, and U. Topcu · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
E. Abadi and R. I. Brafman · 2020
Cited alongside, same era.
Reinforcement learning with non-markovian rewards
M. Gaon and R. I. Brafman · 2020
Cited alongside, same era.
Closest in time.
Reward machines: Exploiting reward function structure in reinforcement learning
R. T. Icarte, T. Q. Klassen, R. A. Valenzano, and S. A. McIlraith · 2022
Closest in time.