Fetching the paper…
Reading the bibliography…
We study the problem of reinforcement learning for a task encoded by a reward machine.
1903
Earlier work this paper cites.
R. Bellman, A markovian decision process, Journal of mathematics and mechanics (1957) 679–684
1957
Earlier work this paper cites.
E. M. Gold, Complexity of automaton identification from given data, Information and control 37 (3) (1978) 302–320
1978
Earlier work this paper cites.
S. P. Singh, Reinforcement learning with a hierarchy of abstract models, in: Proceedings of the National Conference on Artificial Intelligence, no. 10, Citeseer, 1992, p. 202
1992
Earlier work this paper cites.
C. J. Watkins, P. Dayan, Q-learning, Machine learning 8 (3-4) (1992) 279–292
1992
Earlier work this paper cites.
2002
Earlier work this paper cites.
M. E. Taylor, P. Stone, Cross-domain transfer for reinforcement learning, Proceedings of the 24th international conference on Machine learning (2007) 879–886
2007
Earlier work this paper cites.
2009
Earlier work this paper cites.
M. J. Heule, S. Verwer, Exact dfa identification using sat solvers, International Colloquium on Grammatical Inference (2010) 66–79
2010
Earlier work this paper cites.
S. C. Livingston, R. M. Murray, J. W. Burdick, Backtracking temporal logic synthesis for uncertain environments, IEEE International Conference on Robotics and Automation (2012) 5163–5170
2012
Earlier work this paper cites.
M. Guo, K. H. Johansson, D. V. Dimarogonas, Revising motion planning under linear temporal logic specifications in partially known workspaces, 2013 IEEE International Conference on Robotics and Automation (2013) 5025–5032
2013
Earlier work this paper cites.
A. M. Ayala, S. B. Andersson, C. Belta, Temporal logic motion planning in unknown environments, IEEE/RSJ International Conference on Intelligent Robots and Systems (2013) 5279–5284
2013
Earlier work this paper cites.
D. Neider, N. Jansen, Regular model checking using solver technologies and automata learning, NASA Formal Methods Symposium (2013) 16–31
2013
Earlier work this paper cites.
D. Sadigh, E. S. Kim, S. Coogan, S. S. Sastry, S. A. Seshia, A learning based approach to control synthesis of markov decision processes for linear temporal logic specifications, IEEE Conference on Decision and Control (2014) 1091–1096
2014
Earlier work this paper cites.
A.-A. Agha-Mohammadi, S. Chakravorty, N. M. Amato, Firm: Sampling-based feedback motion-planning under motion uncertainty and imperfect measurements, The International Journal of Robotics Research 33 (2) (2014) 268–304
2014
Cited alongside, same era.
M. L. Puterman, Markov decision processes: discrete stochastic dynamic programming, John Wiley & Sons, 2014
2014
Cited alongside, same era.
D. Neider, Applications of automata learning in verification and synthesis, Ph.D. thesis, Hochschulbibliothek der Rheinisch-Westfälischen Technischen Hochschule Aachen (2014)
2014
Cited alongside, same era.
A. Morgado, C. Dodaro, J. Marques-Silva, Core-guided maxsat with soft cardinality constraints (2014) 564–573
2014
Cited alongside, same era.
M. Wen, R. Ehlers, U. Topcu, Correct-by-synthesis reinforcement learning with temporal logic constraints, IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) (2015) 4983–4990
R. Toro Icarte, T. Q. Klassen, R. Valenzano, S. A. McIlraith, Teaching multiple tasks to an rl agent using ltl, Proceedings of the 17th International Conference on Autonomous Agents and MultiAgent Systems (2018) 452–461
2018
Later among the works it cites.
A. Ignatiev, A. Morgado, J. Marques-Silva, PySAT: A Python toolkit for prototyping with SAT oracles (2018) 428–437 doi:10.1007/978-3-319-94144-8_26
2018
Later among the works it cites.
G. Konidaris, On the necessity of abstraction, Current opinion in behavioral sciences 29 (2019) 1–7
2019
Later among the works it cites.
E. M. Hahn, M. Perez, S. Schewe, F. Somenzi, A. Trivedi, D. Wojtczak, Omega-regular objectives in model-free reinforcement learning, International Conference on Tools and Algorithms for the Construction and Analysis of Systems (2019) 395–412
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2015
Cited alongside, same era.
X. Zhang, B. Wu, H. Lin, Learning based supervisor synthesis of pomdp for pctl specifications, 54th IEEE Conference on Decision and Control (CDC) (2015) 7470–7475
2015
Cited alongside, same era.
T. D. Kulkarni, K. Narasimhan, A. Saeedi, J. Tenenbaum, Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation, Advances in neural information processing systems 29 (2016) 3675–3683
2016
Cited alongside, same era.
D. Aksaray, A. Jones, Z. Kong, M. Schwager, C. Belta, Q-learning for robust satisfaction of signal temporal logic specifications, IEEE 55th Conference on Decision and Control (CDC) (2016) 6565–6570
2016
Cited alongside, same era.
M. Lahijanian, M. R. Maly, D. Fried, L. E. Kavraki, H. Kress-Gazit, M. Y. Vardi, Iterative temporal planning in uncertain environments with partial satisfaction guarantees, IEEE Transactions on Robotics 32 (3) (2016) 583–599
2016
Cited alongside, same era.
X. Li, C.-I. Vasile, C. Belta, Reinforcement learning with temporal logic rewards, IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) (2017) 3834–3839
2017
Cited alongside, same era.
J. Andreas, D. Klein, S. Levine, Modular multitask reinforcement learning with policy sketches, International Conference on Machine Learning (2017) 166–175
2017
Cited alongside, same era.
F. J. Montana, J. Liu, T. J. Dodd, Sampling-based reactive motion planning with temporal logic constraints and imperfect state information, Critical Systems: Formal Methods and Automated Verification (2017) 134–149
2017
Cited alongside, same era.
R. R. da Silva, V. Kurtz, H. Lin, Active perception and control from temporal logic specifications, IEEE Control Systems Letters 3 (4) (2019) 1068–1073
2019
Later among the works it cites.
B. Ramasubramanian, A. Clark, L. Bushnell, R. Poovendran, Secure control under partial observability with temporal logic constraints, American Control Conference (ACC) (2019) 1181–1188
2019
Later among the works it cites.
R. Toro Icarte, E. Waldie, T. Klassen, R. Valenzano, M. Castro, S. McIlraith, Learning reward machines for partially observable reinforcement learning, Advances in Neural Information Processing Systems 32 (2019) 15523–15534
2019
Later among the works it cites.
Z. Xu, I. Gavran, Y. Ahmad, R. Majumdar, D. Neider, U. Topcu, B. Wu, Joint inference of reward machines and policies for reinforcement learning, Proceedings of the International Conference on Automated Planning and Scheduling 30 (2020) 590–598
2020
Later among the works it cites.
A. K. Bozkurt, Y. Wang, M. M. Zavlanos, M. Pajic, Control synthesis from linear temporal logic specifications using model-free reinforcement learning, IEEE International Conference on Robotics and Automation (ICRA) (2020) 10349–10355
2020
Later among the works it cites.
M. Ghasemi, E. Bulgur, U. Topcu, Task-oriented active perception and planning in environments with partially known semantics, International Conference on Machine Learning (2020) 3484–3493
2020
Later among the works it cites.
D. Furelos-Blanco, M. Law, A. Russo, K. Broda, A. Jonsson, Induction of subgoal automata for reinforcement learning, in: Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 34, 2020, pp. 3890–3897
2020
Later among the works it cites.
M. Hasanbeig, N. Y. Jeppu, A. Abate, T. Melham, D. Kroening, Deepsynth: Automata synthesis for automatic task segmentation in deep reinforcement learning, in: Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 35, 2021, pp. 7647–7656
2021
Later among the works it cites.
T. Dohmen, N. Topper, G. Atia, A. Beckus, A. Trivedi, A. Velasquez, Inferring probabilistic reward machines from non-markovian reward signals for reinforcement learning, in: Proceedings of the International Conference on Automated Planning and Scheduling, Vol. 32, 2022, pp. 574–582
2022
Closest in time.
M. Jin, Z. Ma, K. Jin, H. H. Zhuo, C. Chen, C. Yu, Creativity of ai: Automatic symbolic option discovery for facilitating deep reinforcement learning, in: Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 36, 2022, pp. 7042–7050
2022
Closest in time.