Fetching the paper…
Reading the bibliography…
Reward Machines provide an automaton-inspired structure for specifying instructions, safety constraints, and other temporally extended reward-worthy behaviour.
Optimal control of Markov processes with incomplete state information
Karl Johan Åström · 1965
Earlier work this paper cites.
The temporal logic of programs
Amir Pnueli · 1977
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Q-learning with linear function approximation
Francisco S Melo and M Isabel Ribeiro · 2007
Earlier work this paper cites.
Temporal logic motion planning for dynamic robots
Georgios E Fainekos, Antoine Girard, Hadas Kress-Gazit, and George J Pappas · 2009
Earlier work this paper cites.
Temporal-logic-based reactive mission and motion planning
Hadas Kress-Gazit, Georgios E Fainekos, and George J Pappas · 2009
Earlier work this paper cites.
LTLMoP: Experimenting with language, temporal logic and robot control
Cameron Finucane, Gangyuan Jing, and Hadas Kress-Gazit · 2010
Earlier work this paper cites.
LTL control in uncertain environments with probabilistic satisfaction guarantees
Xu Chu Dennis Ding, Stephen L Smith, Calin Belta, and Daniela Rus · 2011
Earlier work this paper cites.
Logic programming with simulation-based temporal projection for everyday robot object manipulation
Lars Kunze, Mihai Emanuel Dolha, and Michael Beetz · 2011
Earlier work this paper cites.
Deep recurrent Q-learning for partially observable MDPs
Matthew Hausknecht and Peter Stone · 2015
Earlier work this paper cites.
Towards manipulation planning with temporal logic specifications
Keliang He, Morteza Lahijanian, Lydia E Kavraki, and Moshe Y Vardi · 2015
Earlier work this paper cites.
Q-learning for Robust Satisfaction of Signal Temporal Logic Specifications
Derya Aksaray, Austin Jones, Zhaodan Kong, Mac Schwager, and Calin Belta · 2016
Earlier work this paper cites.
Modular multitask reinforcement learning with policy sketches
Jacob Andreas, Dan Klein, and Sergey Levine · 2017
Earlier work this paper cites.
Reinforcement learning with temporal logic rewards
Xiao Li, Cristian Ioan Vasile, and Calin Belta · 2017
Earlier work this paper cites.
Environment-independent task specifications via GLTL
Michael L Littman, Ufuk Topcu, Jie Fu, Charles Isbell, Min Wen, and James MacGlashan · 2017
Earlier work this paper cites.
Zero-shot Task Generalization with Multi-Task Deep Reinforcement Learning
Junhyuk Oh, Satinder Singh, Honglak Lee, and Pushmeet Kohli · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Logically-constrained reinforcement learning
Mohammadhosein Hasanbeig, Alessandro Abate, and Daniel Kroening · 2018
Earlier work this paper cites.
A policy search method for temporal logic specified reinforcement learning tasks
Xiao Li, Yao Ma, and Calin Belta · 2018
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Richard S. Sutton and Andrew G. Barto · 2018
Earlier work this paper cites.
Using reward machines for high-level task specification and decomposition in reinforcement learning
Rodrigo Toro Icarte, Toryn Q. Klassen, Richard Valenzano, and Sheila A. McIlraith · 2018
Earlier work this paper cites.
Reinforcement learning with probabilistic guarantees for autonomous driving
Maxime Bouton, Jesper Karlsson, Alireza Nakhaei, Kikuo Fujimura, Mykel J Kochenderfer, and Jana Tumova · 2019
Earlier work this paper cites.
LTL and beyond: Formal languages for reward function specification in reinforcement learning
Alberto Camacho, Rodrigo Toro Icarte, Toryn Q. Klassen, Richard Valenzano, and Sheila A. McIlraith · 2019
Cited alongside, same era.
Using natural language for reward shaping in reinforcement learning
Prasoon Goyal, Scott Niekum, and Raymond J Mooney · 2019
Cited alongside, same era.
A composable specification language for reinforcement learning tasks
Kishor Jothimurugan, Rajeev Alur, and Osbert Bastani · 2019
Cited alongside, same era.
Learning reward machines for partially observable reinforcement learning
Rodrigo Toro Icarte, Ethan Waldie, Toryn Q. Klassen, Rick Valenzano, Margarita P. Castro, and Sheila A. McIlraith · 2019
Cited alongside, same era.
Transfer of temporal logic formulas in reinforcement learning
Zhe Xu and Ufuk Topcu · 2019
Cited alongside, same era.
Active finite reward automaton inference and reinforcement learning using queries and counterexamples
Zhe Xu, Bo Wu, Aditya Ojha, Daniel Neider, and Ufuk Topcu · 2021
Later among the works it cites.
Video pretraining (VPT): Learning to act by watching unlabeled online videos
Bowen Baker, Ilge Akkaya, Peter Zhokov, Joost Huizinga, Jie Tang, Adrien Ecoffet, Brandon Houghton, Raul Sampedro, and Jeff Clune · 2022
Later among the works it cites.
Reinforcement learning with stochastic reward machines
Jan Corazza, Ivan Gavran, and Daniel Neider · 2022
Later among the works it cites.
Collaborative rover-copter path planning and exploration with temporal logic specifications based on Bayesian update under uncertain environments
Kazumune Hashimoto, Natsuko Tsumagari, and Toshimitsu Ushio · 2022
Later among the works it cites.
Noisy symbolic abstractions for deep RL: A case study with reward machines
Andrew C Li, Zizhao Chen, Pashootan Vaezipoor, Toryn Q. Klassen, Rodrigo Toro Icarte, and Sheila A. McIlraith · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lim Zun Yuan, Mohammadhosein Hasanbeig, Alessandro Abate, and Daniel Kroening · 2019
Cited alongside, same era.
Temporal logic monitoring rewards via transducers
Giuseppe De Giacomo, Marco Favorito, Luca Iocchi, Fabio Patrizi, and Alessandro Ronca · 2020
Cited alongside, same era.
Induction of subgoal automata for reinforcement learning
Daniel Furelos-Blanco, Mark Law, Alessandra Russo, Krysia Broda, and Anders Jonsson · 2020
Cited alongside, same era.
Task-oriented active perception and planning in environments with partially known semantics
Mahsa Ghasemi, Erdem Arinc Bulgur, and Ufuk Topcu · 2020
Cited alongside, same era.
Deep Reinforcement Learning with temporal logics
Mohammadhosein Hasanbeig, Daniel Kroening, and Alessandro Abate · 2020
Cited alongside, same era.
Encoding formulas as deep networks: Reinforcement learning for zero-shot execution of LTL formulas
Yen-Ling Kuo, Boris Katz, and Andrei Barbu · 2020
Cited alongside, same era.
Learning non-Markovian Reward Models in MDPs
Gavin Rens and Jean-François Raskin · 2020
Cited alongside, same era.
Skill transfer for temporally-extended task specifications
Jason Xinyu Liu, Ankit Shah, Eric Rosen, George Konidaris, and Stefanie Tellex · 2022
Later among the works it cites.
Formalization of intersection traffic rules in temporal logic
Sebastian Maierhofer, Paul Moosbrugger, and Matthias Althoff · 2022
Later among the works it cites.
Reward machines: Exploiting reward function structure in reinforcement learning
Rodrigo Toro Icarte, Toryn Q. Klassen, Richard Valenzano, and Sheila A. McIlraith · 2022
Later among the works it cites.
Learning to follow instructions in text-based games
Mathieu Tuli, Andrew C. Li, Pashootan Vaezipoor, Toryn Q. Klassen, Scott Sanner, and Sheila A. McIlraith · 2022
Later among the works it cites.
Joint learning of reward machines and policies in environments with partially known semantics
Christos Verginis, Cevahir Koprulu, Sandeep Chinchali, and Ufuk Topcu · 2022
Later among the works it cites.
Lifelong reinforcement learning with temporal logic formulas and reward machines
Xuejing Zheng, Chao Yu, and Minjie Zhang · 2022
Later among the works it cites.
Maxime Chevalier-Boisvert, Bolun Dai, Mark Towers, Rodrigo de Lazcano, Lucas Willems, Salem Lahlou, Suman Pal, Pablo Samuel Castro, and Jordan Terry · 2023
Later among the works it cites.
Vision-language models as success detectors
Yuqing Du, Ksenia Konyushkova, Misha Denil, Akhil Raju, Jessica Landon, Felix Hill, Nando de Freitas, and Serkan Cabi · 2023
Later among the works it cites.
Hierarchies of reward machines
Daniel Furelos-Blanco, Mark Law, Anders Jonsson, Krysia Broda, and Alessandra Russo · 2023
Later among the works it cites.
Reinforcement learning of action and query policies with LTL instructions under uncertain event detector
Wataru Hatanaka, Ryota Yamashina, and Takamitsu Matsubara · 2023
Later among the works it cites.
STEVE-1: A generative model for text-to-behavior in Minecraft
Shalev Lifshitz, Keiran Paster, Harris Chan, Jimmy Ba, and Sheila McIlraith · 2023
Later among the works it cites.
Assessing the robustness of intelligence-driven reinforcement learning
Lorenzo Nodari and Federico Cerutti · 2023
Later among the works it cites.
Learning reward machines: A study in partially observable reinforcement learning
Rodrigo Toro Icarte, Toryn Q. Klassen, Richard Valenzano, Margarita P. Castro, Ethan Waldie, and Sheila A. McIlraith · 2023
Later among the works it cites.
Visual reward machines
Elena Umili, Francesco Argenziano, Aymeric Barbin, Roberto Capobianco, et al · 2023
Later among the works it cites.
Exploiting transformer in sparse reward reinforcement learning for interpretable temporal logic motion planning
Hao Zhang, Hao Wang, and Zhen Kan · 2023
Later among the works it cites.
Deep policy optimization with temporal logic constraints
Ameesh Shah, Cameron Voloshin, Chenxi Yang, Abhinav Verma, Swarat Chaudhuri, and Sanjit A Seshia · 2024
Closest in time.
Neurosymbolic motion and task planning for linear temporal logic tasks
Xiaowu Sun and Yasser Shoukry · 2024
Closest in time.