Fetching the paper…
Reading the bibliography…
Reinforcement Learning (RL) has achieved tremendous development in recent years, but still faces significant obstacles in addressing complex real-life problems due to the issues of poor system generalization, low sample efficiency as well as safety and interpretability concerns.
A. Pnueli, “The temporal logic of programs,” in 18th Annual Symposium on Foundations of Computer Science , 1977, pp. 46–57
1977
Earlier work this paper cites.
C. J. Watkins and P. Dayan, “Q-learning,” Machine learning , vol. 8, no. 3-4, pp. 279–292, 1992
1992
Earlier work this paper cites.
M. Ghallab, C. Knoblock, D. Wilkins, and et al., “Pddl - the planning domain definition language,” 1998
1998
Earlier work this paper cites.
R. Alur, T. A. Henzinger, and et al., “Alternating-time temporal logic,” Journal of the ACM , vol. 49, no. 5, pp. 672–713, 2002
2002
Earlier work this paper cites.
R. Brachman and H. Levesque, Knowledge representation and reasoning . Elsevier, 2004
2004
Earlier work this paper cites.
M. Richardson and P. Domingos, “Markov logic networks,” Machine Learning , vol. 62, no. 1-2, pp. 107–136, 2006
2006
Earlier work this paper cites.
L. Mihalkova, T. Huynh, and R. J. Mooney, “Mapping and revising markov logic networks for transfer learning,” in AAAI , vol. 7, 2007, pp. 608–614
2007
Earlier work this paper cites.
L. Mihalkova and R. J. Mooney, “Transfer learning by mapping with minimal target data,” in Proceedings of the AAAI-08 workshop on transfer learning for complex tasks , 2008, pp. 31–36
2008
Earlier work this paper cites.
L. Torrey and J. Shavlik, “Policy transfer via markov logic networks,” in ILP , 2010, pp. 234–248
2010
Earlier work this paper cites.
J. W. Lloyd, Foundations of logic programming . Springer Science & Business Media, 2012
2012
Earlier work this paper cites.
M. Nickles, “Integrating relational reinforcement learning with reasoning about actions and change,” in Inductive Logic Programming: 21st International Conference, ILP 2011, Windsor Great Park, UK, July 31–August 3, 2011, Revised Selected Papers 21 . Springer, 2012, pp. 255–269
2012
Earlier work this paper cites.
F. Mogavero, A. Murano, G. Perelli, and et al., “Reasoning about strategies: On the model-checking problem,” ACM Transactions on Computational Logic , vol. 15, no. 4, pp. 1–47, 2014
2014
Earlier work this paper cites.
J. Garcıa and F. Fernández, “A comprehensive survey on safe reinforcement learning,” Journal of Machine Learning Research , vol. 16, no. 1, pp. 1437–1480, 2015
2015
Earlier work this paper cites.
L. De Raedt and A. Kimmig, “Probabilistic (logic) programming concepts,” Machine Learning , vol. 100, no. 1, pp. 5–47, 2015
2015
Earlier work this paper cites.
M. Leonetti, L. Iocchi, and P. Stone, “A synthesis of automated planning and reinforcement learning for efficient, robust decision-making,” Artificial Intelligence , vol. 241, pp. 103–130, 2016
2016
Earlier work this paper cites.
D. Aksaray, A. Jones, Z. Kong, M. Schwager, and C. Belta, “Q-learning for robust satisfaction of signal temporal logic specifications,” in IEEE 55th CDC , 2016, pp. 6565–6570
2016
Earlier work this paper cites.
S. Sickert, J. Esparza, S. Jaax, and J. Křetínskỳ, “Limit-deterministic büchi automata for linear temporal logic,” in International Conference on Computer Aided Verification . Springer, 2016, pp. 312–332
2016
Earlier work this paper cites.
S. Junges, N. Jansen, C. Dehnert, U. Topcu, and J.-P. Katoen, “Safety-constrained reinforcement learning for mdps,” in ICTA , 2016, pp. 130–146
2016
Earlier work this paper cites.
L. Xiong and Y. Liu, “Strategy representation and reasoning for incomplete information concurrent games in the situation calculus.” in IJCAI , 2016, pp. 1322–1329
2016
Earlier work this paper cites.
Y. Li, “Deep reinforcement learning: An overview,” arXiv preprint arXiv:1701.07274 , 2017
2017
Earlier work this paper cites.
K. Arulkumaran, M. P. Deisenroth, M. Brundage, and A. A. Bharath, “Deep reinforcement learning: A brief survey,” IEEE Signal Processing Magazine , vol. 34, no. 6, pp. 26–38, 2017
2017
Earlier work this paper cites.
X. Li, C. Vasile, and C. Belta, “Reinforcement learning with temporal logic rewards,” in IROS , 2017, pp. 3834–3839
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
R. F. Mansour and S. Hosni, “Recent advances in markov logic networks,” Indian Journal of Science and Technology , vol. 10, p. 19, 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
R. S. Sutton and A. G. Barto, Reinforcement learning: An introduction . MIT press, 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
R. T. Icarte, T. Klassen, R. Valenzano, and S. McIlraith, “Using reward machines for high-level task specification and decomposition in reinforcement learning,” in ICML , 2018, pp. 2107–2116
2018
Earlier work this paper cites.
D. Muniraj, K. G. Vamvoudakis, and M. Farhood, “Enforcing signal temporal logic specifications in multi-agent adversarial environments: A deep q-learning approach,” in IEEE CDC , 2018, pp. 4141–4146
2018
Earlier work this paper cites.
L. A. Ferreira, R. A. Bianchi, P. E. Santos, and R. L. de Mantaras, “A method for the online construction of the set of states of a markov decision process using answer set programming,” in International Conference on Industrial, Engineering and Other Applications of Applied Intelligent Systems . Springer, 2018, pp. 3–15
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
R. Toro Icarte, T. Q. Klassen, and et al., “Teaching multiple tasks to an rl agent using ltl,” in AAMAS , 2018, pp. 452–461
2018
Earlier work this paper cites.
R. Brafman, G. De Giacomo, and F. Patrizi, “Ltlf/ldlf non-markovian rewards,” in AAAI , vol. 32, no. 1, 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
A. Shah, P. Kamath, and et al., “Bayesian inference of temporal task specifications from demonstrations,” in NeurIPS , 2018, pp. 3808–3817
2018
Earlier work this paper cites.
D. Hein, S. Udluft, and et al, “Interpretable policies for reinforcement learning by genetic programming,” Engineering Applications of Artificial Intelligence , vol. 76, pp. 158–169, 2018
2018
Earlier work this paper cites.
O. Bastani, Y. Pu, and A. Solar-Lezama, “Verifiable reinforcement learning via policy extraction,” Advances in neural information processing systems , vol. 31, 2018
2018
Earlier work this paper cites.
A. Verma, V. Murali, R. Singh, P. Kohli, and S. Chaudhuri, “Programmatically interpretable reinforcement learning,” in ICML , 2018, pp. 5045–5054
2018
Earlier work this paper cites.
M. Alshiekh, R. Bloem, R. Ehlers, and et al., “Safe reinforcement learning via shielding,” in AAAI , 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
S. Pathak, L. Pulina, and A. Tacchella, “Verification and repair of control policies for safe reinforcement learning,” Applied Intelligence , vol. 48, no. 4, pp. 886–908, 2018
2018
Earlier work this paper cites.
K. Jothimurugan, R. Alur, and O. Bastani, “A composable specification language for reinforcement learning tasks,” in NeurIPS , 2019, pp. 13 041–13 051
2019
Earlier work this paper cites.
R. T. Icarte, E. Waldie, T. Klassen, and et al., “Learning reward machines for partially observable reinforcement learning,” in NeurIPS , 2019, pp. 15 523–15 534
2019
Earlier work this paper cites.
A. Camacho, R. T. Icarte, T. Q. Klassen, and et al., “Ltl and beyond: Formal languages for reward function specification in reinforcement learning.” in IJCAI , 2019, pp. 6065–6073
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
D. Lyu, F. Yang, B. Liu, and S. Gustafson, “Sdrl: interpretable and data-efficient deep reinforcement learning leveraging symbolic planning,” in AAAI , 2019, pp. 2970–2977
2019
Earlier work this paper cites.
G. De Giacomo, L. Iocchi, and et al., “Foundations for restraining bolts: Reinforcement learning with ltlf/ldlf restraining specifications,” in ICAPS , 2019, pp. 128–136
2019
Earlier work this paper cites.
M. Hasanbeig and et al., “Reinforcement learning for temporal logic control synthesis with probabilistic satisfaction guarantees,” in CDC , 2019, pp. 5338–5343
2019
Earlier work this paper cites.
J. Kim, C. Muise, A. Shah, and et al., “Bayesian inference of linear temporal logic specifications for contrastive explanations,” in IJCAI , 2019, pp. 5591–5598
2019
Earlier work this paper cites.
A. Camacho and S. A. McIlraith, “Learning interpretable models expressed in linear temporal logic,” in ICAPS , 2019, pp. 621–630
2019
Earlier work this paper cites.
Z. Xu and U. Topcu, “Transfer of temporal logic formulas in reinforcement learning,” in IJCAI , 2019, pp. 4010–4018
2019
Earlier work this paper cites.
B. Van Niekerk, S. James, A. Earle, and B. Rosman, “Composing value functions in reinforcement learning,” in ICML , 2019, pp. 6401–6409
2019
Earlier work this paper cites.
A. Verma, H. Le, and et al., “Imitation-projected programmatic reinforcement learning,” NeurIPS , vol. 32, 2019
2019
Earlier work this paper cites.
Z. Jiang and S. Luo, “Neural logic reinforcement learning,” in International conference on machine learning . PMLR, 2019, pp. 3110–3119
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
X. Li, Z. Serlin, G. Yang, and C. Belta, “A formal methods approach to interpretable reinforcement learning for robotic planning,” Science Robotics , vol. 4, no. 37, 2019
2019
Earlier work this paper cites.
N. Vithayathil Varghese and Q. H. Mahmoud, “A survey of multi-task deep reinforcement learning,” Electronics , vol. 9, no. 9, p. 1363, 2020
2020
Earlier work this paper cites.
2020
Cited alongside, same era.
2020
Cited alongside, same era.
Z. Xu, I. Gavran, and et al., “Joint inference of reward machines and policies for reinforcement learning,” in ICAPS , 2020, pp. 590–598
2020
Cited alongside, same era.
D. Furelos-Blanco and et al., “Induction of subgoal automata for reinforcement learning,” in AAAI , 2020, pp. 3890–3897
2020
Cited alongside, same era.
Z. Wu, C. Yu, C. Chen, J. Hao, and H. H. Zhuo, “Plan to predict: Learning an uncertainty-foreseeing model for model-based reinforcement learning,” in NeurIPS , 2022
2022
Later among the works it cites.
X. Zheng, C. Yu, and M. Zhang, “Lifelong reinforcement learning with temporal logic formulas and reward machines,” Knowledge-Based Systems , vol. 257, p. 109650, 2022
2022
Later among the works it cites.
A. Mumuni and F. Mumuni, “Data augmentation: A comprehensive survey of modern approaches,” Array , vol. 16, p. 100258, 2022
2022
Later among the works it cites.
W. Zhou and W. Li, “Programmatic reward design by example,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 36, no. 8, 2022, pp. 9233–9241
2022
Later among the works it cites.
Y. Cao, Z. Li, T. Yang, H. Zhang, Y. Zheng, Y. Li, J. Hao, and Y. Liu, “Galois: boosting deep reinforcement learning via generalizable logic synthesis,” Advances in Neural Information Processing Systems , vol. 35, pp. 19 930–19 943, 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
F. Djeumou, Z. Xu, and U. Topcu, “Probabilistic swarm guidance subject to graph temporal logic specifications,” in Robotics: Science and Systems (RSS) , 2020
2020
Cited alongside, same era.
2020
Cited alongside, same era.
L. Illanes, X. Yan, R. T. Icarte, and S. A. McIlraith, “Symbolic plans as high-level instructions for reinforcement learning,” in ICAPS , 2020, pp. 540–550
2020
Cited alongside, same era.
A. K. Bozkurt, Y. Wang, and et al., “Control synthesis from linear temporal logic specifications using model-free reinforcement learning,” in IEEE ICRA , 2020, pp. 10 349–10 355
2020
Cited alongside, same era.
J. Schrittwieser, I. Antonoglou, T. Hubert, and et al., “Mastering atari, go, chess and shogi by planning with a learned model,” Nat. , vol. 588, no. 7839, pp. 604–609, 2020. [Online]. Available: https://doi.org/10.1038/s41586-020-03051-4
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2022
Later among the works it cites.
X. Peng, M. Riedl, and P. Ammanabrolu, “Inherently explainable reinforcement learning in natural language,” Advances in Neural Information Processing Systems , vol. 35, pp. 16 178–16 190, 2022
2022
Later among the works it cites.
Z. Zhao, J. Xun, X. Wen, and J. Chen, “Safe reinforcement learning for single train trajectory optimization via shield sarsa,” IEEE Transactions on Intelligent Transportation Systems , vol. 24, no. 1, pp. 412–428, 2022
2022
Later among the works it cites.
J. X. Liu, Z. Yang, B. Schornstein, S. Liang, I. Idrees, S. Tellex, and A. Shah, “Lang2ltl: Translating natural language commands to temporal specification with large language models,” in Workshop on Language and Robotics at CoRL 2022 , 2022
2022
Later among the works it cites.
T. M. Moerland, J. Broekens, A. Plaat, C. M. Jonker et al. , “Model-based reinforcement learning: A survey,” Foundations and Trends® in Machine Learning , vol. 16, no. 1, pp. 1–118, 2023
2023
Closest in time.
R. T. Icarte, T. Q. Klassen, R. Valenzano, M. P. Castro, E. Waldie, and S. A. McIlraith, “Learning reward machines: A study in partially observable reinforcement learning,” Artificial Intelligence , vol. 323, p. 103989, 2023
2023
Closest in time.
G. Varricchione, N. Alechina, M. Dastani, and B. Logan, “Synthesising reward machines for cooperative multi-agent reinforcement learning,” in European Conference on Multi-Agent Systems . Springer, 2023, pp. 328–344
2023
Closest in time.
L. Ardon, D. Furelos-Blanco, and A. Russo, “Learning reward machines in cooperative multi-agent tasks,” in International Conference on Autonomous Agents and Multiagent Systems . Springer, 2023, pp. 43–59
2023
Closest in time.
W. Hatanaka, R. Yamashina, and T. Matsubara, “Reinforcement learning of action and query policies with ltl instructions under uncertain event detector,” IEEE Robotics and Automation Letters , 2023
2023
Closest in time.
D. Furelos-Blanco, M. Law, A. Jonsson, K. Broda, and A. Russo, “Hierarchies of reward machines,” in International Conference on Machine Learning . PMLR, 2023, pp. 10 494–10 541
2023
Closest in time.
H. Bourel, A. Jonsson, O.-A. Maillard, and M. S. Talebi, “Exploration in reward machines with low regret,” in International Conference on Artificial Intelligence and Statistics . PMLR, 2023, pp. 4114–4146
2023
Closest in time.
J. Corazza, H. P. Aria, D. Neider, and Z. Xu, “Expediting reinforcement learning by incorporating temporal causal information,” in Causal Representation Learning Workshop at NeurIPS 2023
2023
Closest in time.
H. Sun and F. Wu, “Less is more: Refining datasets for offline reinforcement learning with reward machines,” in Proceedings of the 2023 International Conference on Autonomous Agents and Multiagent Systems (AAMAS) , 2023, pp. 1239–1247
2023
Closest in time.
C. Koprulu and U. Topcu, “Reward-machine-guided, self-paced reinforcement learning,” in Uncertainty in Artificial Intelligence . PMLR, 2023, pp. 1121–1131
2023
Closest in time.
C. Voloshin, A. Verma, and Y. Yue, “Eventual discounting temporal logic counterfactual experience replay,” in International Conference on Machine Learning . PMLR, 2023, pp. 35 137–35 150
2023
Closest in time.
A. Abate, Y. Almulla, J. Fox, D. Hyland, and M. Wooldridge, “Learning task automata for reinforcement learning using hidden markov models,” in ECAI 2023 . IOS Press, 2023, pp. 3–10
2023
Closest in time.
Z. Wu, C. Yu, and et al., “Models as agents: Optimizing multi-step predictions of interactive local models in model-based multi-agent reinforcement learning,” in AAAI , 2023
2023
Closest in time.
L. Guan, K. Valmeekam, S. Sreedharan, and S. Kambhampati, “Leveraging pre-trained large language models to construct and utilize world models for model-based task planning,” Advances in Neural Information Processing Systems , vol. 36, pp. 79 081–79 094, 2023
2023
Closest in time.
Z. Bing, A. Koch, X. Yao, K. Huang, and A. Knoll, “Meta-reinforcement learning via language instructions,” in 2023 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2023, pp. 5985–5991
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
Z. Yu, Y. Tao, L. Chen, T. Sun, and H. Yang, “B-coder: Value-based deep reinforcement learning for program synthesis.” CoRR , 2023
2023
Closest in time.
2023
Closest in time.
N. Bougie, T. Onishi, and Y. Tsuruoka, “Interpretable imitation learning with symbolic rewards,” ACM Transactions on Intelligent Systems and Technology , vol. 15, no. 1, pp. 1–34, 2023
2023
Closest in time.
G.-T. Liu, E.-P. Hu, P.-J. Cheng, H.-Y. Lee, and S.-H. Sun, “Hierarchical programmatic reinforcement learning via learning to compose programs,” in International Conference on Machine Learning . PMLR, 2023, pp. 21 672–21 697
2023
Closest in time.
2023
Closest in time.
H. Hasanbeig, D. Kroening, and A. Abate, “Certified reinforcement learning with logic guidance,” Artificial Intelligence , vol. 322, p. 103949, 2023
2023
Closest in time.
S. Carr, N. Jansen, S. Junges, and U. Topcu, “Safe reinforcement learning via shielding under partial observability,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 37, no. 12, 2023, pp. 14 748–14 756
2023
Closest in time.
C. Glanois, P. Weng, M. Zimmer, D. Li, T. Yang, J. Hao, and W. Liu, “A survey on interpretable reinforcement learning,” Machine Learning , pp. 1–44, 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
C. Zhu, W. Si, J. Zhu, and Z. Jiang, “Decomposing temporal equilibrium strategy for coordinated distributed multi-agent reinforcement learning,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 38, no. 16, 2024, pp. 17 618–17 627
2024
Closest in time.
K. Terashima, K. Kobayashi, and Y. Yamashita, “On reward distribution in reinforcement learning of multi-agent surveillance systems with temporal logic specifications,” Advanced Robotics , vol. 38, no. 6, pp. 386–397, 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
J. X. Liu, A. Shah, E. Rosen, M. Jia, G. Konidaris, and S. Tellex, “Skill transfer for temporal task specification,” in 2024 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2024, pp. 2535–2541
2024
Closest in time.
G. Azran, M. H. Danesh, S. V. Albrecht, and S. Keren, “Contextual pre-planning on reward machine abstractions for enhanced transfer in deep reinforcement learning,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 38, no. 10, 2024, pp. 10 953–10 961
2024
Closest in time.
W. Qiu, W. Mao, and H. Zhu, “Instructing goal-conditioned reinforcement learning agents with temporal logic objectives,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
D. Xu and F. Fekri, “Generalization of temporal logic tasks via future dependent options,” Machine Learning , pp. 1–32, 2024
2024
Closest in time.
D. Kuric, G. Infante, V. Gómez, A. Jonsson, and H. van Hoof, “Planning with a learned policy basis to optimally solve complex tasks,” in Proceedings of the International Conference on Automated Planning and Scheduling , vol. 34, 2024, pp. 333–341
2024
Closest in time.
L. Cao, C. Wang, J. Qi, and Y. Peng, “Exploring into the unseen: Enhancing language-conditioned policy generalization with behavioral information,” Cyborg and Bionic Systems , vol. 5, p. 0084, 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
A. Rashidi Laleh and M. Nili Ahmadabadi, “A survey on enhancing reinforcement learning in complex environments: Insights from human and llm feedback,” arXiv e-prints , pp. arXiv–2411, 2024
2024
Closest in time.
J. Guo, R. Zhang, S. Peng, Q. Yi, X. Hu, R. Chen, Z. Du, L. Li, Q. Guo, Y. Chen et al. , “Efficient symbolic policy learning with differentiable symbolic expression,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
Q. Delfosse, H. Shindo, D. Dhami, and K. Kersting, “Interpretable and explainable logical policies via neurally guided symbolic abstraction,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.