Fetching the paper…
Reading the bibliography…
Goal-conditioned reinforcement learning is a powerful way to control an AI agent's behavior at runtime.
Language identification in the limit
E Mark Gold · 1967
Earlier work this paper cites.
An n log n algorithm for minimizing states in a finite automaton
John Hopcroft · 1971
Earlier work this paper cites.
Introduction to Automata Theory, Languages and Computation
John E. Hopcroft and Jeffrey D. Ullman · 1979
Earlier work this paper cites.
Learning regular sets from queries and counterexamples
Dana Angluin · 1987
Earlier work this paper cites.
Inferring regular languages in polynomial updated time
José Oncina and Pedro Garcia · 1992
Earlier work this paper cites.
Learning dfa from simple examples
Rajesh Parekh and Vasant Honavar · 2001
Earlier work this paper cites.
Visualizing data using t-sne
Laurens Van der Maaten and Geoffrey Hinton · 2008
Earlier work this paper cites.
From church and prior to psl
Moshe Y Vardi · 2008
Earlier work this paper cites.
Exact DFA identification using SAT solvers
Marijn Heule and Sicco Verwer · 2010
Earlier work this paper cites.
Robust control of uncertain markov decision processes with temporal logic specifications
Eric M Wolff, Ufuk Topcu, and Richard M Murray · 2012
Earlier work this paper cites.
Optimal control of markov decision processes with linear temporal logic constraints
Xuchu Ding, Stephen L Smith, Calin Belta, and Daniela Rus · 2014
Earlier work this paper cites.
Probably approximately correct mdp learning and control with temporal logic constraints
Jie Fu and Ufuk Topcu · 2014
Earlier work this paper cites.
A learning based approach to control synthesis of markov decision processes for linear temporal logic specifications
Dorsa Sadigh, Eric S Kim, Samuel Coogan, S Shankar Sastry, and Sanjit A Seshia · 2014
Earlier work this paper cites.
Universal value function approximators
Tom Schaul, Daniel Horgan, Karol Gregor, and David Silver · 2015
Earlier work this paper cites.
Q-learning for robust satisfaction of signal temporal logic specifications
Derya Aksaray, Austin Jones, Zhaodan Kong, Mac Schwager, and Calin Belta · 2016
Earlier work this paper cites.
Modular multitask reinforcement learning with policy sketches
Jacob Andreas, Dan Klein, and Sergey Levine · 2017
Earlier work this paper cites.
Hindsight experience replay
Marcin Andrychowicz, Filip Wolski, Alex Ray, Jonas Schneider, Rachel Fong, Peter Welinder, Bob McGrew, Josh Tobin, OpenAI Pieter Abbeel, and Wojciech Zaremba · 2017
Cited alongside, same era.
Reinforcement learning with temporal logic rewards
Xiao Li, Cristian-Ioan Vasile, and Calin Belta · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Grammatical Inference
Wojciech Wieczorek · 2017
Cited alongside, same era.
Learning to understand goal specifications by modelling reward
Dzmitry Bahdanau, Felix Hill, Jan Leike, Edward Hughes, Arian Hosseini, Pushmeet Kohli, and Edward Grefenstette · 2018
Cited alongside, same era.
Compositional reinforcement learning from logical specifications
Kishor Jothimurugan, Suguman Bansal, Osbert Bastani, and Rajeev Alur · 2021
Later among the works it cites.
Ltl2action: Generalizing ltl instructions for multi-task rl
Pashootan Vaezipoor, Andrew C Li, Rodrigo A Toro Icarte, and Sheila A Mcilraith · 2021
Later among the works it cites.
Demonstration informed specification search
Marcell Vazquez-Chanlatte, Ameesh Shah, Gil Lederman, and Sanjit A Seshia · 2021
Later among the works it cites.
On the (in) tractability of reinforcement learning for ltl objectives
Cambridge Yang, Michael Littman, and Michael Carbin · 2021
Later among the works it cites.
A framework for transforming specifications in reinforcement learning
Rajeev Alur, Suguman Bansal, Osbert Bastani, and Kishor Jothimurugan · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alberto Camacho, Meghyn Bienvenu, and Sheila A McIlraith · 2018
Cited alongside, same era.
Modeling relational data with graph convolutional networks
Michael Schlichtkrull, Thomas N Kipf, Peter Bloem, Rianne Van Den Berg, Ivan Titov, and Max Welling · 2018
Cited alongside, same era.
Ltl and beyond: Formal languages for reward function specification in reinforcement learning
Alberto Camacho, Rodrigo Toro Icarte, Toryn Q Klassen, Richard Anthony Valenzano, and Sheila A McIlraith · 2019
Cited alongside, same era.
Language as an abstraction for hierarchical deep reinforcement learning
Yiding Jiang, Shixiang Shane Gu, Kevin P Murphy, and Chelsea Finn · 2019
Cited alongside, same era.
A composable specification language for reinforcement learning tasks
Kishor Jothimurugan, Rajeev Alur, and Osbert Bastani · 2019
Cited alongside, same era.
A survey of reinforcement learning informed by natural language
Jelena Luketina, Nantas Nardelli, Gregory Farquhar, Jakob Foerster, Jacob Andreas, Edward Grefenstette, Shimon Whiteson, and Tim Rocktäschel · 2019
Cited alongside, same era.
Benchmarking safe exploration in deep reinforcement learning
Alex Ray, Joshua Achiam, and Dario Amodei · 2019
Cited alongside, same era.
Reward machines: Exploiting reward function structure in reinforcement learning
Rodrigo Toro Icarte, Toryn Q Klassen, Richard Valenzano, and Sheila A McIlraith · 2022
Later among the works it cites.
Learning deterministic finite automata decompositions from examples and demonstrations
Niklas Lauffer, Beyazit Yalcinkaya, Marcell Vazquez-Chanlatte, Ameesh Shah, and Sanjit A Seshia · 2022
Later among the works it cites.
Specifications from Demonstrations: Learning, Teaching, and Control
Marcell Jose Vazquez-Chanlatte · 2022
Later among the works it cites.
Policy optimization with linear temporal logic constraints
Cameron Voloshin, Hoang Le, Swarat Chaudhuri, and Yisong Yue · 2022
Later among the works it cites.
Policy synthesis and reinforcement learning for discounted ltl
Rajeev Alur, Osbert Bastani, Kishor Jothimurugan, Mateo Perez, Fabio Somenzi, and Ashutosh Trivedi · 2023
Later among the works it cites.
Automata conditioned reinforcement learning with experience replay
Beyazit Yalcinkaya, Niklas Lauffer, Marcell Vazquez-Chanlatte, and Sanjit Seshia · 2023
Later among the works it cites.
A pac learning algorithm for ltl and omega-regular objectives in mdps
Mateo Perez, Fabio Somenzi, and Ashutosh Trivedi · 2024
Closest in time.
Instructing goal-conditioned reinforcement learning agents with temporal logic objectives
Wenjie Qiu, Wensen Mao, and He Zhu · 2024
Closest in time.
Deep policy optimization with temporal logic constraints
Ameesh Shah, Cameron Voloshin, Chenxi Yang, Abhinav Verma, Swarat Chaudhuri, and Sanjit A Seshia · 2024
Closest in time.
Roboclip: One demonstration is enough to learn robot policies
Sumedh Sontakke, Jesse Zhang, Séb Arnold, Karl Pertsch, Erdem Bıyık, Dorsa Sadigh, Chelsea Finn, and Laurent Itti · 2024
Closest in time.