Fetching the paper…
Reading the bibliography…
It is desirable for an agent to be able to solve a rich variety of problems that can be specified through language in the same environment.
The temporal logic of programs
Amir Pnueli · 1977
Earlier work this paper cites.
C. Watkins · 1989
Earlier work this paper cites.
Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning
Richard Sutton, Doina Precup, and Satinder Singh · 1999
Earlier work this paper cites.
Model checking of safety properties
Orna Kupferman and Moshe Y Vardi · 2001
Earlier work this paper cites.
Handbook of automated reasoning , volume 1
Alan JA Robinson and Andrei Voronkov · 2001
Earlier work this paper cites.
Skill discovery in continuous reinforcement learning domains using skill chaining
George Konidaris and Andrew Barto · 2009
Earlier work this paper cites.
Transfer learning for reinforcement learning domains: a survey
Matthew. Taylor and Peter Stone · 2009
Earlier work this paper cites.
Compositionality of optimal control laws
Emanuel Todorov · 2009
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei Rusu, Joel Veness, Marc Bellemare, Alex Graves, Martin Riedmiller, Andreas Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Spot 2.0—a framework for ltl and ω \omega -automata manipulation
Alexandre Duret-Lutz, Alexandre Lewkowicz, Amaury Fauchille, Thibaud Michaud, Etienne Renault, and Laurent Xu · 2016
Earlier work this paper cites.
End-to-end training of deep visuomotor policies
Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel · 2016
Earlier work this paper cites.
Reinforcement learning with temporal logic rewards
Xiao Li, Cristian-Ioan Vasile, and Calin Belta · 2017
Earlier work this paper cites.
Environment-independent task specifications via GLTL
Michael Littman, Ufuk Topcu, Jie Fu, Charles Isbell, Min Wen, and James MacGlashan · 2017
Earlier work this paper cites.
Spinning Up in Deep Reinforcement Learning
Joshua Achiam · 2018
Cited alongside, same era.
Transfer in deep reinforcement learning using successor features and generalised policy improvement
Andre Barreto, Diana Borsa, John Quan, Tom Schaul, David Silver, Matteo Hessel, Daniel Mankowitz, Augustin Zidek, and Remi Munos · 2018
Cited alongside, same era.
Addressing function approximation error in actor-critic methods
Scott Fujimoto, Herke Hoof, and David Meger · 2018
Cited alongside, same era.
Composable deep reinforcement learning for robotic manipulation
Tuomas Haarnoja, Vitchyr Pong, Aurick Zhou, Murtaza Dalal, Pieter Abbeel, and Sergey Levine · 2018
Cited alongside, same era.
Rudder: Return decomposition for delayed rewards
Jose Arjona-Medina, Michael Gillhofer, Michael Widrich, Thomas Unterthiner, Johannes Brandstetter, and Sepp Hochreiter · 2019
Cited alongside, same era.
The option keyboard: Combining skills in reinforcement learning
Agent57: Outperforming the Atari human benchmark
Adrià Puigdomènech Badia, Bilal Piot, Steven Kapturowski, Pablo Sprechmann, Alex Vitvitskyi, Zhaohan Daniel Guo, and Charles Blundell · 2020
Later among the works it cites.
Fast reinforcement learning with generalized policy updates
André Barreto, Shaobo Hou, Diana Borsa, David Silver, and Doina Precup · 2020
Later among the works it cites.
A Boolean task algebra for reinforcement learning
Geraud Nangue Tasse, Steven James, and Benjamin Rosman · 2020
Later among the works it cites.
The logical options framework
Brandon Araki, Xiao Li, Kiran Vodrahalli, Jonathan DeCastro, Micah Fry, and Daniela Rus · 2021
Later among the works it cites.
Learning to follow language instructions with compositional policies
Vanya Cohen, Geraud Nangue Tasse, Nakul Gopalan, Steven James, Matthew Gombolay, and Benjamin Rosman · 2021
Later among the works it cites.
Compositional reinforcement learning from logical specifications
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
André Barreto, Diana Borsa, Shaobo Hou, Gheorghe Comanici, Eser Aygün, Philippe Hamel, Daniel Toyama, Shibl Mourad, David Silver, Doina Precup, et al · 2019
Cited alongside, same era.
Ltl and beyond: Formal languages for reward function specification in reinforcement learning
Alberto Camacho, Rodrigo Toro Icarte, Toryn Q Klassen, Richard Anthony Valenzano, and Sheila A McIlraith · 2019
Cited alongside, same era.
Composing entropic policies using divergence correction
Jonathan Hunt, Andre Barreto, Timothy Lillicrap, and Nicolas Heess · 2019
Cited alongside, same era.
A composable specification language for reinforcement learning tasks
Kishor Jothimurugan, Rajeev Alur, and Osbert Bastani · 2019
Cited alongside, same era.
MCP: Learning composable hierarchical control with multiplicative compositional policies
Xue Bin Peng, Michael Chang, Grace Zhang, Pieter Abbeel, and Sergey Levine · 2019
Cited alongside, same era.
Benchmarking safe exploration in deep reinforcement learning.(2019)
Alex Ray, Joshua Achiam, and Dario Amodei · 2019
Cited alongside, same era.
Composing value functions in reinforcement learning
Benjamin Van Niekerk, Steven James, Adam Earle, and Benjamin Rosman · 2019
Cited alongside, same era.
Kishor Jothimurugan, Suguman Bansal, Osbert Bastani, and Rajeev Alur · 2021
Later among the works it cites.
Ltl2action: Generalizing ltl instructions for multi-task rl
Pashootan Vaezipoor, Andrew C Li, Rodrigo A Toro Icarte, and Sheila A Mcilraith · 2021
Later among the works it cites.
Constructing a good behavior basis for transfer using generalized policy updates
Safa Alver and Doina Precup · 2022
Closest in time.
End-to-end learning to follow language instructions with compositional policies
Vanya Cohen, Geraud Nangue Tasse, Nakul Gopalan, Steven James, Ray Mooney, and Benjamin Rosman · 2022
Closest in time.
Hierarchical reinforcement learning: A survey and open research challenges
Matthias Hutsebaut-Buysse, Kevin Mets, and Steven Latré · 2022
Closest in time.
Reward machines: Exploiting reward function structure in reinforcement learning
Rodrigo Toro Icarte, Toryn Q Klassen, Richard Valenzano, and Sheila A McIlraith · 2022
Closest in time.
Skill transfer for temporally-extended task specifications
Jason Xinyu Liu, Ankit Shah, Eric Rosen, George Konidaris, and Stefanie Tellex · 2022
Closest in time.