Fetching the paper…
Reading the bibliography…
A key challenge for reinforcement learning is solving long-horizon planning problems.
Deepsynth: Program synthesis for automatic task segmentation in deep reinforcement learning
Mohammadhosein Hasanbeig, Natasha Yogananda Jeppu, Alessandro Abate, Tom Melham, and Daniel Kroening · 1911
Earlier work this paper cites.
Strips: A new approach to the application of theorem proving to problem solving
Richard E Fikes and Nils J Nilsson · 1971
Earlier work this paper cites.
Inferring lisp programs from examples
David E. Shaw, William R. Swartout, and C. Cordell Green · 1975
Earlier work this paper cites.
The complexity of optimization problems
M W Krentel · 1986
Earlier work this paper cites.
Probabilistic planning with information gathering and contingent execution
Denise Draper, S. Hanks, and Daniel S. Weld · 1994
Earlier work this paper cites.
The focussed dˆ* algorithm for real-time replanning
Anthony Stentz et al · 1995
Earlier work this paper cites.
High-level planning and control with incomplete information using pomdp’s
Blai Bonet · 1998
Earlier work this paper cites.
Actor-critic algorithms
Vijay Konda and John Tsitsiklis · 1999
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Richard S Sutton, Doina Precup, and Satinder Singh · 1999
Earlier work this paper cites.
Learning options in reinforcement learning
Martin Stolle and Doina Precup · 2002
Earlier work this paper cites.
Automated hierarchy discovery for planning in partially observable environments
Laurent Charlin, Pascal Poupart, and Romy Shioda · 2007
Earlier work this paper cites.
Learning differentiable programs with admissible neural heuristics
Ameesh Shah, Eric Zhan, Jennifer J. Sun, Abhinav Verma, Yisong Yue, and Swarat Chaudhuri · 2007
Earlier work this paper cites.
Z3: An efficient smt solver
Leonardo De Moura and Nikolaj Bjørner · 2008
Earlier work this paper cites.
Program Synthesis by Sketching
Armando Solar-Lezama · 2008
Earlier work this paper cites.
Hierarchical pomdp controller optimization by likelihood maximization
Marc Toussaint, Laurent Charlin, and Pascal Poupart · 2008
Earlier work this paper cites.
Inductive program synthesis over noisy data
Shivam Handa and Martin Rinard · 2009
Earlier work this paper cites.
Integrated task and motion planning
Caelan Reed Garrett, Rohan Chitnis, Rachel Holladay, Beomjoon Kim, Tom Silver, Leslie Pack Kaelbling, and Tomás Lozano-Pérez · 2010
Earlier work this paper cites.
Automating string processing in spreadsheets using input-output examples
Sumit Gulwani · 2011
Earlier work this paper cites.
Generalized planning: Synthesizing plans that work for multiple environments
Yuxiao Hu and Giuseppe De Giacomo · 2011
Earlier work this paper cites.
Hierarchical task and motion planning in the now
Leslie Pack Kaelbling and Tomás Lozano-Pérez · 2011
Earlier work this paper cites.
Foundations and applications of generalized planning
Siddharth Srivastava · 2011
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Earlier work this paper cites.
Integrated task and motion planning in belief space
Leslie Pack Kaelbling and Tomás Lozano-Pérez · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2013
Earlier work this paper cites.
Unsupervised learning by program synthesis
Kevin Ellis, Armando Solar-Lezama, and Josh Tenenbaum · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
Learning structured output representation using deep conditional generative models
Kihyuk Sohn, Honglak Lee, and Xinchen Yan · 2015
Cited alongside, same era.
High-dimensional continuous control using generalized advantage estimation
John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, and Pieter Abbeel · 2016
Cited alongside, same era.
A deep hierarchical approach to lifelong learning in minecraft, 2016
Chen Tessler, Shahar Givony, Tom Zahavy, Daniel J. Mankowitz, and Shie Mannor · 2016
Cited alongside, same era.
Modular multitask reinforcement learning with policy sketches
Jacob Andreas, Dan Klein, and Sergey Levine · 2017
Cited alongside, same era.
Deep reinforcement learning: A brief survey
Neural guided constraint logic programming for program synthesis
Lisa Zhang, Gregory Rosenblatt, Ethan Fetaya, Renjie Liao, William E. Byrd, Matthew Might, Raquel Urtasun, and Richard S. Zemel · 2018
Later among the works it cites.
Execution-guided neural program synthesis
Xinyun Chen, Chang Liu, and Dawn Song · 2019
Later among the works it cites.
Write, execute, assess: Program synthesis with a repl
Kevin Ellis, Maxwell Nye, Yewen Pu, Felix Sosa, Josh Tenenbaum, and Armando Solar-Lezama · 2019
Later among the works it cites.
Generalized planning via abstraction: Arbitrary numbers of objects
León Illanes and Sheila A. McIlraith · 2019
Later among the works it cites.
A composable specification language for reinforcement learning tasks
Kishor Jothimurugan, Rajeev Alur, and Osbert Bastani · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kai Arulkumaran, Marc Peter Deisenroth, Miles Brundage, and Anil Anthony Bharath · 2017
Cited alongside, same era.
Deepcoder: Learning to write programs, 2017
Matej Balog, Alexander L. Gaunt, Marc Brockschmidt, Sebastian Nowozin, and Daniel Tarlow · 2017
Cited alongside, same era.
Robustfill: Neural program learning under noisy i/o
Jacob Devlin, Jonathan Uesato, Surya Bhupatiraju, Rishabh Singh, Abdel-rahman Mohamed, and Pushmeet Kohli · 2017
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Cited alongside, same era.
Adam: A method for stochastic optimization, 2017
Diederik P. Kingma and Jimmy Ba · 2017
Cited alongside, same era.
Attention is all you need, 2017
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Synthesis of data completion scripts using finite tree automata
Xinyu Wang, Isil Dillig, and Rishabh Singh · 2017
Cited alongside, same era.
Camille Phiquepal and Marc Toussaint · 2019
Later among the works it cites.
Learning to infer and execute 3d shape programs
Yonglong Tian, Andrew Luo, Xingyuan Sun, Kevin Ellis, William T Freeman, Joshua B Tenenbaum, and Jiajun Wu · 2019
Later among the works it cites.
Verifiable and interpretable reinforcement learning through program synthesis
Abhinav Verma · 2019
Later among the works it cites.
Imitation-projected programmatic reinforcement learning
Abhinav Verma, Hoang M Le, Yisong Yue, and Swarat Chaudhuri · 2019
Later among the works it cites.
Learning neurosymbolic generative models via program synthesis
Halley Young, Osbert Bastani, and Mayur Naik · 2019
Later among the works it cites.
Deep reinforcement learning with relational inductive biases
Vinicius Zambaldi, David Raposo, Adam Santoro, Victor Bapst, Yujia Li, Igor Babuschkin, Karl Tuyls, David Reichert, Timothy Lillicrap, Edward Lockhart, Murray Shanahan, Victoria Langston, Razvan Pascanu, Matthew Botvinick, Oriol Vinyals, and Peter Battaglia · 2019
Later among the works it cites.
Dac: The double actor-critic architecture for learning options, 2019
Shangtong Zhang and Shimon Whiteson · 2019
Later among the works it cites.
Value preserving state-action abstractions
David Abel, Nate Umbanhowar, Khimya Khetarpal, Dilip Arumugam, Doina Precup, and Michael Littman · 2020
Later among the works it cites.
Neurosymbolic reinforcement learning with formally verified exploration
Greg Anderson, Abhinav Verma, Isil Dillig, and Swarat Chaudhuri · 2020
Later among the works it cites.
Option discovery using deep skill chaining
Akhil Bagaria and George Konidaris · 2020
Later among the works it cites.
Program synthesis using deduction-guided reinforcement learning
Yanju Chen, Chenglong Wang, Osbert Bastani, Isil Dillig, and Yu Feng · 2020
Later among the works it cites.
Generating programmatic referring expressions via program synthesis
Jiani Huang, Calvin Smith, Osbert Bastani, Rishabh Singh, Aws Albarghouthi, and Mayur Naik · 2020
Later among the works it cites.
Sub-policy adaptation for hierarchical reinforcement learning, 2020
Alexander C. Li, Carlos Florensa, Ignasi Clavera, and Pieter Abbeel · 2020
Later among the works it cites.
Representing partial programs with blended abstract semantics
Maxwell Nye, Yewen Pu, Matthew Bowers, Jacob Andreas, Joshua B Tenenbaum, and Armando Solar-Lezama · 2020
Later among the works it cites.
Program guided agent
Shao-Hua Sun, Te-Lin Wu, and Joseph J. Lim · 2020
Later among the works it cites.
Program synthesis with large language models
Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie J. Cai, Michael Terry, Quoc V. Le, and Charles Sutton · 2021
Closest in time.
Web question answering with neurosymbolic program synthesis
Qiaochu Chen, Aaron Lamoreaux, Xinyu Wang, Greg Durrett, Osbert Bastani, and Isil Dillig · 2021
Closest in time.
Dreamcoder: bootstrapping inductive program synthesis with wake-sleep library learning
Kevin Ellis, Catherine Wong, Maxwell Nye, Mathias Sablé-Meyer, Lucas Morales, Luke Hewitt, Luc Cary, Armando Solar-Lezama, and Joshua B Tenenbaum · 2021
Closest in time.
Data-efficient hindsight off-policy option learning, 2021
Markus Wulfmeier, Dushyant Rao, Roland Hafner, Thomas Lampe, Abbas Abdolmaleki, Tim Hertweck, Michael Neunert, Dhruva Tirumala, Noah Siegel, Nicolas Heess, and Martin Riedmiller · 2021
Closest in time.
Hierarchical reinforcement learning by discovering intrinsic options, 2021
Jesse Zhang, Haonan Yu, and Wei Xu · 2021
Closest in time.