Fetching the paper…
Reading the bibliography…
Despite achieving superior performance in human-level control problems, unlike humans, deep reinforcement learning (DRL) lacks high-order intelligence (e.g., logic deduction and reuse), thus it behaves ineffectively than humans regarding learning and generalization in complex problems.
The perceptron: a probabilistic model for information storage and organization in the brain
Frank Rosenblatt · 1958
Earlier work this paper cites.
Toward a mathematical semantics for computer languages
Dana S Scott and Christopher Strachey · 1971
Earlier work this paper cites.
The republic of Plato
Francis Macdonald Cornford et al · 1976
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Artificial intelligence
Patrick Henry Winston · 1992
Earlier work this paper cites.
Logic programming and knowledge representation
Chitta Baral and Michael Gelfond · 1994
Earlier work this paper cites.
An introduction to proof theory
Samuel R Buss · 1998
Earlier work this paper cites.
Models, reasoning and inference
Judea Pearl et al · 2000
Earlier work this paper cites.
Introduction to statistical relational learning
Daphne Koller, Nir Friedman, Sašo Džeroski, Charles Sutton, Andrew McCallum, Avi Pfeffer, Pieter Abbeel, Ming-Fai Wong, Chris Meek, Jennifer Neville, et al · 2007
Earlier work this paper cites.
Program synthesis by sketching
Armando Solar-Lezama · 2008
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
Automated feedback generation for introductory programming assignments
Rishabh Singh, Sumit Gulwani, and Armando Solar-Lezama · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Inferring algorithmic patterns with stack-augmented recurrent nets
Armand Joulin and Tomas Mikolov · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation
Tejas D Kulkarni, Karthik Narasimhan, Ardavan Saeedi, and Josh Tenenbaum · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
Robustfill: Neural program learning under noisy i/o
Jacob Devlin, Jonathan Uesato, Surya Bhupatiraju, Rishabh Singh, Abdel-rahman Mohamed, and Pushmeet Kohli · 2017
Cited alongside, same era.
Survey of model-based reinforcement learning: Applications on robotics
Athanasios S Polydoros and Lazaros Nalpantidis · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Neural logic reinforcement learning
Zhengyao Jiang and Shan Luo · 2019
Later among the works it cites.
A composable specification language for reinforcement learning tasks
Kishor Jothimurugan, Rajeev Alur, and Osbert Bastani · 2019
Later among the works it cites.
Learning to infer program sketches
Maxwell Nye, Luke Hewitt, Joshua Tenenbaum, and Armando Solar-Lezama · 2019
Later among the works it cites.
Program guided agent
Shao-Hua Sun, Te-Lin Wu, and Joseph J Lim · 2019
Later among the works it cites.
An inductive synthesis framework for verifiable reinforcement learning
He Zhu, Zikang Xiong, Stephen Magill, and Suresh Jagannathan · 2019
Later among the works it cites.
Explainable reinforcement learning: A survey
Erika Puiutta and Eric Veith · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Computational models of affordance in robotics: a taxonomy and systematic classification
Philipp Zech, Simon Haller, Safoura Rezapour Lakani, Barry Ridge, Emre Ugur, and Justus Piater · 2017
Cited alongside, same era.
Verifiable reinforcement learning via policy extraction
Osbert Bastani, Yewen Pu, and Armando Solar-Lezama · 2018
Cited alongside, same era.
Leveraging grammar and reinforcement learning for neural program synthesis
Rudy Bunel, Matthew Hausknecht, Jacob Devlin, Rishabh Singh, and Pushmeet Kohli · 2018
Cited alongside, same era.
Execution-guided neural program synthesis
Xinyun Chen, Chang Liu, and Dawn Song · 2018
Cited alongside, same era.
Minimalistic gridworld environment for openai gym
Maxime Chevalier-Boisvert, Lucas Willems, and Suman Pal · 2018
Cited alongside, same era.
Learning explanatory rules from noisy data
Richard Evans and Edward Grefenstette · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Cited alongside, same era.
Provenance-guided synthesis of datalog programs
Mukund Raghothaman, Jonathan Mendelson, David Zhao, Mayur Naik, and Bernhard Scholz · 2020
Later among the works it cites.
Kogun: accelerating deep reinforcement learning via integrating human suboptimal knowledge
Peng Zhang, Jianye Hao, Weixun Wang, Hongyao Tang, Yi Ma, Yihai Duan, and Yan Zheng · 2020
Later among the works it cites.
Deep reinforcement learning for trading
Zihao Zhang, Stefan Zohren, and Stephen Roberts · 2020
Later among the works it cites.
Deepsynth: Automata synthesis for automatic task segmentation in deep reinforcement learning
Mohammadhosein Hasanbeig, Natasha Yogananda Jeppu, Alessandro Abate, Tom Melham, and Daniel Kroening · 2021
Later among the works it cites.
Value function spaces: Skill-centric state abstractions for long-horizon reasoning
Dhruv Shah, Peng Xu, Yao Lu, Ted Xiao, Alexander T Toshev, Sergey Levine, et al · 2021
Later among the works it cites.
Program synthesis guided reinforcement learning for partially observed environments
Yichen Yang, Jeevana Priya Inala, Osbert Bastani, Yewen Pu, Armando Solar-Lezama, and Martin Rinard · 2021
Later among the works it cites.
Program synthesis guided reinforcement learning for partially observed environments
Yichen Yang, Jeevana Priya Inala, Osbert Bastani, Yewen Pu, Armando Solar-Lezama, and Martin Rinard · 2021
Later among the works it cites.
Value function spaces: Skill-centric state abstractions for long-horizon reasoning
Dhruv Shah, Alexander T Toshev, Sergey Levine, and brian ichter · 2022
Closest in time.