Fetching the paper…
Reading the bibliography…
Most state of the art decision systems based on Reinforcement Learning (RL) are data-driven black-box neural models, where it is often difficult to incorporate expert knowledge into the models or let experts review and validate the learned decision mechanisms.
Reinforcement learning applications
Yuxi Li · 1908
Earlier work this paper cites.
On computable numbers, with an application to the entscheidungsproblem. a correction
A M Turing · 1938
Earlier work this paper cites.
Partially Observable Markov Processes., 1964
Jr Kramer, J David R · 1964
Earlier work this paper cites.
Optimal control of Markov processes with incomplete state information
K J Åström · 1965
Earlier work this paper cites.
Neuronlike adaptive elements that can solve difficult learning control problems
A. G. Barto, R. S. Sutton, and C. W. Anderson · 1983
Earlier work this paper cites.
Efficient memory-based learning for robot control
Andrew William Moore · 1990
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Brainfuck – an eight-instruction turing-complete programming language
U. Muller · 1993
Earlier work this paper cites.
Reinforcement learning: A survey
Leslie Pack Kaelbling, Michael L Littman, and Andrew W Moore · 1996
Earlier work this paper cites.
Learning to forget: Continual prediction with lstm
Felix A Gers, Jürgen Schmidhuber, and Fred Cummins · 1999
Earlier work this paper cites.
Hierarchical reinforcement learning with the maxq value function decomposition
Thomas G. Dietterich · 2000
Earlier work this paper cites.
Dimensions of neural-symbolic integration-a structured survey
Sebastian Bader and Pascal Hitzler · 2005
Earlier work this paper cites.
A field guide to genetic programming
Riccardo Poli, William B Langdon, Nicholas F McPhee, and John R Koza · 2008
Earlier work this paper cites.
The unicode standard
Julie D Allen, Deborah Anderson, Joe Becker, Richard Cook, Mark Davis, Peter Edberg, Michael Everson, Asmus Freytag, Laurentiu Iancu, Richard Ishida, et al · 2012
Earlier work this paper cites.
Lecture 6.5—RmsProp: Divide the gradient by a running average of its recent magnitude
T. Tieleman and G. Hinton · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2013
Earlier work this paper cites.
Flashmeta: A framework for inductive program synthesis
Oleksandr Polozov and Sumit Gulwani · 2015
Cited alongside, same era.
Reinforcement learning neural turing machines
Wojciech Zaremba and Ilya Sutskever · 2015
Cited alongside, same era.
Jason Weston, Sumit Chopra, and Antoine Bordes · 2015
Cited alongside, same era.
Relational knowledge extraction from neural networks
Manoel Vitor Macedo França, Artur S d’Avila Garcez, and Gerson Zaverucha · 2015
Cited alongside, same era.
Neural random-access machines
Karol Kurach, Marcin Andrychowicz, and Ilya Sutskever · 2016
Cited alongside, same era.
control flow in brainfuck | matslina, 2016
Mats Linander · 2016
Recent Advances in Neural Program Synthesis
Neel Kant · 2018
Later among the works it cites.
Neural-guided deductive search for real-time program synthesis from examples
Ashwin K Vijayakumar, Dhruv Batra, Abhishek Mohta, Prateek Jain, Oleksandr Polozov, and Sumit Gulwani · 2018
Later among the works it cites.
Neural Program Synthesis with Priority Queue Training
Daniel A. Abolafia, Mohammad Norouzi, Jonathan Shen, Rui Zhao, and Quoc V. Le · 2018
Later among the works it cites.
Daqn: Deep auto-encoder and q-network
Daiki Kimura · 2018
Later among the works it cites.
With an eye to AI and autonomous diagnosis
P. A. Keane and E. J. Topol · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Cited alongside, same era.
Reinforcement Learning: An Introduction, Second edition in progress , volume 3
Richard S Sutton and Andrew G Barto · 2017
Cited alongside, same era.
Deep reinforcement learning: A brief survey
K. Arulkumaran, M. P. Deisenroth, M. Brundage, and A. A. Bharath · 2017
Cited alongside, same era.
Program synthesis
Sumit Gulwani, Oleksandr Polozov, and Rishabh Singh · 2017
Cited alongside, same era.
Sqlnet: Generating structured queries from natural language without reinforcement learning
Xiaojun Xu, Chang Liu, and Dawn Song · 2017
Cited alongside, same era.
Neural-symbolic learning and reasoning: A survey and interpretation
Tarek R Besold, Artur d’Avila Garcez, Sebastian Bader, Howard Bowman, Pedro Domingos, Pascal Hitzler, Kai-Uwe Kühnberger, Luis C Lamb, Daniel Lowd, Priscila Machado Vieira Lima, et al · 2017
Cited alongside, same era.
Chao Yu, Jiming Liu, and Shamim Nemati · 2019
Later among the works it cites.
Synthetic datasets for neural program synthesis
Richard Shin, Neel Kant, Kavi Gupta, Christopher Bender, Brandon Trabucco, Rishabh Singh, and Dawn Song · 2019
Later among the works it cites.
A survey of genetic programming and its applications
Milad Taleby Ahvanooey, Qianmu Li, Ming Wu, and Shuo Wang · 2019
Later among the works it cites.
Revisiting neural-symbolic learning cycle
Martin Svatoš, Gustav Šourek, and Filip Železný · 2019
Later among the works it cites.
A neural model for generating natural language summaries of program subroutines
A. LeClair, S. Jiang, and C. McMillan · 2019
Later among the works it cites.
Zooming for efficient model-free reinforcement learning in metric spaces
Ahmed Touati, Adrien Ali Taiga, and Marc G Bellemare · 2020
Later among the works it cites.
vadim0x60/evestop: Early stopping with exponential variance elmination
Vadim Liventsev · 2021
Closest in time.
Bytecode — Wikipedia, the free encyclopedia, 2020
Wikipedia contributors · 2021
Closest in time.
https://vadim.me/posts/unreasonable/
The unreasonable effectiveness of language models for source code · vadim liventsev · 2022
Closest in time.
Competition-level code generation with alphacode, 2022
Yujia Li, David Choi, Junyoung Chung, Nate Kushman, Julian Schrittwieser, Rémi Leblond, Tom Eccles, James Keeling, Felix Gimeno, Agustin Dal Lago, Thomas Hubert, Peter Choy, Cyprien de Masson d’Autume, Igor Babuschkin, Xinyun Chen, Po-Sen Huang, Johannes Welbl, Sven Gowal, Alexey Cherepanov, James Molloy, Daniel J. Mankowitz, Esme Sutherland Robson, Pushmeet Kohli, Nando de Freitas, Koray Kavukcuoglu, and Oriol Vinyals · 2022
Closest in time.