Fetching the paper…
Reading the bibliography…
We present a reinforcement learning framework, called Programmatically Interpretable Reinforcement Learning (PIRL), that is designed to generate interpretable and verifiable agent policies.
Symbolic execution and program testing
King, J. C · 1976
Earlier work this paper cites.
Neuronlike adaptive elements that can solve difficult learning control problems
Barto, A. G., Sutton, R. S., and Anderson, C. W · 1983
Earlier work this paper cites.
Automatic tuning of simple regulators with specifications on phase and amplitude margins
Astrom, K. and Hagglund, T · 1984
Earlier work this paper cites.
Variable resolution dynamic programming: Efficiently learning action maps in multivariate real-valued state-spaces
Moore, A · 1991
Earlier work this paper cites.
PID controllers: Theory, Design, and Tuning
Åström, K. J. and Hägglund, T · 1995
Earlier work this paper cites.
Is imitation learning the route to humanoid robots?
Schaal, S · 1999
Earlier work this paper cites.
Perception as abduction: Turning sensor data into meaningful representation
Shanahan, M · 2005
Earlier work this paper cites.
The sketching approach to program synthesis
Solar-Lezama, A · 2009
Earlier work this paper cites.
Learning to overtake in TORCS using simple reinforcement learning
Loiacono, D., Prete, A., Lanzi, P. L., and Cardamone, L · 2010
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Ross, S., Gordon, G. J., and Bagnell, D · 2011
Earlier work this paper cites.
Practical bayesian optimization of machine learning algorithms
Snoek, J., Larochelle, H., and Adams, R. P · 2012
Earlier work this paper cites.
Symbolic execution for software testing: three decades later
Cadar, C. and Sen, K · 2013
Earlier work this paper cites.
Evolving large-scale neural networks for vision-based TORCS
Koutník, J., Cuccu, G., Schmidhuber, J., and Gomez, F. J · 2013
Cited alongside, same era.
Graves, A., Wayne, G., and Danihelka, I · 2014
Cited alongside, same era.
TORCS, The Open Racing Car Simulator
Wymann, B., Espié, E., Guionneau, C., Dimitrakakis, C., Coulom, R., and Sumner, A · 2014
Cited alongside, same era.
Syntax-guided synthesis
Alur, R., Bodík, R., Dallal, E., Fisman, D., Garg, P., Juniwal, G., Kress-Gazit, H., Madhusudan, P., Martin, M. M. K., Raghothaman, M., Saha, S., Seshia, S. A., Singh, R., Solar-Lezama, A., Torlak, E., and Udupa, A · 2015
Cited alongside, same era.
Synthesizing data structure transformations from input-output examples
Feser, J. K., Chaudhuri, S., and Dillig, I · 2015
Cited alongside, same era.
Rlpy: A value-function-based reinforcement learning framework for education and research
Towards deep symbolic reinforcement learning
Garnelo, M., Arulkumaran, K., and Shanahan, M · 2016
Later among the works it cites.
Building machines that learn and think like people
Lake, B. M., Ullman, T. D., Tenenbaum, J. B., and Gershman, S. J · 2016
Later among the works it cites.
The mythos of model interpretability
Lipton, Z. C · 2016
Later among the works it cites.
Mastering the game of Go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., van den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., Dieleman, S., Grewe, D., Nham, J., Kalchbrenner, N., Sutskever, I., Lillicrap, T. P., Leach, M., Kavukcuoglu, K., Graepel, T., and Hassabis, D · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Geramifard, A., Dann, C., Klein, R. H., Dabney, W., and How, J. P · 2015
Cited alongside, same era.
Inferring algorithmic patterns with stack-augmented recurrent nets
Joulin, A. and Mikolov, T · 2015
Cited alongside, same era.
Kurach, K., Andrychowicz, M., and Sutskever, I · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Cited alongside, same era.
Dueling network architectures for deep reinforcement learning
Wang, Z., de Freitas, N., and Lanctot, M · 2015
Cited alongside, same era.
Graying the black box: Understanding dqns
Zahavy, T., Ben-Zrihem, N., and Mannor, S · 2015
Cited alongside, same era.
Deepcoder: Learning to write programs
Balog, M., Gaunt, A. L., Brockschmidt, M., Nowozin, S., and Tarlow, D · 2016
Cited alongside, same era.
Devlin, J., Uesato, J., Bhupatiraju, S., Singh, R., Mohamed, A., and Kohli, P · 2017
Later among the works it cites.
Reluplex: An efficient smt solver for verifying deep neural networks
Katz, G., Barrett, C. W., Dill, D. L., Julian, K., and Kochenderfer, M. J · 2017
Later among the works it cites.
Methods for interpreting and understanding deep neural networks
Montavon, G., Samek, W., and Müller, K · 2017
Later among the works it cites.
Driving in TORCS using modular fuzzy controllers
Salem, M., Mora, A. M., Merelo, J. J., and García-Sánchez, P · 2017
Later among the works it cites.
Mastering chess and Shogi by self-play with a general reinforcement learning algorithm
Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., et al · 2017
Later among the works it cites.
Neural program synthesis with priority queue training
Abolafia, D. A., Norouzi, M., and Le, Q. V · 2018
Closest in time.
Leveraging grammar and reinforcement learning for neural program synthesis
Bunel, R., Hausknecht, M., Devlin, J., Singh, R., and Kohli, P · 2018
Closest in time.
Neural sketch learning for conditional program generation
Murali, V., Chaudhuri, S., and Jermaine, C · 2018
Closest in time.