Fetching the paper…
Reading the bibliography…
The search for interpretable reinforcement learning policies is of high academic and industrial interest.
Koshiyama, A. S., Escovedo, T., Vellasco, M. M. B. R., Tanscheit, R., July 2014. Gpfis-control: A fuzzy genetic model for control tasks. In: 2014 IEEE International Conference on Fuzzy Systems (FUZZ-IEEE). pp. 1953–1959
1959
Earlier work this paper cites.
Schwefel, H.-P., 1981. Numerical optimization of computer models. John Wiley & Sons, Inc
1981
Earlier work this paper cites.
Breiman, L., Friedman, J., Olshen, R., Stone, C., 1984. Classification and Regression Trees. CRC Press, Boca Raton, FL
1984
Earlier work this paper cites.
Sutton, R., 1988. Learning to predict by the methods of temporal differences. Machine learning 3 (1), 9–44
1988
Earlier work this paper cites.
Moore, A. W., nov 1990. Efficient memory-based learning for robot control. Tech. Rep. UCAM-CL-TR-209, University of Cambridge, Computer Laboratory. URL http://www.cl.cam.ac.uk/techreports/UCAM-CL-TR-209.pdf
1990
Earlier work this paper cites.
Koza, J. R., 1992. Genetic Programming: On the Programming of Computers by Means of Natural Selection. MIT Press, Cambridge, MA, USA
1992
Earlier work this paper cites.
Blickle, T., Thiele, L., 1995. A mathematical analysis of tournament selection. In: ICGA. pp. 9–16
1995
Earlier work this paper cites.
Gordon, G., 1995. Stable function approximation in dynamic programming. In: In Machine Learning: Proceedings of the Twelfth International Conference. Morgan Kaufmann
1995
Earlier work this paper cites.
Schwefel, H.-P., 1995. Evolution and optimum seeking. sixth-generation computer technology series
1995
Earlier work this paper cites.
Sutton, R., Barto, A., 1998. Reinforcement learning: an introduction. A Bradford book
1998
Earlier work this paper cites.
Shimooka, H., Fujimoto, Y., 1999. Generating equations with genetic programming for control of a movable inverted pendulum. In: Selected Papers from the Second Asia-Pacific Conference on Simulated Evolution and Learning on Simulated Evolution and Learning. SEAL’98. Springer-Verlag, London, UK, UK, pp. 179–186
1999
Earlier work this paper cites.
Juang, C.-F., Lin, J.-Y., Lin, C.-T., Apr. 2000. Genetic reinforcement learning through symbiotic evolution for fuzzy controller design. Trans. Sys. Man Cyber. Part B 30 (2), 290–302
2000
Earlier work this paper cites.
Downing, K. L., 2001. Adaptive genetic programs via reinforcement learning. In: Proceedings of the 3rd Annual Conference on Genetic and Evolutionary Computation. GECCO’01. Morgan Kaufmann Publishers Inc., San Francisco, CA, USA, pp. 19–26
2001
Earlier work this paper cites.
Fantoni, I., Lozano, R., 2002. Non-linear control for underactuated mechanical systems. Springer
2002
Earlier work this paper cites.
Katagiri, H., Hirasawa, K., Hu, J., Murata, J., Kosaka, M., 2002. Network structure oriented evolutionary model: Genetic network programming. Transactions of the Society of Instrument and Control Engineers 38 (5), 485–494
2002
Cited alongside, same era.
Keane, M. A., Koza, J. R., Streeter, M. J., 2002. Automatic synthesis using genetic programming of an improved general-purpose controller for industrially representative plants. In: Proceedings of the 2002 NASA/DoD Conference on Evolvable Hardware (EH’02). EH ’02. IEEE Computer Society, Washington, DC, USA, pp. 113–123
2002
Cited alongside, same era.
Mabu, S., Hirasawa, K., Hu, J., Murata, J., 2002. Online learning of genetic network programming (GNP). In: Fogel, D. B., El-Sharkawi, M. A., Yao, X., Greenwood, G., Iba, H., Marrow, P., Shackleton, M. (Eds.), Proceedings of the 2002 Congress on Evolutionary Computation CEC2002. IEEE Press, pp. 321–326
2002
Cited alongside, same era.
Ormoneit, D., Sen, S., 2002. Kernel-based reinforcement learning. Machine learning 49 (2), 161–178
Schäfer, A. M., 2008. Reinforcement learning with recurrent neural networks. Ph.D. thesis, University of Osnabrück, Germany
2008
Later among the works it cites.
Riedmiller, M., Gabel, T., Hafner, R., Lange, S., 2009. Reinforcement learning for robot soccer. Autonomous Robots 27 (1), 55–73
2009
Later among the works it cites.
Busoniu, L., Babuska, R., De Schutter, B., Ernst, D., 2010. Reinforcement Learning and Dynamic Programming Using Function Approximators. CRC Press
2010
Later among the works it cites.
Koza, J. R., 2010. Human-competitive results produced by genetic programming. Genetic Programming and Evolvable Machines 11 (3), 251–284
2010
Later among the works it cites.
Dubčáková, R., 2011. Eureqa: software review. Genetic programming and evolvable machines 12 (2), 173–178
2011
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2002
Cited alongside, same era.
Gearhart, C., 2003. Genetic programming as policy search in markov decision processes. In: Koza, J. R. (Ed.), Genetic Algorithms and Genetic Programming at Stanford 2003. Stanford Bookstore, Stanford, California, 94305-3079 USA, pp. 61–67
2003
Cited alongside, same era.
Lagoudakis, M., Parr, R., 2003. Least-squares policy iteration. Journal of Machine Learning Research, 1107–1149
2003
Cited alongside, same era.
Bakker, B., 2004. The state of mind: Reinforcement learning with recurrent neural networks. Ph.D. thesis, Leiden University, Netherlands
2004
Cited alongside, same era.
Mabu, S., Hirasawa, K., Hu, J., 2004. Genetic Network Programming with Reinforcement Learning and Its Performance Evaluation. Springer Berlin Heidelberg, Berlin, Heidelberg, pp. 710–711
2004
Cited alongside, same era.
Ernst, D., Geurts, P., Wehenkel, L., Littman, L., 2005. Tree-based batch mode reinforcement learning. Journal of Machine Learning Research 6, 503–556
2005
Cited alongside, same era.
Kamio, S., Iba, H., Jun. 2005. Adaptation technique for integrating genetic programming and reinforcement learning for real robots. Trans. Evol. Comp 9 (3), 318–333
2005
Cited alongside, same era.
Riedmiller, M., 2005a. Neural fitted Q iteration — first experiences with a data efficient neural reinforcement learning method. In: Machine Learning: ECML 2005. Vol. 3720. Springer, pp. 317–328
2005
Cited alongside, same era.
Riedmiller, M., 2005b. Neural reinforcement learning to swing-up and balance a real pole. In: Systems, Man and Cybernetics, 2005 IEEE International Conference on. Vol. 4. pp. 3191–3196
2005
Cited alongside, same era.
Duell, S., Udluft, S., Sterzing, V., 2012. Solving partially observable reinforcement learning problems with recurrent neural networks. In: Neural Networks: Tricks of the Trade. Springer, pp. 709–733
2012
Later among the works it cites.
Maes, F., Fonteneau, R., Wehenkel, L., Ernst, D., 2012. Policy search in a space of simple closed-form formulas: towards interpretability of reinforcement learning. Discovery Science, 37–50
2012
Later among the works it cites.
Neuneier, R., Zimmermann, H.-G., 2012. How to train neural networks. In: Montavon, G., Orr, G., Müller, K.-R. (Eds.), Neural Networks: Tricks of the Trade, Second Edition. Springer, pp. 369–418
2012
Later among the works it cites.
2016
Later among the works it cites.
Le, N., Xuan, H. N., Brabazon, A., Thi, T. P., 2016. Complexity measures in genetic programming learning: A brief review. In: Evolutionary Computation (CEC), 2016 IEEE Congress on. IEEE, pp. 2409–2416
2016
Later among the works it cites.
Hein, D., Depeweg, S., Tokic, M., Udluft, S., Hentschel, A., Runkler, T. A., Sterzing, V., 2017a. A benchmark environment motivated by industrial control problems. In: 2017 IEEE Symposium Series on Computational Intelligence (SSCI). pp. 1–8
2017
Closest in time.
Hein, D., Udluft, S., Tokic, M., Hentschel, A., Runkler, T. A., Sterzing, V., 2017c. Batch reinforcement learning on the industrial benchmark: First experiences. In: 2017 International Joint Conference on Neural Networks (IJCNN). pp. 4214–4221
2017
Closest in time.
Hein, D., Hentschel, A., Runkler, T. A., Udluft, S., 2018. Particle swarm optimization for model predictive control in reinforcement learning environments. In: Shi, Y. (Ed.), Critical Developments and Applications of Swarm Intelligence. IGI Global, Hershey, PA, USA, Ch. 16, pp. 401–427
2018
Closest in time.