Fetching the paper…
Reading the bibliography…
In the research area of reinforcement learning (RL), frequently novel and promising methods are developed and introduced to the RL community.
R. S. Sutton and A. G. Barto, Reinforcement Learning: An Introduction . Cambridge, MA: MIT Press, 1998
1998
Earlier work this paper cites.
J. Randløv and P. Alstrøm, “Learning to drive a bicycle using reinforcement learning and shaping,” in Proceedings of the Fifteenth International Conference on Machine Learning (ICML 1998) , J. W. Shavlik, Ed. San Francisco, CA, USA: Morgan Kauffman, 1998, pp. 463–471
1998
Earlier work this paper cites.
M. Schlang, B. Feldkeller, B. Lang, P. T., and R. T. A., “Neural computation in steel industry,” in 1999 European Control Conference (ECC) , 1999, pp. 2922–2927
1999
Earlier work this paper cites.
T. A. Runkler, E. Gerstorfer, M. Schlang, E. Jünnemann, and J. Hollatz, “Modelling and optimisation of a refining process for fibre board production,” Control engineering practice , vol. 11, no. 11, pp. 1229–1241, 2003
2003
Earlier work this paper cites.
A. M. Schaefer, D. Schneegass, V. Sterzing, and S. Udluft, “A neural reinforcement learning approach to gas turbine control,” in 2007 International Joint Conference on Neural Networks , 2007, pp. 1691–1696
2007
Earlier work this paper cites.
S. A. Hartmann and T. A. Runkler, “Online optimization of a color sorting assembly buffer using ant colony optimization,” Operations Research Proceedings 2007 , pp. 415–420, 2008
2008
Earlier work this paper cites.
A. Hans, D. Schneegass, A. M. Schaefer, and S. Udluft, “Safe exploration for reinforcement learning,” in 2008 European Symposium on Artificial Neural Networks (ESANN) , 2008, pp. 143–148
2008
Earlier work this paper cites.
A. Hans and S. Udluft, “Efficient uncertainty propagation for reinforcement learning with limited data,” in 2009 Proceedings of the International Conference on Artificial Neural Networks (ICANN 2009) , 2009, pp. 70–79
2009
Cited alongside, same era.
P. Abbeel, A. Coates, and A. Y. Ng, “Autonomous helicopter aerobatics through apprenticeship learning,” The International Journal of Robotics Research , vol. 29, no. 13, pp. 1608–1639, 2010
2010
Cited alongside, same era.
E. Todorov, T. Erez, and Y. Tassa, “Mujoco: A physics engine for model-based control,” in 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, 2012, pp. 5026–5033
2012
Cited alongside, same era.
S. Lange, T. Gabel, and M. Riedmiller, “Batch reinforcement learning,” in Reinforcement Learning . Springer, 2012, pp. 45–73
2012
Cited alongside, same era.
S. Spieckermann, S. Düll, S. Udluft, A. Hentschel, and T. A. Runkler, “Exploiting similarity in system identification tasks with recurrent neural networks,” Neurocomputing , vol. 169, pp. 343–349, 2015
2015
Later among the works it cites.
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. van den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, S. Dieleman, D. Grewe, J. Nham, N. Kalchbrenner, I. Sutskever, T. P. Lillicrap, M. Leach, K. Kavukcuoglu, T. Graepel, and D. Hassabis, “Mastering the game of Go with deep neural networks and tree search,” Nature , vol. 529, no. 7587, pp. 484–489, 2016
2016
Later among the works it cites.
2016
Later among the works it cites.
D. Hein, S. Udluft, M. Tokic, A. Hentschel, T. A. Runkler, and V. Sterzing, “Batch reinforcement learning on the industrial benchmark: First experiences,” in 2017 International Joint Conference on Neural Networks (IJCNN) , 2017, pp. 4214–4221
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling, “The arcade learning environment: An evaluation platform for general agents,” Journal of Artificial Intelligence Research , vol. 47, pp. 253–279, 2013
2013
Cited alongside, same era.
2015
Cited alongside, same era.
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis, “Human-level control through deep reinforcement learning,” Nature , vol. 518, no. 7540, pp. 529–533, 2015
2015
Cited alongside, same era.
2017
Closest in time.
2017
Closest in time.