Fetching the paper…
Reading the bibliography…
This paper focuses on reinforcement learning (RL) with limited prior knowledge.
Duda, R., Hart, P.: Pattern Classification and scene analysis. John Wiley and sons, Menlo Park, CA (1973)
1973
Earlier work this paper cites.
Cortes, C., Vapnik, V.: Support-vector networks. Machine Learning 20(3), 273–297 (1995)
1995
Earlier work this paper cites.
Jones, D., Schonlau, M., Welch, W.: Efficient global optimization of expensive black-box functions. Journal of Global Optimization 13(4), 455–492 (1998)
1998
Earlier work this paper cites.
Sutton, R., Barto, A.: Reinforcement Learning: An Introduction. MIT Press, Cambridge (1998)
1998
Earlier work this paper cites.
Ng, A., Russell, S.: Algorithms for inverse reinforcement learning. In: Langley, P. (ed.) Proc. of the Seventeenth International Conference on Machine Learning (ICML-00). pp. 663–670. Morgan Kaufmann (2000)
2000
Earlier work this paper cites.
Hansen, N., Ostermeier, A.: Completely derandomized self-adaptation in evolution strategies. Evolutionary Computation 9(2), 159–195 (2001)
2001
Earlier work this paper cites.
Herbrich, R., Graepel, T., Campbell, C.: Bayes point machines. Journal of Machine Learning Research 1, 245–279 (2001)
2001
Earlier work this paper cites.
O’Regan, J., Noë, A.: A sensorimotor account of vision and visual consciousness. Behavioral and Brain Sciences 24, 939–973 (2001)
2001
Earlier work this paper cites.
Littman, M.L., Sutton, R.S., Singh, S.: Predictive representations of state. In: Neural Information Processing Systems 14. p. 1555–1561 (2002)
2002
Earlier work this paper cites.
Lagoudakis, M., Parr, R.: Least-squares policy iteration. Journal of Machine Learning Research (JMLR) 4, 1107–1149 (2003)
2003
Earlier work this paper cites.
Abbeel, P., Ng, A.: Apprenticeship learning via inverse reinforcement learning. In: Brodley, C.E. (ed.) ICML. ACM International Conference Proceeding Series, vol. 69. ACM (2004)
2004
Cited alongside, same era.
Dasgupta, S.: Coarse sample complexity bounds for active learning. In: Advances in Neural Information Processing Systems 18 (2005)
2005
Cited alongside, same era.
Joachims, T.: A support vector method for multivariate performance measures. In: Raedt, L.D., Wrobel, S. (eds.) ICML. pp. 377–384 (2005)
2005
Cited alongside, same era.
Tsochantaridis, I., Joachims, T., Hofmann, T., Altun, Y.: Large margin methods for structured and interdependent output variables. Journal of Machine Learning Research 6, 1453–1484 (2005)
2005
Cited alongside, same era.
Joachims, T.: Training linear svms in linear time. In: Eliassi-Rad, T., Ungar, L.H., Craven, M., Gunopulos, D. (eds.) KDD. pp. 217–226. ACM (2006)
Heidrich-Meisner, V., Igel, C.: Hoeffding and bernstein races for selecting policies in evolutionary direct policy search. In: ICML. p. 51 (2009)
2009
Later among the works it cites.
Zhao, Kosorok, M.R., Zeng, D.: Reinforcement learning design for cancer clinical trials. Stat Med (Sep 2009)
2009
Later among the works it cites.
Hachiya, H., Sugiyama, M.: Feature selection for reinforcement learning: Evaluating implicit state-reward dependency via conditional mutual information. In: Proc. ECML/PKDD (1). Lecture Notes in Computer Science, vol. 6321, pp. 474–489 (2010)
2010
Later among the works it cites.
Konidaris, G., Kuindersma, S., Barto, A., Grupen, R.: Constructing skill trees for reinforcement learning agents from demonstration trajectories. In: Advances in Neural Information Processing Systems 23. pp. 1162–1170 (2010)
2010
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2006
Cited alongside, same era.
Calinon, S., Guenter, F., Billard, A.: On Learning, Representing and Generalizing a Task in a Humanoid Robot. IEEE transactions on systems, man and cybernetics, Part B. Special issue on robot learning by observation, demonstration and imitation 37(2), 286–298 (2007)
2007
Cited alongside, same era.
Kolter, J.Z., Abbeel, P., Ng, A.Y.: Hierarchical apprenticeship learning with application to quadruped locomotion. In: NIPS. MIT Press (2007)
2007
Cited alongside, same era.
Bergeron, C., Zaretzki, J., Breneman, C.M., Bennett, K.P.: Multiple instance ranking. In: ICML. pp. 48–55 (2008)
2008
Cited alongside, same era.
Brochu, E., de Freitas, N., Ghosh, A.: Active preference learning with discrete choice data. In: Advances in Neural Information Processing Systems 20. pp. 409–416 (2008)
2008
Cited alongside, same era.
Peters, J., Schaal, S.: Reinforcement learning of motor skills with policy gradients. Neural Networks 21(4), 682–697 (2008)
2008
Cited alongside, same era.
Akrour, R., Schoenauer, M., Sebag, M.: Preference-based policy learning. In: Gunopulos et al. [ 10 ] , pp. 12–27
Cited in the paper.
Cheng, W., Fürnkranz, J., Hüllermeier, E., Park, S.H.: Preference-based policy iteration: Leveraging preference learning for reinforcement learning. In: Gunopulos et al. [ 10 ] , pp. 312–327
Cited in the paper.
2010
Later among the works it cites.
Viappiani, P., Boutilier, C.: Optimal Bayesian recommendation sets and myopically optimal choice query sets. In: NIPS. pp. 2352–2360 (2010)
2010
Later among the works it cites.
Whiteson, S., Taylor, M.E., Stone, P.: Critical factors in the empirical performance of temporal difference and evolutionary methods for reinforcement learning. Journal of Autonomous Agents and Multi-Agent Systems 21(1), 1–27 (2010)
2010
Later among the works it cites.
Gunopulos, D., Hofmann, T., Malerba, D., Vazirgiannis, M. (eds.): Proc. European Conf. on Machine Learning and Knowledge Discovery in Databases, ECML PKDD, Part I. LNCS 6911, Springer Verlag (2011)
2011
Later among the works it cites.
Liu, C., Chen, Q., Wang, D.: Locomotion control of quadruped robots based on cpg-inspired workspace trajectory generation. In: Proc. ICRA. pp. 1250–1255. IEEE (2011)
2011
Later among the works it cites.
Viappiani, P.: Monte-Carlo methods for preference learning. In: Hamadi, Y., Schoenauer, M. (eds.) Proc. Learning and Intelligent OptimizatioN, LION 6. LNCS, Springer Verlag (2012), to appear
2012
Closest in time.