Fetching the paper…
Reading the bibliography…
Continuous action policy search is currently the focus of intensive research, driven both by the recent success of deep reinforcement learning algorithms and the emergence of competitors based on evolutionary algorithms.
Rastrigin, L., 1963. The convergence of the random search method in the extremal control of a many parameter system. Automation and Remote Control 24 (10), 1337–1342
1963
Earlier work this paper cites.
Gill, P. E., Murray, W., Wright, M. H., 1981. Practical optimization. Academic press
1981
Earlier work this paper cites.
Sutton, R. S., 1988. Learning to Predict by the Method of Temporal Differences. Machine Learning 3, 9–44
1988
Earlier work this paper cites.
Goldberg, D. E., 1989. Genetic Algorithms in Search, Optimization, and Machine Learning. Addison Wesley, Reading, MA
1989
Earlier work this paper cites.
Koza, J. R., 1992. Genetic Programming: On the Programming of Computers by Means of Natural Selection. MIT Press, Cambridge, MA
1992
Earlier work this paper cites.
Williams, R. J., May 1992. Simple statistical gradient-following algorithms for connectionist reinforcement learning. Machine Learning 8 (3-4), 229–256
1992
Earlier work this paper cites.
Baird, L. C., 1994. Reinforcement learning in continuous time: Advantage updating. In: Proceedings of the International Conference on Neural Networks. Orlando, FL
1994
Earlier work this paper cites.
Thrun, S., Mitchell, T. M., 1995. Lifelong robot learning. Robotics and autonomous systems 15 (1-2), 25–46
1995
Earlier work this paper cites.
Back, T., 1996. Evolutionary algorithms in theory and practice: evolution strategies, evolutionary programming, genetic algorithms. Oxford university press
1996
Earlier work this paper cites.
Aha, D. W., 1997. Editorial. In: Lazy learning. Springer, pp. 7–10
1997
Earlier work this paper cites.
Sutton, R. S., Barto, A. G., 1998. Reinforcement Learning: An Introduction. MIT Press
1998
Earlier work this paper cites.
Williams, J. K., Singh, S. P., 1998. Experimental results on learning stochastic memoryless policies for partially observable markov decision processes. In: NIPS. pp. 1073–1080
1998
Earlier work this paper cites.
Pelikan, M., Goldberg, D. E., Cantú-Paz, E., 1999. Boa: The Bayesian optimization algorithm. In: Proceedings of the 1st Annual Conference on Genetic and Evolutionary Computation-Volume 1. Morgan Kaufmann Publishers Inc., pp. 525–532
1999
Earlier work this paper cites.
Kearns, M. J., Singh, S. P., 2000. Bias-variance error bounds for temporal difference updates. In: COLT. pp. 142–147
2000
Earlier work this paper cites.
Baxter, J., Bartlett, P. L., 2001. Infinite-horizon policy-gradient estimation. Journal of Artificial Intelligence Research 15, 319–350
2001
Earlier work this paper cites.
Hansen, N., Ostermeier, A., 2001. Completely derandomized self-adaptation in evolution strategies. Evolutionary Computation 9 (2), 159–195
2001
Earlier work this paper cites.
Larrañaga, P., Lozano, J. A., 2001. Estimation of distribution algorithms: A new tool for evolutionary computation. Vol. 2. Springer Science & Business Media
2001
Earlier work this paper cites.
Stanley, K. O., Miikkulainen, R., 2002. Efficient evolution of neural network topologies. In: Evolutionary Computation, 2002. CEC’02. Proceedings of the 2002 Congress on. Vol. 2. IEEE, pp. 1757–1762
2002
Earlier work this paper cites.
Rubinstein, R., Kroese, D., 2004. The Cross-Entropy Method: A Unified Approach to Combinatorial Optimization, Monte-Carlo Simulation, and Machine Learning. Springer-Verlag
2004
Earlier work this paper cites.
Bhatnagar, S., Sutton, R. S., Ghavamzadeh, M., Lee, M., 2007. Incremental natural actor-critic algorithms. In: Advances in Neural Information Processing Systems. MIT Press
2007
Earlier work this paper cites.
Lizotte, D. J., Wang, T., Bowling, M. H., Schuurmans, D., 2007. Automatic gait optimization with gaussian process regression. In: IJCAI. Vol. 7. pp. 944–949
2007
Earlier work this paper cites.
Floreano, D., Dürr, P., Mattiussi, C., 2008. Neuroevolution: from architectures to learning. Evolutionary Intelligence 1 (1), 47–62
2008
Earlier work this paper cites.
Riedmiller, M., Peters, J., Schaal, S., 2008. Evaluation of policy gradient methods and variants on the cart-pole benchmark. In: IEEE International Symposium on Approximate Dynamic Programming and Reinforcement Learning (ADPRL)
2008
Earlier work this paper cites.
Wierstra, D., Schaul, T., Peters, J., Schmidhuber, J., 2008. Natural evolution strategies. In: IEEE Congress on Evolutionary Computation. IEEE, pp. 3381–3387
2008
Earlier work this paper cites.
Argall, B. D., Chernova, S., Veloso, M., Browning, B., 2009. A survey of robot learning from demonstration. Robotics and Autonomous Systems 57, 469–483
2009
Earlier work this paper cites.
Kober, J., Peters, J., 2009. Learning motor primitives for robotics. In: IEEE International Conference on Robotics and Automation. IEEE, pp. 2112–2118
2009
Earlier work this paper cites.
Sun, Y., Wierstra, D., Schaul, T., Schmidhuber, J., 2009. Efficient natural evolution strategies. In: Proceedings of the 11th Annual conference on Genetic and evolutionary computation. ACM, pp. 539–546
2009
Earlier work this paper cites.
Togelius, J., Schaul, T., Wierstra, D., Igel, C., Gomez, F., Schmidhuber, J., 2009. Ontogenetic and phylogenetic reinforcement learning. Künstliche Intelligenz 23 (3), 30–33
2009
Earlier work this paper cites.
Akimoto, Y., Nagata, Y., Ono, I., Kobayashi, S., 2010. Bidirectional relation between cma evolution strategies and natural evolution strategies. In: International Conference on Parallel Problem Solving from Nature. Springer, pp. 154–163
2010
Earlier work this paper cites.
Baranes, A., Oudeyer, P.-Y., 2010. Intrinsically motivated goal exploration for active motor learning in robots: A case study. In: IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2010). IEEE, Taipei, Taiwan, Province Of China
2010
Earlier work this paper cites.
2010
Earlier work this paper cites.
Glasmachers, T., Schaul, T., Yi, S., Wierstra, D., Schmidhuber, J., 2010. Exponential natural evolution strategies. In: Proceedings of the 12th annual conference on Genetic and evolutionary computation. ACM, pp. 393–400
2010
Earlier work this paper cites.
Peters, J., Mülling, K., Altun, Y., 2010. Relative entropy policy search. In: AAAI. Atlanta, pp. 1607–1612
2010
Earlier work this paper cites.
Sehnke, F., Osendorfer, C., Rückstieß, T., Graves, A., Peters, J., Schmidhuber, J., 2010. Parameter-exploring policy gradients. Neural Networks 23 (4), 551–559
2010
Earlier work this paper cites.
Sigaud, O., Buffet, O., 2010. Markov Decision Processes in Artificial Intelligence. iSTE - Wiley
2010
Earlier work this paper cites.
Theodorou, E., Buchli, J., Schaal, S., 2010. A generalized path integral control approach to reinforcement learning. Journal of Machine Learning Research 11, 3137–3181
2010
Earlier work this paper cites.
Arnold, L., Auger, A., Hansen, N., Ollivier, Y., 2011. Information-geometric optimization algorithms: A unifying picture via invariance principles. Tech. rep., INRIA Saclay
2011
Earlier work this paper cites.
Cuccu, G., Gomez, F., 2011. When novelty is not enough. In: European Conference on the Applications of Evolutionary Computation. Springer, pp. 234–243
2011
Earlier work this paper cites.
Deisenroth, M., Rasmussen, C. E., 2011. Pilco: A model-based and data-efficient approach to policy search. In: Proceedings of the 28th International Conference on machine learning. pp. 465–472
2011
Earlier work this paper cites.
Lehman, J., Stanley, K. O., 2011. Abandoning objectives: Evolution through the search for novelty alone. Evolutionary computation 19 (2), 189–223
2011
Earlier work this paper cites.
Neumann, G., 2011. Variational inference for policy search in changing situations. In: Proceedings of the 28th international conference on machine learning. pp. 817–824
2011
Earlier work this paper cites.
Bottou, L., 2012. Stochastic gradient descent tricks. In: Neural networks: Tricks of the trade. Springer, pp. 421–436
2012
Cited alongside, same era.
Grondman, I., Busoniu, L., Lopes, G. A., Babuska, R., 2012. A survey of actor-critic reinforcement learning: Standard and natural policy gradients. IEEE Transactions on Systems, Man, and Cybernetics, Part C (Applications and Reviews) 42 (6), 1291–1307
2012
Cited alongside, same era.
Baranes, A., Oudeyer, P.-Y., 2013. Active learning of inverse models with intrinsically motivated goal exploration in robots. Robotics and Autonomous Systems 61 (1), 49–73
2013
Cited alongside, same era.
Deisenroth, M. P., Neumann, G., Peters, J., et al., 2013. A survey on policy search for robotics. Foundations and Trends® in Robotics 2 (1–2), 1–142
2013
Cited alongside, same era.
2017
Later among the works it cites.
Cully, A., Demiris, Y., 2017. Quality and diversity optimization: A unifying modular framework. IEEE Transactions on Evolutionary Computation
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2013
Cited alongside, same era.
Kober, J., Bagnell, J. A., Peters, J., 2013. Reinforcement learning in robotics: A survey. The International Journal of Robotics Research 32 (11), 1238–1274
2013
Cited alongside, same era.
Levine, S., Koltun, V., 2013. Guided policy search. In: Proceedings of the 30th International Conference on Machine Learning. pp. 1–9
2013
Cited alongside, same era.
Stulp, F., Sigaud, O., august 2013. Robot skill learning: From reinforcement learning to evolution strategies. Paladyn Journal of Behavioral Robotics 4 (1), 49–61
2013
Cited alongside, same era.
Baranes, A. F., Oudeyer, P.-Y., Gottlieb, J., 2014. The effects of task difficulty, novelty and the size of the search space on intrinsically motivated exploration. Frontiers in neuroscience 8, 317
2014
Cited alongside, same era.
Calandra, R., Gopalan, N., Seyfarth, A., Peters, J., Deisenroth, M. P., 2014. Bayesian gait optimization for bipedal locomotion. In: International Conference on Learning and Intelligent Optimization. Springer, pp. 274–290
2014
Cited alongside, same era.
Doncieux, S., Mouret, J.-B., 2014. Beyond black-box optimization: a review of selective pressures for evolutionary robotics. Evolutionary Intelligence 7 (2), 71–93
2014
Cited alongside, same era.
Hwangbo, J., Gehring, C., Sommer, H., Siegwart, R., Buchli, J., 2014. ROCK ∗
2014
Cited alongside, same era.
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
Islam, R., Henderson, P., Gomrokchi, M., Precup, D., 2017. Reproducibility of benchmarked deep reinforcement learning tasks for continuous control. In: Proceedings of the ICML 2017 workshop on Reproducibility in Machine Learning (RML)
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
Martinez-Cantin, R., Tee, K., McCourt, M., 2017. Policy search using robust Bayesian optimization. In: Neural Information Processing Systems (NIPS) Workshop on Acting and Interacting in the Real World: Challenges in Robot Learning
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
Tang, Y., Kucukelbir, A., 2017. Variational deep q network. arXiv preprint arXiv:1711.11225
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
2018
Closest in time.
Barth-maron, G., Hoffman, M., Budden, D., Dabney, W., Horgan, D., TB, D., Muldal, A., Heess, N., Lillicrap, T. P., 2018. Distributional policy gradient. In: ICLR. pp. 1–16
2018
Closest in time.
2018
Closest in time.
2018
Closest in time.
2018
Closest in time.
Gangwani, T., Peng, J., 2018. Policy optimization by genetic distillation. In: ICLR 2018
2018
Closest in time.
2018
Closest in time.
Khadka, S., Tumer, K., 2018. Evolutionary reinforcement learning. arXiv preprint arXiv:1805.07917
2018
Closest in time.
2018
Closest in time.
2018
Closest in time.
2018
Closest in time.
2018
Closest in time.
2018
Closest in time.
2018
Closest in time.
2018
Closest in time.
2018
Closest in time.
2018
Closest in time.
2018
Closest in time.