Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) is a powerful approach for acquiring a good-performing policy.
Completely derandomized self-adaptation in evolution strategies
Hansen, N. and Ostermeier, A · 2001
Earlier work this paper cites.
The cross entropy method for fast policy search
Mannor, S., Rubinstein, R. Y., and Gat, Y · 2003
Earlier work this paper cites.
Pattern recognition and machine learning
Bishop, C · 2006
Earlier work this paper cites.
Dynamic movement primitives-a framework for motor control in humans and humanoid robotics
Schaal, S · 2006
Earlier work this paper cites.
Relative entropy policy search
Peters, J., Mulling, K., and Altun, Y · 2010
Earlier work this paper cites.
Parameter-exploring policy gradients
Sehnke, F., Osendorfer, C., Rückstieß, T., Graves, A., Peters, J., and Schmidhuber, J · 2010
Earlier work this paper cites.
Hierarchical relative entropy policy search
Daniel, C., Neumann, G., and Peters, J · 2012
Earlier work this paper cites.
Data-efficient generalization of robot skills with contextual policy search
Kupcsik, A., Deisenroth, M., Peters, J., and Neumann, G · 2013
Earlier work this paper cites.
Probabilistic movement primitives
Paraschos, A., Daniel, C., Peters, J. R., and Neumann, G · 2013
Earlier work this paper cites.
Natural evolution strategies
Wierstra, D., Schaul, T., Glasmachers, T., Sun, Y., Peters, J., and Schmidhuber, J · 2014
Earlier work this paper cites.
Model-based relative entropy stochastic search
Abdolmaleki, A., Lioutikov, R., Peters, J. R., Lau, N., Pualo Reis, L., and Neumann, G · 2015
Earlier work this paper cites.
Robots that can adapt like animals
Cully, A., Clune, J., Tarapore, D., and Mouret, J.-B · 2015
Earlier work this paper cites.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P · 2015
Earlier work this paper cites.
Openai gym, 2016
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Earlier work this paper cites.
Layered direct policy search for learning hierarchical skills
End, F., Akrour, R., Peters, J., and Neumann, G · 2017
Earlier work this paper cites.
Reverse curriculum generation for reinforcement learning
Florensa, C., Held, D., Wulfmeier, M., Zhang, M., and Abbeel, P · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Policy search with high-dimensional context variables
Tangkaratt, V., Van Hoof, H., Parisi, S., Neumann, G., Peters, J., and Sugiyama, M · 2017
Earlier work this paper cites.
Automatic goal generation for reinforcement learning agents
Florensa, C., Held, D., Geng, X., and Abbeel, P · 2018
Cited alongside, same era.
Intrinsic motivation and automatic curricula via asymmetric self-play
Sukhbaatar, S., Lin, Z., Kostrikov, I., Synnaeve, G., Szlam, A., and Fergus, R · 2018
Cited alongside, same era.
Contextual direct policy search: With regularized covariance matrix estimation
Abdolmaleki, A., Simoes, D., Lau, N., Reis, L. P., and Neumann, G · 2019
Cited alongside, same era.
Diversity is all you need: Learning skills without a reward function
Eysenbach, B., Gupta, A., Ibarz, J., and Levine, S · 2019
Cited alongside, same era.
Efficient exploration via state marginal matching
Lee, L., Eysenbach, B., Parisotto, E., Xing, E., Levine, S., and Salakhutdinov, R · 2019
Cited alongside, same era.
Differentiable trust region layers for deep reinforcement learning
Otto, F., Becker, P., Vien, N. A., Ziesche, H. C., and Neumann, G · 2021
Later among the works it cites.
Normalizing flows for probabilistic modeling and inference
Papamakarios, G., Nalisnick, E., Rezende, D. J., Mohamed, S., and Lakshminarayanan, B · 2021
Later among the works it cites.
Probabilistic mixture-of-experts for efficient deep reinforcement learning
Ren, J., Li, Y., Ding, Z., Pan, W., and Dong, H · 2021
Later among the works it cites.
Contextual latent-movements off-policy optimization for robotic manipulation skills
Tosatto, S., Chalvatzaki, G., and Peters, J · 2021
Later among the works it cites.
Specializing versatile skill libraries using local mixture of experts
Celik, O., Zhou, D., Li, G., Becker, P., and Neumann, G · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Explore, discover and learn: Unsupervised discovery of state-covering skills
Campos, V., Trott, A., Xiong, C., Socher, R., i Nieto, X. G., and Torres, J · 2020
Cited alongside, same era.
Emergent complexity and zero-shot transfer via unsupervised environment design
Dennis, M., Jaques, N., Vinitsky, E., Bayen, A., Russell, S., Critch, A., and Levine, S · 2020
Cited alongside, same era.
Implementation matters in deep policy gradients: A case study on ppo and trpo
Engstrom, L., Ilyas, A., Santurkar, S., Tsipras, D., Janoos, F., Rudolph, L., and Madry, A · 2020
Cited alongside, same era.
Model-based quality-diversity search for efficient robot learning
Keller, L., Tanneberg, D., Stark, S., and Peters, J · 2020
Cited alongside, same era.
Automated curriculum generation through setter-solver interactions
Racaniere, S., Lampinen, A., Santoro, A., Reichert, D., Firoiu, V., and Lillicrap, T · 2020
Cited alongside, same era.
A performance-based start state curriculum framework for reinforcement learning
Wöhlke, J., Schmitt, F., and van Hoof, H · 2020
Cited alongside, same era.
Automatic curriculum learning through value disagreement
Zhang, Y., Abbeel, P., and Pinto, L · 2020
Cited alongside, same era.
Implicit behavioral cloning
Florence, P., Lynch, C., Zeng, A., Ramirez, O. A., Wahid, A., Downs, L., Wong, A., Lee, J., Mordatch, I., and Tompson, J · 2022
Later among the works it cites.
Behavior transformers: Cloning k k modes with one stone
Shafiullah, N. M., Cui, Z., Altanzaya, A. A., and Pinto, L · 2022
Later among the works it cites.
Information maximizing curriculum: A curriculum-based approach for learning versatile skills
Blessing, D., Celik, O., Jia, X., Reuss, M., Li, M. X., Lioutikov, R., and Neumann, G · 2023
Later among the works it cites.
Diffusion policy: Visuomotor policy learning via action diffusion
Chi, C., Feng, S., Du, Y., Xu, Z., Cousineau, E., Burchfiel, B., and Song, S · 2023
Later among the works it cites.
Hierarchical policy blending as inference for reactive robot control
Hansel, K., Urain, J., Peters, J., and Chalvatzaki, G · 2023
Later among the works it cites.
Reparameterized policy learning for multimodal trajectory optimization
Huang, Z., Liang, L., Ling, Z., Li, X., Gan, C., and Su, H · 2023
Later among the works it cites.
Hierarchical policy blending as optimal transport
Le, A. T., Hansel, K., Peters, J., and Chalvatzaki, G · 2023
Later among the works it cites.
Deep black-box reinforcement learning with movement primitives
Otto, F., Celik, O., Zhou, H., Ziesche, H., Ngo, V. A., and Neumann, G · 2023
Later among the works it cites.
Multi-task reinforcement learning with mixture of orthogonal experts
Hendawy, A., Peters, J., and D’Eramo, C · 2024
Closest in time.
Towards diverse behaviors: A benchmark for imitation learning with human demonstrations
Jia, X., Blessing, D., Jiang, X., Reuss, M., Donat, A., Lioutikov, R., and Neumann, G · 2024
Closest in time.
On the benefit of optimal transport for curriculum reinforcement learning
Klink, P., D’Eramo, C., Peters, J., and Pajarinen, J · 2024
Closest in time.
Reverse forward curriculum learning for extreme sample and demonstration efficiency in rl
Tao, S., Shukla, A., Chan, T.-k., and Su, H · 2024
Closest in time.
Task adaptation from skills: Information geometry, disentanglement, and new objectives for unsupervised reinforcement learning
Yang, Y., Zhou, T., He, Q., Han, L., Pechenizkiy, M., and Fang, M · 2024
Closest in time.