Motor skill learning with local trajectory methods
Levine, S. (2014) · 2014
Later among the works it cites.
Learning neural network policies with guided policy search under unknown dynamics
Levine, S. and Abbeel, P. (2014) · 2014
Later among the works it cites.
Learning complex neural network policies with trajectory optimization
Levine, S. and Koltun, V. (2014) · 2014
Later among the works it cites.
Approximate MaxEnt inverse optimal control and its application for mental simulation of human interactions
Huang, D., Farahmand, A., Kitani, K. M., and Bagnell, J. A. (2015) · 2015
Later among the works it cites.
Shared autonomy via hindsight optimization
Javdani, S., Srinivasa, S., and Bagnell, J. A. (2015) · 2015
Later among the works it cites.
Trust region policy optimization
Schulman, J., Levine, S., Moritz, P., Jordan, M. I., and Abbeel, P. (2015) · 2015
Later among the works it cites.
Maximum entropy deep inverse reinforcement learning
Wulfmeier, M., Ondruska, P., and Posner, I. (2015) · 2015
Later among the works it cites.
Generative adversarial imitation learning
Ho, J. and Ermon, S. (2016) · 2016
Later among the works it cites.
End-to-end training of deep visuomotor policies
Levine, S., Finn, C., Darrell, T., and Abbeel, P. (2016) · 2016
Later among the works it cites.
High-dimensional continuous control using generalized advantage estimation
Schulman, J., Moritz, P., Levine, S., Jordan, M., and Abbeel, P. (2016) · 2016
Later among the works it cites.
Reinforcement learning with deep energy-based policies
Haarnoja, T., Tang, H., Abbeel, P., and Levine, S. (2017) · 2017
Later among the works it cites.
Pgq: Combining policy gradient and q-learning
O’Donoghue, B., Munos, R., Kavukcuoglu, K., and Mnih, V. (2017) · 2017
Later among the works it cites.
Equivalence between policy gradients and soft q-learning
Schulman, J., Chen, X., and Abbeel, P. (2017) · 2017
Later among the works it cites.
Maximum a posteriori policy optimisation
Abdolmaleki, A., Springenberg, J. T., Tassa, Y., Munos, R., Heess, N., and Riedmiller, M. (2018) · 2018
Closest in time.
Learning robust rewards with adversarial inverse reinforcement learning
Fu, J., Luo, K., and Levine, S. (2018) · 2018
Closest in time.
Meta-reinforcement learning of structured exploration strategies
Original
Gupta, A., Mendonca, R., Liu, Y., Abbeel, P., and Levine, S. (2018) · 2018
Closest in time.
Composable deep reinforcement learning for robotic manipulation
Haarnjoa, T., Pong, V., Zhou, A., Dalal, M., Abbeel, P., and Levine, S. (2018) · 2018
Closest in time.
Learning an embedding space for transferable robot skills
Hausman, K., Springenberg, J. T., Wang, Z., Heess, N., and Riedmiller, M. (2018) · 2018
Closest in time.