Fetching the paper…
Reading the bibliography…
Reinforcement learning algorithms are typically limited to learning a single solution for a specified task, even though diverse solutions often exist.
R. J. Williams, Simple statistical gradient-following algorithms for connectionist reinforcement learning, Machine Learning 8 (1992) 229–256
1992
Earlier work this paper cites.
M. I. Jordan, Z. Ghahramani, T. S. Jaakkola, L. K. Saul, An introduction to variational methods for graphical models, Machine Learning 37 (1999) 183–233
1999
Earlier work this paper cites.
R. S. Sutton, D. McAllester, S. Singh, Y. Mansour, Policy gradient methods for reinforcement learning with function approximation, in: Advances in Neural Information Processing Systems, Vol. 12, 1999, pp. 1057–1063
1999
Earlier work this paper cites.
D. Barber, F. V. Agakov, The IM algorithm: A variational approach to information maximization, in: Advances in Neural Information Processing Systems, Vol. 16, 2003, pp. 201–208
2003
Earlier work this paper cites.
E. Todorov, T. Erez, Y. Tassa, Mujoco: A physics engine for model-based control, in: 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems, 2012, pp. 5026–5033
2012
Earlier work this paper cites.
D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, M. Riedmiller, Deterministic policy gradient algorithms, in: Proceedings of the International Conference on Machine Learning, 2014
2014
Earlier work this paper cites.
A. Cully, J. Clune, D. Tarapore, J. B. Mouret, Robots that can adapt like animals, Nature (2015)
2015
Earlier work this paper cites.
S. Levine, C. Finn, T. Darrell, P. Abbeel, End-to-end training of deep visuomotor policies, Journal of Machine Learning Research 17 (39) (2016) 1–40
2016
Earlier work this paper cites.
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, S. Dieleman, D. Grewe, J. Nham, N. Kalchbrenner, I. Sutskever, T. Lillicrap, M. Leach, K. Kavukcuoglu, T. Graepel, D. Hassabis, Mastering the game of go with deep neural networks and tree search, Nature 529 (7587) (2016) 484–489
2016
Earlier work this paper cites.
R. Munos, T. Stepleton, A. Harutyunyan, M. G. Bellemare, Safe and efficient off-policy reinforcement learning, in: Advances in Neural Information Processing Systems, Vol. 29, 2016, pp. 1054–1062
2016
Earlier work this paper cites.
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, D. Wierstra, Continuous control with deep reinforcement learning, in: Proceedings of the International Conference on Learning Representations, 2016
2016
Earlier work this paper cites.
J. K. Pugh, L. B. Soros, K. O. Stanley, Quality diversity: A new frontier for evolutionary computation, Frontiers in Robotics and AI (2016)
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
Z. Wang, J. Merel, S. Reed, G. Wayne, N. de Freitas, N. Heess, Robust imitation of diverse behaviors, in: Advances in Neural Information Processing Systems, Vol. 30, 2017, pp. 5320–5329
2017
Earlier work this paper cites.
Y. Li, J. Song, S. Ermon, InfoGAIL: Interpretable imitation learning fromvisual demonstrations, in: Advances in Neural Information Processing Systems, Vol. 30, 2017, pp. 3812–3822
2017
Cited alongside, same era.
P. L. Bacon, J. Harb, D. Precup, The option-critic architecture, in: Proceedings of the AAAI Conference on Artificial Intelligence (AAAI), 2017, pp. 1726–1734
2017
Cited alongside, same era.
C. Florensa, Y. Duan, P. Abbeel, Stochastic neural networks for hierarchical reinforcement learning, in: Proceedings of the International Conference on Learning Representations, 2017
2017
Cited alongside, same era.
R. S. Sutton, A. G. Barto, Reinforcement Learning: An Introduction, 2nd Edition, MIT Press, 2018
2018
Cited alongside, same era.
S. Fujimoto, H. van Hoof, D. Meger, Addressing function approximation error in actor-critic methods, in: Proceedings of the International Conference on Machine Learning, Vol. 80, 2018, pp. 1587–1596
J. Merel, L. Hasenclever, A. Galashov, A. Ahuja, V. Pham, Y. W. T. G. Wayne and, N. Heess, Neural probabilistic motor primitives for humanoid control, in: Proceedings of the International Conference on Learning Representations, 2019
2019
Later among the works it cites.
O. Nachum, S. Gu, H. Lee, S. Levine, Near optimal representation learning for hierarchical reinforcement learning, in: Proceedings of the International Conference on Learning Representations, 2019
2019
Later among the works it cites.
T. Osa, V. Tangkaratt, M. Sugiyama, Hierarchical reinforcement learning via advantage-weighted information maximization, in: Proceedings of the International Conference on Learning Representations, 2019
2019
Later among the works it cites.
C. Bodnar, A. Li, K. Hausman, P. Pastor, M. Kalakrishnan, Quantile qt-opt for risk-awarevision-based robotic grasping, in: Robotics and Science and Systems, 2020
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
M. Toussaint, K. R. Allen, K. A. Smith, J. B. Tenenbaum, Differentiable physics and stable modes for tool-use and manipulation planning, in: Proceedings of Robotics: Sciences and Systems, 2018
2018
Cited alongside, same era.
T. Haarnoja, A. Zhou, P. Abbeel, S. Levine, Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor, in: Proceedings of the International Conference on Machine Learning, Vol. 80, 2018, pp. 1861–1870
2018
Cited alongside, same era.
X. B. Peng, P. Abbeel, S. Levine, M. van de Panne, Deepmimic: Example-guided deep reinforcement learning of physics-based character skills, ACM Transactions on Graphics 37 (4) (2018) 143:1–143:14
2018
Cited alongside, same era.
O. Nachum, S. Gu, H. Lee, S. Levine, Data-efficient hierarchical reinforcement learning, in: Advances in Neural Information Processing Systems, 2018, pp. 3303–3313
2018
Cited alongside, same era.
J. Achiam, Spinning Up in Deep Reinforcement Learning (2018)
2018
Cited alongside, same era.
B. Eysenbach, A. Gupta, J. Ibarz, S. Levine, Diversity is all you need: Learning skills without a reward function, in: Proceedings of the International Conference on Learning Representations, 2019
2019
Cited alongside, same era.
Y. Burda, H. Edwards, A. Storkey, O. Klimov, Exploration by random network distillation, in: Proceedings of the International Conference on Learning Representations, 2019
2019
Cited alongside, same era.
D. Sorokin, A. Ulanov, E. Sazhina, A. Lvovsky, Interferobot: aligning an optical interferometer by a reinforcement learning agent, in: Advances in Neural Information Processing Systems, 2020
2020
Later among the works it cites.
S. Kumar, A. Kumar, S. Levine, C. Finn, One solution is not all you need:few-shot extrapolation via structured maxent rl, in: 33 (Ed.), Advances in Neural Information Processing Systems, 2020, pp. 8198–8210
2020
Later among the works it cites.
A. Sharma, S. Gu, S. Levine, V. Kumar, K. Hausman, Dynamics-aware unsupervised discovery of skills, in: Proceedings of the International Conference on Learning Representations, 2020
2020
Later among the works it cites.
A. P. Badia, B. Piot, S. Kapturowski, P. Sprechmann, A. Vitvitskyi, D. Guo, C. Blundell, Agent57: Outperforming the atari human benchmark, in: Proceedings of the International Conference on Machine Learning, Vol. 119, 2020, pp. 507–517
2020
Later among the works it cites.
J. Parker-Holder, A. Pacchiano, K. Choromanski, S. Roberts, Effective diversity in population based reinforcement learning, in: Advances in Neural Information Processing Systems, Vol. 33, 2020, pp. 18050–18062
2020
Later among the works it cites.
A. Orthey, B. Frész, M. Toussaint, Motion planning explorer: Visualizing localminima using a local-minima tree, IEEE Robotics and Automation Letters 5 (2) (2020) 346–353
2020
Later among the works it cites.
T. Osa, Multimodal trajectory optimization for motion planning, The International Journal of Robotics Research 39 (8) (2020) 983–1001
2020
Later among the works it cites.
T. Gangwani, J. Peng, Y. Zhou, Harnessing distribution ratio estimators forlearning agents with quality and diversity, in: Proceedings of Conference on Robot Learning, 2020
2020
Later among the works it cites.
A. Puigdomènech Badia, P. Sprechmann, D. Vitvitskyi, A. andGuo, B. Piot, S. Kapturowski, O. Tieleman, M. Arjovsky, A. Pritzel, A. Bolt, C. Blundell, Nevergive up: Learning directed exploration strategies, in: Proceedings of the International Conference on Learning Representations, 2020
2020
Later among the works it cites.
T. Osa, Motion planning by learning the solution manifold in trajectory optimization, arXiv (2021)
2021
Closest in time.