Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Later among the works it cites.
The option-critic architecture
Pierre-Luc Bacon, Jean Harb, and Doina Precup · 2017
Later among the works it cites.
Stochastic neural networks for hierarchical reinforcement learning
Original
Carlos Florensa, Yan Duan, and Pieter Abbeel · 2017
Later among the works it cites.
Noisy networks for exploration
Original
Meire Fortunato, Mohammad Gheshlaghi Azar, Bilal Piot, Jacob Menick, Ian Osband, Alex Graves, Vlad Mnih, Remi Munos, Demis Hassabis, Olivier Pietquin, et al · 2017
Later among the works it cites.
Emergence of locomotion behaviours in rich environments
Original
Nicolas Heess, Srinivasan Sriram, Jay Lemmon, Josh Merel, Greg Wayne, Yuval Tassa, Tom Erez, Ziyu Wang, Ali Eslami, Martin Riedmiller, et al · 2017
Later among the works it cites.
Hierarchical actor-critic
Original
Andrew Levy, Robert Platt, and Kate Saenko · 2017
Later among the works it cites.
Parameter space noise for exploration
Original
Matthias Plappert, Rein Houthooft, Prafulla Dhariwal, Szymon Sidor, Richard Y Chen, Xi Chen, Tamim Asfour, Pieter Abbeel, and Marcin Andrychowicz · 2017
Later among the works it cites.
Meta learning shared hierarchies
Kevin Frans, Jonathan Ho, Xi Chen, Pieter Abbeel, and John Schulman · 2018
Later among the works it cites.
Addressing function approximation error in actor-critic methods
Original
Scott Fujimoto, Herke van Hoof, and Dave Meger · 2018
Later among the works it cites.
Rainbow: Combining improvements in deep reinforcement learning
Matteo Hessel, Joseph Modayil, Hado Van Hasselt, Tom Schaul, Georg Ostrovski, Will Dabney, Dan Horgan, Bilal Piot, Mohammad Azar, and David Silver · 2018
Later among the works it cites.
Multi-agent manipulation via locomotion using hierarchical sim2real
Original
Ofir Nachum, Michael Ahn, Hugo Ponte, Shixiang Gu, and Vikash Kumar · 2019
Closest in time.