Dopamine: A Research Framework for Deep Reinforcement Learning
Pablo Samuel Castro, Subhodeep Moitra, Carles Gelada, Saurabh Kumar, and Marc G. Bellemare · 2018
Later among the works it cites.
Revisiting the arcade learning environment: Evaluation protocols and open problems for general agents
Marlos C. Machado, Marc G. Bellemare, Erik Talvitie, Joel Veness, Matthew Hausknecht, and Michael Bowling · 2018
Later among the works it cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clement Hongler · 2018
Later among the works it cites.
The effects of memory replay in reinforcement learning
Ruishan Liu and James Zou · 2018
Later among the works it cites.
Soft actor-critic algorithms and applications
Kristian Hartikainen George Tucker Sehoon Ha Jie Tan Vikash Kumar Henry Zhu Abhishek Gupta Pieter Abbeel Tuomas Haarnoja, Aurick Zhou and Sergey Levine · 2018
Later among the works it cites.
Diagnosing bottlenecks in deep q-learning algorithms
Justin Fu, Aviral Kumar, Matthew Soh, and Sergey Levine · 2019
Later among the works it cites.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
Tianhe Yu, Deirdre Quillen, Zhanpeng He, Ryan Julian, Karol Hausman, Chelsea Finn, and Sergey Levine · 2019
Later among the works it cites.
Combating label noise in deep learning using abstention
Sunil Thulasidasan, Tanmoy Bhattacharya, Jeff Bilmes, Gopinath Chennupati, and Jamal Mohd-Yusof · 2019
Later among the works it cites.
Towards characterizing divergence in deep q-learning
Original
Joshua Achiam, Ethan Knight, and Pieter Abbeel · 2019
Later among the works it cites.
Stabilizing off-policy q-learning via bootstrapping error reduction
Aviral Kumar, Justin Fu, George Tucker, and Sergey Levine · 2019
Later among the works it cites.
Ray interference: a source of plateaus in deep reinforcement learning
Original
Tom Schaul, Diana Borsa, Joseph Modayil, and Razvan Pascanu · 2019
Later among the works it cites.
Provably efficient q q -learning with function approximation via distribution shift error checking oracle
Simon Du, Yuping Luo, Ruosong Wang, and Hanrui Zhang · 2019
Later among the works it cites.
Provably efficient maximum entropy exploration
Elad Hazan, Sham Kakade, Karan Singh, and Abby Van Soest · 2019
Later among the works it cites.
Understanding the impact of entropy on policy optimization
Zafarali Ahmed, Nicolas Le Roux, Mohammad Norouzi, and Dale Schuurmans · 2019
Later among the works it cites.
Off-policy deep reinforcement learning without exploration
Scott Fujimoto, David Meger, and Doina Precup · 2019
Later among the works it cites.
Behavior regularized offline reinforcement learning
Original
Yifan Wu, George Tucker, and Ofir Nachum · 2019
Later among the works it cites.
Is a good representation sufficient for sample efficient reinforcement learning?
Simon S. Du, Sham M. Kakade, Ruosong Wang, and Lin F. Yang · 2020
Closest in time.
Gradient surgery for multi-task learning
Original
Tianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine, Karol Hausman, and Chelsea Finn · 2020
Closest in time.