Fetching the paper…
Reading the bibliography…
Deep Reinforcement Learning (RL) is proven powerful for decision making in simulated environments.
Distributional reinforcement learning for efficient exploration
Borislav Mavrin, Shangtong Zhang, Hengshuai Yao, Linglong Kong, Kaiwen Wu, and Yaoliang Yu · 1905
Earlier work this paper cites.
Learning dynamic algorithm portfolios
Matteo Gagliolo and Jürgen Schmidhuber · 2006
Earlier work this paper cites.
Algorithm selection as a bandit problem with unbounded losses
Matteo Gagliolo and Jürgen Schmidhuber · 2008
Earlier work this paper cites.
Analyzing bandit-based adaptive operator selection mechanisms
Álvaro Fialho, Luís Da Costa, Marc Schoenauer, and Michèle Sebag · 2010
Earlier work this paper cites.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
Sébastien Bubeck and Nicolò Cesa-Bianchi · 2012
Earlier work this paper cites.
Regret bounds for reinforcement learning with policy advice
Mohammad Gheshlaghi Azar, Alessandro Lazaric, and Emma Brunskill · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Cited alongside, same era.
Doubly robust off-policy evaluation for reinforcement learning
Nan Jiang and Lihong Li · 2015
Cited alongside, same era.
Trust region policy optimization
John Schulman, Sergey Levine, Philipp Moritz, Michael I. Jordan, and Pieter Abbeel · 2015
Cited alongside, same era.
Algorithm selection for combinatorial search problems: A survey
Lars Kotthoff · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adrià Puigdomènech Badia, Mehdi Mirza, Alex Graves, Timothy P. Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
A distributional perspective on reinforcement learning
Marc G. Bellemare, Will Dabney, and Rémi Munos · 2017
Later among the works it cites.
Population based training of neural networks
Max Jaderberg, Valentin Dalibard, Simon Osindero, Wojciech M Czarnecki, Jeff Donahue, Ali Razavi, Oriol Vinyals, Tim Green, Iain Dunning, Karen Simonyan, et al · 2017
Later among the works it cites.
More robust doubly robust off-policy evaluation
Mehrdad Farajtabar, Yinlam Chow, and Mohammad Ghavamzadeh · 2018
Later among the works it cites.
Reinforcement learning algorithm selection
Romain Laroche and Raphael Feraud · 2018
Later among the works it cites.
A generalized framework for population based training
Ang Li, Ola Spyra, Sagi Perel, Valentin Dalibard, Max Jaderberg, Chenjie Gu, David Budden, Tim Harley, and Pramod Gupta · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deep exploration via bootstrapped dqn
Ian Osband, Charles Blundell, Alexander Pritzel, and Benjamin Van Roy · 2016
Cited alongside, same era.
Later among the works it cites.