Fetching the paper…
Reading the bibliography…
Offline reinforcement learning (offline RL), which aims to find an optimal policy from a previously collected static dataset, bears algorithmic difficulties due to function approximation errors from out-of-distribution (OOD) data points.
Expected normal order statistics(exact and approximate)
JP Royston · 1982
Earlier work this paper cites.
Understanding the qr algorithm
David S Watkins · 1982
Earlier work this paper cites.
Actor-critic algorithms
Vijay R Konda and John N Tsitsiklis · 2000
Earlier work this paper cites.
Exploration–exploitation tradeoff using variance estimates in multi-armed bandits
Jean-Yves Audibert, Rémi Munos, and Csaba Szepesvári · 2009
Earlier work this paper cites.
Practical bayesian optimization of machine learning algorithms
Jasper Snoek, Hugo Larochelle, and Ryan P Adams · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Marc G Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2013
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Deep exploration via bootstrapped dqn
Ian Osband, Charles Blundell, Alexander Pritzel, and Benjamin Van Roy · 2016
Earlier work this paper cites.
Deep reinforcement learning with double q-learning
Hado Van Hasselt, Arthur Guez, and David Silver · 2016
Earlier work this paper cites.
Ucb exploration via q-ensembles
Richard Y Chen, Szymon Sidor, Pieter Abbeel, and John Schulman · 2017
Earlier work this paper cites.
Addressing function approximation error in actor-critic methods
Scott Fujimoto, Herke Hoof, and David Meger · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Cited alongside, same era.
Randomized prior functions for deep reinforcement learning
Ian Osband, John Aslanides, and Albin Cassirer · 2018
Cited alongside, same era.
Generative adversarial user model for reinforcement learning based recommendation system
Xinshi Chen, Shuang Li, Hui Li, Shaohua Jiang, and Le Song · 2019
Cited alongside, same era.
Better exploration with optimistic actor-critic
Kamil Ciosek, Quan Vuong, Robert Loftin, and Katja Hofmann · 2019
Cited alongside, same era.
Estimating risk and uncertainty in deep reinforcement learning
Selective dyna-style planning under limited model capacity
Zaheer Abbas, Samuel Sokota, Erin Talvitie, and Martha White · 2020
Later among the works it cites.
An optimistic perspective on offline reinforcement learning
Rishabh Agarwal, Dale Schuurmans, and Mohammad Norouzi · 2020
Later among the works it cites.
D4rl: Datasets for deep data-driven reinforcement learning
Justin Fu, Aviral Kumar, Ofir Nachum, George Tucker, and Sergey Levine · 2020
Later among the works it cites.
Morel: Model-based offline reinforcement learning
Rahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, and Thorsten Joachims · 2020
Later among the works it cites.
Conservative q-learning for offline reinforcement learning
Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine · 2020
Later among the works it cites.
Maxmin q-learning: Controlling the estimation bias of q-learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
William R Clements, Bastien Van Delft, Benoît-Marie Robaglia, Reda Bahi Slaoui, and Sébastien Toth · 2019
Cited alongside, same era.
Off-policy deep reinforcement learning without exploration
Scott Fujimoto, David Meger, and Doina Precup · 2019
Cited alongside, same era.
Stabilizing off-policy q-learning via bootstrapping error reduction
Aviral Kumar, Justin Fu, George Tucker, and Sergey Levine · 2019
Cited alongside, same era.
Alphastar: Mastering the real-time strategy game starcraft ii
Oriol Vinyals, Igor Babuschkin, Junyoung Chung, Michael Mathieu, Max Jaderberg, Wojciech M Czarnecki, Andrew Dudzik, Aja Huang, Petko Georgiev, Richard Powell, et al · 2019
Cited alongside, same era.
Behavior regularized offline reinforcement learning
Yifan Wu, George Tucker, and Ofir Nachum · 2019
Cited alongside, same era.
Qingfeng Lan, Yangchen Pan, Alona Fyshe, and Martha White · 2020
Later among the works it cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu · 2020
Later among the works it cites.
Mopo: Model-based offline policy optimization
Tianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon, James Zou, Sergey Levine, Chelsea Finn, and Tengyu Ma · 2020
Later among the works it cites.
Sunrise: A simple unified framework for ensemble learning in deep reinforcement learning
Kimin Lee, Michael Laskin, Aravind Srinivas, and Pieter Abbeel · 2021
Closest in time.
Uncertainty weighted actor-critic for offline reinforcement learning, 2021
Yue Wu, Shuangfei Zhai, Nitish Srivastava, Joshua M. Susskind, Jian Zhang, Ruslan Salakhutdinov, and Hanlin Goh · 2021
Closest in time.