Fetching the paper…
Reading the bibliography…
Learning to act from observational data without active environmental interaction is a well-known challenge in Reinforcement Learning (RL).
Benchmarking batch deep reinforcement learning algorithms
Scott Fujimoto, Edoardo Conti, Mohammad Ghavamzadeh, and Joelle Pineau · 1910
Earlier work this paper cites.
Movement-produced stimulation in the development of visually guided behavior
Richard Held and Alan Hein · 1963
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Richard S. Sutton · 1988
Earlier work this paper cites.
On-line Q-learning using connectionist systems
Gavin A. Rummery and Mahesan Niranjan · 1994
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
Leemon Baird · 1995
Earlier work this paper cites.
Analysis of temporal-difference learning with function approximation
John N. Tsitsiklis and Benjamin Van Roy · 1996
Earlier work this paper cites.
Ziyu Wang, Alexander Novikov, Konrad Żołna, Jost Tobias Springenberg, Scott Reed, Bobak Shahriari, Noah Siegel, Josh Merel, Caglar Gulcehre, Nicolas Heess, et al · 2006
Earlier work this paper cites.
Implicit under-parameterization inhibits data-efficient deep reinforcement learning
Aviral Kumar, Rishabh Agarwal, Dibya Ghosh, and Sergey Levine · 2010
Earlier work this paper cites.
Double Q-learning
Hado van Hasselt · 2010
Earlier work this paper cites.
What are the statistical limits of offline RL with linear function approximation?
Ruosong Wang, Dean P. Foster, and Sham M. Kakade · 2010
Earlier work this paper cites.
Lecture 6.5: RMSProp
Tijmen Tieleman and Geoffrey Hinton · 2012
Earlier work this paper cites.
The Arcade Learning Environment: an evaluation platform for general agents
Marc G. Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2013
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeffrey Dean · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Earlier work this paper cites.
High confidence policy improvement
Philip Thomas, Georgios Theocharous, and Mohammad Ghavamzadeh · 2015
Earlier work this paper cites.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Earlier work this paper cites.
Deep reinforcement learning with double Q-learning
Hado van Hasselt, Arthur Guez, and David Silver · 2016
Earlier work this paper cites.
JAX: composable transformations of Python+NumPy programs, 2018
James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, and Qiao Zhang · 2018
Cited alongside, same era.
Dopamine: A research framework for deep reinforcement learning
Pablo Samuel Castro, Subhodeep Moitra, Carles Gelada, Saurabh Kumar, and Marc G. Bellemare · 2018
Cited alongside, same era.
Distributional reinforcement learning with quantile regression
Will Dabney, Mark Rowland, Marc G. Bellemare, and Rémi Munos · 2018
Cited alongside, same era.
Rainbow: combining improvements in deep reinforcement learning
Matteo Hessel, Joseph Modayil, Hado van Hasselt, Tom Schaul, Georg Ostrovski, Will Dabney, Dan Horgan, Bilal Piot, Mohammad Azar, and David Silver · 2018
Cited alongside, same era.
Revisiting the Arcade Learning Environment: Evaluation protocols and open problems for general agents
Marlos C Machado, Marc G. Bellemare, Erik Talvitie, Joel Veness, Matthew Hausknecht, and Michael Bowling · 2018
Optax: Composable gradient transformation and optimisation, in JAX!, 2020
Matteo Hessel, David Budden, Fabio Viola, Mihaela Rosca, Eren Sezener, and Tom Hennigan · 2020
Later among the works it cites.
MOReL: Model-based offline reinforcement learning
Rahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, and Thorsten Joachims · 2020
Later among the works it cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu · 2020
Later among the works it cites.
Provably good batch reinforcement learning without great exploration
Yao Liu, Adith Swaminathan, Alekh Agarwal, and Emma Brunskill · 2020
Later among the works it cites.
Deployment-efficient reinforcement learning via model-based offline optimization
Tatsuya Matsushima, Hiroki Furuta, Yutaka Matsuo, Ofir Nachum, and Shixiang Gu · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Reinforcement Learning: An Introduction
Richard S. Sutton and Andrew G. Barto · 2018
Cited alongside, same era.
Deep reinforcement learning and the deadly triad
Hado van Hasselt, Yotam Doron, Florian Strub, Matteo Hessel, Nicolas Sonnerat, and Joseph Modayil · 2018
Cited alongside, same era.
Towards characterizing divergence in deep Q-learning
Joshua Achiam, Ethan Knight, and Pieter Abbeel · 2019
Cited alongside, same era.
Learning from a learner
Alexis Jacq, Matthieu Geist, Ana Paiva, and Olivier Pietquin · 2019
Cited alongside, same era.
Recurrent experience replay in distributed reinforcement learning
Steven Kapturowski, Georg Ostrovski, John Quan, Remi Munos, and Will Dabney · 2019
Cited alongside, same era.
Stabilizing off-policy Q-learning via bootstrapping error reduction
Aviral Kumar, Justin Fu, Matthew Soh, George Tucker, and Sergey Levine · 2019
Cited alongside, same era.
Behavior regularized offline reinforcement learning
Yifan Wu, George Tucker, and Ofir Nachum · 2019
Cited alongside, same era.
Later among the works it cites.
Accelerating online reinforcement learning with offline datasets
Ashvin Nair, Murtaza Dalal, Abhishek Gupta, and Sergey Levine · 2020
Later among the works it cites.
DQN Zoo: Reference implementations of DQN-based agents, 2020
John Quan and Georg Ostrovski · 2020
Later among the works it cites.
MOPO: Model-based offline policy optimization
Tianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon, James Zou, Sergey Levine, Chelsea Finn, and Tengyu Ma · 2020
Later among the works it cites.
The importance of pessimism in fixed-dataset policy optimization
Jacob Buckman, Carles Gelada, and Marc G. Bellemare · 2021
Closest in time.
Continuous doubly constrained batch reinforcement learning
Rasool Fakoor, Jonas Mueller, Pratik Chaudhari, and Alexander J. Smola · 2021
Closest in time.
The low-rank simplicity bias in deep networks
Minyoung Huh, Hossein Mobahi, Richard Zhang, Brian Cheung, Pulkit Agrawal, and Phillip Isola · 2021
Closest in time.
Revisiting Rainbow: Promoting more insightful and inclusive deep reinforcement learning research
Johan S. Obando-Ceron and Pablo Samuel Castro · 2021
Closest in time.
Online and offline reinforcement learning by planning with a learned model
Julian Schrittwieser, Thomas Hubert, Amol Mandhane, Mohammadamin Barekatain, Ioannis Antonoglou, and David Silver · 2021
Closest in time.
Instabilities of offline RL with pre-trained neural representation
Ruosong Wang, Yifan Wu, Ruslan Salakhutdinov, and Sham M. Kakade · 2021
Closest in time.
Uncertainty weighted actor-critic for offline reinforcement learning
Yue Wu, Shuangfei Zhou, Nitish Srivastava, Joshua Susskind, Jian Zhang, Ruslan Salakhutdinov, and Hanlin Goh · 2021
Closest in time.
On the sample complexity of batch reinforcement learning with policy-induced data
Chenjun Xiao, Ilbin Lee, Bo Dai, Dale Schuurmans, and Csaba Szepesvari · 2021
Closest in time.
Off-policy deep reinforcement learning without exploration
Scott Fujimoto, David Meger, and Doina Precup · 2062
Closest in time.