Fetching the paper…
Reading the bibliography…
We study an approach to offline reinforcement learning (RL) based on optimally solving finitely-represented MDPs derived from a static dataset of experience.
Dynamic programming
Richard Bellman and Samuel G. Kneale · 1958
Earlier work this paper cites.
Markov decision processes: Discrete stochastic dynamic programming
Martin L. Puterman · 1994
Earlier work this paper cites.
Stable function approximation in dynamic programming
G. Gordon · 1995
Earlier work this paper cites.
Nonparametric model-based reinforcement learning
Christopher G Atkeson · 1998
Earlier work this paper cites.
Approximate solutions to markov decision processes
G. Gordon and Tom Michael Mitchell · 1999
Earlier work this paper cites.
Stochastic dynamic programming with factored representations
Craig Boutilier, Richard Dearden, and Moisés Goldszmidt · 2000
Earlier work this paper cites.
Kernel-based reinforcement learning in average-cost problems
Dirk Ormoneit and P. Glynn · 2002
Earlier work this paper cites.
Kernel-based reinforcement learning
Dirk Ormoneit and Ś. Sen · 2002
Earlier work this paper cites.
Tree-based batch mode reinforcement learning
D. Ernst, P. Geurts, and L. Wehenkel · 2005
Earlier work this paper cites.
Model-based function approximation in reinforcement learning
Nicholas K Jong and Peter Stone · 2007
Earlier work this paper cites.
Op-elm: optimally pruned extreme learning machine
Yoan Miche, Antti Sorjamaa, Patrick Bas, Olli Simula, Christian Jutten, and Amaury Lendasse · 2009
Earlier work this paper cites.
Gpu-based markov decision process solver
Ársæll ór Jóhannsson · 2009
Earlier work this paper cites.
Implementation of kd-trees on the gpu to achieve real time graphics processing
W. W. Martin · 2012
Earlier work this paper cites.
Planning in factored action spaces with symbolic dynamic programming
Aswin Raghavan, Saket Joshi, Alan Fern, Prasad Tadepalli, and Roni Khardon · 2012
Earlier work this paper cites.
Batch mode reinforcement learning based on the synthesis of artificial trajectories
R. Fonteneau, S. Murphy, L. Wehenkel, and D. Ernst · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, Ioannis Antonoglou, Daan Wierstra, and Martin A. Riedmiller · 2013
Earlier work this paper cites.
Pac optimal exploration in continuous space markov decision processes
Jason Pazis and Ronald Parr · 2013
Cited alongside, same era.
Massively parallel kd-tree construction and nearest neighbor search algorithms
Linjia Hu, S. Nooshabadi, and M. Ahmadi · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, Andrei A. Rusu, J. Veness, Marc G. Bellemare, A. Graves, Martin A. Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, S. Petersen, C. Beattie, A. Sadik, Ioannis Antonoglou, H. King, D. Kumaran, Daan Wierstra, S. Legg, and Demis Hassabis · 2015
Cited alongside, same era.
A parallel solver for markov decision process in crowd simulations
S. Ruiz and B. Hernández · 2015
Cited alongside, same era.
An optimistic perspective on offline reinforcement learning
Rishabh Agarwal, D. Schuurmans, and Mohammad Norouzi · 2019
Later among the works it cites.
Benchmarking batch deep reinforcement learning algorithms
Scott Fujimoto, Edoardo Conti, Mohammad Ghavamzadeh, and Joelle Pineau · 2019
Later among the works it cites.
Algaedice: Policy gradient from arbitrary experience
Ofir Nachum, Bo Dai, Ilya Kostrikov, Yinlam Chow, L. Li, and D. Schuurmans · 2019
Later among the works it cites.
Behavior regularized offline reinforcement learning
Y. Wu, G. Tucker, and Ofir Nachum · 2019
Later among the works it cites.
Flambe: Structural complexity and representation learning of low rank mdps
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Charles Blundell, B. Uria, A. Pritzel, Y. Li, Avraham Ruderman, Joel Z. Leibo, Jack W. Rae, Daan Wierstra, and Demis Hassabis · 2016
Cited alongside, same era.
An emphatic approach to the problem of off-policy temporal-difference learning
R. Sutton, A. R. Mahmood, and Martha White · 2016
Cited alongside, same era.
Gpu-accelerated value iteration for the computation of reachability probabilities in mdps
Zhimin Wu, E. Hahn, A. Günay, L. Zhang, and Y. Liu · 2016
Cited alongside, same era.
Deep episodic value iteration for model-based meta-reinforcement learning
Steven Stenberg Hansen · 2017
Cited alongside, same era.
Neural episodic control
A. Pritzel, B. Uria, S. Srinivasan, Adrià Puigdomènech Badia, Oriol Vinyals, Demis Hassabis, Daan Wierstra, and Charles Blundell · 2017
Cited alongside, same era.
Handbook of discrete and computational geometry
Csaba D Toth, Joseph O’Rourke, and Jacob E Goodman · 2017
Cited alongside, same era.
gym-miniworld environment for openai gym
Maxime Chevalier-Boisvert · 2018
Cited alongside, same era.
A. Agarwal, Sham M. Kakade, A. Krishnamurthy, and W. Sun · 2020
Closest in time.
The importance of pessimism in fixed-dataset policy optimization
J. Buckman, Carles Gelada, and Marc G. Bellemare · 2020
Closest in time.
Rl unplugged: Benchmarks for offline reinforcement learning
Caglar Gulcehre, Ziyu Wang, A. Novikov, T. L. Paine, Sergio Gomez Colmenarejo, Konrad Zolna, Rishabh Agarwal, Josh Merel, Daniel J. Mankowitz, Cosmin Paduraru, Gabriel Dulac-Arnold, J. Li, Mohammad Norouzi, Matt Hoffman, Ofir Nachum, G. Tucker, Nicolas Heess, and N. D. Freitas · 2020
Closest in time.
Morel : Model-based offline reinforcement learning
R. Kidambi, A. Rajeswaran, Praneeth Netrapalli, and T. Joachims · 2020
Closest in time.
Conservative q-learning for offline reinforcement learning
Aviral Kumar, Aurick Zhou, G. Tucker, and Sergey Levine · 2020
Closest in time.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Sergey Levine, Aviral Kumar, G. Tucker, and Justin Fu · 2020
Closest in time.
Exact (then approximate) dynamic programming for deep reinforcement learning
Henrik Marklund, Suraj Nair, and Chelsea Finn · 2020
Closest in time.
Plannable approximations to mdp homomorphisms: Equivariance under actions
Elise van der Pol, Thomas Kipf, Frans A. Oliehoek, and M. Welling · 2020
Closest in time.
Ziyu Wang, A. Novikov, Konrad Zolna, Jost Tobias Springenberg, Scott Reed, B. Shahriari, N. Siegel, Josh Merel, Caglar Gulcehre, Nicolas Heess, and N. D. Freitas · 2020
Closest in time.
Plan2vec: Unsupervised representation learning by latent plans
G. Yang, A. Zhang, Ari S. Morcos, Joelle Pineau, P. Abbeel, and R. Calandra · 2020
Closest in time.
Mopo: Model-based offline policy optimization
Tianhe Yu, G. Thomas, Lantao Yu, S. Ermon, J. Zou, Sergey Levine, Chelsea Finn, and Tengyu Ma · 2020
Closest in time.
Episodic reinforcement learning with associative memory
Guangxiang Zhu, Zichuan Lin, G. Yang, and C. Zhang · 2020
Closest in time.