Fetching the paper…
Reading the bibliography…
Model-based offline reinforcement learning approaches generally rely on bounds of model error.
The jackknife, the bootstrap and other resampling plans
Bradley Efron · 1982
Earlier work this paper cites.
Riemannian geometry
Manfredo Perdigao do Carmo · 1992
Earlier work this paper cites.
A framework for behavioural cloning
Michael Bain and Claude Sammut · 1995
Earlier work this paper cites.
Learning embedded maps of markov processes
Yaakov Engel and Shie Mannor · 2001
Earlier work this paper cites.
Predictive representations of state
Michael L Littman and Richard S Sutton · 2002
Earlier work this paper cites.
Tree-based batch mode reinforcement learning
Damien Ernst, Pierre Geurts, and Louis Wehenkel · 2005
Earlier work this paper cites.
Neural fitted q iteration–first experiences with a data efficient neural reinforcement learning method
Martin Riedmiller · 2005
Earlier work this paper cites.
Visualizing data using t-sne
Laurens Van der Maaten and Geoffrey Hinton · 2008
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Stéphane Ross, Geoffrey Gordon, and Drew Bagnell · 2011
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Batch mode reinforcement learning based on the synthesis of artificial trajectories
Raphael Fonteneau, Susan A Murphy, Louis Wehenkel, and Damien Ernst · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2013
Earlier work this paper cites.
Reliability of classification and prediction in k-nearest neighbours
Joe Luis Villa Medina et al · 2013
Earlier work this paper cites.
Nice: Non-linear independent components estimation
Laurent Dinh, David Krueger, and Yoshua Bengio · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Martin L Puterman · 2014
Earlier work this paper cites.
Stochastic backpropagation and approximate inference in deep generative models
Danilo Jimenez Rezende, Shakir Mohamed, and Daan Wierstra · 2014
Earlier work this paper cites.
Xi Chen, Diederik P Kingma, Tim Salimans, Yan Duan, Prafulla Dhariwal, John Schulman, Ilya Sutskever, and Pieter Abbeel · 2016
Earlier work this paper cites.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Yarin Gal and Zoubin Ghahramani · 2016
Earlier work this paper cites.
Deep learning , volume 1
Ian Goodfellow, Yoshua Bengio, Aaron Courville, and Yoshua Bengio · 2016
Earlier work this paper cites.
Latent constraints: Learning to generate conditionally from unconditional generative models
Jesse Engel, Matthew Hoffman, and Adam Roberts · 2017
Earlier work this paper cites.
Unsupervised learning of disentangled and interpretable representations from sequential data
Wei-Ning Hsu, Yu Zhang, and James Glass · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
A hierarchical latent variable encoder-decoder model for generating dialogues
Iulian Serban, Alessandro Sordoni, Ryan Lowe, Laurent Charlin, Joelle Pineau, Aaron Courville, and Yoshua Bengio · 2017
Cited alongside, same era.
Neural discrete representation learning
Aaron Van Den Oord, Oriol Vinyals, et al · 2017
Cited alongside, same era.
Latent space oddity: On the curvature of deep generative models
Georgios Arvanitidis, Lars Kai Hansen, and Søren Hauberg · 2018
Cited alongside, same era.
Optimizing the latent space of generative networks
Piotr Bojanowski, Armand Joulin, David Lopez-Pas, and Arthur Szlam · 2018
Cited alongside, same era.
Metrics for deep generative models
Disentangled skill embeddings for reinforcement learning
Janith C Petangoda, Sergio Pascual-Diaz, Vincent Adam, Peter Vrancx, and Jordi Grau-Moya · 2019
Later among the works it cites.
The natural language of actions
Guy Tennenholtz and Shie Mannor · 2019
Later among the works it cites.
An optimistic perspective on offline reinforcement learning
Rishabh Agarwal, Dale Schuurmans, and Mohammad Norouzi · 2020
Later among the works it cites.
Geometrically enriched latent spaces
Georgios Arvanitidis, Søren Hauberg, and Bernhard Schölkopf · 2020
Later among the works it cites.
The importance of pessimism in fixed-dataset policy optimization
Jacob Buckman, Carles Gelada, and Marc G Bellemare · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nutan Chen, Alexej Klushyn, Richard Kurle, Xueyan Jiang, Justin Bayer, and Patrick Smagt · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Cited alongside, same era.
Learning an embedding space for transferable robot skills
Karol Hausman, Jost Tobias Springenberg, Ziyu Wang, Nicolas Heess, and Martin Riedmiller · 2018
Cited alongside, same era.
An environment for autonomous driving decision-making
Edouard Leurent · 2018
Cited alongside, same era.
Rllib: Abstractions for distributed reinforcement learning
Eric Liang, Richard Liaw, Robert Nishihara, Philipp Moritz, Roy Fox, Ken Goldberg, Joseph Gonzalez, Michael Jordan, and Ion Stoica · 2018
Cited alongside, same era.
Randomized prior functions for deep reinforcement learning
Ian Osband, John Aslanides, and Albin Cassirer · 2018
Cited alongside, same era.
Learning action representations for reinforcement learning
Yash Chandak, Georgios Theocharous, James Kostas, Scott Jordan, and Philip Thomas · 2019
Cited alongside, same era.
Guided variational autoencoder for disentanglement learning
Zheng Ding, Yifan Xu, Weijian Xu, Gaurav Parmar, Yang Yang, Max Welling, and Zhuowen Tu · 2020
Later among the works it cites.
Revisiting fundamentals of experience replay
William Fedus, Prajit Ramachandran, Rishabh Agarwal, Yoshua Bengio, Hugo Larochelle, Mark Rowland, and Will Dabney · 2020
Later among the works it cites.
D4rl: Datasets for deep data-driven reinforcement learning
Justin Fu, Aviral Kumar, Ofir Nachum, George Tucker, and Sergey Levine · 2020
Later among the works it cites.
Mastering atari with discrete world models
Danijar Hafner, Timothy Lillicrap, Mohammad Norouzi, and Jimmy Ba · 2020
Later among the works it cites.
Is pessimism provably efficient for offline rl?
Ying Jin, Zhuoran Yang, and Zhaoran Wang · 2020
Later among the works it cites.
Variational autoencoders with riemannian brownian motion priors
Dimitris Kalatzis, David Eklund, Georgios Arvanitidis, and Søren Hauberg · 2020
Later among the works it cites.
Morel: Model-based offline reinforcement learning
Rahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, and Thorsten Joachims · 2020
Later among the works it cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu · 2020
Later among the works it cites.
Simple and effective vae training with calibrated decoders
Oleh Rybkin, Kostas Daniilidis, and Sergey Levine · 2020
Later among the works it cites.
Mastering atari, go, chess and shogi by planning with a learned model
Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert, Karen Simonyan, Laurent Sifre, Simon Schmitt, Arthur Guez, Edward Lockhart, Demis Hassabis, Thore Graepel, et al · 2020
Later among the works it cites.
Decoupling representation learning from reinforcement learning
Adam Stooke, Kimin Lee, Pieter Abbeel, and Michael Laskin · 2020
Later among the works it cites.
Overcoming model bias for robust offline deep reinforcement learning
Phillip Swazinna, Steffen Udluft, and Thomas Runkler · 2020
Later among the works it cites.
What are the statistical limits of offline rl with linear function approximation?
Ruosong Wang, Dean P Foster, and Sham M Kakade · 2020
Later among the works it cites.
Mopo: Model-based offline policy optimization
Tianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon, James Zou, Sergey Levine, Chelsea Finn, and Tengyu Ma · 2020
Later among the works it cites.
Andrea Zanette · 2020
Later among the works it cites.
Masked contrastive representation learning for reinforcement learning
Jinhua Zhu, Yingce Xia, Lijun Wu, Jiajun Deng, Wengang Zhou, Tao Qin, and Houqiang Li · 2020
Later among the works it cites.
Comparison of bayesian, k-nearest neighbor and gaussian process regression methods for quantifying uncertainty of suspended sediment concentration prediction
Aboalhasan Fathabadi, Seyed Morteza Seyedian, and Arash Malekian · 2021
Closest in time.