Fetching the paper…
Reading the bibliography…
Offline Reinforcement Learning promises to learn effective policies from previously-collected, static datasets without the need for exploration.
Neuro-dynamic programming
Bertsekas, D. P. and Tsitsiklis, J. N · 1996
Earlier work this paper cites.
Error bounds for approximate policy iteration
Munos, R · 2003
Earlier work this paper cites.
Error propagation for approximate policy and value iteration
Farahmand, A.-m., Szepesvári, C., and Munos, R · 2010
Earlier work this paper cites.
A kernel two-sample test
Gretton, A., Borgwardt, K. M., Rasch, M. J., Schölkopf, B., and Smola, A · 2012
Earlier work this paper cites.
Batch reinforcement learning
Lange, S., Gabel, T., and Riedmiller, M · 2012
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M · 2013
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R · 2014
Earlier work this paper cites.
Taming the noise in reinforcement learning via soft updates
Fox, R., Pakman, A., and Tishby, N · 2015
Earlier work this paper cites.
Variational dropout and the local reparameterization trick
Kingma, D. P., Salimans, T., and Welling, M · 2015
Earlier work this paper cites.
Approximate modified policy iteration and its application to the game of tetris
Scherrer, B., Ghavamzadeh, M., Gabillon, V., Lesner, B., and Geist, M · 2015
Earlier work this paper cites.
Uncertainty in deep learning
Gal, Y · 2016
Earlier work this paper cites.
Safe and efficient off-policy reinforcement learning
Munos, R., Stepleton, T., Harutyunyan, A., and Bellemare, M · 2016
Earlier work this paper cites.
Deep exploration via bootstrapped dqn
Osband, I., Blundell, C., Pritzel, A., and Van Roy, B · 2016
Earlier work this paper cites.
Averaged-dqn: Variance reduction and stabilization for deep reinforcement learning
Anschel, O., Baram, N., and Shimkin, N · 2017
Earlier work this paper cites.
Noisy networks for exploration
Fortunato, M., Azar, M. G., Piot, B., Menick, J., Osband, I., Graves, A., Mnih, V., Munos, R., Hassabis, D., Pietquin, O., et al · 2017
Earlier work this paper cites.
Simple and scalable predictive uncertainty estimation using deep ensembles
Lakshminarayanan, B., Pritzel, A., and Blundell, C · 2017
Cited alongside, same era.
Learning complex dexterous manipulation with deep reinforcement learning and demonstrations
Rajeswaran, A., Kumar, V., Gupta, A., Vezzani, G., Schulman, J., Todorov, E., and Levine, S · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Cited alongside, same era.
Mastering the game of go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., et al · 2017
Cited alongside, same era.
Tang, Y. and Kucukelbir, A · 2017
Off-policy deep reinforcement learning without exploration
Fujimoto, S., Meger, D., and Precup, D · 2019
Later among the works it cites.
Way off-policy batch deep reinforcement learning of implicit human preferences in dialog
Jaques, N., Ghandeharioun, A., Shen, J. H., Ferguson, C., Lapedriza, A., Jones, N., Gu, S., and Picard, R · 2019
Later among the works it cites.
Stabilizing off-policy q-learning via bootstrapping error reduction
Kumar, A., Fu, J., Soh, M., Tucker, G., and Levine, S · 2019
Later among the works it cites.
Safe policy improvement with baseline bootstrapping
Laroche, R., Trichelair, P., and Des Combes, R. T · 2019
Later among the works it cites.
Practical deep learning with bayesian principles
Osawa, K., Swaroop, S., Khan, M. E. E., Jain, A., Eschenhagen, R., Turner, R. E., and Yokota, R · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Distributional reinforcement learning with quantile regression
Dabney, W., Rowland, M., Bellemare, M. G., and Munos, R · 2018
Cited alongside, same era.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
Espeholt, L., Soyer, H., Munos, R., Simonyan, K., Mnih, V., Ward, T., Doron, Y., Firoiu, V., Harley, T., Dunning, I., et al · 2018
Cited alongside, same era.
Addressing function approximation error in actor-critic methods
Fujimoto, S., van Hoof, H., Meger, D., et al · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Cited alongside, same era.
Variational bayesian dropout: pitfalls and fixes
Hron, J., Matthews, A., and Ghahramani, Z · 2018
Cited alongside, same era.
Scalable deep reinforcement learning for vision-based robotic manipulation
Kalashnikov, D., Irpan, A., Pastor, P., Ibarz, J., Herzog, A., Jang, E., Quillen, D., Holly, E., Kalakrishnan, M., Vanhoucke, V., et al · 2018
Cited alongside, same era.
Bbq-networks: Efficient exploration in deep reinforcement learning for task-oriented dialogue systems
Lipton, Z., Li, X., Gao, J., Li, L., Ahmed, F., and Deng, L · 2018
Cited alongside, same era.
Peng, X. B., Kumar, A., Zhang, G., and Levine, S · 2019
Later among the works it cites.
Behavior regularized offline reinforcement learning
Wu, Y., Tucker, G., and Nachum, O · 2019
Later among the works it cites.
Mopo: Model-based offline policy optimization
Yu, T., Thomas, G., Yu, L., Ermon, S., Zou, J., Levine, S., Finn, C., and Ma, T · 2019
Later among the works it cites.
An optimistic perspective on offline reinforcement learning
Agarwal, R., Schuurmans, D., and Norouzi, M · 2020
Later among the works it cites.
D4rl: Datasets for deep data-driven reinforcement learning
Fu, J., Kumar, A., Nachum, O., Tucker, G., and Levine, S · 2020
Later among the works it cites.
Reinforcement learning with uncertainty estimation for tactical decision-making in intersections
Hoel, C.-J., Tram, T., and Sjöberg, J · 2020
Later among the works it cites.
Morel: Model-based offline reinforcement learning
Kidambi, R., Rajeswaran, A., Netrapalli, P., and Joachims, T · 2020
Later among the works it cites.
Conservative q-learning for offline reinforcement learning
Kumar, A., Zhou, A., Tucker, G., and Levine, S · 2020
Later among the works it cites.
Accelerating online reinforcement learning with offline datasets
Nair, A., Dalal, M., Gupta, A., and Levine, S · 2020
Later among the works it cites.
Randomized value functions via multiplicative normalizing flows
Touati, A., Satija, H., Romoff, J., Pineau, J., and Vincent, P · 2020
Later among the works it cites.