Fetching the paper…
Reading the bibliography…
Motivated by the success of ensembles for uncertainty estimation in supervised learning, we take a renewed look at how ensembles of $Q$-functions can be leveraged as the primary source of pessimism for offline reinforcement learning (RL).
Robust impulsive synchronization of coupled delayed neural networks with uncertainties
Li, P., Cao, J., and Wang, Z · 2007
Earlier work this paper cites.
Reinforcement learning in finite mdps: Pac analysis
Strehl, A. L., Li, L., and Littman, M. L · 2009
Earlier work this paper cites.
Batch reinforcement learning
Lange, S., Gabel, T., and Riedmiller, M · 2012
Earlier work this paper cites.
Bayesian learning for neural networks , volume 118
Neal, R. M · 2012
Earlier work this paper cites.
Batch mode reinforcement learning based on the synthesis of artificial trajectories
Fonteneau, R., Murphy, S. A., Wehenkel, L., and Ernst, D · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Why m heads are better than one: Training a diverse ensemble of deep networks
Lee, S., Purushwalkam, S., Cogswell, M., Crandall, D., and Batra, D · 2015
Earlier work this paper cites.
High confidence off-policy evaluation
Thomas, P., Theocharous, G., and Ghavamzadeh, M · 2015
Earlier work this paper cites.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Earlier work this paper cites.
Bayesian reinforcement learning: A survey
Ghavamzadeh, M., Mannor, S., Pineau, J., and Tamar, A · 2016
Earlier work this paper cites.
Simple and scalable predictive uncertainty estimation using deep ensembles
Lakshminarayanan, B., Pritzel, A., and Blundell, C · 2016
Earlier work this paper cites.
Deep exploration via bootstrapped dqn
Osband, I., Blundell, C., Pritzel, A., and Van Roy, B · 2016
Earlier work this paper cites.
Ucb exploration via q-ensembles
Chen, R. Y., Sidor, S., Abbeel, P., and Schulman, J · 2017
Earlier work this paper cites.
Count-based exploration with neural density models
Ostrovski, G., Bellemare, M. G., Oord, A., and Munos, R · 2017
Earlier work this paper cites.
Implicit weight uncertainty in neural networks
Pawlowski, N., Brock, A., Lee, M. C., Rajchl, M., and Glocker, B · 2017
Earlier work this paper cites.
Addressing function approximation error in actor-critic methods
Fujimoto, S., Hoof, H., and Meger, D · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C · 2018
Earlier work this paper cites.
Deep reinforcement learning and the deadly triad
Van Hasselt, H., Doron, Y., Strub, F., Hessel, M., Sonnerat, N., and Modayil, J · 2018
Earlier work this paper cites.
Off-policy deep reinforcement learning without exploration
Fujimoto, S., Meger, D., and Precup, D · 2019
Earlier work this paper cites.
Smile: Scalable meta inverse reinforcement learning through context-conditional policies
Ghasemipour, S. K. S., Gu, S., and Zemel, R · 2019
Cited alongside, same era.
Stabilizing off-policy q-learning via bootstrapping error reduction
Kumar, A., Fu, J., Tucker, G., and Levine, S · 2019
Cited alongside, same era.
Wide neural networks of any depth evolve as linear models under gradient descent
Lee, J., Xiao, L., Schoenholz, S., Bahri, Y., Novak, R., Sohl-Dickstein, J., and Pennington, J · 2019
Cited alongside, same era.
Neural tangents: Fast and easy infinite neural networks in python
Novak, R., Xiao, L., Hron, J., Lee, J., Alemi, A. A., Sohl-Dickstein, J., and Schoenholz, S. S · 2019
Cited alongside, same era.
Deep exploration via randomized value functions
Osband, I., Van Roy, B., Russo, D. J., Wen, Z., et al · 2019
Cited alongside, same era.
Conservative q-learning for offline reinforcement learning
Kumar, A., Zhou, A., Tucker, G., and Levine, S · 2020
Later among the works it cites.
Bandit algorithms
Lattimore, T. and Szepesvári, C · 2020
Later among the works it cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Levine, S., Kumar, A., Tucker, G., and Fu, J · 2020
Later among the works it cites.
Awac: Accelerating online reinforcement learning with offline datasets
Nair, A., Dalal, M., Gupta, A., and Levine, S · 2020
Later among the works it cites.
Hydra: Preserving ensemble diversity for model distillation
Tran, L., Veeling, B. S., Roth, K., Swiatkowski, J., Dillon, J. V., Snoek, J., Mandt, S., Salimans, T., Nowozin, S., and Jenatton, R · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Can you trust your model’s uncertainty? evaluating predictive uncertainty under dataset shift
Ovadia, Y., Fertig, E., Ren, J., Nado, Z., Sculley, D., Nowozin, S., Dillon, J. V., Lakshminarayanan, B., and Snoek, J · 2019
Cited alongside, same era.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning
Peng, X. B., Kumar, A., Zhang, G., and Levine, S · 2019
Cited alongside, same era.
Behavior regularized offline reinforcement learning
Wu, Y., Tucker, G., and Nachum, O · 2019
Cited alongside, same era.
Meta-inverse reinforcement learning with probabilistic context variables
Yu, L., Yu, T., Finn, C., and Ermon, S · 2019
Cited alongside, same era.
Binary ensemble neural network: More bits per network or more networks per bit?
Zhu, S., Dong, X., and Su, H · 2019
Cited alongside, same era.
An optimistic perspective on offline reinforcement learning
Agarwal, R., Schuurmans, D., and Norouzi, M · 2020
Cited alongside, same era.
Opal: Offline primitive discovery for accelerating offline reinforcement learning
Ajay, A., Kumar, A., Agrawal, P., Levine, S., and Nachum, O · 2020
Cited alongside, same era.
Wang, Z., Novikov, A., Zolna, K., Springenberg, J. T., Reed, S., Shahriari, B., Siegel, N., Merel, J., Gulcehre, C., Heess, N., et al · 2020
Later among the works it cites.
Batchensemble: an alternative approach to efficient ensemble and lifelong learning
Wen, Y., Tran, D., and Ba, J · 2020
Later among the works it cites.
Feature learning in infinite-width neural networks
Yang, G. and Hu, E. J · 2020
Later among the works it cites.
Offline policy selection under uncertainty
Yang, M., Dai, B., Nachum, O., Tucker, G., and Schuurmans, D · 2020
Later among the works it cites.
Uncertainty-based offline reinforcement learning with diversified q-ensemble
An, G., Moon, S., Kim, J.-H., and Song, H. O · 2021
Later among the works it cites.
A minimalist approach to offline reinforcement learning
Fujimoto, S. and Gu, S. S · 2021
Later among the works it cites.
Emaq: Expected-max q-learning operator for simple yet effective offline and online rl
Ghasemipour, S. K. S., Schuurmans, D., and Gu, S. S · 2021
Later among the works it cites.
Is pessimism provably efficient for offline rl?
Jin, Y., Yang, Z., and Wang, Z · 2021
Later among the works it cites.
Confident off-policy evaluation and selection through self-normalized importance weighting
Kuzborskij, I., Vernade, C., Gyorgy, A., and Szepesvári, C · 2021
Later among the works it cites.
Sunrise: A simple unified framework for ensemble learning in deep reinforcement learning
Lee, K., Laskin, M., Srinivas, A., and Abbeel, P · 2021
Later among the works it cites.
Online and offline reinforcement learning by planning with a learned model
Schrittwieser, J., Hubert, T., Mandhane, A., Barekatain, M., Antonoglou, I., and Silver, D · 2021
Later among the works it cites.
Pebl: Pessimistic ensembles for offline deep reinforcement learning
Smit, J., Ponnambalam, C. T., Spaan, M. T., and Oliehoek, F. A · 2021
Later among the works it cites.
Bellman-consistent pessimism for offline reinforcement learning
Xie, T., Cheng, C.-A., Jiang, N., Mineiro, P., and Agarwal, A · 2021
Later among the works it cites.
Offline-to-online reinforcement learning via balanced replay and pessimistic q-ensemble
Lee, S., Seo, Y., Lee, K., Abbeel, P., and Shin, J · 2022
Closest in time.