Fetching the paper…
Reading the bibliography…
Using a model of the environment and a value function, an agent can construct many estimates of a state's value, by unrolling the model for different lengths and bootstrapping with its value function.
Dream to control: Learning behaviors by latent imagination
Hafner, D., Lillicrap, T., Ba, J., and Norouzi, M · 1912
Earlier work this paper cites.
Games against nature
Milnor, J · 1951
Earlier work this paper cites.
A Markovian decision process
Bellman, R · 1957
Earlier work this paper cites.
The foundations of statistics
Savage, L. J · 1972
Earlier work this paper cites.
Model predictive heuristic control
Richalet, J., Rault, A., Testud, J., and Papon, J · 1978
Earlier work this paper cites.
Learning how the world works: Specifications for predictive networks in robots and brains
Werbos, P. J · 1987
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Sutton, R. S · 1988
Earlier work this paper cites.
Model predictive control: Theory and practice—A survey
Garcia, C. E., Prett, D. M., and Morari, M · 1989
Earlier work this paper cites.
An on-line algorithm for dynamic reinforcement learning and planning in reactive environments
Schmidhuber, J · 1990
Earlier work this paper cites.
Dyna, an integrated architecture for learning, planning, and reacting
Sutton, R. S · 1991
Earlier work this paper cites.
Q-learning
Watkins, C. J. and Dayan, P · 1992
Earlier work this paper cites.
Python reference manual
Van Rossum, G. and Drake Jr, F. L · 1995
Earlier work this paper cites.
A comparison of some error estimates for neural network models
Tibshirani, R · 1996
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Bayesian Q-learning
Dearden, R., Friedman, N., and Russell, S · 1998
Earlier work this paper cites.
Model-based Bayesian exploration
Dearden, R., Friedman, N., and Andre, D · 1999
Earlier work this paper cites.
A Bayesian framework for reinforcement learning
Strens, M · 2000
Earlier work this paper cites.
Building ensembles with heterogeneous models , 2003
Wichard, J., Merkwirth, C., and Ogorzalek, M · 2003
Earlier work this paper cites.
Universal artificial intelligence: Sequential decisions based on algorithmic probability
Hutter, M · 2004
Earlier work this paper cites.
Efficient selectivity and backup operators in Monte-Carlo tree search
Coulom, R · 2006
Earlier work this paper cites.
Matplotlib: A 2D graphics environment
Hunter, J. D · 2007
Earlier work this paper cites.
A theoretical and empirical analysis of Expected Sarsa
Van Seijen, H., Van Hasselt, H., Whiteson, S., and Wiering, M · 2009
Earlier work this paper cites.
Formal theory of creativity, fun, and intrinsic motivation (1990–2010)
Schmidhuber, J · 2010
Earlier work this paper cites.
Improving translation model by monolingual data
Bojar, O. and Tamchyna, A · 2011
Earlier work this paper cites.
Exploration in model-based reinforcement learning by empirically estimating learning progress
Lopes, M., Lang, T., Toussaint, M., and Oudeyer, P.-Y · 2012
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Puterman, M. L · 2014
Earlier work this paper cites.
Weight uncertainty in neural network
Blundell, C., Cornebise, J., Kavukcuoglu, K., and Wierstra, D · 2015
Earlier work this paper cites.
Fast and accurate deep network learning by exponential linear units (elus)
Clevert, D.-A., Unterthiner, T., and Hochreiter, S · 2015
Earlier work this paper cites.
Neural network ensembles in reinforcement learning
Faußer, S. and Schwenker, F · 2015
Earlier work this paper cites.
Stochastic systems: Estimation, identification, and adaptive control
Kumar, P. R. and Varaiya, P · 2015
Earlier work this paper cites.
Incentivizing exploration in reinforcement learning with deep predictive models
Stadie, B. C., Levine, S., and Abbeel, P · 2015
Cited alongside, same era.
Embed to control: A locally linear latent dynamics model for control from raw images
Watter, M., Springenberg, J. T., Boedecker, J., and Riedmiller, M · 2015
Cited alongside, same era.
Tensorflow: A system for large-scale machine learning
Abadi, M., Barham, P., Chen, J., Chen, Z., Davis, A., Dean, J., Devin, M., Ghemawat, S., Irving, G., Isard, M., et al · 2016
Cited alongside, same era.
Unifying count-based exploration and intrinsic motivation
Bellemare, M., Srinivasan, S., Ostrovski, G., Schaul, T., Saxton, D., and Munos, R · 2016
Cited alongside, same era.
Stochastic latent actor-critic: Deep reinforcement learning with a latent variable model
Lee, A. X., Nagabandi, A., Abbeel, P., and Levine, S · 2019
Later among the works it cites.
A simple baseline for bayesian uncertainty in deep learning
Maddox, W. J., Izmailov, P., Garipov, T., Vetrov, D. P., and Wilson, A. G · 2019
Later among the works it cites.
Self-supervised exploration via disagreement
Pathak, D., Gandhi, D., and Gupta, A · 2019
Later among the works it cites.
Model-based active exploration
Shyam, P., Jaśkowski, W., and Gomez, F · 2019
Later among the works it cites.
Minatar: An atari-inspired testbed for thorough and reproducible reinforcement learning experiments
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Huang, G., Sun, Y., Liu, Z., Sedra, D., and Weinberger, K. Q · 2016
Cited alongside, same era.
Reinforcement learning with unsupervised auxiliary tasks
Jaderberg, M., Mnih, V., Czarnecki, W. M., Schaul, T., Leibo, J. Z., Silver, D., and Kavukcuoglu, K · 2016
Cited alongside, same era.
Simple and scalable predictive uncertainty estimation using deep ensembles
Lakshminarayanan, B., Pritzel, A., and Blundell, C · 2016
Cited alongside, same era.
Safe and efficient off-policy reinforcement learning
Munos, R., Stepleton, T., Harutyunyan, A., and Bellemare, M. G · 2016
Cited alongside, same era.
Deep exploration via bootstrapped DQN
Osband, I., Blundell, C., Pritzel, A., and Van Roy, B · 2016
Cited alongside, same era.
Averaged-dqn: Variance reduction and stabilization for deep reinforcement learning
Anschel, O., Baram, N., and Shimkin, N · 2017
Cited alongside, same era.
A distributional perspective on reinforcement learning
Bellemare, M. G., Dabney, W., and Munos, R · 2017
Cited alongside, same era.
Treeqn and atreec: Differentiable tree-structured models for deep reinforcement learning
Farquhar, G., Rocktäschel, T., Igl, M., and Whiteson, S · 2017
Cited alongside, same era.
Young, K. and Tian, T · 2019
Later among the works it cites.
An optimistic perspective on offline reinforcement learning
Agarwal, R., Schuurmans, D., and Norouzi, M · 2020
Later among the works it cites.
Depth uncertainty in neural networks
Antorán, J., Allingham, J. U., and Hernández-Lobato, J. M · 2020
Later among the works it cites.
The DeepMind JAX Ecosystem , 2020
Babuschkin, I., Baumli, K., Bell, A., Bhupatiraju, S., Bruce, J., Buchlovsky, P., Budden, D., Cai, T., Clark, A., Danihelka, I., Fantacci, C., Godwin, J., Jones, C., Hennigan, T., Hessel, M., Kapturowski, S., Keck, T., Kemaev, I., King, M., Martens, L., Mikulik, V., Norman, T., Quan, J., Papamakarios, G., Ring, R., Ruiz, F., Sanchez, A., Schneider, R., Sezener, E., Spencer, S., Srinivasan, S., Stokowiec, W., and Viola, F · 2020
Later among the works it cites.
Ready policy one: World building through active learning
Ball, P., Parker-Holder, J., Pacchiano, A., Choromanski, K., and Roberts, S · 2020
Later among the works it cites.
Imagined value gradients: Model-based policy optimization with tranferable latent dynamics models
Byravan, A., Springenberg, J. T., Abdolmaleki, A., Hafner, R., Neunert, M., Lampe, T., Siegel, N., Heess, N., and Riedmiller, M · 2020
Later among the works it cites.
The value-improvement path: Towards better representations for reinforcement learning
Dabney, W., Barreto, A., Rowland, M., Dadashi, R., Quan, J., Bellemare, M. G., and Silver, D · 2020
Later among the works it cites.
Efficient and scalable bayesian neural nets with rank-1 factors
Dusenberry, M., Jerfel, G., Wen, Y., Ma, Y., Snoek, J., Heller, K., Lakshminarayanan, B., and Tran, D · 2020
Later among the works it cites.
Can autonomous vehicles identify, recover from, and adapt to distribution shifts?
Filos, A., Tigkas, P., McAllister, R., Rhinehart, N., Levine, S., and Gal, Y · 2020
Later among the works it cites.
Temporal Difference Uncertainties as a Signal for Exploration
Flennerhag, S., Wang, J. X., Sprechmann, P., Visin, F., Galashov, A., Kapturowski, S., Borsa, D. L., Heess, N., Barreto, A., and Pascanu, R · 2020
Later among the works it cites.
Value-driven hindsight modelling
Guez, A., Viola, F., Weber, T., Buesing, L., Kapturowski, S., Precup, D., Silver, D., and Heess, N · 2020
Later among the works it cites.
Contrastive variational model-based reinforcement learning for complex observations
Ma, X., Chen, S., Hsu, D., and Lee, W. S · 2020
Later among the works it cites.
Uncertainty in neural networks: Approximately bayesian ensembling
Pearce, T., Leibfried, F., and Brintrup, A · 2020
Later among the works it cites.
RIDE: Rewarding impact-driven exploration for procedurally-generated environments
Raileanu, R. and Rocktäschel, T · 2020
Later among the works it cites.
Off-policy actor-critic with shared experience replay
Schmitt, S., Hessel, M., and Simonyan, K · 2020
Later among the works it cites.
Mastering atari, go, chess and shogi by planning with a learned model
Schrittwieser, J., Antonoglou, I., Hubert, T., Simonyan, K., Sifre, L., Schmitt, S., Guez, A., Lockhart, E., Hassabis, D., Graepel, T., et al · 2020
Later among the works it cites.
Planning to explore via self-supervised world models
Sekar, R., Rybkin, O., Daniilidis, K., Abbeel, P., Hafner, D., and Pathak, D · 2020
Later among the works it cites.
dm_control: Software and tasks for continuous control
Tunyasuvunakool, S., Muldal, A., Doron, Y., Liu, S., Bohez, S., Merel, J., Erez, T., Lillicrap, T., Heess, N., and Tassa, Y · 2020
Later among the works it cites.
Hyperparameter ensembles for robustness and uncertainty quantification
Wenzel, F., Snoek, J., Tran, D., and Jenatton, R · 2020
Later among the works it cites.
Bayesian deep learning and a probabilistic perspective of generalization
Wilson, A. G. and Izmailov, P · 2020
Later among the works it cites.
Self-consistent models and values
Farquhar, G., Baumli, K., Marinho, Z., Filos, A., Hessel, M., van Hasselt, H., and Silver, D · 2021
Closest in time.
Filos, A., Lyle, C., Gal, Y., Levine, S., Jaques, N., and Farquhar, G · 2021
Closest in time.
Grimm, C., Barreto, A., Farquhar, G., Silver, D., and Singh, S · 2021
Closest in time.
Muesli: Combining improvements in policy optimization
Hessel, M., Danihelka, I., Viola, F., Guez, A., Schmitt, S., Sifre, L., Weber, T., Silver, D., and van Hasselt, H · 2021
Closest in time.
Reinforcement Learning, Bit by Bit
Lu, X., Van Roy, B., Dwaracherla, V., Ibrahimi, M., Osband, I., and Wen, Z · 2021
Closest in time.
On The Effect of Auxiliary Tasks on Representation Dynamics
Lyle, C., Rowland, M., Ostrovski, G., and Dabney, W · 2021
Closest in time.
Control-Oriented Model-Based Reinforcement Learning with Implicit Differentiation
Nikishin, E., Abachi, R., Agarwal, R., and Bacon, P.-L · 2021
Closest in time.
Playvirtual: Augmenting cycle-consistent virtual trajectories for reinforcement learning
Yu, T., Lan, C., Zeng, W., Feng, M., Zhang, Z., and Chen, Z · 2021
Closest in time.