Fetching the paper…
Reading the bibliography…
Randomized value functions offer a promising approach towards the challenge of efficient exploration in complex environments with high dimensional state and action spaces.
On the theory of the brownian motion
Uhlenbeck, G. E. and Ornstein, L. S. (1930) · 1930
Earlier work this paper cites.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
Thompson, W. R. (1933) · 1933
Earlier work this paper cites.
Dropout: A simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R. (2014) · 1958
Earlier work this paper cites.
Efficient memory-based learning for robot control
Moore, A. W. (1990) · 1990
Earlier work this paper cites.
Keeping neural networks simple by minimising the description length of weights
Hinton, G. and van Camp, D. (1993) · 1993
Earlier work this paper cites.
Introduction to reinforcement learning
Sutton, R. S. and Barto, A. G. (1998) · 1998
Earlier work this paper cites.
Approximate solutions to markov decision processes
Gordon, G. J. (1999) · 1999
Earlier work this paper cites.
A bayesian framework for reinforcement learning
Strens, M. (2000) · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S. P., and Mansour, Y. (2000) · 2000
Earlier work this paper cites.
R-max-a general polynomial time algorithm for near-optimal reinforcement learning
Brafman, R. I. and Tennenholtz, M. (2002) · 2002
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Kearns, M. and Singh, S. (2002) · 2002
Earlier work this paper cites.
Neural fitted q iteration–first experiences with a data efficient neural reinforcement learning method
Riedmiller, M. (2005) · 2005
Earlier work this paper cites.
A theoretical analysis of model-based interval estimation
Strehl, A. L. and Littman, M. L. (2005) · 2005
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Jaksch, T., Ortner, R., and Auer, P. (2010) · 2010
Earlier work this paper cites.
Parameter-exploring policy gradients
Sehnke, F., Osendorfer, C., Rückstieß, T., Graves, A., Peters, J., and Schmidhuber, J. (2010) · 2010
Cited alongside, same era.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M. (2013) · 2013
Cited alongside, same era.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M. (2013) · 2013
Cited alongside, same era.
(more) efficient reinforcement learning via posterior sampling
Osband, I., Russo, D., and Van Roy, B. (2013) · 2013
Cited alongside, same era.
Stochastic backpropagation and approximate inference in deep generative models
Rezende, D. J., Mohamed, S., and Wierstra, D. (2014) · 2014
Cited alongside, same era.
Weight uncertainty in neural network
Density estimation using real nvp
Dinh, L., Sohl-Dickstein, J., and Bengio, S. (2016) · 2016
Later among the works it cites.
Vime: Variational information maximizing exploration
Houthooft, R., Chen, X., Duan, Y., Schulman, J., De Turck, F., and Abbeel, P. (2016) · 2016
Later among the works it cites.
Improved variational inference with inverse autoregressive flow
Kingma, D. P., Salimans, T., Jozefowicz, R., Chen, X., Sutskever, I., and Welling, M. (2016) · 2016
Later among the works it cites.
Efficient exploration for dialogue policy learning with bbq networks & replay buffer spiking
Lipton, Z. C., Gao, J., Li, L., Li, X., Ahmed, F., and Deng, L. (2016) · 2016
Later among the works it cites.
Hierarchical variational models
Ranganath, R., Tran, D., and Blei, D. (2016) · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Blundell, C., Cornebise, J., Kavukcuoglu, K., and Wierstra, D. (2015) · 2015
Cited alongside, same era.
Variational dropout and the local reparameterization trick
Kingma, D. P., Salimans, T., and Welling, M. (2015) · 2015
Cited alongside, same era.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D. (2015) · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al. (2015) · 2015
Cited alongside, same era.
Variational inference with normalizing flows
Rezende, D. and Mohamed, S. (2015) · 2015
Cited alongside, same era.
Markov chain monte carlo and variational inference: Bridging the gap
Salimans, T., Kingma, D., and Welling, M. (2015) · 2015
Cited alongside, same era.
Incentivizing exploration in reinforcement learning with deep predictive models
Stadie, B. C., Levine, S., and Abbeel, P. (2015) · 2015
Cited alongside, same era.
Deep reinforcement learning with double q-learning
van Hasselt, H., Guez, A., and Silver, D. (2016) · 2016
Later among the works it cites.
Noisy networks for exploration
Fortunato, M., Azar, M. G., Piot, B., Menick, J., Osband, I., Graves, A., Mnih, V., Munos, R., Hassabis, D., Pietquin, O., et al. (2017) · 2017
Later among the works it cites.
Multiplicative normalizing flows for variational bayesian neural networks
Louizos, C. and Welling, M. (2017) · 2017
Later among the works it cites.
Deep exploration via randomized value functions
Osband, I., Russo, D., Wen, Z., and Van Roy, B. (2017) · 2017
Later among the works it cites.
Parameter space noise for exploration
Plappert, M., Houthooft, R., Dhariwal, P., Sidor, S., Chen, R. Y., Chen, X., Asfour, T., Abbeel, P., and Andrychowicz, M. (2017) · 2017
Later among the works it cites.
Playing hard exploration games by watching youtube
Aytar, Y., Pfaff, T., Budden, D., Paine, T. L., Wang, Z., and de Freitas, N. (2018) · 2018
Closest in time.
Efficient exploration through bayesian deep q-networks
Azizzadenesheli, K., Brunskill, E., and Anandkumar, A. (2018) · 2018
Closest in time.
Hierarchical imitation and reinforcement learning
Le, H. M., Jiang, N., Agarwal, A., Dudík, M., Yue, Y., and Daumé III, H. (2018) · 2018
Closest in time.
Randomized prior functions for deep reinforcement learning
Osband, I., Aslanides, J., and Cassirer, A. (2018) · 2018
Closest in time.