Fetching the paper…
Reading the bibliography…
Applying probabilistic models to reinforcement learning (RL) enables the application of powerful optimisation tools such as variational inference to RL.
Maven: Multi-agent variational exploration, 2019
Mahajan, A., Rashid, T., Samvelyan, M., and Whiteson, S · 1910
Earlier work this paper cites.
Deterministic Policy Gradient Algorithms
Silver, D., Lever, G., Heess, N., Degris, T., Wierstra, D., and Riedmiller, M · 1938
Earlier work this paper cites.
Algorithms for Reinforcement Learning
Szepesvári, C · 1939
Earlier work this paper cites.
Maximum likelihood from incomplete data via the em algorithm
Dempster, A. P., Laird, N. M., and Rubin, D. B · 1977
Earlier work this paper cites.
On the Convergence Properties of the EM Algorithm’
Wu, C. F. J · 1983
Earlier work this paper cites.
Function optimization using connectionist reinforcement learning algorithms
Williams, R. J. and Peng, J · 1991
Earlier work this paper cites.
Q-learning
Watkins, C. J. C. H. and Dayan, P · 1992
Earlier work this paper cites.
Analysis of some incremental variants of policy iteration: First steps toward understanding actor-critic learning systems, 1993
Williams, R. J., Baird, L. C., and III · 1993
Earlier work this paper cites.
Fast exact multiplication by the hessian
Pearlmutter, B. A · 1994
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
Baird, L · 1995
Earlier work this paper cites.
Constrained Optimization and Lagrange Multiplier Methods
Bertsekas, D · 1996
Earlier work this paper cites.
Using expectation-maximization for reinforcement learning
Dayan, P. and Hinton, G. E · 1997
Earlier work this paper cites.
An analysis of temporal-difference learning with function approximation
Tsitsiklis, J. N. and Van Roy, B · 1997
Earlier work this paper cites.
Sutton & Barto Book: Reinforcement Learning: An Introduction
Sutton, R. S. and Barto, A. G · 1998
Earlier work this paper cites.
Learning in Graphical Models
Jordan, M. I. (ed.) · 1999
Earlier work this paper cites.
Policy Gradient Methods for Reinforcement Learning with Function Approximation
Sutton, R. S., Mcallester, D., Singh, S., and Mansour, Y · 1999
Earlier work this paper cites.
Policy iteration for factored mdps
Koller, D. and Parr, R · 2000
Earlier work this paper cites.
Variational algorithms for approximate Bayesian inference
Beal, M. J · 2003
Earlier work this paper cites.
Reinforcement learning with factored states and actions
Sallans, B. and Hinton, G. E · 2004
Earlier work this paper cites.
Convergence theorems for generalized alternating minimization procedures
Gunawardana, A. and Byrne, W · 2005
Earlier work this paper cites.
Pattern Recognition and Machine Learning (Information Science and Statistics)
Bishop, C. M · 2006
Earlier work this paper cites.
Probabilistic inference for solving discrete and continuous state Markov Decision Processes
Toussaint, M. and Storkey, A · 2006
Cited alongside, same era.
Reinforcement learning by reward-weighted regression for operational space control
Peters, J. and Schaal, S · 2007
Cited alongside, same era.
Linearly-solvable markov decision problems
Todorov, E · 2007
Cited alongside, same era.
Generalized Functions , chapter 4, pp. 111–124
Kelly, J · 2008
Cited alongside, same era.
Maximum entropy inverse reinforcement learning
Ziebart, B. D., Maas, A., Bagnell, J. A., and Dey, A. K · 2008
Cited alongside, same era.
Convergent temporal-difference learning with arbitrary smooth function approximation
Bhatnagar, S., Precup, D., Silver, D., Sutton, R. S., Maei, H. R., and Szepesvári, C · 2009
Cited alongside, same era.
Motor Skill Learning with Trajectory Methods
Levine, S · 2014
Later among the works it cites.
Bias in natural actor-critic algorithms
Thomas, P · 2014
Later among the works it cites.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2015
Later among the works it cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., and Hassabis, D · 2015
Later among the works it cites.
Deep Reinforcement Learning with Double Q-learning
van Hasselt, H., Guez, A., and Silver, D · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Efficient sample reuse in em-based policy search
Hachiya, H., Peters, J., and Sugiyama, M · 2009
Cited alongside, same era.
A tutorial on variational Bayesian inference
Fox, C. W. and Roberts, S. J · 2010
Cited alongside, same era.
Variational Methods For Reinforcement Learning
Furmston, T. and Barber, D · 2010
Cited alongside, same era.
Approximate inference and stochastic optimal control
Rawlik, K., Toussaint, M., and Vijayakumar, S · 2010
Cited alongside, same era.
Modeling Purposeful Adaptive Behavior with the Principle of Maximum Causal Entropy
Ziebart, B. D · 2010
Cited alongside, same era.
Modeling interaction via the principle of maximum causal entropy
Ziebart, B. D., Bagnell, J. A., and Dey, A. K · 2010
Cited alongside, same era.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Later among the works it cites.
Benchmarking deep reinforcement learning for continuous control
Duan, Y., Chen, X., Houthooft, R., Schulman, J., and Abbeel, P · 2016
Later among the works it cites.
Q-Prop: Sample-Efficient Policy Gradient with An Off-Policy Critic
Gu, S., Lillicrap, T., Ghahramani, Z., Turner, R. E., and Levine, S · 2016
Later among the works it cites.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K · 2016
Later among the works it cites.
Variational Inference: A Review for Statisticians, 2017
Blei, D. M., Kucukelbir, A., and McAuliffe, J. D · 2017
Later among the works it cites.
Reinforcement learning with deep energy-based policies
Haarnoja, T., Tang, H., Abbeel, P., and Levine, S · 2017
Later among the works it cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Later among the works it cites.
Introduction to Reinforcement Learning
Sutton, R. S. and Barto, A. G · 2017
Later among the works it cites.
Maximum a posteriori policy optimisation
Abdolmaleki, A., Springenberg, J. T., Tassa, Y., Munos, R., Heess, N., and Riedmiller, M · 2018
Closest in time.
Fourier policy gradients
Fellows, M., Ciosek, K., and Whiteson, S · 2018
Closest in time.
DiCE: The infinitely differentiable Monte Carlo estimator
Foerster, J., Farquhar, G., Al-Shedivat, M., Rocktäschel, T., Xing, E., and Whiteson, S · 2018
Closest in time.
Addressing function approximation error in actor-critic methods
Fujimoto, S., van Hoof, H., and Meger, D · 2018
Closest in time.
Recall traces: Backtracking models for efficient reinforcement learning
Goyal, A., Brakel, P., Fedus, W., Lillicrap, T. P., Levine, S., Larochelle, H., and Bengio, Y · 2018
Closest in time.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Closest in time.
Reinforcement Learning and Control as Probabilistic Inference: Tutorial and Review
Levine, S · 2018
Closest in time.