Fetching the paper…
Reading the bibliography…
We propose a method for learning expressive energy-based policies for continuous states and actions, which has been feasible only in tabular domains before.
On the theory of the brownian motion
Uhlenbeck, G. E. and Ornstein, L. S · 1930
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
Lai, T. L. and Robbins, H · 1985
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, Ronald J · 1992
Earlier work this paper cites.
Reinforcement learning: A survey
Kaelbling, L. P., Littman, M. L., and Moore, A. W · 1996
Earlier work this paper cites.
Reinforcement learning: An introduction , volume 1
Sutton, R. S. and Barto, A. G · 1998
Earlier work this paper cites.
A natural policy gradient
Kakade, S · 2002
Earlier work this paper cites.
Reinforcement learning with factored states and actions
Sallans, B. and Hinton, G. E · 2004
Earlier work this paper cites.
Path integrals and symmetry breaking for optimal control theory
Kappen, H. J · 2005
Earlier work this paper cites.
Linearly-solvable Markov decision problems
Todorov, E · 2007
Earlier work this paper cites.
General duality between optimal control and estimation
Todorov, E · 2008
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Ziebart, B. D., Maas, A. L., Bagnell, J. A., and Dey, A. K · 2008
Earlier work this paper cites.
Linear Bellman combination for control of character animation
Da Silva, M., Durand, F., and Popović, J · 2009
Earlier work this paper cites.
Compositionality of optimal control laws
Todorov, E · 2009
Earlier work this paper cites.
Robot trajectory optimization using approximate inference
Toussaint, M · 2009
Earlier work this paper cites.
Free-energy based reinforcement learning for vision-based navigation with high-dimensional sensory inputs
Elfwing, S., Otsuka, M., Uchibe, E., and Doya, K · 2010
Earlier work this paper cites.
Free-energy-based reinforcement learning in a partially observable environment
Otsuka, M., Yoshimoto, J., and Doya, K · 2010
Earlier work this paper cites.
Relative entropy policy search
Peters, J., Mülling, K., and Altun, Y · 2010
Cited alongside, same era.
Modeling purposeful adaptive behavior with the principle of maximum causal entropy
Ziebart, B. D · 2010
Cited alongside, same era.
Reinforcement learning in feedback control
Hafner, R. and Riedmiller, M · 2011
Cited alongside, same era.
Variational inference for policy search in changing situations
Neumann, G · 2011
Cited alongside, same era.
Hierarchical relative entropy policy search
Daniel, C., Neumann, G., and Peters, J · 2012
Cited alongside, same era.
Actor-critic reinforcement learning with energy-based policies
Heess, N., Silver, D., and Teh, Y. W · 2012
Cited alongside, same era.
Deep learning
Goodfellow, Ian, Bengio, Yoshua, and Courville, Aaron · 2016
Later among the works it cites.
Learning and transfer of modulated locomotor controllers
Heess, N., Wayne, G., Tassa, Y., Lillicrap, T., Riedmiller, M., and Silver, D · 2016
Later among the works it cites.
Reinforcement learning with unsupervised auxiliary tasks
Jaderberg, M., Mnih, V., Czarnecki, W. M., Schaul, T., Leibo, J. Z., Silver, D., and Kavukcuoglu, K · 2016
Later among the works it cites.
Deep directed generative models with energy-based probability estimation
Kim, T. and Bengio, Y · 2016
Later among the works it cites.
End-to-end training of deep visuomotor policies
Levine, S., Finn, C., Darrell, T., and Abbeel, P · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On stochastic optimal control and reinforcement learning by approximate inference
Rawlik, K., Toussaint, M., and Vijayakumar, S · 2012
Cited alongside, same era.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M · 2013
Cited alongside, same era.
Learning neural network policies with guided policy search under unknown dynamics
Levine, S. and Abbeel, P · 2014
Cited alongside, same era.
Deterministic policy gradient algorithms
Silver, D., Lever, G., Heess, N., Degris, T., Wierstra, D., and Riedmiller, M · 2014
Cited alongside, same era.
Bias in natural actor-critic algorithms
Thomas, P · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. and Ba, J · 2015
Cited alongside, same era.
Stein variational gradient descent: A general purpose bayesian inference algorithm
Liu, Q. and Wang, D · 2016
Later among the works it cites.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T. P., Harley, T., Silver, D., and Kavukcuoglu, K · 2016
Later among the works it cites.
PGQ: Combining policy gradient and Q-learning
O’Donoghue, B., Munos, R., Kavukcuoglu, K., and Mnih, V · 2016
Later among the works it cites.
Loss is its own reward: Self-supervision for reinforcement learning
Shelhamer, E., Mahmoudieh, P., Argus, M., and Darrell, T · 2016
Later among the works it cites.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., van den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., Dieleman, S., Grewe, D., Nham, J., Kalchbrenner, N., Sutskever, I., Lillicrap, T., Leach, M., Kavukcuoglu, K., Graepel, T., and Hassabis, D · 2016
Later among the works it cites.
Learning to draw samples: With application to amortized mle for generative adversarial learning
Wang, D. and Liu, Q · 2016
Later among the works it cites.
Energy-based generative adversarial network
Zhao, J., Mathieu, M., and LeCun, Y · 2016
Later among the works it cites.
Stochastic neural networks for hierarchical reinforcement learning
Florensa, C., Duan, Y., and P., Abbeel · 2017
Closest in time.
Stein variational policy gradient
Liu, Y., Ramachandran, P., Liu, Q., and Peng, J · 2017
Closest in time.
Equivalence between policy gradients and soft Q-learning
Schulman, J., Abbeel, P., and Chen, X · 2017
Closest in time.