Understand
Embed-to-control (E2C) is a model for solving high-dimensional optimal control problems by combining variational auto-encoders with locally-optimal controllers.
- However, the E2C model suffers from two major drawbacks: 1) its objective function does not correspond to the likelihood of the data sequence and 2) the variational encoder used for embedding typically has large variational approximation error, especially when there is noise in the system dynamics.
- In this paper, we present a new model for learning robust locally-linear controllable embedding (RCE).
- Our model directly estimates the predictive conditional density of the future observation given the current one, while introducing the bottleneck between the current and future observations.
Built on
Differential Dynamic Programming
D. Jacobson and D. Mayne · 1970
Earlier work this paper cites.
An approach to fuzzy control of nonlinear systems; stability and design issues
H. Wang, K. Tanaka, and M. Griffin · 1996
Earlier work this paper cites.
Introduction to Reinforcement Learning
R. Sutton and A. Barto · 1998
Earlier work this paper cites.
Non-parametric representations of policies and value functions: A trajectory-based approach
C. Atkeson and J. Murimoto · 2002
Earlier work this paper cites.
Iterative linear quadratic regulator design for nonlinear biological movement systems
W. Li and E. Todorov · 2004
Earlier work this paper cites.
A generalized iterative LQG method for locally-optimal feedback control of constrained non-linear stochastic systems
E. Todorov and W. Li · 2005
Earlier work this paper cites.
Similar
Receding horizon differential dynamic programming
Y. Tassa, T. Erez, and W. Smart · 2008
Cited alongside, same era.
Deep auto-encoder neural networks in reinforcement learning
S. Lange and M. Riedmiller · 2010
Cited alongside, same era.
Variational policy search via trajectory optimization
S. Levine and V. Koltun · 2013
Cited alongside, same era.
Auto-encoding variational Bayes
D. Kingma and M. Welling · 2014
Cited alongside, same era.
Probabilistic differential dynamic programming
Y. Pan and E. Theodorou · 2014
Cited alongside, same era.
Stochastic backpropagation and approximate inference in deep generative models
D. Rezende, S. Mohamed, and D. Wierstra · 2014
Cited alongside, same era.
Then
Autonomous learning of state representations for control: An emerging field aims to autonomously learn state representations for reinforcement learning agents from their real-world sensor observations
W. Böhmer, J. Springenberg, J. Boedecker, M. Riedmiller, and K. Obermayer · 2015
Later among the works it cites.
From pixels to torques: Policy learning with deep dynamical models
N. Wahlström, T. Schön, and M. Desienroth · 2015
Later among the works it cites.
Embed to control: A locally linear latent dynamics model for control from raw images
M. Watter, J. Springenberg, J. Boedecker, and M. Riedmiller · 2015
Later among the works it cites.
Deep variational bayes filters: Unsupervised learning of state space models from raw data
M. Karl, M. Soelch, J. Bayer, and P. van der Smagt · 2017
Closest in time.
Bottleneck conditional density estimation
R. Shu, H. Bui, and M. Ghavamzadeh · 2017
Closest in time.
Beyond the bibliography
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…