2017

Robust Locally-Linear Controllable Embedding

Banijamali, Ershad, Shu, Rui, Ghavamzadeh, Mohammad et al.

Understand

Embed-to-control (E2C) is a model for solving high-dimensional optimal control problems by combining variational auto-encoders with locally-optimal controllers.

  • However, the E2C model suffers from two major drawbacks: 1) its objective function does not correspond to the likelihood of the data sequence and 2) the variational encoder used for embedding typically has large variational approximation error, especially when there is noise in the system dynamics.
  • In this paper, we present a new model for learning robust locally-linear controllable embedding (RCE).
  • Our model directly estimates the predictive conditional density of the future observation given the current one, while introducing the bottleneck between the current and future observations.

Built on

  • Differential Dynamic Programming

    D. Jacobson and D. Mayne · 1970

    Earlier work this paper cites.

  • An approach to fuzzy control of nonlinear systems; stability and design issues

    H. Wang, K. Tanaka, and M. Griffin · 1996

    Earlier work this paper cites.

  • Introduction to Reinforcement Learning

    R. Sutton and A. Barto · 1998

    Earlier work this paper cites.

  • Non-parametric representations of policies and value functions: A trajectory-based approach

    C. Atkeson and J. Murimoto · 2002

    Earlier work this paper cites.

  • Iterative linear quadratic regulator design for nonlinear biological movement systems

    W. Li and E. Todorov · 2004

    Earlier work this paper cites.

  • A generalized iterative LQG method for locally-optimal feedback control of constrained non-linear stochastic systems

    E. Todorov and W. Li · 2005

    Earlier work this paper cites.

Similar

  • Receding horizon differential dynamic programming

    Y. Tassa, T. Erez, and W. Smart · 2008

    Cited alongside, same era.

  • Deep auto-encoder neural networks in reinforcement learning

    S. Lange and M. Riedmiller · 2010

    Cited alongside, same era.

  • Variational policy search via trajectory optimization

    S. Levine and V. Koltun · 2013

    Cited alongside, same era.

  • Auto-encoding variational Bayes

    D. Kingma and M. Welling · 2014

    Cited alongside, same era.

  • Probabilistic differential dynamic programming

    Y. Pan and E. Theodorou · 2014

    Cited alongside, same era.

  • Stochastic backpropagation and approximate inference in deep generative models

    D. Rezende, S. Mohamed, and D. Wierstra · 2014

    Cited alongside, same era.

Then

  • Autonomous learning of state representations for control: An emerging field aims to autonomously learn state representations for reinforcement learning agents from their real-world sensor observations

    W. Böhmer, J. Springenberg, J. Boedecker, M. Riedmiller, and K. Obermayer · 2015

    Later among the works it cites.

  • From pixels to torques: Policy learning with deep dynamical models

    Original

    N. Wahlström, T. Schön, and M. Desienroth · 2015

    Later among the works it cites.

  • Embed to control: A locally linear latent dynamics model for control from raw images

    M. Watter, J. Springenberg, J. Boedecker, and M. Riedmiller · 2015

    Later among the works it cites.

  • Deep variational bayes filters: Unsupervised learning of state space models from raw data

    M. Karl, M. Soelch, J. Bayer, and P. van der Smagt · 2017

    Closest in time.

  • Bottleneck conditional density estimation

    R. Shu, H. Bui, and M. Ghavamzadeh · 2017

    Closest in time.

Beyond the bibliography

alphaXiv searches the wider corpus for related work and actual follow-ups.

Open on alphaXiv

alphaXiv is searching for related work…