Fetching the paper…
Reading the bibliography…
In this paper, we introduce Path Integral Networks (PI-Net), a recurrent network representation of the Path Integral optimal control algorithm.
An on-line algorithm for dynamic reinforcement learning and planning in reactive environments
J. Schmidhuber · 1990
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. J. Williams · 1992
Earlier work this paper cites.
An approach to fuzzy control of nonlinear systems: Stability and design issues
H. O. Wang, K. Tanaka, and M. F. Griffin · 1996
Earlier work this paper cites.
Vlsi implementation of locally connected neural network for solving partial differential equations
R. Yentis and M. Zaghloul · 1996
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
A. Y. Ng, D. Harada, and S. Russell · 1999
Earlier work this paper cites.
Movement imitation with nonlinear dynamical systems in humanoid robots
A. J. Ijspeert, J. Nakanishi, and S. Schaal · 2002
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
P. Abbeel and A. Y. Ng · 2004
Earlier work this paper cites.
Linear theory for control of nonlinear stochastic systems
H. J. Kappen · 2005
Earlier work this paper cites.
A generalized iterative LQG method for locally-optimal feedback control of constrained nonlinear stochastic systems
E. Todorov and W. Li · 2005
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
B. D. Ziebart, A. L. Maas, J. A. Bagnell, and A. K. Dey · 2008
Earlier work this paper cites.
A generalized path integral control approach to reinforcement learning
E. Theodorou, J. Buchli, and S. Schaal · 2010
Cited alongside, same era.
Relative entropy inverse reinforcement learning
A. Boularias, J. Kober, and J. Peters · 2011
Cited alongside, same era.
PILCO: A model-based and data-efficient approach to policy search
M. Deisenroth and C. E. Rasmussen · 2011
Cited alongside, same era.
A reduction of imitation learning and structured prediction to no-regret online learning
S. Ross, G. J. Gordon, and D. Bagnell · 2011
Cited alongside, same era.
Learning objective functions for manipulation
M. Kalakrishnan, P. Pastor, L. Righetti, and S. Schaal · 2013
Cited alongside, same era.
DeepDriving: Learning affordance for direct perception in autonomous driving
C. Chen, A. Seff, A. Kornhauser, and J. Xiao · 2015
Trust region policy optimization
J. Schulman, S. Levine, P. Abbeel, M. I. Jordan, and P. Moritz · 2015
Later among the works it cites.
End to end learning for self-driving cars
M. Bojarski, D. Del Testa, D. Dworakowski, B. Firner, B. Flepp, et al · 2016
Later among the works it cites.
Path integral guided policy search
Y. Chebotar, M. Kalakrishnan, A. Yahya, A. Li, S. Schaal, and S. Levine · 2016
Later among the works it cites.
Guided cost learning: Deep inverse optimal control via policy optimization
C. Finn, S. Levine, and P. Abbeel · 2016
Later among the works it cites.
End-to-end training of deep visuomotor policies
S. Levine, C. Finn, T. Darrell, and P. Abbeel · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
TensorFlow: Large-scale machine learning on heterogeneous distributed systems
J. Dean and R. Monga · 2015
Cited alongside, same era.
Continuous control with deep reinforcement learning
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, et al · 2015
Cited alongside, same era.
Value iteration networks
A. Tamar, S. Levine, P. Abbeel, Y. Wu, and G. Thomas
Cited in the paper.
Learning from the hindsight plan–episodic MPC improvement
A. Tamar, G. Thomas, T. Zhang, S. Levine, and P. Abbeel
Cited in the paper.
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. P. Lillicrap, et al · 2016
Later among the works it cites.
Aggressive driving with model predictive path integral control
G. Williams, P. Drews, B. Goldfain, J. M. Rehg, and E. A. Theodorou · 2016
Later among the works it cites.
PLATO: Policy learning using adaptive trajectory optimization
G. Kahn, T. Zhang, S. Levine, and P. Abbeel · 2017
Closest in time.
Model predictive path integral control: From theory to parallel computation
G. Williams, A. Aldrich, and E. A. Theodorou · 2017
Closest in time.