Fetching the paper…
Reading the bibliography…
Policy search can in principle acquire complex strategies for control of robots and other autonomous systems.
D. Pollard, “Asymptopia: an exposition of statistical asymptotic theory,” 2000. [Online]. Available: stat.yale.edu/~pollard/Books/Asymptopia/
2000
Earlier work this paper cites.
D. Mayne, M. M. Seron, and S. V. Rakovic, “Robust model predictive control of constrained linear systems with bounded disturbances,” in Automatica , 2005
2005
Earlier work this paper cites.
E. Todorov and W. Li, “A generalized iterative LQG method for locally-optimal feedback control of constrained nonlinear stochastic systems,” in American Control Conference , 2005
2005
Earlier work this paper cites.
X. Nguyen, M. J. Wainwright, and M. I. Jordan, “Divergences, surrogate loss functions and experimental design,” in NIPS , 2005
2005
Earlier work this paper cites.
B. Williams, G. Klein, and I. Reid, “Real-time slam relocalisation,” in ICCV , 2007
2007
Earlier work this paper cites.
V. Nair and G. Hinton, “Rectified linear units improve restricted boltzmann machines,” in ICML , 2010
2010
Earlier work this paper cites.
P. Martin and E. Salaun, “The true role of accelerometer feedback in quadrotor control,” in ICRA , 2010
2010
Earlier work this paper cites.
M. P. Deisenroth, G. Neumann, and J. Peters, “A survey on policy search for robotics,” in Foundations and Trends in Robotics , 2011
2011
Earlier work this paper cites.
S. Ross, G. Gordon, and J. Bagnell, “A reduction of imitation learning and structured prediction to no-regret online learning,” in AISTATS , 2011
2011
Earlier work this paper cites.
H. He, J. Eisner, and H. Daume, “Imitation learning by coaching,” in NIPS , 2012
2012
Earlier work this paper cites.
H. Kappen, V. Gomez, and M. Opper, “Optimal control as a graphical model inference problem,” in Machine Learning , 2012
2012
Cited alongside, same era.
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and R. M., “Playing atari with deep reinforcement learning,” in Workshop on Deep Learning, NIPS , 2013
2013
Cited alongside, same era.
S. Levine and V. Koltun, “Guided policy search,” in ICML , 2013
2013
Cited alongside, same era.
S. Ross, N. Melik-Barkhudarov, K. S. Shankar, A. Wendel, D. Dey, J. A. Bagnell, and M. Hebert, “Learning monocular reactive uav control in cluttered natural environments,” in ICRA , 2013
2013
Cited alongside, same era.
S. Levine and V. Koltun, “Variational policy search via trajectory optimization,” in NIPS , 2013
2013
Cited alongside, same era.
Y. LeCun, Y. Bengio, and G. Hinton, “Deep learning,” Nature , vol. 521, no. 7553, pp. 436–444, 2015
2015
Later among the works it cites.
C. Chen, A. Seff, A. Kornhauser, and J. Xiao, “Deepdriving: learning affordance for direct perception in autonomous driving,” in ICCV , 2015
2015
Later among the works it cites.
N. Heess, G. Wayne, D. Silver, T. Lillicrap, Y. Tassa, and T. Erez, “Learning continuous control policies by stochastic value gradients,” in NIPS , 2015
2015
Later among the works it cites.
J. Schulman, S. Levine, P. Moritz, M. I. Jordan, and P. Abbeel, “Trust region policy optimization,” in ICML , 2015
2015
Later among the works it cites.
D. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in ICLR , 2015
2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Mueller and R. D’Andrea, “A model predictive controller for quadrotor state interception,” in ECC , 2013
2013
Cited alongside, same era.
S. Levine and P. Abbeel, “Learning neural network policies with guided policy search under unknown dynamics,” in NIPS , 2014
2014
Cited alongside, same era.
X. Guo, S. Singh, H. Lee, R. L. Lewis, and X. Wang, “Deep learning for real-time Atari game play using offline Monte-Carlo tree search planning,” in NIPS , 2014
2014
Cited alongside, same era.
2014
Cited alongside, same era.
A. Giusti, J. Guzzi, D. C. Ciresan, F. He, J. P. Rodriguez, F. Fontana, M. Faessler, C. Forster, J. Schmidhuber, G. Caro, D. Scaramuzza, and L. Gambardella, “A machine learning approach to visual perception of forest trails for mobile robots,” in IEEE Robotics and Automation Letters , 2016
2016
Closest in time.
S. Levine, C. Finn, T. Darrell, and P. Abbeel, “End-to-end training of deep visuomotor policies,” JMLR , 2016
2016
Closest in time.
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra, “Continuous control with deep reinforcement learning,” in ICLR , 2016
2016
Closest in time.
T. Zhang, G. Kahn, S. Levine, and P. Abbeel, “Learning deep control policies for autonomous aerial vehicles with mpc-guided policy search,” in ICRA , 2016
2016
Closest in time.