Fetching the paper…
Reading the bibliography…
The majority of methods used to compute approximations to the Hamilton-Jacobi-Isaacs partial differential equation (HJI PDE) rely on the discretization of the state space to perform dynamic programming updates.
Y. LeCun, B. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. Hubbard, and L. D. Jackel, “Backpropagation Applied to Handwritten Zip Code Recognition,” vol. 1, no. 4, pp. 541–551, 1989
1989
Earlier work this paper cites.
K. Hornik, M. Stinchcombe, and H. White, “Multilayer feedforward networks are universal approximators,” Neural Networks , vol. 2, no. 5, pp. 359–366, 1989
1989
Earlier work this paper cites.
C. J. C. H. Watkins and P. Dayan, “Technical Note: Q-Learning,” Machine Learning , vol. 8, no. 3, pp. 279–292, 1992
1992
Earlier work this paper cites.
I. E. Lagaris, A. Likas, and D. I. Fotiadis, “Artificial neural networks for solving ordinary and partial differential equations,” IEEE Transactions on Neural Networks , vol. 9, no. 5, pp. 987–1000, 1998
1998
Earlier work this paper cites.
N. Qian, “On the momentum term in gradient descent learning algorithms,” pp. 145–151, 1999
1999
Earlier work this paper cites.
J. Lygeros, C. Tomlin, and S. Sastry, “Controllers for reachability specifications for hybrid systems,” Automatica , vol. 35, no. 3, pp. 349–370, 1999
1999
Earlier work this paper cites.
I. M. Mitchell, A. M. Bayen, and C. J. Tomlin, “A time-dependent Hamilton-Jacobi formulation of reachable sets for continuous dynamic games,” IEEE Transactions on Automatic Control , vol. 50, no. 7, pp. 947–957, 2005
2005
Cited alongside, same era.
B. Djeridane and J. Lygeros, “Neural approximation of PDE solutions: An application to reachability computations,” Proceedings of the 45th IEEE Conference on Decision and Control , pp. 3034–3039, 2006
2006
Cited alongside, same era.
I. Mitchell, “A toolbox of level set methods,” Tech. Rep., 2007
2007
Cited alongside, same era.
W. B. Powell, “What you should know about approximate dynamic programming,” Naval Research Logistics , vol. 56, no. 3, pp. 239–249, 2009
2009
Cited alongside, same era.
L. Bottou, “Large-Scale Machine Learning with Stochastic Gradient Descent,” Proceedings of COMPSTAT’2010 , pp. 177–186, 2010
——, “Approximate Dynamic Programming: Solving the Curses of Dimensionality: Second Edition,” pp. 1–638, 2011
2011
Later among the works it cites.
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller, “Playing Atari with Deep Reinforcement Learning,” arXiv preprint arXiv: … , pp. 1–9, 2013
2013
Later among the works it cites.
J. Schulman, S. Levine, M. Jordan, and P. Abbeel, “Trust Region Policy Optimization,” Icml-2015 , p. 16, 2015
2015
Later among the works it cites.
M. Chen, J. Fisac, S. Sastry, and C. J. Tomlin, “Safe sequential path planning of multi-vehicle systems via double-obstacle Hamilton-Jacobi-Isaacs variational inequality,” 2015 European Control Conference, ECC 2015 , pp. 3304–3309, 2015
2015
Later among the works it cites.
S. Levine, C. Finn, T. Darrell, and P. Abbeel, “End-to-End Training of Deep Visuomotor Policies,” Journal of Machine Learning Research , vol. 17, pp. 1–40, 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2010
Cited alongside, same era.
J. H. Gillula, H. Huang, M. P. Vitus, and C. J. Tomlin, “Design of guaranteed safe maneuvers using reachable sets: Autonomous quadrotor aerobatics in theory and practice,” in Proceedings - IEEE International Conference on Robotics and Automation , 2010, pp. 1649–1654
2010
Cited alongside, same era.
2016
Closest in time.
M. Chen, Q. Hu, C. Mackin, J. F. Fisac, and C. J. Tomlin, “Safe platooning of unmanned aerial vehicles via reachability,” Proceedings of the IEEE Conference on Decision and Control , vol. 2016-Febru, pp. 4695–4701, 2016
2016
Closest in time.