Fetching the paper…
Reading the bibliography…
In this paper, we propose Q-learning algorithms for continuous-time deterministic optimal control problems with Lipschitz continuous controls.
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu, “Asynchronous methods for deep reinforcement learning,” in International Conference on Machine Learning , 2016, pp. 1928–1937
1937
Earlier work this paper cites.
H. Robbins and S. Monro, “A stochastic approximation method,” Annals of Mathematical Statistics , vol. 22, pp. 400–407, 1951
1951
Earlier work this paper cites.
M. Crandall and P.-L. Lions, “Viscosity solutions of Hamilton–Jacobi equations,” Transactions of the American Mathematical Society , vol. 277, pp. 1–42, 1983
1983
Earlier work this paper cites.
C. J. Watkins and P. Dayan, “Q-learning,” Machine Learning , vol. 8, pp. 279–292, 1992
1992
Earlier work this paper cites.
L. C. Baird, “Reinforcement learning in continuous time: advantage updating,” in IEEE International Conference on Neural Networks , 1994, pp. 2448–2453
1994
Earlier work this paper cites.
G. J. Gordon, “Stable function approximation in dynamic programming,” in International Conference on Machine Learning , 1995, pp. 261–268
1995
Earlier work this paper cites.
S. J. Bradtke and M. O. Duff, “Reinforcement learning methods for continuous-time Markov decision problems,” in Advances in Neural Information Processing Systems , 1995, pp. 393–400
1995
Earlier work this paper cites.
D. P. Bertsekas and J. N. Tsitsiklis, Neuro-Dynamic Programming . Belmont, MA: Athena Scientific, 1996
1996
Earlier work this paper cites.
P. Dayan and S. P. Singh, “Improving policies without measuring merits,” in Advances in Neural Information Processing Systems , 1996, pp. 1059–1065
1996
Earlier work this paper cites.
M. Bardi and I. Capuzzo-Dolcetta, Optimal Control and Viscosity Solutions of Hamilton–Jacobi–Bellman Equations . Boston, MA: Birkhäuser, 1997
1997
Earlier work this paper cites.
R. S. Sutton and A. G. Barto, Reinforcement Learning: An Introduction . Cambridge, MA: MIT Press, 1998
1998
Earlier work this paper cites.
L. Ljung, System Identification: Theory for the User , 2nd ed. Pearson, 1998
1998
Earlier work this paper cites.
R. Munos and A. W. Moore, “Barycentric interpolators for continuous space and time reinforcement learning,” in Advances in Neural Information Processing Systems , 1999, pp. 1024–1030
1999
Earlier work this paper cites.
K. Doya, “Reinforcement learning in continuous time and space,” Neural Computation , vol. 12, pp. 219–245, 2000
2000
Earlier work this paper cites.
——, “A study of reinforcement learning in the continuous case by the means of viscosity solutions,” Machine Learning , vol. 40, pp. 265–299, 2000
2000
Earlier work this paper cites.
H. Kushner and G. G. Yin, Stochastic Approximation and Recursive Algorithms and Applications . New York: Springer Science & Business Media, 2003
2003
Earlier work this paper cites.
N. Kohl and P. Stone, “Policy gradient reinforcement learning for fast quadrupedal locomotion,” in IEEE International Conference on Robotics and Automation , 2004, pp. 2619–2624
2004
Earlier work this paper cites.
M. Abu-Khalaf and F. L. Lewis, “Nearly optimal control laws for nonlinear systems with saturating actuators using a neural network HJB approach,” Automatica , vol. 41, pp. 779–791, 2005
2005
Earlier work this paper cites.
R. Munos, “Policy gradient in continuous time,” Journal of Machine Learning Research , vol. 7, pp. 771–791, 2006
2006
Earlier work this paper cites.
Y. Tassa and T. Erez, “Least squares solutions of the HJB equation with neural network value-function approximators,” IEEE Transactions on Neural Networks , vol. 18, pp. 1031–1041, 2007
2007
Cited alongside, same era.
P. Mehta and S. Meyn, “Q-learning and pontryagin’s minimum principle,” in IEEE Conference on Decision and Control , 2009, pp. 3598–3605
2009
Cited alongside, same era.
C. Szepesvari, Algorithms for Reinforcement Learning . San Rafael, CA: Morgan and Claypool Publishers, 2010
2010
Cited alongside, same era.
K. G. Vamvoudakis and F. Lewis, “Online actor-critic algorithm to solve the continuous-time infinite horizon optimal control problem,” Automatica , vol. 46, pp. 878–888, 2010
2010
Cited alongside, same era.
E. Theodorou, J. Buchli, and S. Schaal, “A generalized path integral control approach to reinforcement learning,” Journal of Machine Learning Research , vol. 11, pp. 3137–3181, 2010
T. Bian and Z.-P. Jiang, “Value iteration and adaptive dynamic programming for data-driven adaptive optimal control design,” Automatica , vol. 71, pp. 348–360, 2016
2016
Later among the works it cites.
2017
Later among the works it cites.
K. G. Vamvoudakis, “Q-learning for continuous-time linear systems: A model-free infinite horizon optimal control approach,” Systems & Control Letters , vol. 100, pp. 14–20, 2017
2017
Later among the works it cites.
Y. Yang, D. Wunsch, and Y. Yin, “Hamiltonian-driven adaptive dynamic programming for continuous nonlinear dynamical systems,” IEEE Transactions on Neural Networks and Learning Systems , vol. 28, pp. 1929–1940, 2017
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2010
Cited alongside, same era.
2010
Cited alongside, same era.
E. Todorov, T. Erez, and Y. Tassa, “MuJoCo: A physics engine for model-based control,” in IEEE/RSJ International Conference on Intelligent Robots and Systems , 2012, pp. 5026–5033
2012
Cited alongside, same era.
S. Bhasin, R. Kamalapurkar, M. Johnson, K. G. Vamvoudakis, F. L. Lewis, and W. E. Dixon, “A novel actor-critic-identifier architecture for approximate optimal control of uncertain nonlinear systems,” Automatica , vol. 49, pp. 82–92, 2013
2013
Cited alongside, same era.
S. Levine and V. Koltun, “Guided policy search,” in International Conference on Machine Learning , 2013, pp. 1–9
2013
Cited alongside, same era.
H. Modares and F. L. Lewis, “Optimal tracking control of nonlinear partially-unknown constrained-input systems using integral reinforcement learning,” Automatica , vol. 50, pp. 1780–1792, 2014
2014
Cited alongside, same era.
D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. Riedmiller, “Deterministic policy gradient algorithms,” in International Conference on Machine Learning , 2014, pp. 387–395
2014
Cited alongside, same era.
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, and S. Petersen, “Human-level control through deep reinforcement learning,” Nature , vol. 518, pp. 529–533, 2015
2015
Cited alongside, same era.
K. Rajagopal, S. N. Balakrishnan, and J. R. Busemeyer, “Neural network-based solutions for stochastic optimal control using path integrals,” IEEE Transactions on Neural Networks and Learning Systems , vol. 28, pp. 534–545, 2017
2017
Later among the works it cites.
2018
Later among the works it cites.
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine, “Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor,” in International Conference on Machine Learning , 2018, pp. 1861–1870
2018
Later among the works it cites.
S. Fujimoto, H. Hoof, and D. Meger, “Addressing function approximation error in actor-critic methods,” in International Conference on Machine Learning , 2018, pp. 1587–1596
2018
Later among the works it cites.
M. Ohnishi, M. Yukawa, M. Johansson, and M. Sugiyama, “Continuous-time value function approximation in reproducing kernel Hilbert spaces,” in Advances in Neural Information Processing Systems , 2018, pp. 2813–2824
2018
Later among the works it cites.
2018
Later among the works it cites.
D. Quillen, E. Jang, O. Nachum, C. Finn, J. Ibarz, and S. Levine, “Deep reinforcement learning for vision-based robotic grasping: A simulated comparative evaluation of off-policy methods,” in IEEE International Conference on Robotics and Automation , 2018, pp. 6284–6291
2018
Later among the works it cites.
C. Tessler, G. Tennenholtz, and S. Mannor, “Distributional policy optimization: An alternative approach for continuous control,” in Advances in Neural Information Processing Systems , 2019, pp. 1350–1360
2019
Later among the works it cites.
G. P. Kontoudis and K. G. Vamvoudakis, “Kinodynamic motion planning with continuous-time Q-learning: An online, model-free, and safe navigation framework,” IEEE Transactions on Neural Networks and Learning Systems , vol. 30, pp. 3803–3817, 2019
2019
Later among the works it cites.
C. Tallec, L. Blier, and Y. Ollivier, “Making deep Q-learning methods robust to time discretization,” in International Conference on Machine Learning , 2019, pp. 6096–6104
2019
Later among the works it cites.
2020
Closest in time.
2020
Closest in time.
M. Lutter, B. Belousov, K. Listmann, D. Clever, and J. Peters, “HJB optimal feedback control with deep differential value functions and action constraints,” in Conference on Robot Learning , 2020, pp. 640–650
2020
Closest in time.
J. Kim and I. Yang, “Hamilton–Jacobi–Bellman for Q-learning in continuous time,” in Learning for Dynamics and Control (L4DC) , 2020, pp. 739–748
2020
Closest in time.
H. Van Hasselt, A. Guez, and D. Silver, “Deep reinforcement learning with double Q-learning,” in Thirtieth AAAI Conference on Artificial Intelligence , 2016, pp. 2094–2100
2094
Closest in time.