Fetching the paper…
Reading the bibliography…
Techniques based on Reinforcement Learning (RL) are increasingly being used to design control policies for robotic systems.
1903
Earlier work this paper cites.
1904
Earlier work this paper cites.
1909
Earlier work this paper cites.
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu, “Asynchronous Methods for Deep Reinforcement Learning,” in International Conference on Machine Learning , June 2016, pp. 1928–1937. [Online]. Available: http://proceedings.mlr.press/v48/mniha16.html
1937
Earlier work this paper cites.
A. G. Barto, R. S. Sutton, and C. W. Anderson, “Neuronlike adaptive elements that can solve difficult learning control problems,” IEEE Transactions on Systems, Man, and Cybernetics , vol. SMC-13, no. 5, pp. 834–846, Sept. 1983
1983
Earlier work this paper cites.
A. W. Moore, “Efficient Memory-based Learning for Robot Control,” Tech. Rep., 1990
1990
Earlier work this paper cites.
S. B. Thrun, “Efficient Exploration In Reinforcement Learning,” Tech. Rep., 1992
1992
Earlier work this paper cites.
C. J. C. H. Watkins and P. Dayan, “Q-learning,” Machine Learning , vol. 8, no. 3, pp. 279–292, May 1992
1992
Earlier work this paper cites.
M. V. Kothare, V. Balakrishnan, and M. Morari, “Robust constrained model predictive control using linear matrix inequalities,” Automatica , vol. 32, no. 10, pp. 1361–1379, Oct. 1996
1996
Earlier work this paper cites.
C. Atkeson and J. Santamaria, “A comparison of direct and model-based reinforcement learning,” in Proceedings of International Conference on Robotics and Automation , vol. 4, Apr. 1997, pp. 3557–3564 vol.4
1997
Earlier work this paper cites.
D. Q. Mayne, J. B. Rawlings, C. V. Rao, and P. O. M. Scokaert, “Constrained model predictive control: Stability and optimality,” Automatica , vol. 36, no. 6, pp. 789–814, June 2000
2000
Earlier work this paper cites.
J. Kocijan, R. Murray-Smith, C. Rasmussen, and A. Girard, “Gaussian process model based predictive control,” in Proceedings of the 2004 American Control Conference , vol. 3, June 2004, pp. 2214–2219 vol.3
2004
Earlier work this paper cites.
A. Richards and J. How, “Mixed-integer programming for control,” in Proceedings of the 2005, American Control Conference, 2005. , June 2005, pp. 2676–2683 vol. 4
2005
Earlier work this paper cites.
N. Hansen, The CMA Evolution Strategy: A Comparing Review , 2006
2006
Earlier work this paper cites.
C. E. Rasmussen and C. K. I. Williams, Gaussian Processes for Machine Learning , ser. Adaptive Computation and Machine Learning. Cambridge, Mass: MIT Press, 2006
2006
Earlier work this paper cites.
D. Nguyen-tuong, J. R. Peters, and M. Seeger, “Local Gaussian Process Regression for Real Time Online Model Learning,” in Advances in Neural Information Processing Systems 21 , D. Koller, D. Schuurmans, Y. Bengio, and L. Bottou, Eds. Curran Associates, Inc., 2009, pp. 1193–1200. [Online]. Available: http://papers.nips.cc/paper/3403-local-gaussian-process-regression-for-real-time-online-model-learning.pdf
2009
Earlier work this paper cites.
A. Donzé and O. Maler, “Robust Satisfaction of Temporal Logic over Real-Valued Signals,” in Formal Modeling and Analysis of Timed Systems , ser. Lecture Notes in Computer Science, K. Chatterjee and T. A. Henzinger, Eds. Berlin, Heidelberg: Springer, 2010, pp. 92–106
2010
Earlier work this paper cites.
M. Lahijanian, J. Wasniewski, S. B. Andersson, and C. Belta, “Motion planning and control from temporal logic specifications with probabilistic satisfaction guarantees,” in 2010 IEEE International Conference on Robotics and Automation , May 2010, pp. 3227–3232
2010
Earlier work this paper cites.
M. P. Deisenroth and C. E. Rasmussen, “PILCO: A Model-Based and Data-Efficient Approach to Policy Search,” in In Proceedings of the International Conference on Machine Learning , 2011
2011
Earlier work this paper cites.
F. Allgöwer and A. Zheng, Nonlinear model predictive control . Birkhäuser, 2012, vol. 26
2012
Earlier work this paper cites.
E. F. Camacho and C. B. Alba, Model Predictive Control . Springer Science & Business Media, Jan. 2013
2013
Cited alongside, same era.
Z. I. Botev, D. P. Kroese, R. Y. Rubinstein, and P. L’Ecuyer, “The Cross-Entropy Method for Optimization,” in Handbook of Statistics , ser. Handbook of Statistics, C. R. Rao and V. Govindaraju, Eds. Elsevier, Jan. 2013, vol. 31, pp. 35–59
2013
Cited alongside, same era.
J. Fu and U. Topcu, “Probably Approximately Correct MDP Learning and Control With Temporal Logic Constraints,” in Robotics: Science and Systems X , vol. 10, July 2014. [Online]. Available: http://www.roboticsproceedings.org/rss10/p39.html
2014
Cited alongside, same era.
D. Wierstra, T. Schaul, T. Glasmachers, Y. Sun, J. Peters, and J. Schmidhuber, “Natural evolution strategies,” The Journal of Machine Learning Research , vol. 15, no. 1, pp. 949–980, Jan. 2014
2014
Cited alongside, same era.
2017
Later among the works it cites.
F. Berkenkamp, M. Turchetta, A. Schoellig, and A. Krause, “Safe Model-based Reinforcement Learning with Stability Guarantees,” in Advances in Neural Information Processing Systems 30 , I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, Eds. Curran Associates, Inc., 2017, pp. 908–918. [Online]. Available: http://papers.nips.cc/paper/6692-safe-model-based-reinforcement-learning-with-stability-guarantees.pdf
2017
Later among the works it cites.
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. Riedmiller, “Deterministic Policy Gradient Algorithms,” in ICML , June 2014. [Online]. Available: https://hal.inria.fr/hal-00938992
2014
Cited alongside, same era.
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis, “Human-level control through deep reinforcement learning,” Nature , vol. 518, no. 7540, pp. 529–533, Feb. 2015
2015
Cited alongside, same era.
S. S. Farahani, V. Raman, and R. M. Murray, “Robust Model Predictive Control for Signal Temporal Logic Synthesis,” IFAC-PapersOnLine , vol. 48, no. 27, pp. 323–328, Jan. 2015
2015
Cited alongside, same era.
M. P. Deisenroth, D. Fox, and C. E. Rasmussen, “Gaussian Processes for Data-Efficient Learning in Robotics and Control,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 37, no. 2, pp. 408–423, Feb. 2015
2015
Cited alongside, same era.
A. Donzé, X. Jin, J. V. Deshmukh, and S. A. Seshia, “Automotive systems requirement mining using breach,” in 2015 American Control Conference (ACC) , July 2015, pp. 4097–4097
2015
Cited alongside, same era.
X. Jin, A. Donzé, J. V. Deshmukh, and S. A. Seshia, “Mining Requirements From Closed-Loop Control Models,” IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems , vol. 34, no. 11, pp. 1704–1717, Nov. 2015
2015
Cited alongside, same era.
2016
Cited alongside, same era.
D. Aksaray, A. Jones, Z. Kong, M. Schwager, and C. Belta, “Q-Learning for robust satisfaction of signal temporal logic specifications,” in 2016 IEEE 55th Conference on Decision and Control (CDC) , Dec. 2016, pp. 6565–6570
2016
Cited alongside, same era.
X. Li, C.-I. Vasile, and C. Belta, “Reinforcement learning with temporal logic rewards,” in 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , Sept. 2017, pp. 3834–3839
2017
Later among the works it cites.
2017
Later among the works it cites.
M. Grześ, “Reward Shaping in Episodic Reinforcement Learning,” in Proceedings of the 16th Conference on Autonomous Agents and MultiAgent Systems , ser. AAMAS ’17. Richland, SC: International Foundation for Autonomous Agents and Multiagent Systems, May 2017, pp. 565–573
2017
Later among the works it cites.
R. S. Sutton and A. G. Barto, Reinforcement Learning: An Introduction , 2nd ed., ser. Adaptive Computation and Machine Learning Series. Cambridge, Massachusetts: The MIT Press, 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
K. Chua, R. Calandra, R. McAllister, and S. Levine, “Deep Reinforcement Learning in a Handful of Trials using Probabilistic Dynamics Models,” in Advances in Neural Information Processing Systems 31 , S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett, Eds. Curran Associates, Inc., 2018, pp. 4754–4765
2018
Later among the works it cites.
2018
Later among the works it cites.
E. Leurent, “An environment for autonomous driving decision-making,” GitHub repository , 2018. [Online]. Available: https://github.com/eleurent/highway-env
2018
Later among the works it cites.
S. Depeweg, J.-M. Hernandez-Lobato, F. Doshi-Velez, and S. Udluft, “Decomposition of Uncertainty in Bayesian Deep Learning for Efficient and Risk-sensitive Learning,” in International Conference on Machine Learning , July 2018, pp. 1184–1193. [Online]. Available: http://proceedings.mlr.press/v80/depeweg18a.html
2018
Later among the works it cites.
E. Bartocci, J. Deshmukh, A. Donzé, G. Fainekos, O. Maler, D. Ničković, and S. Sankaranarayanan, “Specification-Based Monitoring of Cyber-Physical Systems: A Survey on Theory, Tools and Applications,” in Lectures on Runtime Verification: Introductory and Advanced Topics , ser. Lecture Notes in Computer Science, E. Bartocci and Y. Falcone, Eds. Cham: Springer International Publishing, 2018, pp. 135–175
2018
Later among the works it cites.
Y. V. Pant, H. Abbas, R. Quaye, and R. Mangharam, “Fly-by-Logic: Control of Multi-Drone Fleets with Temporal Logic Objectives,” The 9th ACM/IEEE International Conference on Cyber-Physical Systems (ICCPS), 2018 , Mar. 2018. [Online]. Available: https://repository.upenn.edu/mlab_papers/107
2018
Later among the works it cites.
2019
Later among the works it cites.
A. Balakrishnan and J. V. Deshmukh, “Structured Reward Shaping using Signal Temporal Logic specifications,” in 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , Nov. 2019, pp. 3481–3486
2019
Later among the works it cites.
L. Lindemann and D. V. Dimarogonas, “Control Barrier Functions for Signal Temporal Logic Tasks,” IEEE Control Systems Letters , vol. 3, no. 1, pp. 96–101, Jan. 2019
2019
Later among the works it cites.
M. Ohnishi, L. Wang, G. Notomista, and M. Egerstedt, “Barrier-Certified Adaptive Reinforcement Learning with Applications to Brushbot Navigation,” IEEE Transactions on Robotics , vol. 35, no. 5, pp. 1186–1205, Oct. 2019
2019
Later among the works it cites.
S. Moarref and H. Kress-Gazit, “Automated synthesis of decentralized controllers for robot swarms from high-level temporal logic specifications,” Autonomous Robots , vol. 44, no. 3, pp. 585–600, Mar. 2020
2020
Closest in time.