Fetching the paper…
Reading the bibliography…
Reinforcement learning algorithms discover policies that maximize reward, but do not necessarily guarantee safety during learning or execution phases.
Pnueli, A.: The temporal logic of programs. In: 18th Annual Symposium on Foundations of Computer Science, Providence, Rhode Island, USA, 31 October - 1 November 1977. pp. 46–57 (1977), https://doi.org/10.1109/SFCS.1977.32
1977
Earlier work this paper cites.
Valiant, L.G.: A theory of the learnable. Commun. ACM 27(11), 1134–1142 (1984)
1984
Earlier work this paper cites.
Emerson, E.A.: Handbook of theoretical computer science (vol. b). chap. Temporal and Modal Logic, pp. 995–1072. MIT Press, Cambridge, MA, USA (1990)
1990
Earlier work this paper cites.
Manna, Z., Pnueli, A.: Temporal verification of reactive systems - safety. Springer (1995)
1995
Earlier work this paper cites.
Clouse, J.: On integrating apprentice learning and reinforcement learning title2:. Tech. rep., Amherst, MA, USA (1997)
1997
Earlier work this paper cites.
Sutton, R.S., Barto, A.G.: Reinforcement learning: An introduction. IEEE Trans. Neural Networks 9(5), 1054–1054 (1998)
1998
Earlier work this paper cites.
Henzinger, T.A., Kopke, P.W.: Discrete-time control for rectangular hybrid automata. Theor. Comput. Sci. 221(1-2), 369–392 (1999)
1999
Earlier work this paper cites.
Kupferman, O., Vardi, M.Y.: Model checking of safety properties. Formal Methods in System Design 19(3), 291–314 (2001), https://doi.org/10.1023/A:1011254632723
2001
Earlier work this paper cites.
Thomaz, A.L., Breazeal, C.: Reinforcement learning with human teachers: Evidence of feedback and guidance with implications for learning performance. In: Proceedings of the 21st National Conference on Artificial Intelligence - Vol. 1. pp. 1000–1005. AAAI Press (2006)
2006
Cited alongside, same era.
Baier, C., Katoen, J.P.: Principles of Model Checking (Representation and Mind Series). The MIT Press (2008)
2008
Cited alongside, same era.
Thomaz, A.L., Breazeal, C.: Teachable robots: Understanding human teaching behavior to build more effective robot learners. Artif. Intell. 172(6-7), 716–737 (2008)
2008
Cited alongside, same era.
2012
Cited alongside, same era.
Pecka, M., Svoboda, T.: Safe Exploration Techniques for Reinforcement Learning – An Overview, pp. 357–375. Springer International Publishing, Cham (2014)
2014
Later among the works it cites.
Bloem, R., Könighofer, B., Könighofer, R., Wang, C.: Shield synthesis: - runtime enforcement for reactive systems. In: Tools and Algorithms for the Construction and Analysis of Systems - 21st Int. Conf, TACAS 2015, London, UK, April 11-18, 2015. pp. 533–548 (2015)
2015
Later among the works it cites.
García, J., Fernández, F.: A comprehensive survey on safe reinforcement learning. Journal of Machine Learning Research 16, 1437–1480 (2015)
2015
Later among the works it cites.
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A.A., Veness, J., Bellemare, M.G., Graves, A., Riedmiller, M., Fidjeland, A.K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., Hassabis, D.: Human-level control through deep reinforcement learning. Nature 518(7540), 529–533 (02 2015)
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2012
Cited alongside, same era.
Bellemare, M.G., Naddaf, Y., Veness, J., Bowling, M.: The arcade learning environment: An evaluation platform for general agents. Journal of Art. Intell. Research 47, 253–279 (2013)
2013
Cited alongside, same era.
Sohail, S., Somenzi, F.: Safety first: a two-stage algorithm for the synthesis of reactive systems. STTT 15(5-6), 433–454 (2013), https://doi.org/10.1007/s10009-012-0224-3
2013
Cited alongside, same era.
Vidal, P.Q., Rodríguez, R.I., González, M.R., Regueiro, C.V.: Learning on real robots from experience and simple user feedback. Journal of Physical Agents 7(1) (2013)
2013
Cited alongside, same era.
2015
Later among the works it cites.
Wen, M., Ehlers, R., Topcu, U.: Correct-by-synthesis reinforcement learning with temporal logic constraints. In: IROS 2015, Germany, Sep. 28 - Oct. 2, 2015. pp. 4983–4990 (2015)
2015
Later among the works it cites.
DeCastro, J.A., Kress-Gazit, H.: Nonlinear controller synthesis and automatic workspace partitioning for reactive high-level behaviors. In: Proceedings of the 19th International Conference on Hybrid Systems: Computation and Control, HSCC 2016. pp. 225–234 (2016)
2016
Later among the works it cites.
Fu, J., Topcu, U.: Synthesis of shared autonomy policies with temporal logic specifications. IEEE Trans. Automation Science and Engineering 13(1), 7–17 (2016)
2016
Later among the works it cites.
Junges, S., Jansen, N., Dehnert, C., Topcu, U., Katoen, J.: Safety-constrained reinforcement learning for mdps. In: Tools and Algorithms for the Construction and Analysis of Systems - 22nd Int. Conference, TACAS 2016, The Netherlands, April 2-8, 2016. pp. 130–146 (2016)
2016
Later among the works it cites.