Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) depends critically on the choice of reward functions used to capture the de- sired behavior and constraints of a robot.
“Policy invariance under reward transformations : Theory and application to reward shaping,” Sixteenth International Conference on Machine Learning , vol. 3, pp. 278–287, 1999
1999
Earlier work this paper cites.
A. Ng and S. Russell, “Algorithms for inverse reinforcement learning,” Proceedings of the Seventeenth International Conference on Machine Learning , vol. 0, pp. 663–670, 2000. [Online]. Available: http://www-cs.stanford.edu/people/ang/papers/icml00-irl.pdf
2000
Earlier work this paper cites.
T. G. Dietterich, “Hierarchical reinforcement learning with the maxq value function decomposition,” J. Artif. Intell. Res.(JAIR) , vol. 13, pp. 227–303, 2000
2000
Earlier work this paper cites.
T. Latvala, A. Biere, K. Heljanko, and T. Junttila, “Simple bounded LTL model checking,” Formal Methods in Computer-Aided Design , vol. 3312, no. LCNS, pp. 186–200, 2004. [Online]. Available: http://www.springerlink.com/index/A1JNFCB7Q9KNC1Q1.pdf
2004
Earlier work this paper cites.
A. Donzé and O. Maler, “Robust satisfaction of temporal logic over real-valued signals,” Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) , vol. 6246 LNCS, pp. 92–106, 2010
2010
Earlier work this paper cites.
M. P. Deisenroth, “A Survey on Policy Search for Robotics,” Foundations and Trends in Robotics , vol. 2, no. 1, pp. 1–142, 2011. [Online]. Available: http://www.nowpublishers.com/articles/foundations-and-trends-in-robotics/ROB-021
2011
Earlier work this paper cites.
2012
Earlier work this paper cites.
G. De Giacomo and M. Y. Vardi, “Linear temporal logic and Linear Dynamic Logic on finite traces,” IJCAI International Joint Conference on Artificial Intelligence , pp. 854–860, 2013
2013
Cited alongside, same era.
V. Gómez, H. J. Kappen, J. Peters, and G. Neumann, “Policy search for path integral control,” in Joint European Conference on Machine Learning and Knowledge Discovery in Databases . Springer, 2014, pp. 482–497
2014
Cited alongside, same era.
D. Sadigh, E. S. Kim, S. Coogan, S. S. Sastry, S. Seshia, and Others, “A learning based approach to control synthesis of markov decision processes for linear temporal logic specifications,” Decision and Control (CDC), 2014 IEEE 53rd Annual Conference on , pp. 1091–1096, 2014
2014
Cited alongside, same era.
2014
Cited alongside, same era.
2016
Closest in time.
D. Aksaray, A. Jones, Z. Kong, M. Schwager, and C. Belta, “Q -Learning for Robust Satisfaction of Signal Temporal Logic Specifications,” 2016
2016
Closest in time.
Y. Duan, X. Chen, R. Houthooft, J. Schulman, and P. Abbeel, “Benchmarking deep reinforcement learning for continuous control,” in Proceedings of the 33rd International Conference on Machine Learning (ICML) , 2016
2016
Closest in time.
2016
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2015
Cited alongside, same era.
2016
Cited alongside, same era.
2016
Closest in time.