Fetching the paper…
Reading the bibliography…
Reinforcement learning requires interaction with an environment, which is expensive for robots.
J. MacQueen, “Some methods for classification and analysis of multivariate observations,” in Proc. 5th Berkeley Symposium on Math., Stat., and Prob
1965
Earlier work this paper cites.
A. Pnueli, “The temporal logic of programs,” in Proceedings of the 18th Annual Symposium on Foundations of Computer Science
1977
Earlier work this paper cites.
S. Lloyd, “Least squares quantization in pcm,” IEEE transactions on information theory
1982
Earlier work this paper cites.
L.-J. Lin, “Self-improving reactive agents based on reinforcement learning, planning and teaching,” Machine learning
1992
Earlier work this paper cites.
C. J. Watkins and P. Dayan, “Q-learning,” Machine learning
1992
Earlier work this paper cites.
S. Thrun, “Lifelong learning algorithms.,” Learning to learn
1998
Earlier work this paper cites.
F. Bacchus and F. Kabanza, “Using temporal logics to express search control knowledge for planning,” Artificial intelligence
2000
Earlier work this paper cites.
P. Auer, “Using confidence bounds for exploitation-exploration trade-offs,” Journal of Machine Learning Research
2002
Earlier work this paper cites.
S. Thiébaux, C. Gretton, J. Slaney, D. Price, and F. Kabanza, “Decision-theoretic planning with non-markovian rewards,” Journal of Artificial Intelligence Research
2006
Earlier work this paper cites.
D. Arthur and S. Vassilvitskii, “k-means++: The advantages of careful seeding,” tech. rep., Stanford, 2006
2006
Earlier work this paper cites.
C. Diuk, A. Cohen, and M. L. Littman, “An object-oriented representation for efficient reinforcement learning,” in Proceedings of the 25th international conference on Machine learning
2008
Earlier work this paper cites.
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al
2015
Cited alongside, same era.
2015
Cited alongside, same era.
S. Narvekar, J. Sinapov, M. Leonetti, and P. Stone, “Source task creation for curriculum learning,” in Proceedings of the 2016 international conference on autonomous agents & multiagent systems
2016
Cited alongside, same era.
2017
Cited alongside, same era.
A. Radford, K. Narasimhan, T. Salimans, I. Sutskever, et al
2018
Later among the works it cites.
2018
Later among the works it cites.
2019
Later among the works it cites.
S. Narvekar, B. Peng, M. Leonetti, J. Sinapov, M. E. Taylor, and P. Stone, “Curriculum learning for reinforcement learning domains: A framework and survey,” The Journal of Machine Learning Research
2020
Later among the works it cites.
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” The Journal of Machine Learning Research
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
2017
Cited alongside, same era.
A. Camacho, O. Chen, S. Sanner, and S. A. McIlraith, “Non-markovian rewards expressed in ltl: guiding search via reward shaping,” in Tenth annual symposium on combinatorial search
2017
Cited alongside, same era.
R. T. Icarte, T. Klassen, R. Valenzano, and S. McIlraith, “Using reward machines for high-level task specification and decomposition in reinforcement learning,” in International Conference on Machine Learning
2018
Cited alongside, same era.
R. Toro Icarte, T. Q. Klassen, R. Valenzano, and S. A. McIlraith, “Teaching multiple tasks to an rl agent using ltl,” in Proceedings of the 17th International Conference on Autonomous Agents and MultiAgent Systems
2018
Cited alongside, same era.
F. L. D. Silva and A. H. R. Costa, “Object-oriented curriculum generation for reinforcement learning,” in Proceedings of the 17th international conference on autonomous agents and multiagent systems
2018
Cited alongside, same era.
R. Brafman, G. De Giacomo, and F. Patrizi, “Ltlf/ldlf non-markovian rewards,” in Proceedings of the AAAI conference on artificial intelligence
2018
Cited alongside, same era.
2020
Later among the works it cites.
2021
Later among the works it cites.
P. Vaezipoor, A. C. Li, R. A. T. Icarte, and S. A. Mcilraith, “Ltl2action: Generalizing ltl instructions for multi-task rl,” in International Conference on Machine Learning
2021
Later among the works it cites.
R. T. Icarte, T. Q. Klassen, R. Valenzano, and S. A. McIlraith, “Reward machines: Exploiting reward function structure in reinforcement learning,” Journal of Artificial Intelligence Research
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.