Fetching the paper…
Reading the bibliography…
To achieve the ambitious goals of artificial intelligence, reinforcement learning must include planning with a model of the world that is abstract in state and time.
1974
Earlier work this paper cites.
Sutton, R. S. (1984). Temporal Credit Assignment in Reinforcement Learning
1984
Earlier work this paper cites.
Sutton, R. S. (1988). Learning to predict by the methods of temporal differences. Machine Learning 3
1988
Earlier work this paper cites.
Sutton, R. S. (1995). TD models: Modeling the world at a mixture of time scales. In Proceedings of the International Conference on Machine Learning
1995
Earlier work this paper cites.
Sutton, R. S., Barto, A. (2018). Reinforcement Learning: An Introduction
2000
Earlier work this paper cites.
McGovern, A., Barto, A. G. (2001). Automatic discovery of subgoals in reinforcement learning using diverse density. In Proceedings of the International Conference on Machine Learning
2001
Earlier work this paper cites.
LeCun, Y., Bottou, L., Bengio, Y., Haffner, P. (1998). Gradient-based learning applied to document recognition. Proceedings of the IEEE 86 (11):2278–2324. Littman, M., Sutton, R. S., Singh, S. (2002). Predictive representations of state. Advances in Neural Information Processing Systems 14
2002
Earlier work this paper cites.
Jaderberg, M., Mnih, V., Czarnecki, W. M., Schaul, T., Leibo, J. Z., Silver, D., Kavukcuoglu, K. (2017). Reinforcement learning with unsupervised auxiliary tasks. In Proceedings of the International Conference on Learning Representations
2003
Earlier work this paper cites.
Simsek, Ö., Barto, A. G. (2004). Using relative novelty to identify useful temporal abstractions in reinforcement learning. In Proceedings of the International Conference on Machine Learning
2004
Cited alongside, same era.
Singh, S. P., Barto, A. G., Chentanez, N. (2004). Intrinsically motivated reinforcement learning. In Advances in Neural Information Processing Systems
2004
Cited alongside, same era.
Simsek, Ö., Wolfe, A. P., Barto, A. G. (2005). Identifying useful subgoals in reinforcement learning by local graph partitioning. In Proceedings of the International Conference on Machine Learning
2005
Cited alongside, same era.
Sorg, J., Singh, S. P. (2010). Linear options. In Proceedings of the International Conference on Autonomous Agents and Multiagent Systems
2010
Cited alongside, same era.
Solway, A., Diuk, C., Cordova, N., Yee, D., Barto, A. G., Niv, Y., Botvinick, M. (2014). Optimal behavioral hierarchy. PLoS Computational Biology 10
2014
Later among the works it cites.
Machado, M. C., Bellemare, M. G., Bowling, M. (2017). A Laplacian framework for option discovery in reinforcement learning. In Proceedings of the International Conference on Machine Learning
2017
Later among the works it cites.
Machado, M. C., Rosenbaum, C., Guo, X., Liu, M., Tesauro, G., Campbell, M. (2018). Eigenoption discovery through the deep successor representation. In Proceedings of the International Conference on Learning Representations
2018
Later among the works it cites.
Deisenroth, M. P., Rasmussen, C. E. (2011). PILCO: A model-based and data-efficient approach to policy search. In Proceedings of the International Conference on Machine Learning . Eysenbach, B., Gupta, A., Ibarz, J., Levine, S. (2019). Diversity is all you need: Learning skills without a reward function. In Proceedings of the International Conference on Learning Representations
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sutton, R. S., Modayil, J., Delp, M., Degris, T., Pilarski, P. M., White, A., Precup, D. (2011). Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction. In Proceedings of the International Conference on Autonomous Agents and Multiagent Systems
2011
Cited alongside, same era.
Silver, D., Ciosek, K. (2012). Compositional planning using optimal option models. In Proceedings of the International Conference on Machine Learning
2012
Cited alongside, same era.
Bacon, P.-L., Harb, J., Precup, D. (2017). The option-critic architecture. In Proceedings of the Association for Advancement of Artificial Intelligence . Baranes, A., Oudeyer, P.-Y. (2013). Active learning of inverse models with intrinsically motivated goal exploration in robots. Robotics and Autonomous Systems 61
2013
Cited alongside, same era.
2019
Later among the works it cites.
Wan, Y., Zaheer, M., White, A., White, M., Sutton, R. S. (2019). Planning with expectation models. In Proceedings of the International Joint Conference on Artificial Intelligence
2019
Later among the works it cites.
Sutton, R. S., Precup, D., Singh, S. (1999). Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning. Artificial Intelligence 112
2022
Closest in time.
Machado, M. C., Barreto, A., Precup, D, Bowling, M. (2023). Temporal abstraction in reinforcement learning with the successor representation. Journal of Machine Learning Research 24
2023
Closest in time.