Fetching the paper…
Reading the bibliography…
A Relational Markov Decision Process (RMDP) is a first-order representation to express all instances of a single probabilistic planning domain with possibly unbounded number of objects.
A Markovian Decision Process
Bellman, R · 1957
Earlier work this paper cites.
Proceedings of the 11th International Joint Conference on Artificial Intelligence. Detroit, MI, USA, August 1989 , 1989. Morgan Kaufmann
Sridharan, N. S. (ed.) · 1989
Earlier work this paper cites.
Markov Decision Processes
Puterman, M · 1994
Earlier work this paper cites.
Robot learning from demonstration
Atkeson, C. G. and Schaal, S · 1997
Earlier work this paper cites.
Symbolic dynamic programming for first-order mdps
Boutilier, C., Reiter, R., and Price, B · 2001
Earlier work this paper cites.
Generalizing plans to new environments in relational mdps
Guestrin, C., Koller, D., Gearhart, C., and Kanodia, N · 2003
Earlier work this paper cites.
Solving relational MDPs with first-order machine learning
Mausam and Weld, D. S · 2003
Earlier work this paper cites.
Approximate linear programming for first-order mdps
Sanner, S. and Boutilier, C · 2005
Earlier work this paper cites.
Approximate policy iteration with a policy language bias: Solving relational markov decision processes
Fern, A., Yoon, S. W., and Givan, R · 2006
Earlier work this paper cites.
Practical solution techniques for first-order mdps
Sanner, S. and Boutilier, C · 2008
Earlier work this paper cites.
Transfer via soft homomorphisms
Sorg, J. and Singh, S. P · 2009
Earlier work this paper cites.
Transfer learning for reinforcement learning domains: A survey
Taylor, M. E. and Stone, P · 2009
Earlier work this paper cites.
Relational Dynamic Influence Diagram Language (RDDL): Language Description
Sanner, S · 2010
Cited alongside, same era.
Probabilistic relational planning with first order decision diagrams
Joshi, S. and Khardon, R · 2011
Cited alongside, same era.
PROST: probabilistic planning based on UCT
Keller, T. and Eyerich, P · 2012
Cited alongside, same era.
A theory of goal-oriented mdps with dead ends
Kolobov, A., Mausam, and Weld, D. S · 2012
Cited alongside, same era.
Planning with Markov Decision Processes: An AI Perspective
Mausam and Kolobov, A · 2012
Cited alongside, same era.
International Probabilistic Planning Competition (IPPC) 2014
Grzes, M., Hoey, J., and Sanner, S · 2014
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., van den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., Dieleman, S., Grewe, D., Nham, J., Kalchbrenner, N., Sutskever, I., Lillicrap, T. P., Leach, M., Kavukcuoglu, K., Graepel, T., and Hassabis, D · 2016
Later among the works it cites.
DARLA: improving zero-shot transfer in reinforcement learning
Higgins, I., Pal, A., Rusu, A. A., Matthey, L., Burgess, C., Pritzel, A., Botvinick, M., Blundell, C., and Lerchner, A · 2017
Later among the works it cites.
Semi-supervised classification with graph convolutional networks
Kipf, T. N. and Welling, M · 2017
Later among the works it cites.
Teacher-student curriculum learning
Matiisen, T., Oliver, A., Cohen, T., and Schulman, J · 2017
Later among the works it cites.
Velickovic, P., Cucurull, G., Casanova, A., Romero, A., Liò, P., and Bengio, Y · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Younes, H. L. S., Littman, M. L., Weissman, D., and Asmuth, J · 2014
Cited alongside, same era.
Actor-mimic: Deep multitask and transfer reinforcement learning
Parisotto, E., Ba, L. J., and Salakhutdinov, R · 2015
Cited alongside, same era.
Empirical evaluation of rectified activations in convolutional network
Xu, B., Wang, N., Chen, T., and Li, M · 2015
Cited alongside, same era.
Towards deep symbolic reinforcement learning
Garnelo, M., Arulkumaran, K., and Shanahan, M · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T. P., Harley, T., Silver, D., and Kavukcuoglu, K · 2016
Cited alongside, same era.
An overview of gradient descent optimization algorithms, 2016
Ruder, S · 2016
Cited alongside, same era.
Later among the works it cites.
Transfer of deep reactive policies for mdp planning
Bajpai, A., Garg, S., and Mausam · 2018
Later among the works it cites.
Learning generalized reactive policies using deep neural networks
Groshev, E., Tamar, A., Goldstein, M., Srivastava, S., and Abbeel, P · 2018
Later among the works it cites.
Training deep reactive policies for probabilistic planning problems
Issakkimuthu, M., Fern, A., and Tadepalli, P · 2018
Later among the works it cites.
Action schema networks: Generalised policies with deep learning
Toyer, S., Trevizan, F. W., Thiébaux, S., and Xie, L · 2018
Later among the works it cites.
Size independent neural transfer for rddl planning
Garg, S., Bajpai, A., and Mausam · 2019
Later among the works it cites.
Guiding Search with Generalized Policies for Probabilistic Planning
Shen, W., Trevizan, F., Toyer, S., Thiébaux, S., and Xie, L · 2019
Later among the works it cites.