Fetching the paper…
Reading the bibliography…
In the field of Sequential Decision Making (SDM), two paradigms have historically vied for supremacy: Automated Planning (AP) and Reinforcement Learning (RL).
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. (2020) · 1901
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K. (2016) · 1937
Earlier work this paper cites.
A formal basis for the heuristic determination of minimum cost paths
Hart, P. E., Nilsson, N. J., and Raphael, B. (1968) · 1968
Earlier work this paper cites.
Learning and executing generalized robot plans
Fikes, R. E., Hart, P. E., and Nilsson, N. J. (1972) · 1972
Earlier work this paper cites.
Depth-first search and linear graph algorithms
Tarjan, R. (1972) · 1972
Earlier work this paper cites.
The nonlinear nature of plans
Sacerdoti, E. D. (1975) · 1975
Earlier work this paper cites.
Generating project networks
Tate, A. (1977) · 1977
Earlier work this paper cites.
Breadth-first search
Bundy, A. and Wallen, L. (1984) · 1984
Earlier work this paper cites.
Macro-operators: A weak method for learning
Korf, R. E. (1985) · 1985
Earlier work this paper cites.
Rule creation and rule learning through environmental exploration
Shen, W. M. and Simon, H. A. (1989) · 1989
Earlier work this paper cites.
Learning from delayed rewards
Watkins, C. J. C. H. (1989) · 1989
Earlier work this paper cites.
A survey of algorithmic methods for partially observed Markov decision processes
Lovejoy, W. S. (1991) · 1991
Earlier work this paper cites.
Dyna, an integrated architecture for learning, planning, and reacting
Sutton, R. S. (1991) · 1991
Earlier work this paper cites.
Reinforcement learning with perceptual aliasing: The perceptual distinctions approach
Chrisman, L. (1992) · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J. (1992) · 1992
Earlier work this paper cites.
Probabilistic planning with information gathering and contingent execution
Draper, D., Hanks, S., and Weld, D. S. (1994) · 1994
Earlier work this paper cites.
Consideration of risk in reinforcement learning
Heger, M. (1994) · 1994
Earlier work this paper cites.
Reinforcement learning with soft state aggregation
Singh, S., Jaakkola, T., and Jordan, M. (1994) · 1994
Earlier work this paper cites.
Structural regression trees
Kramer, S. (1996) · 1996
Earlier work this paper cites.
Algorithms for sequential decision-making
Littman, M. L. (1996) · 1996
Earlier work this paper cites.
Searching for planning operators with context-dependent and probabilistic effects
Oates, T. and Cohen, P. R. (1996) · 1996
Earlier work this paper cites.
Learning planning operators by observation and practice
Wang, X. (1996) · 1996
Earlier work this paper cites.
Fast planning through planning graph analysis
Blum, A. L. and Furst, M. L. (1997) · 1997
Earlier work this paper cites.
The automatic inference of state invariants in tim
Fox, M. and Long, D. (1998) · 1998
Earlier work this paper cites.
Inferring state constraints for domain-independent planning
Gerevini, A. and Schubert, L. (1998) · 1998
Earlier work this paper cites.
Macro-actions in reinforcement learning: An empirical analysis
McGovern, A. and Sutton, R. S. (1998) · 1998
Earlier work this paper cites.
Spudd: stochastic planning using decision diagrams
Hoey, J., St-Aubin, R., Hu, A., and Boutilier, C. (1999) · 1999
Earlier work this paper cites.
Efficient implementation of the plan graph in stan
Long, D. and Fox, M. (1999) · 1999
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Sutton, R. S., Precup, D., and Singh, S. (1999) · 1999
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Ng, A. Y., Russell, S., et al. (2000) · 2000
Earlier work this paper cites.
Relational reinforcement learning
Džeroski, S., De Raedt, L., and Driessens, K. (2001) · 2001
Earlier work this paper cites.
Planning with pattern databases
Edelkamp, S. (2001) · 2001
Earlier work this paper cites.
Lao*: A heuristic search algorithm that finds solutions with loops
Hansen, E. A. and Zilberstein, S. (2001) · 2001
Earlier work this paper cites.
Ff: The fast-forward planning system
Hoffmann, J. (2001) · 2001
Earlier work this paper cites.
Symbolic heuristic search for factored Markov decision processes
Feng, Z. and Hansen, E. A. (2002) · 2002
Earlier work this paper cites.
Learning options in reinforcement learning
Stolle, M. and Precup, D. (2002) · 2002
Earlier work this paper cites.
Labeled rtdp: Improving the convergence of real-time dynamic programming
Bonet, B. and Geffner, H. (2003) · 2003
Earlier work this paper cites.
Shop2: An htn planning system
Nau, D. S., Au, T.-C., Ilghami, O., Kuter, U., Murdock, J. W., Wu, D., and Yaman, F. (2003) · 2003
Earlier work this paper cites.
Learning first-order Markov models for control
Abbeel, P. and Ng, A. (2004) · 2004
Earlier work this paper cites.
Learning domain-specific control knowledge from random walks
Fern, A., Yoon, S. W., and Givan, R. (2004) · 2004
Earlier work this paper cites.
Ordered landmarks in planning
Hoffmann, J., Porteous, J., and Sebastia, L. (2004) · 2004
Earlier work this paper cites.
Val: Automatic plan validation, continuous effects and mixed initiative planning using pddl
Howey, R., Long, D., and Fox, M. (2004) · 2004
Earlier work this paper cites.
Relational reinforcement learning: An overview
Tadepalli, P., Givan, R., and Driessens, K. (2004) · 2004
Earlier work this paper cites.
Ppddl1. 0: An extension to pddl for expressing planning domains with probabilistic effects
Younes, H. L. and Littman, M. L. (2004) · 2004
Earlier work this paper cites.
Macro-ff: Improving ai planning with automatically learned macro-operators
Botea, A., Enzenberger, M., Müller, M., and Schaeffer, J. (2005) · 2005
Earlier work this paper cites.
Efficiently handling temporal knowledge in an htn planner
Castillo, L. A., Fernández-Olivares, J., Garcia-Perez, O., and Palao, F. (2006) · 2006
Earlier work this paper cites.
Optimizations of data structures, heuristics and algorithms for path-finding on maps
Cazenave, T. (2006) · 2006
Earlier work this paper cites.
The fast downward planning system
Helmert, M. (2006) · 2006
Earlier work this paper cites.
Learning hierarchical task networks by observation
Nejati, N., Langley, P., and Konik, T. (2006) · 2006
Earlier work this paper cites.
Reinforcement learning with human teachers: Evidence of feedback and guidance with implications for learning performance
Thomaz, A. L., Breazeal, C., et al. (2006) · 2006
Earlier work this paper cites.
Learning heuristic functions from relaxed plans
Yoon, S. W., Fern, A., and Givan, R. (2006) · 2006
Earlier work this paper cites.
Marvin: A heuristic search planner with online macro-action learning
Coles, A. I. and Smith, A. J. (2007) · 2007
Earlier work this paper cites.
Automated creation of pattern database search heuristics
Edelkamp, S. (2007) · 2007
Earlier work this paper cites.
Introduction to statistical relational learning
Koller, D., Friedman, N., Džeroski, S., Sutton, C., McCallum, A., Pfeffer, A., Abbeel, P., Wong, M.-F., Meek, C., Neville, J., et al. (2007) · 2007
Earlier work this paper cites.
Building portable options: Skill transfer in reinforcement learning
Konidaris, G. D. and Barto, A. G. (2007) · 2007
Earlier work this paper cites.
Learning symbolic models of stochastic domains
Pasula, H. M., Zettlemoyer, L. S., and Kaelbling, L. P. (2007) · 2007
Earlier work this paper cites.
Incremental learning of planning operators in stochastic domains
Safaei, J. and Ghassem-Sani, G. (2007) · 2007
Earlier work this paper cites.
Learning action models from plan examples using weighted max-sat
Yang, Q., Wu, K., and Jiang, Y. (2007) · 2007
Earlier work this paper cites.
Towards model-lite planning: A proposal for learning & planning with incomplete domain models
Yoon, S. and Kambhampati, S. (2007) · 2007
Earlier work this paper cites.
Ff-replan: A baseline for probabilistic planning
Yoon, S. W., Fern, A., and Givan, R. (2007) · 2007
Cited alongside, same era.
An object-oriented representation for efficient reinforcement learning
Diuk, C., Cohen, A., and Littman, M. L. (2008) · 2008
Cited alongside, same era.
Safe exploration for reinforcement learning
Hans, A., Schneegaß, D., Schäfer, A. M., and Udluft, S. (2008) · 2008
Cited alongside, same era.
Htn-maker: Learning htns with minimal additional knowledge engineering required
Hogg, C., Munoz-Avila, H., and Kuter, U. (2008) · 2008
Cited alongside, same era.
The PELA architecture: integrating planning and learning to improve execution
Jiménez, S., Fernández, F., and Borrajo, D. (2008) · 2008
Cited alongside, same era.
Using kernel perceptrons to learn action effects for planning
Mourao, K., Petrick, R. P., and Steedman, M. (2008) · 2008
A review of learning planning action models
Arora, A., Fiorino, H., Pellier, D., Métivier, M., and Pesty, S. (2018) · 2018
Later among the works it cites.
Towards a simple approach to multi-step model-based reinforcement learning
Asadi, K., Cater, E., Misra, D., and Littman, M. L. (2018) · 2018
Later among the works it cites.
Relational inductive biases, deep learning, and graph networks
Battaglia, P. W., Hamrick, J. B., Bapst, V., Sanchez-Gonzalez, A., Zambaldi, V., Malinowski, M., Tacchetti, A., Raposo, D., Santoro, A., Faulkner, R., et al. (2018) · 2018
Later among the works it cites.
Treeqn and atreec: Differentiable tree planning for deep reinforcement learning
Farquhar, G., Rockt aschel, T., Igl, M., and Whiteson, S. (2018) · 2018
Later among the works it cites.
Learning to search with MCTSnets
Guez, A., Weber, T., Antonoglou, I., Simonyan, K., Vinyals, O., Wierstra, D., Munos, R., and Silver, D. (2018) · 2018
Later among the works it cites.
Curiosity driven exploration of learned disentangled goal spaces
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Reinforcement learning and automated planning: A survey
Partalas, I., Vrakas, D., and Vlahavas, I. (2008) · 2008
Cited alongside, same era.
Regression for classical and nondeterministic planning
Rintanen, J. (2008) · 2008
Cited alongside, same era.
Dyna-style planning with linear function approximation and prioritized sweeping
Sutton, R. S., Szepesvári, C., Geramifard, A., and Bowling, M. H. (2008) · 2008
Cited alongside, same era.
Efficient learning of action schemas and web-service descriptions
Walsh, T. J. and Littman, M. L. (2008) · 2008
Cited alongside, same era.
Concise finite-domain representations for pddl planning tasks
Helmert, M. (2009) · 2009
Cited alongside, same era.
Cost-optimal planning with landmarks
Karpas, E. and Domshlak, C. (2009) · 2009
Cited alongside, same era.
Laversanne-Finot, A., Pere, A., and Oudeyer, P.-Y. (2018) · 2018
Later among the works it cites.
Deep reinforcement learning
Li, Y. (2018) · 2018
Later among the works it cites.
Deep learning: A critical appraisal
Marcus, G. (2018) · 2018
Later among the works it cites.
Learning-driven goal generation
Pozanco, A., Fernández, S., and Borrajo, D. (2018) · 2018
Later among the works it cites.
Asp-based time-bounded planning for logistics robots
Schäpers, B., Niemueller, T., Lakemeyer, G., Gebser, M., and Schaub, T. (2018) · 2018
Later among the works it cites.
A general reinforcement learning algorithm that masters chess, shogi, and go through self-play
Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., et al. (2018) · 2018
Later among the works it cites.
Universal planning networks: Learning generalizable representations for visuomotor control
Srinivas, A., Jabri, A., Abbeel, P., Levine, S., and Finn, C. (2018) · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G. (2018) · 2018
Later among the works it cites.
Action schema networks: Generalised policies with deep learning
Toyer, S., Trevizan, F., Thiébaux, S., and Xie, L. (2018) · 2018
Later among the works it cites.
Deep reinforcement learning for nlp
Wang, W. Y., Li, J., and He, X. (2018) · 2018
Later among the works it cites.
Reinforcement learning and optimal control
Bertsekas, D. (2019) · 2019
Later among the works it cites.
Neural logic machines
Dong, H., Mao, J., Lin, T., Wang, C., Li, L., and Zhou, D. (2019) · 2019
Later among the works it cites.
Reconciling deep learning with symbolic artificial intelligence: representing objects and relations
Garnelo, M. and Shanahan, M. (2019) · 2019
Later among the works it cites.
An investigation of model-free planning
Guez, A., Mirza, M., Gregor, K., Kabra, R., Racanière, S., Weber, T., Raposo, D., Santoro, A., Orseau, L., Eccles, T., et al. (2019) · 2019
Later among the works it cites.
Darpa’s explainable artificial intelligence (xai) program
Gunning, D. and Aha, D. (2019) · 2019
Later among the works it cites.
An introduction to the planning domain definition language
Haslum, P., Lipovetzky, N., Magazzeni, D., and Muise, C. (2019) · 2019
Later among the works it cites.
A review of generalized planning
Jiménez, S., Segovia-Aguas, J., and Jonsson, A. (2019) · 2019
Later among the works it cites.
Sdrl: interpretable and data-efficient deep reinforcement learning leveraging symbolic planning
Lyu, D., Yang, F., Liu, B., and Gustafson, S. (2019) · 2019
Later among the works it cites.
Counterexample-guided abstraction refinement for pattern selection in optimal classical planning
Rovner, A., Sievers, S., and Helmert, M. (2019) · 2019
Later among the works it cites.
Deep reinforcement learning with relational inductive biases
Zambaldi, V., Raposo, D., Santoro, A., Bapst, V., Li, Y., Babuschkin, I., Tuyls, K., Reichert, D., Lillicrap, T., Lockhart, E., et al. (2019) · 2019
Later among the works it cites.
The emerging landscape of explainable automated planning & decision making
Chakraborti, T., Sreedharan, S., and Kambhampati, S. (2020) · 2020
Later among the works it cites.
The value equivalence principle for model-based reinforcement learning
Grimm, C., Barreto, A., Singh, S., and Silver, D. (2020) · 2020
Later among the works it cites.
Model based reinforcement learning for atari
Kaiser, Ł., Babaeizadeh, M., Miłos, P., Osiński, B., Campbell, R. H., Czechowski, K., Erhan, D., Finn, C., Kozakowski, P., Levine, S., et al. (2020) · 2020
Later among the works it cites.
Generating data in planning: Sas planning tasks of a given causal structure
Katz, M. and Sohrabi, S. (2020) · 2020
Later among the works it cites.
Contrastive learning of structured world models
Kipf, T. N., van der Pol, E., and Welling, M. (2020) · 2020
Later among the works it cites.
A framework for reinforcement learning and planning
Moerland, T. M., Broekens, J., and Jonker, C. M. (2020) · 2020
Later among the works it cites.
Landmark-based approaches for goal recognition as planning
Pereira, R. F., Oren, N., and Meneguzzi, F. (2020) · 2020
Later among the works it cites.
Artificial Intelligence: A Modern Approach (4th Edition)
Russell, S. J. and Norvig, P. (2020) · 2020
Later among the works it cites.
Mastering atari, go, chess and shogi by planning with a learned model
Schrittwieser, J., Antonoglou, I., Hubert, T., Simonyan, K., Sifre, L., Schmitt, S., Guez, A., Lockhart, E., Hassabis, D., Graepel, T., et al. (2020) · 2020
Later among the works it cites.
Learning domain-independent planning heuristics with hypergraph networks
Shen, W., Trevizan, F., and Thiébaux, S. (2020) · 2020
Later among the works it cites.
A survey of inverse reinforcement learning: Challenges, methods and progress
Arora, S. and Doshi, P. (2021) · 2021
Later among the works it cites.
Combinatorial optimization and reasoning with graph neural networks
Cappart, Q., Chételat, D., Khalil, E. B., Lodi, A., Morris, C., and Velickovic, P. (2021) · 2021
Later among the works it cites.
Reinforcement learning in economics and finance
Charpentier, A., Elie, R., and Remlinger, C. (2021) · 2021
Later among the works it cites.
Learning hierarchical task networks with preferences from unannotated demonstrations
Chen, K., Srikanth, N. S., Kent, D., Ravichandar, H., and Chernova, S. (2021) · 2021
Later among the works it cites.
Landmark generation in htn planning
Höller, D. and Bercher, P. (2021) · 2021
Later among the works it cites.
Scenario planning in the wild: A neuro-symbolic approach
Katz, M., Srinivas, K., Sohrabi, S., Feblowitz, M., Udrea, O., and Hassanzadeh, O. (2021) · 2021
Later among the works it cites.
Discovering symbolic policies with deep reinforcement learning
Landajuela, M., Petersen, B. K., Kim, S., Santiago, C. P., Glatt, R., Mundhenk, N., Pettit, J. F., and Faissol, D. (2021) · 2021
Later among the works it cites.
Discovering relational and numerical expressions from plan traces for learning action models
Segura-Muros, J. Á., Pérez, R., and Fernández-Olivares, J. (2021) · 2021
Later among the works it cites.
Automatic instance generation for classical planning
Torralba, A., Seipp, J., and Sievers, S. (2021) · 2021
Later among the works it cites.
Classical planning in deep latent space
Asai, M., Kajino, H., Fukunaga, A., and Muise, C. (2022) · 2022
Later among the works it cites.
Intrinsically motivated goal exploration processes with automatic curriculum learning
Forestier, S., Portelas, R., Mollard, Y., and Oudeyer, P.-Y. (2022) · 2022
Later among the works it cites.
Neural-symbolic learning and reasoning: a survey and interpretation
Garcez, A. d., Bader, S., Bowman, H., Lamb, L. C., de Penning, L., Illuminoo, B., Poon, H., and Zaverucha, C. G. (2022) · 2022
Later among the works it cites.
Reinforcement learning for classical planning: Viewing heuristics as dense reward generators
Gehring, C., Asai, M., Chitnis, R., Silver, T., Kaelbling, L., Sohrabi, S., and Katz, M. (2022) · 2022
Later among the works it cites.
Creativity of ai: Automatic symbolic option discovery for facilitating deep reinforcement learning
Jin, M., Ma, Z., Jin, K., Zhuo, H. H., Chen, C., and Yu, C. (2022) · 2022
Later among the works it cites.
Planning with Markov decision processes: An AI perspective
Natarajan, M. and Kolobov, A. (2022) · 2022
Later among the works it cites.
Learning to select goals in automated planning with deep-q learning
Núñez-Molina, C., Fernández-Olivares, J., and Pérez, R. (2022) · 2022
Later among the works it cites.
Neurosymbolic reinforcement learning and planning: A survey
Acharya, K., Raza, W., Dourado, C., Velasquez, A., and Song, H. H. (2023) · 2023
Closest in time.
Explainability in deep reinforcement learning: A review into current methods and applications
Hickling, T., Zenati, A., Aouf, N., and Spencer, P. (2023) · 2023
Closest in time.
Model-based reinforcement learning: A survey
Moerland, T. M., Broekens, J., Plaat, A., Jonker, C. M., et al. (2023) · 2023
Closest in time.
Nesig: A neuro-symbolic method for learning to generate planning problems
Núñez-Molina, C., Mesejo, P., and Fernández-Olivares, J. (2023) · 2023
Closest in time.
High-accuracy model-based reinforcement learning, a survey
Plaat, A., Kosters, W., and Preuss, M. (2023) · 2023
Closest in time.
Reinforcement learning algorithms: A brief survey
Shakya, A. K., Pillai, G., and Chakrabarty, S. (2023) · 2023
Closest in time.
Reinforcement learning with knowledge representation and reasoning: A brief survey
Yu, C., Zheng, X., Zhuo, H. H., Wan, H., and Luo, W. (2023) · 2023
Closest in time.