Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) is a central problem in artificial intelligence.
Challenges of real-world reinforcement learning
Dulac-Arnold, G., Mankowitz, D., Hester, T., 2019 · 1904
Earlier work this paper cites.
Deepsynth: Automata synthesis for automatic task segmentation in deep reinforcement learning
Hasanbeig, M., Jeppu, N.Y., Abate, A., Melham, T., Kroening, D., 2021 · 1911
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning, in: Proceedings of the 33rd International Conference on Machine Learning (ICML), pp. 1928–1937
Mnih, V., Badia, A.P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., Kavukcuoglu, K., 2016 · 1937
Earlier work this paper cites.
Active task-inference-guided deep inverse reinforcement learning, in: Proceedings of the 59th IEEE Conference on on Decision and Control (CDC), pp. 1932–1938
Memarian, F., Xu, Z., Wu, B., Wen, M., Topcu, U., 2020 · 1938
Earlier work this paper cites.
Inductive inference: Theory and methods
Angluin, D., Smith, C.H., 1983 · 1983
Earlier work this paper cites.
Learning regular sets from queries and counterexamples
Angluin, D., 1987 · 1987
Earlier work this paper cites.
Q-learning
Watkins, C.J.C.H., Dayan, P., 1992 · 1992
Earlier work this paper cites.
An optimization-based categorization of reinforcement learning environments, in: From Animals to Animats 2: Proceedings of the Second International Conference on Simulation of Adaptive Behavior, pp. 262–270
Littman, M.L., 1993 · 1993
Earlier work this paper cites.
Learning finite state machines with self-clustering recurrent networks
Zeng, Z., Goodman, R.M., Smyth, P., 1993 · 1993
Earlier work this paper cites.
Acting optimally in partially observable stochastic domains, in: Proceedings of the 12th National Conference on Artificial Intelligence (AAAI), pp. 1023–1028
Cassandra, A.R., Kaelbling, L.P., Littman, M.L., 1994 · 1994
Earlier work this paper cites.
Learning without state-estimation in partially observable Markovian decision processes, in: Machine Learning Proceedings 1994. Elsevier, pp. 284–292
Singh, S.P., Jaakkola, T., Jordan, M.I., 1994 · 1994
Earlier work this paper cites.
Reinforcement learning: A survey
Kaelbling, L.P., Littman, M.L., Moore, A.W., 1996 · 1996
Earlier work this paper cites.
Learning finite-state controllers for partially observable environments, in: Proceedings of the 15th Conference on Uncertainty in Artificial Intelligence (UAI), pp. 427–436
Meuleau, N., Peshkin, L., Kim, K.E., Kaelbling, L.P., 1999 · 1999
Earlier work this paper cites.
Learning policies with external memory, in: Proceedings of the 16th International Conference on Machine Learning (ICML), pp. 307–314
Peshkin, L., Meuleau, N., Kaelbling, L.P., 1999 · 1999
Earlier work this paper cites.
Predictive representations of state, in: Proceedings of the 15th Conference on Advances in Neural Information Processing Systems (NIPS), pp. 1555–1561
Littman, M.L., Sutton, R.S., Singh, S., 2002 · 2002
Earlier work this paper cites.
Local search in combinatorial optimization
Aarts, E., Aarts, E.H., Lenstra, J.K., 2003 · 2003
Earlier work this paper cites.
Using rewards for belief state updates in partially observable Markov decision processes, in: Proceedings of the 16th European Conference on Machine Learning (ECML), pp. 593–600
Izadi, M.T., Precup, D., 2005 · 2005
Earlier work this paper cites.
Handbook of constraint programming
Rossi, F., Van Beek, P., Walsh, T., 2006 · 2006
Earlier work this paper cites.
Xu, Z., Wu, B., Neider, D., Topcu, U., 2020b · 2006
Earlier work this paper cites.
50 Years of integer programming 1958-2008: From the early years to the state-of-the-art
Jünger, M., Liebling, T.M., Naddef, D., Nemhauser, G.L., Pulleyblank, W.R., Reinelt, G., Rinaldi, G., Wolsey, L.A., 2009 · 2008
Cited alongside, same era.
Model-based Bayesian reinforcement learning in partially observable domains, in: Proceedings of the 10th International Symposium on Artificial Intelligence and Mathematics (ISAIM), pp. 1–2
Poupart, P., Vlassis, N., 2008 · 2008
Cited alongside, same era.
Online learning of non-Markovian reward models
Rens, G., Raskin, J.F., Reynouad, R., Marra, G., 2020 · 2009
Cited alongside, same era.
Grammatical inference: learning automata and grammars
De la Higuera, C., 2010 · 2010
Cited alongside, same era.
Constructing states for reinforcement learning, in: Proceedings of the 27th International Conference on Machine Learning (ICML), pp. 727–734
Mahmud, M., 2010 · 2010
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., Klimov, O., 2017 · 2017
Later among the works it cites.
Mastering the game of Go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., et al., 2017 · 2017
Later among the works it cites.
Learning dexterous in-hand manipulation
Andrychowicz, M., Baker, B., Chociej, M., Jozefowicz, R., McGrew, B., Pachocki, J., Petron, A., Plappert, M., Powell, G., Ray, A., et al., 2018 · 2018
Later among the works it cites.
Gurobi Optimizer Reference Manual
Gurobi Optimization, LLC, 2018 · 2018
Later among the works it cites.
Optimizing agent behavior over long time scales by transporting value
Hung, C.C., Lillicrap, T., Abramson, J., Wu, Y., Mirza, M., Carnevale, F., Ahuja, A., Wayne, G., 2018 · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Large neighborhood search, in: Handbook of metaheuristics. Springer, pp. 399–419
Pisinger, D., Ropke, S., 2010 · 2010
Cited alongside, same era.
Reward machines: Exploiting reward function structure in reinforcement learning
Toro Icarte, R., Klassen, T.Q., Valenzano, R., McIlraith, S.A., 2020a · 2010
Cited alongside, same era.
The act of remembering: a study in partially observable reinforcement learning
Toro Icarte, R., Valenzano, R., Klassen, T.Q., Christoffersen, P., massoud Farahmand, A., McIlraith, S.A., 2020b · 2010
Cited alongside, same era.
Bayesian nonparametric methods for partially-observable reinforcement learning
Doshi-Velez, F., Pfau, D., Wood, F., Roy, N., 2013 · 2013
Cited alongside, same era.
Bayesian reinforcement learning: A survey
Ghavamzadeh, M., Mannor, S., Pineau, J., Tamar, A., et al., 2015 · 2015
Cited alongside, same era.
Deep recurrent q-learning for partially observable MDPs, in: AAAI Fall Symposium on Sequential Decision Making for Intelligent Agents (AAAI-SDMIA15)
Hausknecht, M., Stone, P., 2015 · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A.A., Veness, J., Bellemare, M.G., Graves, A., Riedmiller, M., Fidjeland, A.K., Ostrovski, G., et al., 2015 · 2015
Cited alongside, same era.
Later among the works it cites.
ILOG CP Optimizer 12.8 Manual
IBM, 2018 · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Sutton, R.S., Barto, A.G., 2018 · 2018
Later among the works it cites.
Using reward machines for high-level task specification and decomposition in reinforcement learning, in: Proceedings of the 35th International Conference on Machine Learning (ICML), pp. 2112–2121
Toro Icarte, R., Klassen, T.Q., Valenzano, R., McIlraith, S.A., 2018 · 2018
Later among the works it cites.
LTL and beyond: Formal languages for reward function specification in reinforcement learning, in: Proceedings of the 28th International Joint Conference on Artificial Intelligence (IJCAI), pp. 6065–6073
Camacho, A., Toro Icarte, R., Klassen, T.Q., Valenzano, R., McIlraith, S.A., 2019 · 2019
Later among the works it cites.
Learning reward machines for partially observable reinforcement learning, in: Proceedings of the 32nd Conference on Advances in Neural Information Processing Systems (NeurIPS), pp. 15497–15508
Toro Icarte, R., Waldie, E., Klassen, T.Q., Valenzano, R., Castro, M.P., McIlraith, S.A., 2019 · 2019
Later among the works it cites.
Temporal logic monitoring rewards via transducers, in: Proceedings of the 17th International Conference on Knowledge Representation and Reasoning (KR), pp. 860–870
De Giacomo, G., Favorito, M., Iocchi, L., Patrizi, F., Ronca, A., 2020 · 2020
Later among the works it cites.
Induction of subgoal automata for reinforcement learning., in: Proceedings of the 34th AAAI Conference on Artificial Intelligence (AAAI), pp. 3890–3897
Furelos-Blanco, D., Law, M., Russo, A., Broda, K., Jonsson, A., 2020 · 2020
Later among the works it cites.
Reinforcement learning with non-Markovian rewards, in: Proceedings of the 34th AAAI Conference on Artificial Intelligence (AAAI), pp. 3980–3987
Gaon, M., Brafman, R., 2020 · 2020
Later among the works it cites.
Challenges of real-world reinforcement learning: definitions, benchmarks and analysis
Dulac-Arnold, G., Levine, N., Mankowitz, D.J., Li, J., Paduraru, C., Gowal, S., Hester, T., 2021 · 2021
Closest in time.
Advice-guided reinforcement learning in a non-Markovian environment, in: Proceedings of the 35th AAAI Conference on Artificial Intelligence (AAAI), pp. 9073–9080
Neider, D., Gaglione, J.R., Gavran, I., Topcu, U., Wu, B., Xu, Z., 2021 · 2021
Closest in time.
Interpretable sequence classification via discrete optimization, in: Proceedings of the 35th AAAI Conference on Artificial Intelligence (AAAI), pp. 9647–9656
Shvo, M., Li, A.C., Toro Icarte, R., McIlraith, S.A., 2021 · 2021
Closest in time.
Tabu search, in: Handbook of combinatorial optimization. Springer, pp. 2093–2229
Glover, F., Laguna, M., 1998 · 2093
Closest in time.
Deep reinforcement learning with Double Q-learning, in: Proceedings of the 30th AAAI Conference on Artificial Intelligence (AAAI), pp. 2094–2100
Van Hasselt, H., Guez, A., Silver, D., 2016 · 2094
Closest in time.