Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) has demonstrated its ability to solve high dimensional tasks by leveraging non-linear function approximators.
L. A. Zadeh, “Fuzzy sets,” Information and control , vol. 8, no. 3, pp. 338–353, 1965
1965
Earlier work this paper cites.
P. Dayan and G. E. Hinton, “Feudal reinforcement learning,” in Advances in neural information processing systems (NIPS) , 1993, pp. 271–278
1993
Earlier work this paper cites.
M. Craven and J. W. Shavlik, “Extracting tree-structured representations of trained networks,” in Advances in Neural Information Processing Systems (NIPS) , 1996
1996
Earlier work this paper cites.
J. L. Bresina, “Heuristic-biased stochastic sampling,” in AAAI/IAAI, Vol. 1 , 1996, pp. 271–278
1996
Earlier work this paper cites.
R. Parr and S. J. Russell, “Reinforcement learning with hierarchies of machines,” in Advances in neural information processing systems (NIPS) , 1997
1997
Earlier work this paper cites.
R. S. Sutton and A. G. Barto, Reinforcement Learning: An Introduction . Boston, MA: MIT Press, 1998
1998
Earlier work this paper cites.
D. Precup, R. S. Sutton, and S. P. Singh, “Theoretical results on reinforcement learning with temporally abstract options,” in European Conference on Machine Learning (ECML) , 1998, pp. 382–393
1998
Earlier work this paper cites.
R. Sutton, D. McAllester, S. Singh, and Y. Mansour, “Policy Gradient Methods for Reinforcement Learning with Function Approximation,” in Advances in Neural Information Processing Systems (NIPS) , 1999
1999
Earlier work this paper cites.
T. G. Dietterich, “State abstraction in maxq hierarchical reinforcement learning,” in Advances in Neural Information Processing Systems (NIPS) , 1999
1999
Earlier work this paper cites.
R. S. Sutton, D. Precup, and S. P. Singh, “Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning,” Artif. Intell. , vol. 112, pp. 181–211, 1999
1999
Earlier work this paper cites.
D. P. Bertsekas, Nonlinear Programming . Belmont, MA: Athena Scientific, 1999
1999
Earlier work this paper cites.
C. W. Anderson, “Approximating a policy can be easier than approximating a value function,” Computer Science Technical Report , 2000
2000
Earlier work this paper cites.
T. G. Dietterich, “Hierarchical reinforcement learning with the maxq value function decomposition,” J. Artif. Intell. Res. , vol. 13, pp. 227–303, 2000
2000
Earlier work this paper cites.
S. M. Kakade, “A natural policy gradient,” in Advances in neural information processing systems (NIPS) , 2002, pp. 1531–1538
2002
Earlier work this paper cites.
A. G. Barto and S. Mahadevan, “Recent advances in hierarchical reinforcement learning,” Discrete event dynamic systems , vol. 13, no. 1-2, pp. 41–77, 2003
2003
Earlier work this paper cites.
M. Lagoudakis and R. Parr, “Least-squares policy iteration,” Journal of Machine Learning Research (JMLR) , vol. 4, pp. 1107–1149, 2003
2003
Earlier work this paper cites.
D. Ernst, P. Geurts, and L. Wehenkel, “Tree-based batch mode reinforcement learning,” Journal of Machine Learning Research , 2005
2005
Earlier work this paper cites.
L. Li, T. J. Walsh, and M. L. Littman, “Towards a unified theory of state abstraction for mdps,” in International Symposium on Artificial Intelligence and Mathematics (ISAIM) , 2006
2006
Earlier work this paper cites.
X. Xu, D. Hu, and X. Lu, “Kernel-based least squares policy iteration for reinforcement learning,” IEEE Transactions on Neural Networks , 2007
2007
Earlier work this paper cites.
I. Rexakis and M. G. Lagoudakis, “Classifier-based policy representation,” in Seventh International Conference on Machine Learning and Applications . IEEE, 2008, pp. 91–98
2008
Earlier work this paper cites.
J. Peters and S. Schaal, “Natural Actor-Critic,” Neurocomputation , vol. 71, no. 7-9, pp. 1180–1190, 2008
2008
Earlier work this paper cites.
G. Konidaris and A. Barto, “Skill discovery in continuous reinforcement learning domains using skill chaining,” in Advances in Neural Information Processing Systems (NIPS) , 2009, pp. 1015–1023
2009
Earlier work this paper cites.
C. Szepesvari, Algorithms for Reinforcement Learning . Morgan & Claypool, 2010
2010
Earlier work this paper cites.
E. Strumbelj and I. Kononenko, “An efficient explanation of individual classifications using game theory,” Journal of Machine Learning Resource (JMLR) , 2010
2010
Earlier work this paper cites.
S. Ross and D. Bagnell, “Efficient reductions for imitation learning,” in AISTATS , 2010, pp. 661–668
2010
Earlier work this paper cites.
C. D’Eramo, D. Tateo, A. Bonarini, M. Restelli, and J. Peters, “Mushroomrl: Simplifying reinforcement learning research,” https://github.com/MushroomRL/mushroom-rl , 2020
2010
Earlier work this paper cites.
D. P. Bertsekas, “Approximate policy iteration: a survey and some new methods,” Journal of Control Theory and Applications , vol. 9, no. 3, pp. 310–335, Aug 2011
2011
Cited alongside, same era.
B. da Silva, G. Konidaris, and A. Barto, “Learning Parameterized Skills,” in International Conference on Machine Learning (ICML) , 2012
2012
Cited alongside, same era.
E. Todorov, T. Erez, and Y. Tassa, “Mujoco: A physics engine for model-based control.” in International Conference on Intelligent Robots andSystems (IROS) . IEEE, 2012, pp. 5026–5033
2012
Cited alongside, same era.
C. Gentile, S. Li, and G. Zappella, “Online clustering of bandits,” in International Conference on Machine Learning (ICML) , 2014
2014
Cited alongside, same era.
B. Scherrer, “Approximate policy iteration schemes: A comparison,” in Proceedings of the 31th International Conference on Machine Learning, ICML 2014, Beijing, China, 21-26 June 2014 , 2014, pp. 1314–1322
C. Florensa, Y. Duan, and P. Abbeel, “Stochastic neural networks for hierarchical reinforcement learning,” in International Conference in Learning Representations (ICLR) , 2017
2017
Later among the works it cites.
C. Tessler, S. Givony, T. Zahavy, D. J. Mankowitz, and S. Mannor, “A deep hierarchical approach to lifelong learning in minecraft.” in AAAI Conference on Artificial Intelligence (AAAI) , vol. 3, 2017, p. 6
2017
Later among the works it cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in Neural Information Processing Systems (NIPS) , 2017
2017
Later among the works it cites.
B. Hayes and J. A. Shah, “Improving robot controller transparency through autonomous policy explanation,” in International Conference on Human-Robot Interaction (HRI) , 2017
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2014
Cited alongside, same era.
S. Masoudnia and R. Ebrahimpour, “Mixture of experts: A literature survey,” Artif. Intell. Rev. , 2014
2014
Cited alongside, same era.
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis, “Human-level control through deep reinforcement learning,” Nature , vol. 518, no. 7540, pp. 529–533, 02 2015
2015
Cited alongside, same era.
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra, “Continuous control with deep reinforcement learning,” CoRR , 2015
2015
Cited alongside, same era.
J. Schulman, S. Levine, M. Jordan, and P. Abbeel, “Trust Region Policy Optimization,” International Conference on Machine Learning (ICML) , p. 16, 2015
2015
Cited alongside, same era.
2015
Cited alongside, same era.
B. Ustun and C. Rudin, “Supersparse linear integer models for optimized medical scoring systems,” Machine Learning , 2015
2015
Cited alongside, same era.
B. Letham, C. Rudin, T. H. McCormick, and D. Madigan, “Interpretable classifiers using rules and bayesian analysis: Building a better stroke prediction model,” The Annals of Applied Statistics , 2015
2015
Cited alongside, same era.
A. Rajeswaran, K. Lowrey, E. Todorov, and S. M. Kakade, “Towards generalization and simplicity in continuous control,” in Conference on Neural Information Processing Systems (NIPS) , 2017
2017
Later among the works it cites.
2017
Later among the works it cites.
P. Dhariwal, C. Hesse, O. Klimov, A. Nichol, M. Plappert, A. Radford, J. Schulman, S. Sidor, and Y. Wu, “Openai baselines,” https://github.com/openai/baselines , 2017
2017
Later among the works it cites.
R. Akrour, F. Veiga, J. Peters, and G. Neumann, “Regularizing reinforcement learning with state abstraction,” in International Conference on Intelligent Robots and Systems (IROS) , 2018
2018
Later among the works it cites.
O. Nachum, S. S. Gu, H. Lee, and S. Levine, “Data-efficient hierarchical reinforcement learning,” in Advances in neural information processing systems (NIPS) , 2018, pp. 3303–3313
2018
Later among the works it cites.
A. Levy, G. Konidaris, R. Platt, and K. Saenko, “Learning multi-level hierarchies with hindsight,” in International Conference in Learning Representations (ICLR) , 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
A. Verma, V. Murali, R. Singh, P. Kohli, and S. Chaudhuri, “Programmatically interpretable reinforcement learning,” in International Conference on Machine Learning (ICML) , 2018
2018
Later among the works it cites.
O. Bastani, Y. Pu, and A. Solar-Lezama, “Verifiable reinforcement learning via policy extraction,” in Advances in Neural Information Processing Systems (NeurIPS) , 2018
2018
Later among the works it cites.
N. Rajaraman and R. Vaze, “Submodular maximization under a matroid constraint: Asking more from an old friend, the greedy algorithm,” 2018
2018
Later among the works it cites.
A. Marjaninejad, D. Urbina-Meléndez, B. A. Cohn, and F. J. Valero-Cuevas, “Autonomous functional movements in a tendon-driven limb via limited experience,” Nature machine intelligence , vol. 1, no. 3, pp. 144–154, 2019
2019
Later among the works it cites.
D. Tateo, İ. S. Erdenliğ, and A. Bonarini, “Graph-based design of hierarchical reinforcement learning agents,” in 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, 2019, pp. 1003–1009
2019
Later among the works it cites.
P. Madumal, T. Miller, L. Sonenberg, and F. Vetere, “Explainable reinforcement learning through a causal lens,” in Conference on Artificial Intelligence (AAAI) , 2019
2019
Later among the works it cites.
X. Nie, E. Brunskill, and S. Wager, “Learning When-to-Treat Policies,” arXiv e-prints , 2019
2019
Later among the works it cites.
Y. Coppens, K. Efthymiadis, T. Lenaerts, and A. Nowe, “Distilling deep reinforcement learning policies in soft decision trees,” in IJCAI Workshop on Explainable Artificial Intelligence , 2019
2019
Later among the works it cites.
R. Akrour, J. Pajarinen, J. Peters, and G. Neumann, “Projections for approximate policy iteration algorithms,” in International Conference on Machine Learning (ICML) , 2019
2019
Later among the works it cites.
E. Coumans and Y. Bai, “Pybullet, a python module for physics simulation for games, robotics and machine learning,” http://pybullet.org , 2016–2019
2019
Later among the works it cites.
P. Sundaresan, J. Grannen, B. Thananjeyan, A. Balakrishna, M. Laskey, K. Stone, J. E. Gonzalez, and K. Goldberg, “Learning rope manipulation policies using dense object descriptors trained on synthetic depth data,” in 2020 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2020, pp. 9411–9418
2020
Closest in time.
M. Riemer, I. Cases, C. Rosenbaum, M. Liu, and G. Tesauro, “On the role of weight sharing during deep option learning,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 34, no. 04, 2020, pp. 5519–5526
2020
Closest in time.
K. Mahadik, Q. Wu, S. Li, and A. Sabne, “Fast distributed bandits for online recommendation systems,” in International Conference on Supercomputing (ICS) , 2020
2020
Closest in time.