Fetching the paper…
Reading the bibliography…
Deep reinforcement learning (DRL) has gained great success by learning directly from high-dimensional sensory inputs, yet is notorious for the lack of interpretability.
A reinforcement learning method for maximizing undiscounted rewards
Schwartz, A. (1993) · 1993
Earlier work this paper cites.
Average reward reinforcement learning: Foundations, algorithms, and empirical results
Mahadevan, S. (1996) · 1996
Earlier work this paper cites.
Action languages
Gelfond, M. and Lifschitz, V. (1998) · 1998
Earlier work this paper cites.
Pddl-the planning domain definition language
McDermott, D., Ghallab, M., Howe, A., Knoblock, C., Ram, A., Veloso, M., Weld, D., and Wilkins, D. (1998) · 1998
Earlier work this paper cites.
Reinforcement learning with hierarchies of machines
Parr, R. and Russell, S. J. (1998) · 1998
Earlier work this paper cites.
Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning
Sutton, R. S., Precup, D., and Singh, S. (1999) · 1999
Earlier work this paper cites.
Using abstract models of behaviours to automatically generate reinforcement learning hierarchies
Ryan, M. R. K. (2002) · 2002
Earlier work this paper cites.
Recent advances in hierarchical reinforcement learning
Barto, A. and Mahadevan, S. (2003) · 2003
Earlier work this paper cites.
Intrinsically motivated reinforcement learning
Chentanez, N., Barto, A. G., and Singh, S. P. (2005) · 2005
Earlier work this paper cites.
Actions as special cases
Erdoğan, S. T. and Lifschitz, V. (2006) · 2006
Earlier work this paper cites.
The fast downward planning system
Helmert, M. (2006) · 2006
Earlier work this paper cites.
Automated planning
Cimatti, A., Pistore, M., and Traverso, P. (2008) · 2008
Earlier work this paper cites.
What is answer set programming?
Lifschitz, V. (2008) · 2008
Earlier work this paper cites.
What is intrinsic motivation? a typology of computational approaches
Oudeyer, P.-Y. and Kaplan, F. (2009) · 2009
Cited alongside, same era.
Learning methods to generate good plans: Integrating htn learning and reinforcement learning
Hogg, C., Kuter, U., and Munoz-Avila, H. (2010) · 2010
Cited alongside, same era.
Conflict-driven answer set solving: From theory to practice
Gebser, M., Kaufmann, B., and Schaub, T. (2012) · 2012
Cited alongside, same era.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M. (2013) · 2013
Cited alongside, same era.
Action Language ℬ 𝒞 \mathcal{BC} : A Preliminary Report
Lee, J., Lifschitz, V., and Yang, F. (2013) · 2013
Cited alongside, same era.
Robot task planning and explanation in open and uncertain worlds
Hanheide, M., Göbelbecker, M., Horn, G. S., et al. (2015) · 2015
Why should i trust you?: Explaining the predictions of any classifier
Ribeiro, M. T., Singh, S., and Guestrin, C. (2016) · 2016
Later among the works it cites.
Towards a rigorous science of interpretable machine learning
Doshi-Velez, F. and Kim, B. (2017) · 2017
Later among the works it cites.
Bwibots: A platform for bridging the gap between ai and human–robot interaction research
Khandelwal, P., Zhang, S., Sinapov, J., Leonetti, M., Thomason, J., Yang, F., Gori, I., Svetlik, M., Khante, P., Lifschitz, V., and Stone, P. (2017) · 2017
Later among the works it cites.
Advantages and limitations of using successor features for transfer in reinforcement learning
Lehnert, L., Tellex, S., and Littman, M. L. (2017) · 2017
Later among the works it cites.
Human learning in atari
Tsividis, P. A., Pouncy, T., Xu, J. L., Tenenbaum, J. B., and Gershman, S. J. (2017) · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al. (2015) · 2015
Cited alongside, same era.
Schaul, T., Quan, J., Antonoglou, I., and Silver, D. (2015) · 2015
Cited alongside, same era.
Planning with task-oriented knowledge acquisition for a service robot
Chen, K., Yang, F., and Chen, X. (2016) · 2016
Cited alongside, same era.
Modular action language alm
Inclezan, D. and Gelfond, M. (2016) · 2016
Cited alongside, same era.
Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation
Kulkarni, T. D., Narasimhan, K., Saeedi, A., and Tenenbaum, J. (2016) · 2016
Cited alongside, same era.
A synthesis of automated planning and reinforcement learning for efficient, robust decision-making
Leonetti, M., Iocchi, L., and Stone, P. (2016) · 2016
Cited alongside, same era.
Explaining explanations: An approach to evaluating interpretability of machine learning
Gilpin, L. H., Bau, D., Yuan, B. Z., Bajwa, A., Specter, M., and Kagal, L. (2018) · 2018
Closest in time.
Strategic object oriented reinforcement learning
Keramati, R., Whang, J., Cho, P., and Brunskill, E. (2018) · 2018
Closest in time.
Hierarchical imitation and reinforcement learning
Le, H. M., Jiang, N., Agarwal, A., Dud\́hbox{\it\/}k, M., Yue, Y., and Daumé III, H. (2018) · 2018
Closest in time.
Robot represention and reasoning with knowledge from reinforcement learning
Lu, K., Zhang, S., Stone, P., and Chen, X. (2018) · 2018
Closest in time.
Programmatically interpretable reinforcement learning
Verma, A., Murali, V., Singh, R., Kohli, P., and Chaudhuri, S. (2018) · 2018
Closest in time.
Peorl: Integrating symbolic planning and hierarchical reinforcement learning for robust decision-making
Yang, F., Lyu, D., Liu, B., and Gustafson, S. (2018) · 2018
Closest in time.
Deep reinforcement learning with double q-learning
Van Hasselt, H., Guez, A., and Silver, D. (2016) · 2094
Closest in time.