Fetching the paper…
Reading the bibliography…
Deep reinforcement learning (deep RL) excels in various domains but lacks generalizability and interpretability.
Karel the robot: a gentle introduction to the art of programming
Pattis, R. E · 1981
Earlier work this paper cites.
Neocognitron: A self-organizing neural network model for a mechanism of visual pattern recognition
Fukushima, K. and Miyake, S · 1982
Earlier work this paper cites.
Optimization of computer simulation models with rare events
Rubinstein, R. Y · 1997
Earlier work this paper cites.
Learning from demonstration
Schaal, S · 1997
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Sutton, R. S., Precup, D., and Singh, S · 1999
Earlier work this paper cites.
Hierarchical reinforcement learning with the MAXQ value function decomposition
Dietterich, T. G · 2000
Earlier work this paper cites.
Crafting papers on machine learning
Langley, P · 2000
Earlier work this paper cites.
Recent advances in hierarchical reinforcement learning
Barto, A. G. and Mahadevan, S · 2003
Earlier work this paper cites.
Distill: Learning domain-specific planners by example
Winner, E. and Veloso, M · 2003
Earlier work this paper cites.
Hierarchical neural program synthesis
Zhong, L., Lindeborg, R., Zhang, J., Lim, J. J., and Sun, S.-H · 2003
Earlier work this paper cites.
Learning teleoreactive logic programs from problem solving
Choi, D. and Langley, P · 2005
Earlier work this paper cites.
Learning phrase representations using rnn encoder-decoder for statistical machine translation, 2014
Cho, K., van Merrienboer, B., Gulcehre, C., Bahdanau, D., Bougares, F., Schwenk, H., and Bengio, Y · 2014
Earlier work this paper cites.
beta-vae: Learning basic visual concepts with a constrained variational framework
Higgins, I., Matthey, L., Pal, A., Burgess, C., Glorot, X., Botvinick, M., Mohamed, S., and Lerchner, A · 2016
Earlier work this paper cites.
Generative adversarial imitation learning
Ho, J. and Ermon, S · 2016
Earlier work this paper cites.
Vizdoom: A doom-based ai research platform for visual reinforcement learning
Kempka, M., Wydmuch, M., Runc, G., Toczek, J., and Jaśkowski, W · 2016
Earlier work this paper cites.
The mythos of model interpretability
Lipton, Z. C · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al · 2016
Earlier work this paper cites.
Modular multitask reinforcement learning with policy sketches
Andreas, J., Klein, D., and Levine, S · 2017
Earlier work this paper cites.
The option-critic architecture
Bacon, P.-L., Harb, J., and Precup, D · 2017
Earlier work this paper cites.
Deepcoder: Learning to write programs
Balog, M., Gaunt, A. L., Brockschmidt, M., Nowozin, S., and Tarlow, D · 2017
Earlier work this paper cites.
Robustfill: Neural program learning under noisy i/o
Devlin, J., Uesato, J., Bhupatiraju, S., Singh, R., Mohamed, A.-r., and Kohli, P · 2017
Earlier work this paper cites.
Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates
Gu, S., Holly, E., Lillicrap, T., and Levine, S · 2017
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2017
Cited alongside, same era.
Survey of model-based reinforcement learning: Applications on robotics
Polydoros, A. S. and Nalpantidis, L · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Cited alongside, same era.
Mastering the game of go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., Chen, Y., Lillicrap, T., Hui, F., Sifre, L., van den Driessche, G., Graepel, T., and Hassabis, D · 2017
Cited alongside, same era.
Feudal networks for hierarchical reinforcement learning
Vezhnevets, A. S., Osindero, S., Schaul, T., Heess, N., Jaderberg, M., Silver, D., and Kavukcuoglu, K · 2017
Cited alongside, same era.
Synthesizing programmatic policies that inductively generalize
Inala, J. P., Bastani, O., Tavares, Z., and Solar-Lezama, A · 2020
Later among the works it cites.
Interpretability in ML: A broad overview, 2020
Shen, O · 2020
Later among the works it cites.
Program guided agent
Sun, S.-H., Wu, T.-L., and Lim, J. J · 2020
Later among the works it cites.
Joint inference of reward machines and policies for reinforcement learning
Xu, Z., Gavran, I., Ahmad, Y., Majumdar, R., Neider, D., Topcu, U., and Wu, B · 2020
Later among the works it cites.
Deepsynth: Automata synthesis for automatic task segmentation in deep reinforcement learning
Hasanbeig, M., Jeppu, N. Y., Abate, A., Melham, T. F., and Kroening, D · 2021
Later among the works it cites.
Latent programmer: Discrete latent codes for program synthesis
Hong, J., Dohan, D., Singh, R., Sutton, C., and Zaheer, M · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Verifiable reinforcement learning via policy extraction
Bastani, O., Pu, Y., and Solar-Lezama, A · 2018
Cited alongside, same era.
Leveraging grammar and reinforcement learning for neural program synthesis
Bunel, R. R., Hausknecht, M., Devlin, J., Singh, R., and Kohli, P · 2018
Cited alongside, same era.
Meta learning shared hierarchies
Frans, K., Ho, J., Chen, X., Abbeel, P., and Schulman, J · 2018
Cited alongside, same era.
Using reward machines for high-level task specification and decomposition in reinforcement learning
Icarte, R. T., Klassen, T., Valenzano, R., and McIlraith, S · 2018
Cited alongside, same era.
Improving neural program synthesis with inferred execution traces
Shin, E. C., Polosukhin, I., and Song, D · 2018
Cited alongside, same era.
Neural program synthesis from diverse demonstration videos
Sun, S.-H., Noh, H., Somasundaram, S., and Lim, J · 2018
Cited alongside, same era.
Programmatically interpretable reinforcement learning
Verma, A., Murali, V., Singh, R., Kohli, P., and Chaudhuri, S · 2018
Cited alongside, same era.
Later among the works it cites.
How to train your robot with deep reinforcement learning: lessons we have learned
Ibarz, J., Tan, J., Finn, C., Kalakrishnan, M., Pastor, P., and Levine, S · 2021
Later among the works it cites.
Flexible option learning
Klissarov, M. and Precup, D · 2021
Later among the works it cites.
Discovering symbolic policies with deep reinforcement learning
Landajuela, M., Petersen, B. K., Kim, S., Santiago, C. P., Glatt, R., Mundhenk, N., Pettit, J. F., and Faissol, D · 2021
Later among the works it cites.
Generalizable imitation learning from observation via inferring goal proximity
Lee, Y., Szot, A., Sun, S.-H., and Lim, J. J · 2021
Later among the works it cites.
Learning to synthesize programs as interpretable and generalizable policies
Trivedi, D., Zhang, J., Sun, S.-H., and Lim, J. J · 2021
Later among the works it cites.
Leveraging approximate symbolic models for reinforcement learning via skill diversity
Guan, L., Sreedharan, S., and Kambhampati, S · 2022
Later among the works it cites.
Skill-based meta-reinforcement learning
Nam, T., Sun, S.-H., Pertsch, K., Hwang, S. J., and Lim, J. J · 2022
Later among the works it cites.
Outracing champion gran turismo drivers with deep reinforcement learning
Wurman, P. R., Barrett, S., Kawamoto, K., MacGlashan, J., Subramanian, K., Walsh, T. J., Capobianco, R., Devlic, A., Eckert, F., Fuchs, F., et al · 2022
Later among the works it cites.
Show me the way! bilevel search for synthesizing programmatic strategies
Aleixo, D. S. and Lelis, L. H · 2023
Closest in time.
LEAGUE: Guided skill learning and abstraction for long-horizon manipulation
Cheng, S. and Xu, D · 2023
Closest in time.
Chevalier-Boisvert, M., Dai, B., Towers, M., de Lazcano, R., Willems, L., Lahlou, S., Pal, S., Castro, P. S., and Terry, J · 2023
Closest in time.
Hierarchies of reward machines
Furelos-Blanco, D., Law, M., Jonsson, A., Broda, K., and Russo, A · 2023
Closest in time.
Code as policies: Language model programs for embodied control
Liang, J., Huang, W., Xia, F., Xu, P., Hausman, K., Ichter, B., Florence, P., and Zeng, A · 2023
Closest in time.
Hierarchical programmatic reinforcement learning via learning to compose programs
Liu, G.-T., Hu, E.-P., Cheng, P.-J., Lee, H.-Y., and Sun, S.-H · 2023
Closest in time.
Bootstrap your own skills: Learning to solve new tasks with large language model guidance
Zhang, J., Jiahui Zhang, K. P., Liu, Z., Ren, X., Chang, M., Sun, S.-H., and Lim, J. J · 2023
Closest in time.
Integrating planning and deep reinforcement learning via automatic induction of task substructures
Liu, J.-C., Chang, C.-H., Sun, S.-H., and Yu, T.-L · 2024
Closest in time.