Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) algorithms typically start tabula rasa, without any prior knowledge of the environment, and without any prior skills.
Learning from delayed rewards
Watkins, C. J. C. H · 1989
Earlier work this paper cites.
Transfer Learning for Reinforcement Learning Domains: A Survey
Taylor, M. E. and Stone, P · 2009
Earlier work this paper cites.
Learning to Interpret Natural Language Navigation Instructions from Observation
Chen, D. L. and Mooney, R. J · 2011
Earlier work this paper cites.
Efficient Estimation of Word Representations in Vector Space
Mikolov, T., Chen, K., Corrado, G., and Dean, J · 2013
Earlier work this paper cites.
Lifelong Machine Learning Systems: Beyond Learning Algorithms
Silver, D. L., Yang, Q., and Li, L · 2013
Earlier work this paper cites.
OntoNotes: A Large Training Corpus for Enhanced Processing
Weischedel, R., Palmer, M., Marcus, M., Hovy, E., Pradhan, S., Ramshaw, L., Xue, N., Taylor, A., Kaufman, J., Franchini, M., El-Bachouti, M., Belvin, R., and Houston, A · 2013
Earlier work this paper cites.
How transferable are features in deep neural networks?
Yosinski, J., Clune, J., Bengio, Y., and Lipson, H · 2014
Earlier work this paper cites.
Deep Recurrent Q-Learning for Partially Observable MDPs
Hausknecht, M. and Stone, P · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., and Hassabis, D · 2015
Earlier work this paper cites.
Transfer Deep Reinforcement Learning in 3D Environments: An Empirical Study
Chaplot, D. S., Sathyendra, K. M., Lample, G., and Salakhutdinov, R · 2016
Cited alongside, same era.
Listen, Attend, and Walk: Neural Mapping of Navigational Instructions to Action Sequences
Mei, H., Bansal, M., and Walter, M. R · 2016
Cited alongside, same era.
The Option-Critic Architecture
Bacon, P.-L., Harb, J., and Precup, D · 2017
Cited alongside, same era.
Grounded Language Learning in a Simulated 3D World
Hermann, K. M., Hill, F., Green, S., Wang, F., Faulkner, R., Soyer, H., Szepesvari, D., Czarnecki, W. M., Jaderberg, M., Teplyashin, D., Wainwright, M., Apps, C., Hassabis, D., and Blunsom, P · 2017
Cited alongside, same era.
spaCy 2: Natural language understanding with Bloom embeddings, convolutional neural networks and incremental parsing
Honnibal, M. and Montani, I · 2017
Cited alongside, same era.
Building machines that learn and think like people
Diversity is All You Need: Learning Skills without a Reward Function
Eysenbach, B., Gupta, A., Ibarz, J., and Levine, S · 2019
Later among the works it cites.
Recurrent Experience Replay in Distributed Reinforcement Learning
Kapturowski, S., Ostrovski, G., Quan, J., Munos, R., and Dabney, W · 2019
Later among the works it cites.
A Survey of Reinforcement Learning Informed by Natural Language
Luketina, J., Nardelli, N., Farquhar, G., Foerster, J., Andreas, J., Grefenstette, E., Whiteson, S., and Rocktäschel, T · 2019
Later among the works it cites.
Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model
Schrittwieser, J., Antonoglou, I., Hubert, T., Simonyan, K., Sifre, L., Schmitt, S., Guez, A., Lockhart, E., Hassabis, D., Graepel, T., Lillicrap, T., and Silver, D · 2019
Later among the works it cites.
RTFM: Generalising to Novel Environment Dynamics via Reading
Zhong, V., Rocktäschel, T., and Grefenstette, E · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lake, B. M., Ullman, T. D., Tenenbaum, J. B., and Gershman, S. J · 2017
Cited alongside, same era.
gym-miniworld environment for openai gym
Chevalier-Boisvert, M · 2018
Cited alongside, same era.
Grounding Language for Transfer in Deep Reinforcement Learning
Narasimhan, K., Barzilay, R., and Jaakkola, T · 2018
Cited alongside, same era.
Learning to Understand Goal Specifications by Modelling Reward
Bahdanau, D., Hill, F., Leike, J., Hughes, E., Hosseini, A., Kohli, P., and Grefenstette, E · 2019
Cited alongside, same era.
Agent57: Outperforming the Atari Human Benchmark
Badia, A. P., Piot, B., Kapturowski, S., Sprechmann, P., Vitvitskyi, A., Guo, D., and Blundell, C · 2020
Closest in time.
Fast Task-Adaptation for tasks labeled using Natural Language in Reinforcement Learning
Hutsebaut-Buysse, M., Mets, K., and Latré, S · 2020
Closest in time.
Exploration in Reinforcement Learning with Deep Covering Options
Jinnai, Y., Park, J. W., Machado, M. C., and Konidaris, G · 2020
Closest in time.
Curriculum Learning for Reinforcement Learning Domains: A Framework and Survey
Narvekar, S., Peng, B., Leonetti, M., Sinapov, J., Taylor, M. E., and Stone, P · 2020
Closest in time.