Fetching the paper…
Reading the bibliography…
Sparse reward is one of the most challenging problems in reinforcement learning (RL).
Understanding natural language
Winograd, T · 1972
Earlier work this paper cites.
The symbol grounding problem
Harnad, S · 1990
Earlier work this paper cites.
Shaping as a method for accelerating reinforcement learning
Gullapalli, V. and Barto, A. G · 1992
Earlier work this paper cites.
Incorporating advice into agents that learn from reinforcements
Maclin, R. and Shavlik, J. W · 1994
Earlier work this paper cites.
Grounding language in perception
Siskind, J. M · 1994
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Ng, A. Y., Harada, D., and Russell, S. J · 1999
Earlier work this paper cites.
Guiding a reinforcement learner with natural language advice: Initial results in robocup soccer
Kuhlmann, G., Stone, P., Mooney, R., and Shavlik, J · 2004
Earlier work this paper cites.
Visualizing data using t-sne
Maaten, L. v. d. and Hinton, G · 2008
Earlier work this paper cites.
Devise: A deep visual-semantic embedding model
Frome, A., Corrado, G. S., Shlens, J., Bengio, S., Dean, J., Mikolov, T., et al · 2013
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Mikolov, T., Sutskever, I., Chen, K., Corrado, G. S., and Dean, J · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M. A · 2013
Earlier work this paper cites.
Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling
Chung, J., Gulcehre, C., Cho, K., and Bengio, Y · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. and Ba, J · 2014
Cited alongside, same era.
A large annotated corpus for learning natural language inference
Bowman, S. R., Angeli, G., Potts, C., and Manning, C. D · 2015
Cited alongside, same era.
Effective approaches to attention-based neural machine translation
Luong, T., Pham, H., and Manning, C. D · 2015
Cited alongside, same era.
Universal value function approximators
Schaul, T., Horgan, D., Gregor, K., and Silver, D · 2015
Cited alongside, same era.
Faulty reward functions in the wild, 2016
Clark, J. and Amodei, D · 2016
Cited alongside, same era.
Deep reinforcement learning with double q-learning
Hasselt, H. v., Guez, A., and Silver, D · 2016
Cited alongside, same era.
Beating atari with natural language guided reinforcement learning
Kaplan, R., Sauer, C., and Sosa, A · 2017
Later among the works it cites.
Playing fps games with deep reinforcement learning
Lample, G. and Chaplot, D. S · 2017
Later among the works it cites.
Teaching machines to describe images via natural language feedback
Ling, H. and Fidler, S · 2017
Later among the works it cites.
Mapping instructions and visual observations to actions with reinforcement learning
Misra, D., Langford, J., and Artzi, Y · 2017
Later among the works it cites.
Data-efficient deep reinforcement learning for dexterous manipulation
Popov, I., Heess, N., Lillicrap, T. P., Hafner, R., Barth-Maron, G., Vecerik, M., Lampe, T., Tassa, Y., Erez, T., and Riedmiller, M. A · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
ViZDoom: A Doom-based AI research platform for visual reinforcement learning
Kempka, M., Wydmuch, M., Runc, G., Toczek, J., and Jaśkowski, W · 2016
Cited alongside, same era.
Asynchronous Methods for Deep Reinforcement Learning
Mnih, V., Puigdomènech Badia, A., Mirza, M., Graves, A., Lillicrap, T. P., Harley, T., Silver, D., and Kavukcuoglu, K · 2016
Cited alongside, same era.
Learning with Latent Language
Andreas, J., Klein, D., and Levine, S · 2017
Cited alongside, same era.
Hindsight experience replay
Andrychowicz, M., Crow, D., Ray, A. K., Schneider, J., Fong, R., Welinder, P., McGrew, B., Tobin, J., Abbeel, P., and Zaremba, W · 2017
Cited alongside, same era.
Supervised learning of universal sentence representations from natural language inference data
Conneau, A., Kiela, D., Schwenk, H., Barrault, L., and Bordes, A · 2017
Cited alongside, same era.
Grounded Language Learning in a Simulated 3D World
Hermann, K. M., Hill, F., Green, S., Wang, F., Faulkner, R., Soyer, H., Szepesvari, D., Czarnecki, W. M., Jaderberg, M., Teplyashin, D., Wainwright, M., Apps, C., Hassabis, D., and Blunsom, P · 2017
Cited alongside, same era.
Kiros, J. R. and Chan, W · 2018
Later among the works it cites.
The importance of sampling inmeta-reinforcement learning
Stadie, B., Yang, G., Houthooft, R., Chen, P., Duan, Y., Wu, Y., Abbeel, P., and Sutskever, I · 2018
Later among the works it cites.
Reinforcement Learning: An Introduction
Sutton, R. S. and Barto, A. G · 2018
Later among the works it cites.
A broad-coverage challenge corpus for sentence understanding through inference
Williams, A., Nangia, N., and Bowman, S · 2018
Later among the works it cites.
Interactive grounded language acquisition and generalization in a 2d world
Yu, H., Zhang, H., and Xu, W · 2018
Later among the works it cites.
Learning to understand goal specifications by modelling reward
Bahdanau, D., Hill, F., Leike, J., Hughes, E., Kohli, P., and Grefenstette, E · 2019
Closest in time.
Meta-learning language-guided policy learning
Co-Reyes, J. D., Gupta, A., Sanjeev, S., Altieri, N., DeNero, J., Abbeel, P., and Levine, S · 2019
Closest in time.