Fetching the paper…
Reading the bibliography…
The BabyAI platform is designed to measure the sample efficiency of training an agent to follow grounded-language instructions.
Backpropagation through time: what it does and how to do it
Werbos, P. J. (1990) · 1990
Earlier work this paper cites.
Long Short-Term Memory
Hochreiter, S. and Schmidhuber, J. (1997) · 1997
Earlier work this paper cites.
Combating false negatives in adversarial imitation learning
Żołna, K., Saharia, C., Boussioux, L., Hui, D. Y.-T., Chevalier-Boisvert, M., Bahdanau, D., and Bengio, Y. (2020) · 2002
Earlier work this paper cites.
Gaussian Processes for Machine Learning
Rasmussen, C. E. and Williams, C. K. I. (2005) · 2005
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Mikolov, T., Chen, K., Corrado, G., and Dean, J. (2013) · 2013
Earlier work this paper cites.
Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation
Cho, K., van Merrienboer, B., Gulcehre, C., Bougares, F., Schwenk, H., and Bengio, Y. (2014) · 2014
Earlier work this paper cites.
Neural Machine Translation by Jointly Learning to Align and Translate
Bahdanau, D., Cho, K., and Bengio, Y. (2015) · 2015
Cited alongside, same era.
Adam: A Method for Stochastic Optimization
Kingma, D. P. and Ba, J. (2015) · 2015
Cited alongside, same era.
High-Dimensional Continuous Control Using Generalized Advantage Estimation
Schulman, J., Moritz, P., Levine, S., Jordan, M., and Abbeel, P. (2015) · 2015
Cited alongside, same era.
Leike, J., Martic, M., Krakovna, V., Ortega, P. A., Everitt, T., Lefrancq, A., Orseau, L., and Legg, S. (2017) · 2017
Cited alongside, same era.
FiLM: Visual Reasoning with a General Conditioning Layer
Perez, E., Strub, F., de Vries, H., Dumoulin, V., and Courville, A. (2017) · 2017
Cited alongside, same era.
Attend, adapt and transfer: Attentive deep architecture for adaptive transfer from multiple sources in the same domain
Rajendran, J., Lakshminarayanan, A. S., Khapra, M. M., Prasanna, P., and Ravindran, B. (2015) · 2017
Later among the works it cites.
Proximal Policy Optimization Algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. (2017) · 2017
Later among the works it cites.
Scalable trust-region method for deep reinforcement learning using kronecker-factored approximation
Wu, Y., Mansimov, E., Liao, S., Grosse, R. B., and Ba, J. (2017) · 2017
Later among the works it cites.
gym-sokoban
Schrader, M.-P. B. (2018) · 2018
Later among the works it cites.
BabyAI: First steps towards grounded language learning with a human in the loop
Chevalier-Boisvert, M., Bahdanau, D., Lahlou, S., Willems, L., Saharia, C., Nguyen, T. H., and Bengio, Y. (2019) · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…