Fetching the paper…
Reading the bibliography…
Imitation learning is a powerful family of techniques for learning sensorimotor coordination in immersive environments.
Fixup initialization: Residual learning without normalization
Zhang, H., Dauphin, Y. N., and Ma, T. (2019) · 1901
Earlier work this paper cites.
The minerl competition on sample efficient reinforcement learning using human priors
Guss, W. H., Codel, C., Hofmann, K., Houghton, B., Kuno, N., Milani, S., Mohanty, S. P., Liebana, D. P., Salakhutdinov, R., Topin, N., Veloso, M., and Wang, P. (2019a) · 1904
Earlier work this paper cites.
Boosted bellman residual minimization handling expert demonstrations
Piot, B., Geist, M., and Pietquin, O. (2014) · 2014
Earlier work this paper cites.
Deep apprenticeship learning for playing video games
Bogdanovic, M., Markovikj, D., Denil, M., and De Freitas, N. (2015) · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al. (2015) · 2015
Earlier work this paper cites.
The Malmo platform for artificial intelligence experimentation
Johnson, M., Hofmann, K., Hutton, T., and Bignell, D. (2016) · 2016
Earlier work this paper cites.
Vizdoom: A doom-based ai research platform for visual reinforcement learning
Kempka, M., Wydmuch, M., Runc, G., Toczek, J., and Jaśkowski, W. (2016) · 2016
Earlier work this paper cites.
The game imitation: Deep supervised convolutional networks for quick video game AI
Chen, Z. and Yi, D. (2017) · 2017
Cited alongside, same era.
Learning macromanagement in starcraft from replays using deep learning
Justesen, N. and Risi, S. (2017) · 2017
Cited alongside, same era.
Hierarchical and interpretable skill acquisition in multi-task reinforcement learning
Shu, T., Xiong, C., and Socher, R. (2017) · 2017
Cited alongside, same era.
Mastering chess and shogi by self-play with a general reinforcement learning algorithm
Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., et al. (2017) · 2017
Cited alongside, same era.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
Human-level performance in 3D multiplayer games with population-based reinforcement learning
Jaderberg, M., Czarnecki, W. M., Dunning, I., Marris, L., Lever, G., Castañeda, A. G., Beattie, C., Rabinowitz, N. C., Morcos, A. S., Ruderman, A., Sonnerat, N., Green, T., Deason, L., Leibo, J. Z., Silver, D., Hassabis, D., Kavukcuoglu, K., and Graepel, T. (2019) · 2019
Later among the works it cites.
Teacher-student curriculum learning
Matiisen, T., Oliver, A., Cohen, T., and Schulman, J. (2019) · 2019
Later among the works it cites.
Retrospective analysis of the 2019 minerl competition on sample efficient reinforcement learning
Milani, S., Topin, N., Houghton, B., Guss, W. H., Mohanty, S. P., Nakata, K., Vinyals, O., and Kuno, N. S. (2020) · 2019
Later among the works it cites.
Competing in the obstacle tower challenge
Nichol, A. (2019) · 2019
Later among the works it cites.
Grandmaster level in StarCraft II using multi-agent reinforcement learning
Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., Choi, D. H., Powell, R., Ewalds, T., Georgiev, P., et al. (2019) · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Espeholt, L., Soyer, H., Munos, R., Simonyan, K., Mnih, V., Ward, T., Doron, Y., Firoiu, V., Harley, T., Dunning, I., et al. (2018) · 2018
Cited alongside, same era.
Rainbow: Combining improvements in deep reinforcement learning
Hessel, M., Modayil, J., Van Hasselt, H., Schaul, T., Ostrovski, G., Dabney, W., Horgan, D., Piot, B., Azar, M., and Silver, D. (2018) · 2018
Cited alongside, same era.
Deep q-learning from demonstrations
Hester, T., Vecerik, M., Pietquin, O., Lanctot, M., Schaul, T., Piot, B., Horgan, D., Quan, J., Sendonaris, A., Osband, I., et al. (2018) · 2018
Cited alongside, same era.
MineRL: a large-scale dataset of Minecraft demonstrations
Guss, W. H., Houghton, B., Topin, N., Wang, P., Codel, C., Veloso, M., and Salakhutdinov, R. (2019b)
Cited in the paper.
Sample factory: Egocentric 3D control from pixels at 100000 FPS with asynchronous reinforcement learning
Petrenko, A., Huang, Z., Kumar, T., Sukhatme, G., and Koltun, V. (2020) · 2020
Closest in time.
DD-PPO: Learning near-perfect PointGoal navigators from 2.5 billion frames
Wijmans, E., Kadian, A., Morcos, A., Lee, S., Essa, I., Parikh, D., Savva, M., and Batra, D. (2020) · 2020
Closest in time.