Fetching the paper…
Reading the bibliography…
Recent advances in artificial intelligence have been driven by the presence of increasingly realistic and complex simulated environments.
Wang, R., Lehman, J., Clune, J., and Stanley, K. O. (2019) · 1901
Earlier work this paper cites.
Winning isn’t everything: Enhancing game development with intelligent agents
Zhao, Y., Borovikov, I., de Mesentier Silva, F., Beirami, A., Rupert, J., Somers, C., Harder, J., Kolen, J., Pinto, J., Pourabolghasem, R., Pestrak, J., Chaput, H., Sardari, M., Lin, L., Narravula, S., Aghdaie, N., and Zaman, K. (2019) · 1903
Earlier work this paper cites.
Clune, J. (2019) · 1905
Earlier work this paper cites.
Emergent tool use from multi-agent autocurricula
Baker, B., Kanitscheider, I., Markov, T., Wu, Y., Powell, G., McGrew, B., and Mordatch, I. (2019) · 1909
Earlier work this paper cites.
Making efficient use of demonstrations to solve hard exploration problems
Gulcehre, C., Paine, T. L., Shriari, B., Denil, M., Hoffman, M., Soyer, H., Tanburn, R., Kapturowski, S., Rabinowitz, N., Williams, D., Barth-Maron, G., Wang, Z., de Freitas, N., and Worlds Team (2019) · 1909
Earlier work this paper cites.
Leveraging procedural generation to benchmark reinforcement learning
Cobbe, K., Hesse, C., Hilton, J., and Schulman, J. (2019a) · 1912
Earlier work this paper cites.
Xxii. programming a computer for playing chess
Shannon, C. E. (1950) · 1950
Earlier work this paper cites.
Some studies in machine learning using the game of checkers
Samuel, A. L. (1959) · 1959
Earlier work this paper cites.
Continual learning in reinforcement environments
Ring, M. B. (1994) · 1994
Earlier work this paper cites.
Temporal difference learning and td-gammon
Tesauro, G. (1995) · 1995
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J. (1997) · 1997
Earlier work this paper cites.
Robotic grasping and contact: A review
Bicchi, A. and Kumar, V. (2000) · 2000
Earlier work this paper cites.
Human-level AI’s killer application: Interactive computer games
Laird, J. and VanLent, M. (2001) · 2001
Earlier work this paper cites.
Computer go
Müller, M. (2002) · 2002
Earlier work this paper cites.
Agent57: Outperforming the atari human benchmark
Puigdomènech Badia, A., Piot, B., Kapturowski, S., Sprechmann, P., Vitvitskyi, A., Guo, D., and Blundell, C. (2020) · 2003
Earlier work this paper cites.
Primate vocalization, gesture, and the evolution of human language
Arbib, M. A., Liebal, K., and Pika, S. (2008) · 2008
Earlier work this paper cites.
Hierarchical models of behavior and prefrontal function
Botvinick, M. M. (2008) · 2008
Earlier work this paper cites.
TAMER: Training an agent manually via evaluative reinforcement
Knox, W. B. and Stone, P. (2008) · 2008
Earlier work this paper cites.
Curriculum learning
Bengio, Y., Louradour, J., Collobert, R., and Weston, J. (2009) · 2009
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y. (2012) · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M. (2013) · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M. (2013) · 2013
Earlier work this paper cites.
Superintelligence: Paths, Dangers, Strategies
Bostrom, N. (2014) · 2014
Earlier work this paper cites.
Deep learning
LeCun, Y., Bengio, Y., and Hinton, G. (2015) · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, Georg, P. S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., and Hassabis, D. (2015) · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., Berg, A., and Fei-Fei, L. (2015) · 2015
Earlier work this paper cites.
Schmidhuber, J. (2015) · 2015
Earlier work this paper cites.
Agent-agnostic human-in-the-loop reinforcement learning
Abel, D., Salvatier, J., Stuhlmüller, A., and Evans, O. (2016) · 2016
Cited alongside, same era.
Beattie, C., Leibo, J. Z., Teplyashin, D., Ward, T., Wainwright, M., Küttler, H., Lefrancq, A., Green, S., Valdés, V., Sadik, Amir Schrittwieser, J., Anderson, K., York, S., Cant, M., Cain, A., Bolton, A., Gaffney, S., King, H., Hassabis, D., Legg, S., and Petersen, S. (2016) · 2016
Cited alongside, same era.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W. (2016) · 2016
Cited alongside, same era.
Pybullet, a python module for physics simulation for games, robotics and machine learning
Coumans, E. and Bai, Y. (2016) · 2016
Cited alongside, same era.
Mastering the game of go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., Chen, Y., Lillicrap, T., Hui, F., Sifre, L., van den Driessche, G., Graepel, T., and Demis, H. (2017) · 2017
Later among the works it cites.
A deep hierarchical approach to lifelong learning in minecraft
Tessler, C., Givony, S., Zahavy, T., Mankowitz, D. J., and Mannor, S. (2017) · 2017
Later among the works it cites.
Domain randomization for transferring deep neural networks from simulation to the real world
Tobin, J., Fong, R., Ray, A., Schneider, J., Zaremba, W., and Abbeel, P. (2017) · 2017
Later among the works it cites.
Training agent for first-person shooter game with actor-critic curriculum learning
Wu, Y. and Tian, Y. (2017) · 2017
Later among the works it cites.
Target-driven visual navigation in indoor scenes using deep reinforcement learning
Zhu, Y., Mottaghi, R., Kolve, E., Lim, J. J., Gupta, A., Fei-Fei, L., and Farhadi, A. (2017) · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dosovitskiy, A. and Koltun, V. (2016) · 2016
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J. (2016) · 2016
Cited alongside, same era.
Generative adversarial imitation learning
Ho, J. and Ermon, S. (2016) · 2016
Cited alongside, same era.
The malmo platform for artificial intelligence experimentation
Johnson, M., Hofmann, K., Hutton, T., and Bignell, D. (2016) · 2016
Cited alongside, same era.
Vizdoom: A doom-based AI research platform for visual reinforcement learning
Kempka, M., Wydmuch, M., Runc, G., Toczek, J., and Jaśkowski, W. (2016) · 2016
Cited alongside, same era.
End-to-end training of deep visuomotor policies
Levine, S., Finn, C., Darrell, T., and Abbeel, P. (2016) · 2016
Cited alongside, same era.
Control of memory, active perception, and action in minecraft
Oh, J., Chockalingam, V., Singh, S., and Lee, H. (2016) · 2016
Cited alongside, same era.
Safely interruptible agents
Orseau, L. and Armstrong, S. (2016) · 2016
Cited alongside, same era.
Later among the works it cites.
Learning dexterous in-hand manipulation
Andrychowicz, M., Baker, B., Chociej, M., Jozefowicz, R., McGrew, B., Pachocki, J., Petron, A., Plappert, M., Powell, G., Ray, A., Schneider, J., Sidor, S., Tobin, J., Welinder, P., Weng, L., and Zaremba, W. (2018) · 2018
Closest in time.
IMPALA: Scalable distributed deep-rl with importance weighted actor-learner architectures
Espeholt, L., Soyer, H., Munos, R., Simonyan, K., Mnih, V., Ward, T., Doron, Y., Firoiu, V., Harley, T., Dunning, I., Legg, S., and Kavukcuoglu, K. (2018) · 2018
Closest in time.
Designing efficient neural attention systems towards achieving human-level sharp vision
Ghani, A. R. A., Koganti, N., Solano, A., Iwasawa, Y., Nakayama, K., and Matsuo, Y. (2018) · 2018
Closest in time.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S. (2018) · 2018
Closest in time.
Semiparametric reinforcement learning
Jain, M. S. and Lindsey, J. (2018) · 2018
Closest in time.
Illuminating generalization in deep reinforcement learning through procedural level generation
Justesen, N., Rodriguez Torrado, R., Bontrager, P., Khalifa, A., Togelius, J., and Risi, S. (2018) · 2018
Closest in time.
Psychlab: a psychology laboratory for deep reinforcement learning agents
Leibo, J. Z., d’Autume, C. d. M., Zoran, D., Amos, D., Beattie, C., Anderson, K., Castañeda, A. G., Sanchez, M., Green, S., Gruslys, A., Legg, S., Hassabis, D., and Botvinick, M. (2018) · 2018
Closest in time.
Gotta learn fast: A new benchmark for generalization in rl
Nichol, A., Pfau, V., Hesse, C., Klimov, O., and Schulman, J. (2018) · 2018
Closest in time.
Schmidhuber, J. (2018) · 2018
Closest in time.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G. (2018) · 2018
Closest in time.
Sim-to-real: Learning agile locomotion for quadruped robots
Tan, J., Zhang, T., Coumans, E., Iscen, A., Bai, Y., Hafner, D., Bohez, S., and Vanhoucke, V. (2018) · 2018
Closest in time.
CHALET: Cornell house agent learning environment
Yan, C., Misra, D., Bennett, A., Walsman, A., Bisk, Y., and Artzi, Y. (2018) · 2018
Closest in time.
Artificial Intelligence and Games
Yannakakis, G. N. and Togelius, J. (2018) · 2018
Closest in time.
Learning from demonstration in the wild
Behbahani, F., Shiarlis, K., Chen, X., Kurin, V., Kasewa, S., Stirbu, C., Gomes, J., Paul, S., Oliehoek, F. A., Messias, J., and Whiteson, S. (2019) · 2019
Closest in time.
Large-scale study of curiosity-driven learning
Burda, Y., Edwards, H., Pathak, D., Storkey, A., Darrell, T., and Efros, A. A. (2019) · 2019
Closest in time.
Obstacle tower: A generalization challenge in vision, control, and planning
Juliani, A., Khalifa, A., Berges, V.-P., Harper, J., Henry, H., Crespi, A., Togelius, J., and Lange, D. (2019) · 2019
Closest in time.
Competing in the obstacle tower challenge
Nichol, A. (2019) · 2019
Closest in time.
Habitat: A Platform for Embodied AI Research
Savva, M., Kadian, A., Maksymets, O., Zhao, Y., Wijmans, E., Jain, B., Straub, J., Liu, J., Koltun, V., Malik, J., Parikh, D., and Batra, D. (2019) · 2019
Closest in time.
Alphastar: Mastering the real-time strategy game starcraft ii
Vinyals, O., Babuschkin, I., Chung, J., Mathieu, M., and Jaderberg, M. (2019) · 2019
Closest in time.
Leveraging human guidance for deep reinforcement learning tasks
Zhang, R., Torabi, F., Guan, L., H. Ballard, D., and Stone, P. (2019) · 2019
Closest in time.
Arena: A general evaluation platform and building toolkit for multi-agent intelligence
Song, Y., Wang, J., Lukasiewicz, T., Xu, Z., Xu, M., Ding, Z., and Wu, L. (2020) · 2020
Closest in time.