Fetching the paper…
Reading the bibliography…
Neural architectures inspired by our own human cognitive system, such as the recently introduced world models, have been shown to outperform traditional deep reinforcement learning (RL) methods in a variety of different domains.
An on-line algorithm for dynamic reinforcement learning and planning in reactive environments. In 1990 IJCNN international joint conference on neural networks . IEEE, 253–258
Jürgen Schmidhuber. 1990 · 1990
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Evolving neural networks through augmenting topologies
Kenneth O Stanley and Risto Miikkulainen. 2002 · 2002
Earlier work this paper cites.
Neuroevolution: from architectures to learning
Dario Floreano, Peter Dürr, and Claudio Mattiussi. 2008 · 2008
Earlier work this paper cites.
Exploiting open-endedness to solve problems through the search for novelty.. In ALIFE . 329–336
Joel Lehman and Kenneth O Stanley. 2008 · 2008
Earlier work this paper cites.
Visualizing data using t-SNE
Laurens van der Maaten and Geoffrey Hinton. 2008 · 2008
Earlier work this paper cites.
A hypercube-based encoding for evolving large-scale neural networks
Kenneth O Stanley, David B D’Ambrosio, and Jason Gauci. 2009 · 2009
Earlier work this paper cites.
PILCO: A model-based and data-efficient approach to policy search. In Proceedings of the 28th International Conference on machine learning (ICML-11) . 465–472
Marc Deisenroth and Carl E Rasmussen. 2011 · 2011
Earlier work this paper cites.
Search-based procedural content generation: A taxonomy and survey
Julian Togelius, Georgios N Yannakakis, Kenneth O Stanley, and Cameron Browne. 2011 · 2011
Earlier work this paper cites.
An enhanced hypercube-based encoding for evolving the placement, density, and connectivity of neurons
Sebastian Risi and Kenneth O Stanley. 2012 · 2012
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Yoshua Bengio, Nicholas Léonard, and Aaron Courville. 2013 · 2013
Earlier work this paper cites.
Generating sequences with recurrent neural networks
Alex Graves. 2013a · 2013
Earlier work this paper cites.
Evolving deep unsupervised convolutional networks for vision-based reinforcement learning. In Proceedings of the 2014 Annual Conference on Genetic and Evolutionary Computation . ACM, 541–548
Jan Koutník, Jürgen Schmidhuber, and Faustino Gomez. 2014 · 2014
Earlier work this paper cites.
Techniques for learning binary stochastic feedforward neural networks
Tapani Raiko, Mathias Berglund, Guillaume Alain, and Laurent Dinh. 2014 · 2014
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification. In Proceedings of the IEEE international conference on computer vision . 1026–1034
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2015 · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Trust region policy optimization. In International Conference on Machine Learning . 1889–1897
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz. 2015 · 2015
Cited alongside, same era.
From pixels to torques: Policy learning with deep dynamical models
Niklas Wahlström, Thomas B Schön, and Marc Peter Deisenroth. 2015 · 2015
Cited alongside, same era.
Embed to control: A locally linear latent dynamics model for control from raw images. In Advances in neural information processing systems . 2746–2754
Manuel Watter, Jost Springenberg, Joschka Boedecker, and Martin Riedmiller. 2015 · 2015
Cited alongside, same era.
Carracing-v0
Oleg Klimov. 2016 · 2016
Cited alongside, same era.
Neuromodulation improves the evolution of forward models. In Proceedings of the Genetic and Evolutionary Computation Conference 2016 . ACM, 157–164
Mohammad Sadegh Norouzzadeh and Jeff Clune. 2016 · 2016
Evolution strategies as a scalable alternative to reinforcement learning
Tim Salimans, Jonathan Ho, Xi Chen, Szymon Sidor, and Ilya Sutskever. 2017 · 2017
Later among the works it cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017 · 2017
Later among the works it cites.
Felipe Petroski Such, Vashisht Madhavan, Edoardo Conti, Joel Lehman, Kenneth O Stanley, and Jeff Clune. 2017 · 2017
Later among the works it cites.
Neural discrete representation learning. In Advances in Neural Information Processing Systems . 6306–6315
Aaron van den Oord, Oriol Vinyals, et al · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Discrete variational autoencoders
Jason Tyler Rolfe. 2016 · 2016
Cited alongside, same era.
Autoencoder-augmented neuroevolution for visual doom playing. In Computational Intelligence and Games (CIG), 2017 IEEE Conference on . IEEE, 1–8
Samuel Alvernaz and Julian Togelius. 2017 · 2017
Cited alongside, same era.
Black-box data-efficient policy search for robotics. In 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, 51–58
Konstantinos Chatzilygeroudis, Roberto Rama, Rituraj Kaushik, Dorian Goepp, Vassilis Vassiliades, and Jean-Baptiste Mouret. 2017 · 2017
Cited alongside, same era.
Game Engine Learning from Video.. In IJCAI . 3707–3713
Matthew Guzdial, Boyang Li, and Mark O Riedl. 2017 · 2017
Cited alongside, same era.
A neural representation of sketch drawings
David Ha and Douglas Eck. 2017 · 2017
Cited alongside, same era.
Car racing with A3C
Min J. Jang, S. and C. Lee. 2017 · 2017
Cited alongside, same era.
Car racing using reinforcement learning
M. Khan and O. Elibol. 2018 · 2017
Cited alongside, same era.
Classical planning in deep latent space: Bridging the subsymbolic-symbolic boundary. In Thirty-Second AAAI Conference on Artificial Intelligence
Masataro Asai and Alex Fukunaga. 2018 · 2018
Later among the works it cites.
Quantifying generalization in reinforcement learning
Karl Cobbe, Oleg Klimov, Chris Hesse, Taehoon Kim, and John Schulman. 2018 · 2018
Later among the works it cites.
Solving OpenAI’s Car Racing Environment with Deep Reinforcement Learning and Dropout
P. Gerber, J. Guan, E. Nunez, K. Phamdo, T. Monsoor, and N. Malaya. 2018 · 2018
Later among the works it cites.
Recurrent world models facilitate policy evolution. In Advances in Neural Information Processing Systems . 2455–2467
David Ha and Jürgen Schmidhuber. 2018 · 2018
Later among the works it cites.
Illuminating Generalization in Deep Reinforcement Learning through Procedural Level Generation
Niels Justesen, Ruben Rodriguez Torrado, Philip Bontrager, Ahmed Khalifa, Julian Togelius, and Sebastian Risi. 2018 · 2018
Later among the works it cites.
An atari model zoo for analyzing, visualizing, and comparing deep reinforcement learning agents
Felipe Petroski Such, Vashisht Madhavan, Rosanne Liu, Rui Wang, Pablo Samuel Castro, Yulun Li, Ludwig Schubert, Marc Bellemare, Jeff Clune, and Joel Lehman. 2018 · 2018
Later among the works it cites.
Reimplementation of World-Models (Ha and Schmidhuber 2018) in pytorch
Corentin Tallec, Léonard Blier, and Diviyan Kalainathan. 2018a · 2018
Later among the works it cites.
Unsupervised Predictive Memory in a Goal-Directed Agent
Greg Wayne, Chia-Chun Hung, David Amos, Mehdi Mirza, Arun Ahuja, Agnieszka Grabska-Barwinska, Jack Rae, Piotr Mirowski, Joel Z Leibo, Adam Santoro, et al · 2018
Later among the works it cites.
A study on overfitting in deep reinforcement learning
Chiyuan Zhang, Oriol Vinyals, Remi Munos, and Samy Bengio. 2018 · 2018
Later among the works it cites.
Deep learning for video game playing
Niels Justesen, Philip Bontrager, Julian Togelius, and Sebastian Risi. 2019 · 2019
Closest in time.
Evolving deep neural networks
Risto Miikkulainen, Jason Liang, Elliot Meyerson, Aditya Rawal, Daniel Fink, Olivier Francon, Bala Raju, Hormoz Shahrzad, Arshak Navruzyan, Nigel Duffy, et al · 2019
Closest in time.