Fetching the paper…
Reading the bibliography…
We explore building generative neural network models of popular reinforcement learning environments.
Gradient theory of optimal flight paths
Kelley, H. J · 1960
Earlier work this paper cites.
The representation of the cumulative rounding error of an algorithm as a taylor expansion of the local rounding errors
Linnainmaa, S · 1970
Earlier work this paper cites.
Evolutionsstrategie: optimierung technischer systeme nach prinzipien der biologischen evolution
Rechenberg, I · 1973
Earlier work this paper cites.
Numerical Optimization of Computer Models
Schwefel, H · 1977
Earlier work this paper cites.
Applications of advances in nonlinear sensitivity analysis
Werbos, Paul J · 1982
Earlier work this paper cites.
A dual back-propagation scheme for scalar reinforcement learning
Munro, P. W · 1987
Earlier work this paper cites.
Learning how the world works: Specifications for predictive networks in robots and brains
Werbos, P. J · 1987
Earlier work this paper cites.
The truck backer-upper: An example of self learning in neural networks
Nguyen, N. and Widrow, B · 1989
Earlier work this paper cites.
Dynamic reinforcement driven error propagation networks with application to game playing
Robinson, T. and Fallside, F · 1989
Earlier work this paper cites.
Neural networks for control and system identification
Werbos, P. J · 1989
Earlier work this paper cites.
Connectionist models of recognition memory: constraints imposed by learning and forgetting functions
Ratcliff, Rodney Mark · 1990
Earlier work this paper cites.
An on-line algorithm for dynamic reinforcement learning and planning in reactive environments
Schmidhuber, J · 1990
Earlier work this paper cites.
Learning complex, extended sequences using the principle of history compression
Schmidhuber, J · 1991
Earlier work this paper cites.
Learning to generate artificial fovea trajectories for target detection
Schmidhuber, J. and Huber, R · 1991
Earlier work this paper cites.
Understanding Comics: The Invisible Art
McCloud, Scott · 1993
Earlier work this paper cites.
Mixture density networks
Bishop, Christopher M · 1994
Earlier work this paper cites.
Catastrophic interference in connectionist networks: Can it be predicted, can it be prevented?
French, Robert M · 1994
Earlier work this paper cites.
Reinforcement driven information acquisition in nondeterministic environments
Schmidhuber, J., Storck, J., and Hochreiter, S · 1994
Earlier work this paper cites.
Reinforcement learning: a survey
Kaelbling, L. P., Littman, M. L., and Moore, A. W · 1996
Earlier work this paper cites.
Long short-term memory
Hochreiter, Sepp and Schmidhuber, Juergen · 1997
Earlier work this paper cites.
Learning to forget: Continual prediction with lstm
Gers, F., Schmidhuber, J., and Cummins, F · 2000
Earlier work this paper cites.
Completely derandomized self-adaptation in evolution strategies
Hansen, Nikolaus and Ostermeier, Andreas · 2001
Earlier work this paper cites.
Akiyoshi’s illusion pages, 2002
Kitaoka, Akiyoshi · 2002
Earlier work this paper cites.
Optimal ordered problem solver
Schmidhuber, J · 2002
Earlier work this paper cites.
Evolving neural networks through augmenting topologies
Stanley, Kenneth O. and Miikkulainen, Risto · 2002
Earlier work this paper cites.
Co-evolving recurrent neurons learn deep memory pomdps
Gomez, F. and Schmidhuber, J · 2005
Earlier work this paper cites.
Invariant visual representation by single neurons in the human brain
Quiroga, R., Reddy, L., Kreiman, G., Koch, C., and Fried, I · 2005
Earlier work this paper cites.
Intrinsic motivation systems for autonomous mental development
Oudeyer, P., Kaplan, F., and Hafner, V · 2006
Earlier work this paper cites.
Developmental robotics, optimal artificial curiosity, creativity, music, and the fine arts
Schmidhuber, J · 2006
Earlier work this paper cites.
Accelerated neural evolution through cooperatively coevolved synapses
Gomez, F., Schmidhuber, J., and Miikkulainen, R · 2008
Earlier work this paper cites.
Parameter-exploring policy gradients
Sehnke, F., Osendorfer, C., Ruckstieb, T., Graves, A., Peters, J., and Schmidhuber, J · 2009
Earlier work this paper cites.
Autonomous evolution of topographic regularities in artificial neural networks
Gauci, Jason and Stanley, Kenneth O · 2010
Earlier work this paper cites.
Formal theory of creativity, fun, and intrinsic motivation (1990-2010)
Schmidhuber, J · 2010
Earlier work this paper cites.
Pilco: A model-based and data-efficient approach to policy search
Deisenroth, M. and Rasmussen, C · 2011
Earlier work this paper cites.
Abandoning objectives: Evolution through the search for novelty alone
Lehman, Joel and Stanley, Kenneth · 2011
Earlier work this paper cites.
More thoughts from understanding comics by scott mccloud, 2012
E, M · 2012
Earlier work this paper cites.
Sensorimotor mismatch signals in primary visual cortex of the behaving mouse
Keller, Georg B., Bonhoeffer, Tobias, and Hübener, Mark · 2012
Earlier work this paper cites.
Neuro-visual control in the quake ii environment
Parker, M. and Bryant, B · 2012
Earlier work this paper cites.
First experiments with powerplay
Srivastava, R., Steunebrink, B., and Schmidhuber, J · 2012
Earlier work this paper cites.
Reinforcement Learning
Wiering, Marco and van Otterlo, Martijn · 2012
Cited alongside, same era.
Motion-dependent representation of space in area mt+
Gerrit, M., Fischer, J., and Whitney, D · 2013
Cited alongside, same era.
Information-seeking, curiosity, and attention: computational and neural mechanisms
Gottlieb, J., Oudeyer, P., Lopes, M., and Baranes, A · 2013
Cited alongside, same era.
Generating sequences with recurrent neural networks
Graves, Alex · 2013
Cited alongside, same era.
A neuroevolution approach to general atari game playing
Hausknecht, M., Lehman, J., Miikkulainen, R., and Stone, P · 2013
Cited alongside, same era.
Tracking fastballs, 2013
Hirshon, B · 2013
Cited alongside, same era.
Wavenet: A generative model for raw audio
van den Oord, A., Dieleman, S., Zen, H., Simonyan, K., Vinyals, O., Graves, A., Kalchbrenner, N., Senior, A., and Kavukcuoglu, K · 2016
Later among the works it cites.
Continuous adaptation via meta-learning in nonstationary and competitive environments
Al-Shedivat, M., Bansal, T., Burda, Y., Sutskever, I., Mordatch, I., and Abbeel, P · 2017
Later among the works it cites.
Autoencoder-augmented neuroevolution for visual doom playing
Alvernaz, S. and Togelius, J · 2017
Later among the works it cites.
Deep reinforcement learning: A brief survey
Arulkumaran, K., Deisenroth, M. P., Brundage, M., and Bharath, A. A · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Auto-encoding variational bayes
Kingma, D. and Welling, M · 2013
Cited alongside, same era.
Evolving large-scale neural networks for vision-based reinforcement learning
Koutnik, J., Cuccu, G., Schmidhuber, J., and Gomez, F · 2013
Cited alongside, same era.
Evolving neural networks
Miikkulainen, R · 2013
Cited alongside, same era.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M · 2013
Cited alongside, same era.
Cortical interneurons that specialize in disinhibitory control
Pi, H., Hangya, B., Kvitsiani, D., Sanders, J., Huang, Z., and Kepecs, A · 2013
Cited alongside, same era.
Powerplay: Training an increasingly general problem solver by continually searching for the simplest still unsolvable problem
Schmidhuber, J · 2013
Cited alongside, same era.
Bansal, T., Pachocki, J., Sidor, S., Sutskever, I., and Mordatch, I · 2017
Later among the works it cites.
Using simulation and domain adaptation to improve efficiency of deep robotic grasping
Bousmalis, K., Irpan, A., Wohlhart, P., Bai, Y., Kelcey, M., Kalakrishnan, M., Downs, L., Ibarz, J., Pastor, P., Konolige, K., Levine, S., and Vanhoucke, V · 2017
Later among the works it cites.
The code for facial identity in the primate brain
Cheang, L. and Tsao, D · 2017
Later among the works it cites.
Recurrent environment simulators
Chiappa, S., Racaniere, S., Wierstra, D., and Mohamed, S · 2017
Later among the works it cites.
Unsupervised learning of disentangled representations from video
Denton, E. and Birodkar, V · 2017
Later among the works it cites.
Pathnet: Evolution channels gradient descent in super neural networks
Fernando, C., Banarse, D., Blundell, C., Zwols, Y., Ha, D., Rusu, A., Pritzel, A., and Wierstra, D · 2017
Later among the works it cites.
Model-based rl lecture at deep rl bootcamp 2017, 2017
Finn, Chelsea · 2017
Later among the works it cites.
Counterintuitive behavior of social systems, 1971
Forrester, Jay Wright · 2017
Later among the works it cites.
Replay comes of age
Foster, David J · 2017
Later among the works it cites.
Generative temporal models with memory
Gemici, M., Hung, C., Santoro, A., Wayne, G., Mohamed, S., Rezende, D., Amos, D., and Lillicrap, T · 2017
Later among the works it cites.
Recurrent neural network tutorial for artists
Ha, D · 2017
Later among the works it cites.
Evolving stable strategies
Ha, D · 2017
Later among the works it cites.
A neural representation of sketch drawings
Ha, D. and Eck, D · 2017
Later among the works it cites.
A benchmark environment motivated by industrial control problems
Hein, D., Depeweg, S., Tokic, M., Udluft, S., Hentschel, A., Runkler, T., and Sterzing, V · 2017
Later among the works it cites.
Darla: Improving zero-shot transfer in reinforcement learning
Higgins, I., Pal, A., Rusu, A., Matthey, L., Burgess, C., Pritzel, A., Botvinick, M., Blundell, C., and Lerchner, A · 2017
Later among the works it cites.
Self-driving cars in the browser, 2017
Hünermann, Jan · 2017
Later among the works it cites.
Reinforcement car racing with a3c
Jang, S., Min, J., and Lee, C · 2017
Later among the works it cites.
Car racing using reinforcement learning
Khan, M. and Elibol, O · 2017
Later among the works it cites.
A sensorimotor circuit in mouse cortex for visual flow predictions
Leinweber, Marcus, Ward, Daniel R., Sobczak, Jan M., Attinger, Alexander, and Keller, Georg B · 2017
Later among the works it cites.
Game engine learning from video
Matthew Guzdial, Boyang Li, Mark O. Riedl · 2017
Later among the works it cites.
Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuning
Nagabandi, A., Kahn, G., Fearing, R., and Levine, S · 2017
Later among the works it cites.
Curiosity-driven exploration by self-supervised prediction
Pathak, D., Agrawal, P., A., Efros, and Darrell, T · 2017
Later among the works it cites.
Deep-q learning for box2d racecar rl problem., 2017
Prieur, Luc · 2017
Later among the works it cites.
Evolution strategies as a scalable alternative to reinforcement learning
Salimans, T., Ho, J., Chen, X., Sidor, S., and Sutskever, I · 2017
Later among the works it cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q., Hinton, G., and Dean, J · 2017
Later among the works it cites.
David silver’s lecture on integrating learning and planning, 2017
Silver, David · 2017
Later among the works it cites.
Welcoming the era of deep neuroevolution, 2017
Stanley, Kenneth and Clune, Jeff · 2017
Later among the works it cites.
Language modeling with recurrent highway hypernetworks
Suarez, Joseph · 2017
Later among the works it cites.
Intrinsic motivation and automatic curricula via asymmetric self-play
Sukhbaatar, S., Lin, Z., Kostrikov, I., Synnaeve, G., Szlam, A., and Fergus, R · 2017
Later among the works it cites.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A, Kaiser, L., and Polosukhin, I · 2017
Later among the works it cites.
Watters, N., Tacchetti, A., Weber, T., Pascanu, R., Battaglia, P., and Zoran, D · 2017
Later among the works it cites.
Imagination-augmented agents for deep reinforcement learning
Weber, T., Racanière, S., Reichert, D., Buesing, L., Guez, A., Rezende, D., Badia, A., Vinyals, O., Heess, N., Li, Y., Pascanu, R., Battaglia, P., Silver, D., and Wierstra, D · 2017
Later among the works it cites.
Video game exploits, 2017
Wikipedia, Authors · 2017
Later among the works it cites.
Schmidhuber, J · 2018
Closest in time.
Illusory motion reproduced by deep neural networks trained for prediction
Watanabe, Eiji, Kitaoka, Akiyoshi, Sakamoto, Kiwako, Yasugi, Masaki, and Tanaka, Kenta · 2018
Closest in time.