Fetching the paper…
Reading the bibliography…
Much of model-based reinforcement learning involves learning a model of an agent's world, and training an agent to leverage this model to perform a task more efficiently.
Evolutionsstrategie–optimierung technisher systeme nach prinzipien der biologischen evolution
Ingo Rechenberg · 1973
Earlier work this paper cites.
Adaptation in natural and artificial systems: an introductory analysis with applications to biology, control, and artificial intelligence
John Henry Holland et al · 1975
Earlier work this paper cites.
Numerische Optimierung von Computer-Modellen mittels der Evolutionsstrategie.(Teil 1, Kap. 1-5)
H-P Schwefel · 1977
Earlier work this paper cites.
Planning using a temporal world model
James F Allen and Johannes A Koomen · 1983
Earlier work this paper cites.
Neuronlike adaptive elements that can solve difficult learning control problems
Andrew G Barto, Richard S Sutton, and Charles W Anderson · 1983
Earlier work this paper cites.
Learning how the world works: Specifications for predictive networks in robots and brains
Paul J Werbos · 1987
Earlier work this paper cites.
Genetic algorithms and machine learning
David E Goldberg and John H Holland · 1988
Earlier work this paper cites.
Making the world differentiable: On using self-supervised fully recurrent neural networks for dynamic reinforcement learning and planning in non-stationary environments
Jürgen Schmidhuber · 1990
Earlier work this paper cites.
Planning with an adaptive world model
Sebastian Thrun, Knut Möller, and Alexander Linden · 1991
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Introduction to reinforcement learning , volume 135
Richard S Sutton, Andrew G Barto, et al · 1998
Earlier work this paper cites.
Reducing the time complexity of the derandomized evolution strategy with covariance matrix adaptation (cma-es)
Nikolaus Hansen, Sibylle D Müller, and Petros Koumoutsakos · 2003
Earlier work this paper cites.
Resilient machines through continuous self-modeling
Josh Bongard, Victor Zykov, and Hod Lipson · 2006
Earlier work this paper cites.
Reducing the dimensionality of data with neural networks
Geoffrey E Hinton and Ruslan R Salakhutdinov · 2006
Earlier work this paper cites.
Sensorless but not senseless: Prediction in evolutionary car racing
Hugo Marques, Julian Togelius, Magdalena Kogutowska, Owen Holland, and Simon M Lucas · 2007
Earlier work this paper cites.
Natural evolution strategies
Daan Wierstra, Tom Schaul, Jan Peters, and Juergen Schmidhuber · 2008
Earlier work this paper cites.
Underactuated robotics: Learning, planning, and control for efficient and agile machines: Course notes for mit 6.832
Russ Tedrake · 2009
Earlier work this paper cites.
Pilco: A model-based and data-efficient approach to policy search
Marc Deisenroth and Carl E Rasmussen · 2011
Earlier work this paper cites.
Building high-level features using large scale unsupervised learning
Quoc V Le, Marc’Aurelio Ranzato, Rajat Monga, Matthieu Devin, Kai Chen, Greg S Corrado, Jeff Dean, and Andrew Y Ng · 2011
Earlier work this paper cites.
The ubiquity of model-based reinforcement learning
Bradley B Doll, Dylan A Simon, and Nathaniel D Daw · 2012
Earlier work this paper cites.
Representation learning: A review and new perspectives
Yoshua Bengio, Aaron Courville, and Pascal Vincent · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
D. Kingma and M. Welling · 2013
Earlier work this paper cites.
Kernel methods in system identification, machine learning and function estimation: A survey
Gianluigi Pillonetto, Francesco Dinuzzo, Tianshi Chen, Giuseppe De Nicolao, and Lennart Ljung · 2014
Earlier work this paper cites.
Stochastic backpropagation and approximate inference in deep generative models
D. Rezende, S. Mohamed, and D. Wierstra · 2014
Cited alongside, same era.
Dropout: a simple way to prevent neural networks from overfitting
N Srivastava, G Hinton, A Krizhevsky, I Sutskever, and R Salakhutdinov · 2014
Cited alongside, same era.
Model regularization for stable sample rollouts
Erik Talvitie · 2014
Cited alongside, same era.
Deepmpc: Learning deep latent features for model predictive control
Ian Lenz, Ross A Knepper, and Ashutosh Saxena · 2015
Cited alongside, same era.
Deep multi-scale video prediction beyond mean square error
Michael Mathieu, Camille Couprie, and Yann LeCun · 2015
Cited alongside, same era.
Felipe Petroski Such, Vashisht Madhavan, Edoardo Conti, Joel Lehman, Kenneth O Stanley, and Jeff Clune · 2017
Later among the works it cites.
Differentiable mpc for end-to-end planning and control
Brandon Amos, Ivan Jimenez, Jacob Sacks, Byron Boots, and J Zico Kolter · 2018
Later among the works it cites.
Lipschitz continuity in model-based reinforcement learning
Kavosh Asadi, Dipendra Misra, and Michael L Littman · 2018
Later among the works it cites.
Vector-based navigation using grid-like representations in artificial agents
Andrea Banino, Caswell Barry, Benigno Uria, Charles Blundell, Timothy Lillicrap, Piotr Mirowski, Alexander Pritzel, Martin J Chadwick, Thomas Degris, Joseph Modayil, et al · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
Action-conditional video prediction using deep networks in atari games
Junhyuk Oh, Xiaoxiao Guo, Honglak Lee, Richard L Lewis, and Satinder Singh · 2015
Cited alongside, same era.
Spatio-temporal video autoencoder with differentiable memory
Viorica Patraucean, Ankur Handa, and Roberto Cipolla · 2015
Cited alongside, same era.
J. Schmidhuber · 2015
Cited alongside, same era.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Cited alongside, same era.
Unsupervised learning of video representations using lstms
Nitish Srivastava, Elman Mansimov, and Ruslan Salakhutdinov · 2015
Cited alongside, same era.
Embed to control: A locally linear latent dynamics model for control from raw images
Manuel Watter, Jost Springenberg, Joschka Boedecker, and Martin Riedmiller · 2015
Cited alongside, same era.
Christopher J Cueva and Xue-Xin Wei · 2018
Later among the works it cites.
Stochastic video generation with a learned prior
Emily Denton and Rob Fergus · 2018
Later among the works it cites.
Visual foresight: Model-based deep reinforcement learning for vision-based robotic control
Frederik Ebert, Chelsea Finn, Sudeep Dasari, Annie Xie, Alex Lee, and Sergey Levine · 2018
Later among the works it cites.
Cmpsci embedded systems 503
Roderic A. Grupen · 2018
Later among the works it cites.
Reinforcement learning for improving agent design
David Ha · 2018
Later among the works it cites.
Recurrent World Models Facilitate Policy Evolution
David Ha and Jürgen Schmidhuber · 2018
Later among the works it cites.
Learning latent dynamics for planning from pixels
Danijar Hafner, Timothy Lillicrap, Ian Fischer, Ruben Villegas, David Ha, Honglak Lee, and James Davidson · 2018
Later among the works it cites.
Towards a definition of disentangled representations
Irina Higgins, David Amos, David Pfau, Sebastien Racaniere, Loic Matthey, Danilo Rezende, and Alexander Lerchner · 2018
Later among the works it cites.
Joel Lehman, Jeff Clune, Dusan Misevic, Christoph Adami, Lee Altenberg, Julie Beaulieu, Peter J Bentley, Samuel Bernard, Guillaume Beslon, David M Bryson, et al · 2018
Later among the works it cites.
Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuning
Anusha Nagabandi, Gregory Kahn, Ronald S Fearing, and Sergey Levine · 2018
Later among the works it cites.
Learning image-conditioned dynamics models for control of underactuated legged millirobots
Anusha Nagabandi, Guangzhao Yang, Thomas Asmar, Ravi Pandya, Gregory Kahn, Sergey Levine, and Ronald S Fearing · 2018
Later among the works it cites.
Representation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Later among the works it cites.
Time-contrastive networks: Self-supervised learning from video
Pierre Sermanet, Corey Lynch, Yevgen Chebotar, Jasmine Hsu, Eric Jang, Stefan Schaal, and Sergey Levine · 2018
Later among the works it cites.
PyTorch implementation of Improving PILCO with Bayesian neural network dynamics models, 2018
Xingdong Zuo · 2018
Later among the works it cites.
Weight agnostic neural networks
Adam Gaier and David Ha · 2019
Closest in time.
Model-based reinforcement learning for atari
Lukasz Kaiser, Mohammad Babaeizadeh, Piotr Milos, Blazej Osinski, Roy H Campbell, Konrad Czechowski, Dumitru Erhan, Chelsea Finn, Piotr Kozakowski, Sergey Levine, et al · 2019
Closest in time.
Videoflow: A flow-based generative model for video
M Kumar, M Babaeizadeh, D Erhan, C Finn, S Levine, L Dinh, and D Kingma · 2019
Closest in time.
Deep neuroevolution of recurrent and discrete world models
Sebastian Risi and Kenneth O. Stanley · 2019
Closest in time.