Fetching the paper…
Reading the bibliography…
Action planning using learned and differentiable forward models of the world is a general approach which has a number of desirable properties, including improved sample complexity over model-free RL methods, reuse of learned models across different tasks, and the ability to perform efficient gradient-based optimization in continuous action spaces.
Dynamic Programming
Bellman, Richard · 1957
Earlier work this paper cites.
An on-line algorithm for dynamic reinforcement learning and planning in reactive environments
Kelley, Henry · 1960
Earlier work this paper cites.
The numerical solution of variational problems
Dreyfus, Stuart · 1962
Earlier work this paper cites.
Neural networks for control
Nguyen, Derrick and Widrow, Bernard · 1990
Earlier work this paper cites.
An on-line algorithm for dynamic reinforcement learning and planning in reactive environments
Schmidhuber, Jurgen · 1990
Earlier work this paper cites.
Efficient training of artificial neural networks for autonomous navigation
Pomerleau, Dean A · 1991
Earlier work this paper cites.
Dyna, an integrated architecture for learning, planning, and reacting
Sutton, Richard S · 1991
Earlier work this paper cites.
Forward models: Supervised learning with a distal teacher
Jordan, Michael I. and Rumelhart, David E · 1992
Earlier work this paper cites.
Crafting papers on machine learning
Langley, P · 2000
Earlier work this paper cites.
A generalized iterative lqg method for locally-optimal feedback control of constrained nonlinear stochastic systems
Todorov, Emanuel and Li, Weiwei · 2005
Earlier work this paper cites.
An application of reinforcement learning to aerobatic helicopter flight
Abbeel, Pieter, Coates, Adam, Quigley, Morgan, and Ng, Andrew Y · 2007
Earlier work this paper cites.
Efficient selectivity and backup operators in monte-carlo tree search
Coulom, Rémi · 2007
Earlier work this paper cites.
Using a monte-carlo approach for bus regulation
Cazenave, Tristan, Balbo, Flavien, and Pinson, Suzanne · 2009
Cited alongside, same era.
On the scalability of parallel uct
Segal, Richard B · 2011
Cited alongside, same era.
A survey of monte carlo tree search methods
Browne, Cameron, Powley, Edward, Whitehouse, Daniel, Lucas, Simon, Cowling, Peter I., Tavener, Stephen, Perez, Diego, Samothrakis, Spyridon, Colton, Simon, and et al · 2012
Cited alongside, same era.
Guiding combinatorial optimization with uct
Sabharwal, Ashish, Samulowitz, Horst, and Reddy, Chandra · 2012
Cited alongside, same era.
Mujoco: A physics engine for model-based control
Todorov, Emanuel, Erez, Tom, and Tassa, Yuval · 2012
Cited alongside, same era.
Deep learning for real-time atari game play using offline monte-carlo tree search planning
Guo, Xiaoxiao, Singh, Satinder, Lee, Honglak, Lewis, Richard L, and Wang, Xiaoshi · 2014
Deep reinforcement learning for robotic manipulation
Gu, Shixiang, Holly, Ethan, Lillicrap, Timothy P., and Levine, Sergey · 2016
Later among the works it cites.
Optimal control with learned local models: Application to dexterous manipulation
Kumar, Vikash, Todorov, Emanuel, and Levine, Sergey · 2016
Later among the works it cites.
The concrete distribution: A continuous relaxation of discrete random variables
Maddison, Chris J., Mnih, Andriy, and Teh, Yee Whye · 2016
Later among the works it cites.
Asynchronous methods for deep reinforcement learning
Mnih, Volodymyr, Badia, Adrià Puigdomènech, Mirza, Mehdi, Graves, Alex, Lillicrap, Timothy P., Harley, Tim, Silver, David, and Kavukcuoglu, Koray · 2016
Later among the works it cites.
Control of memory, active perception, and action in minecraft
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Don’t until the final verb wait: Reinforcement learning for simultaneous machine translation
Ii, Alvin C. Grissom, He, He, Morgan, John, and III, Hal Daume · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, Diederik P. and Ba, Jimmy · 2014
Cited alongside, same era.
Distilling the knowledge in a neural network
Hinton, Geoffrey E., Vinyals, Oriol, and Dean, Jeffrey · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, Volodymyr, Kavukcuoglu, Koray, Silver, David, Rusu, Andrei A., Veness, Joel, Bellemare, Marc G., Graves, Alex, Riedmiller, Martin, Fidjeland, Andreas K., Ostrovski, Georg, Petersen, Stig, Beattie, Charles, Sadik, Amir, Antonoglou, Ioannis, King, Helen, Kumaran, Dharshan, Wierstra, Daan, Legg, Shane, and Hassabis, Demis · 2015
Cited alongside, same era.
Rusu, Andrei A., Colmenarejo, Sergio Gomez, Gülçehre, Çaglar, Desjardins, Guillaume, Kirkpatrick, James, Pascanu, Razvan, Mnih, Volodymyr, Kavukcuoglu, Koray, and Hadsell, Raia · 2015
Cited alongside, same era.
Openai gym, 2016
Brockman, Greg, Cheung, Vicki, Pettersson, Ludwig, Schneider, Jonas, Schulman, John, Tang, Jie, and Zaremba, Wojciech · 2016
Cited alongside, same era.
Oh, Junhyuk, Chockalingam, Valliappa, Singh, Satinder P., and Lee, Honglak · 2016
Later among the works it cites.
Mastering the game of Go with deep neural networks and tree search
Silver, David, Huang, Aja, Maddison, Chris J., Guez, Arthur, Sifre, Laurent, van den Driessche, George, Schrittwieser, Julian, Antonoglou, Ioannis, Panneershelvam, Veda, Lanctot, Marc, Dieleman, Sander, Grewe, Dominik, Nham, John, Kalchbrenner, Nal, Sutskever, Ilya, Lillicrap, Timothy, Leach, Madeleine, Kavukcuoglu, Koray, Graepel, Thore, and Hassabis, Demis · 2016
Later among the works it cites.
Tamar, Aviv, Levine, Sergey, and Abbeel, Pieter · 2016
Later among the works it cites.
Metacontrol for adaptive imagination-based optimization
Hamrick, Jessica, Ballard, Andrew, Pascanu, Razvan, Vinyals, Oriol, Heess, Nicolas, and Battaglia, Peter · 2017
Closest in time.
Deep reinforcement learning that matters
Henderson, Peter, Islam, Riashat, Bachman, Philip, Pineau, Joelle, Precup, Doina, and Meger, David · 2017
Closest in time.
Categorical reparameterization with gumbel-softmax
Jang, Eric, Gu, Shixiang, and Poole, Ben · 2017
Closest in time.
Learning model-based planning from scratch
Pascanu, Razvan, Li, Yujia, Vinyals, Oriol, Heess, Nicolas, Buesing, Lars, Racanière, Sébastien, Reichert, David P., Weber, Theophane, Wierstra, Daan, and Battaglia, Peter · 2017
Closest in time.
Imagination-augmented agents for deep reinforcement learning
Weber, Theophane, Racanière, Sébastien, Reichert, David P., Buesing, Lars, Guez, Arthur, Rezende, Danilo Jimenez, Badia, Adrià Puigdomènech, Vinyals, Oriol, Heess, Nicolas, Li, Yujia, Pascanu, Razvan, Battaglia, Peter, Silver, David, and Wierstra, Daan · 2017
Closest in time.