Fetching the paper…
Reading the bibliography…
Learning a policy using only observational data is challenging because the distribution of states it induces at execution time may differ from the distribution observed during training.
The truck backer-upper: an example of self-learning in neural networks
Nguyen and Widrow · 1989
Earlier work this paper cites.
Neural networks for control
Derrick Nguyen and Bernard Widrow · 1990
Earlier work this paper cites.
An on-line algorithm for dynamic reinforcement learning and planning in reactive environments
Jurgen Schmidhuber · 1990
Earlier work this paper cites.
Efficient training of artificial neural networks for autonomous navigation
Dean A. Pomerleau · 1991
Earlier work this paper cites.
Forward models: Supervised learning with a distal teacher
Michael I. Jordan and David E. Rumelhart · 1992
Earlier work this paper cites.
Bayesian Learning for Neural Networks
Radford M. Neal · 1995
Earlier work this paper cites.
A comparison of direct and model-based reinforcement learning
C. G. Atkeson and J. C. Santamaria · 1997
Earlier work this paper cites.
An introduction to variational methods for graphical models
Michael I. Jordan, Zoubin Ghahramani, Tommi S. Jaakkola, and Lawrence K. Saul · 1999
Earlier work this paper cites.
NGSIM interstate 80 freeway dataset
John Halkias and James Colyar · 2006
Earlier work this paper cites.
Off-road obstacle avoidance through end-to-end learning
Yann LeCun, Urs Muller, Jan Ben, Eric Cosatto, and Beat Flepp · 2006
Earlier work this paper cites.
Efficient reductions for imitation learning
Stéphane Ross and J. Andrew Bagnell · 2010
Earlier work this paper cites.
Pilco: A model-based and data-efficient approach to policy search
Marc Peter Deisenroth and Carl Edward Rasmussen · 2011
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Stéphane Ross, Geoffrey J. Gordon, and Drew Bagnell · 2011
Earlier work this paper cites.
Improving neural networks by preventing co-adaptation of feature detectors
Geoffrey E. Hinton, Nitish Srivastava, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2012
Earlier work this paper cites.
Model-based imitation learning by probabilistic trajectory matching
P. Englert, A. Paraschos, J. Peters, and M. P. Deisenroth · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P. Kingma and Max Welling · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Dropout: A simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Cited alongside, same era.
Learning continuous control policies by stochastic value gradients
Nicolas Heess, Gregory Wayne, David Silver, Tim Lillicrap, Tom Erez, and Yuval Tassa · 2015
Cited alongside, same era.
Continuous control with deep reinforcement learning
Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adrià Puigdomènech Badia, Mehdi Mirza, Alex Graves, Timothy P. Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Later among the works it cites.
Planning for autonomous cars that leverage effects on human actions, 06 2016
Dorsa Sadigh, Shankar Sastry, Sanjit A. Seshia, and Anca D. Dragan · 2016
Later among the works it cites.
Query-efficient imitation learning for end-to-end autonomous driving
Jiakai Zhang and Kyunghyun Cho · 2016
Later among the works it cites.
Stochastic variational video prediction
Mohammad Babaeizadeh, Chelsea Finn, Dumitru Erhan, Roy H. Campbell, and Sergey Levine · 2017
Later among the works it cites.
End-to-end differentiable adversarial imitation learning
Nir Baram, Oron Anschel, Itai Caspi, and Shie Mannor · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Language understanding for text-based games using deep reinforcement learning
Karthik Narasimhan, Tejas D. Kulkarni, and Regina Barzilay · 2015
Cited alongside, same era.
Action-conditional video prediction using deep networks in atari games
Junhyuk Oh, Xiaoxiao Guo, Honglak Lee, Richard L. Lewis, and Satinder P. Singh · 2015
Cited alongside, same era.
Learning simple algorithms from examples
Wojciech Zaremba, Tomas Mikolov, Armand Joulin, and Rob Fergus · 2015
Cited alongside, same era.
Towards vision-based deep reinforcement learning for robotic motion control
Fangyi Zhang, Jürgen Leitner, Michael Milford, Ben Upcroft, and Peter I. Corke · 2015
Cited alongside, same era.
Learning to poke by poking: Experiential learning of intuitive physics
Pulkit Agrawal, Ashvin Nair, Pieter Abbeel, Jitendra Malik, and Sergey Levine · 2016
Cited alongside, same era.
End to end learning for self-driving cars
Mariusz Bojarski, Davide Del Testa, Daniel Dworakowski, Bernhard Firner, Beat Flepp, Prasoon Goyal, Lawrence D. Jackel, Mathew Monfort, Urs Muller, Jiakai Zhang, Xin Zhang, Jake Zhao, and Karol Zieba · 2016
Cited alongside, same era.
Openai gym, 2016
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Cited alongside, same era.
Later among the works it cites.
Uncertainty-aware reinforcement learning for collision avoidance
Gregory Kahn, Adam Villaflor, Vitchyr Pong, Pieter Abbeel, and Sergey Levine · 2017
Later among the works it cites.
Dropout inference in Bayesian neural networks with alpha-divergences
Yingzhen Li and Yarin Gal · 2017
Later among the works it cites.
Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuning
Anusha Nagabandi, Gregory Kahn, Ronald S. Fearing, and Sergey Levine · 2017
Later among the works it cites.
Agile off-road autonomous driving using end-to-end deep imitation learning
Yunpeng Pan, Ching-An Cheng, Kamil Saigol, Keuntaek Lee, Xinyan Yan, Evangelos Theodorou, and Byron Boots · 2017
Later among the works it cites.
Learning model-based planning from scratch
Razvan Pascanu, Yujia Li, Oriol Vinyals, Nicolas Heess, Lars Buesing, Sébastien Racanière, David P. Reichert, Theophane Weber, Daan Wierstra, and Peter Battaglia · 2017
Later among the works it cites.
Imagination-augmented agents for deep reinforcement learning
Theophane Weber, Sébastien Racanière, David P. Reichert, Lars Buesing, Arthur Guez, Danilo Jimenez Rezende, Adrià Puigdomènech Badia, Oriol Vinyals, Nicolas Heess, Yujia Li, Razvan Pascanu, Peter Battaglia, David Silver, and Daan Wierstra · 2017
Later among the works it cites.
Information theoretic mpc for model-based reinforcement learning
G. Williams, N. Wagener, B. Goldfain, P. Drews, J. M. Rehg, B. Boots, and E. A. Theodorou · 2017
Later among the works it cites.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
Kurtland Chua, Roberto Calandra, Rowan McAllister, and Sergey Levine · 2018
Later among the works it cites.
Stochastic video generation with a learned prior
Emily Denton and Rob Fergus · 2018
Later among the works it cites.
Decomposition of uncertainty in Bayesian deep learning for efficient and risk-sensitive learning
Stefan Depeweg, Jose-Miguel Hernandez-Lobato, Finale Doshi-Velez, and Steffen Udluft · 2018
Later among the works it cites.
Aravind Srinivas, Allan Jabri, Pieter Abbeel, Sergey Levine, and Chelsea Finn · 2018
Later among the works it cites.