Leveraging procedural generation to benchmark reinforcement learning
Original
K. Cobbe, C. Hesse, J. Hilton, and J. Schulman · 1912
Earlier work this paper cites.
Probabilistic Robotics
S. Thrun, W. Burgard, and D. Fox · 2005
Earlier work this paper cites.
Optimal Control: Linear Quadratic Methods
B. Anderson and J. Moore · 2007
Earlier work this paper cites.
Causality
J. Pearl · 2009
Earlier work this paper cites.
Best response dynamics for continuous games
E. Barron, R. Goebel, and R. Jensen · 2010
Earlier work this paper cites.
Local characterizations of causal bayesian networks
E. Bareinboim, C. Brito, and J. Pearl · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Earlier work this paper cites.
Playing Atari with deep reinforcement learning
Original
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Original
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov · 2014
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Original
S. Ioffe and C. Szegedy · 2015
Earlier work this paper cites.
Causal inference by using invariant prediction: identification and confidence intervals
J. Peters, P. Bühlmann, and N. Meinshausen · 2016
Earlier work this paper cites.
Elements of Causal Inference: Foundations and Learning Algorithms
J. Peters, D. Janzing, and B. Schölkopf · 2017
Earlier work this paper cites.