Backpropagation Applied to Handwritten Zip Code Recognition
Y LeCun, B Boser, JS Denker, D Henderson, RE Howard, W Hubbard, LD Jackel · 1989
Earlier work this paper cites.
Dyna, an Integrated Architecture for Learning, Planning, and Reacting
RS Sutton · 1991
Earlier work this paper cites.
Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning
RJ Williams · 1992
Earlier work this paper cites.
Efficient Selectivity and Backup Operators in Monte-Carlo Tree Search
R Coulom · 2006
Earlier work this paper cites.
Solar: Deep Structured Representations for Model-Based Reinforcement Learning
M Zhang, S Vikram, L Smith, P Abbeel, M Johnson, S Levine · 2010
Earlier work this paper cites.
The Arcade Learning Environment: An Evaluation Platform for General Agents
MG Bellemare, Y Naddaf, J Veness, M Bowling · 2013
Earlier work this paper cites.
Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation
Original
Y Bengio, N Léonard, A Courville · 2013
Earlier work this paper cites.
Auto-Encoding Variational Bayes
Original
DP Kingma M Welling · 2013
Earlier work this paper cites.
Learning Phrase Representations Using Rnn Encoder-Decoder for Statistical Machine Translation
Original
K Cho, B Van Merriënboer, C Gulcehre, D Bahdanau, F Bougares, H Schwenk, Y Bengio · 2014
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Original
DP Kingma J Ba · 2014
Earlier work this paper cites.
Stochastic Backpropagation and Approximate Inference in Deep Generative Models
Original
DJ Rezende, S Mohamed, D Wierstra · 2014
Earlier work this paper cites.
Fast and Accurate Deep Network Learning by Exponential Linear Units (Elus)
Original
DA Clevert, T Unterthiner, S Hochreiter · 2015
Earlier work this paper cites.
Deep Kalman Filters
Original
RG Krishnan, U Shalit, D Sontag · 2015
Earlier work this paper cites.
Human-Level Control Through Deep Reinforcement Learning
V Mnih, K Kavukcuoglu, D Silver, AA Rusu, J Veness, MG Bellemare, A Graves, M Riedmiller, AK Fidjeland, G Ostrovski, et al · 2015
Earlier work this paper cites.
Action-Conditional Video Prediction Using Deep Networks in Atari Games
J Oh, X Guo, H Lee, RL Lewis, S Singh · 2015
Earlier work this paper cites.
Prioritized Experience Replay
Original
T Schaul, J Quan, I Antonoglou, D Silver · 2015
Earlier work this paper cites.
High-Dimensional Continuous Control Using Generalized Advantage Estimation
Original
J Schulman, P Moritz, S Levine, M Jordan, P Abbeel · 2015
Earlier work this paper cites.
Deep Reinforcement Learning With Double Q-Learning
Original
H Van Hasselt, A Guez, D Silver · 2015
Earlier work this paper cites.
Embed to Control: A Locally Linear Latent Dynamics Model for Control From Raw Images
M Watter, J Springenberg, J Boedecker, M Riedmiller · 2015
Earlier work this paper cites.
Openai Gym, 2016
G Brockman, V Cheung, L Pettersson, J Schneider, J Schulman, J Tang, W Zaremba · 2016
Earlier work this paper cites.
Improving Pilco With Bayesian Neural Network Dynamics Models
Y Gal, R McAllister, CE Rasmussen · 2016
Earlier work this paper cites.
Beta-Vae: Learning Basic Visual Concepts With a Constrained Variational Framework
I Higgins, L Matthey, A Pal, C Burgess, X Glorot, M Botvinick, S Mohamed, A Lerchner · 2016
Earlier work this paper cites.
Deep Variational Bayes Filters: Unsupervised Learning of State Space Models From Raw Data
Original
M Karl, M Soelch, J Bayer, P van der Smagt · 2016
Earlier work this paper cites.
Asynchronous Methods for Deep Reinforcement Learning
V Mnih, AP Badia, M Mirza, A Graves, T Lillicrap, T Harley, D Silver, K Kavukcuoglu · 2016
Earlier work this paper cites.
Dueling Network Architectures for Deep Reinforcement Learning
Z Wang, T Schaul, M Hessel, H Hasselt, M Lanctot, N Freitas · 2016
Earlier work this paper cites.
Stochastic Variational Video Prediction
Original
M Babaeizadeh, C Finn, D Erhan, RH Campbell, S Levine · 2017
Earlier work this paper cites.
A Distributional Perspective on Reinforcement Learning
Original
MG Bellemare, W Dabney, R Munos · 2017
Earlier work this paper cites.