Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) algorithms usually require a substantial amount of interaction data and perform well only for specific tasks in a fixed environment.
Mixture density networks
C. M. Bishop · 1994
Earlier work this paper cites.
Convergence of stochastic iterative dynamic programming algorithms
T. Jaakkola, M. I. Jordan, and S. P. Singh · 1994
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Learning from demonstration
S. Schaal · 1997
Earlier work this paper cites.
Feedforward networks with monotone constraints
H. Zhang and Z. Zhang · 1999
Earlier work this paper cites.
Causality: Models, Reasoning, and Inference
J. Pearl · 2000
Earlier work this paper cites.
Monotonic multi-layer perceptron networks as universal approximators
B. Lang · 2005
Earlier work this paper cites.
A linear non-Gaussian acyclic model for causal discovery
S. Shimizu, P.O. Hoyer, A. Hyvärinen, and A.J. Kerminen · 2006
Earlier work this paper cites.
Nonlinear causal discovery with additive noise models
P.O. Hoyer, D. Janzing, J. Mooji, J. Peters, and B. Schölkopf · 2009
Earlier work this paper cites.
On the identifiability of the post-nonlinear causal model
K. Zhang and A. Hyvärinen · 2009
Earlier work this paper cites.
Pilco: A model-based and data-efficient approach to policy search
M. Deisenroth and C. E. Rasmussen · 2011
Earlier work this paper cites.
Playing atari with deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller · 2013
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G.and ⋯ \cdots Bellemare, and S. Petersen · 2015
Cited alongside, same era.
Deep reinforcement learning with double Q-learning
H. Hasselt, A. Guez, and D. Silver · 2015
Cited alongside, same era.
Dueling network architectures for deep reinforcement learning
Z. Wang, N. de Freitas, and M. Lanctot · 2015
Cited alongside, same era.
Clustering longitudinal clinical marker trajectories from electronic health data: Applications to phenotyping and endotype discovery
P. Schulam, F. Wigley, and S. Saria · 2015
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, ⋯ \cdots , and S. Dieleman · 2016
Cited alongside, same era.
Counterfactual data-fusion for online reinforcement learners
A. Forney, J. Pearl, and E. Bareinboim · 2017
Later among the works it cites.
B. Petersen, J. Yang, W. Grathwohl, C. Cockrell, C. Santiago, G. An, and D. Faissol · 2018
Later among the works it cites.
Evaluating reinforcement learning algorithms in observational health settings
O. Gottesman, F. Johansson, J. Meier, J. Dent, D. Lee, S. Srinivasan, et al., and J. Yao · 2018
Later among the works it cites.
Model-based value estimation for efficient model-free reinforcement learning
V. Feinberg, A. Wan, I. Stoica, M. I. Jordan, J. Gonzalez, and S. Levine · 2018
Later among the works it cites.
Sample-efficient reinforcement learning with stochastic ensemble value expansion
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Duan, J. Schulman, X. Chen, P. Bartlett, I. Sutskever, and P. Abbeel · 2016
Cited alongside, same era.
Learning to reinforcement learn
J. Wang, Z. Kurth-Nelson, D. Tirumala, H. Soyer, J. Leibo, R. Munos, C. Blundell, D. Kumaran, and M. Botvinick · 2016
Cited alongside, same era.
On estimation of functional causal models: general results and application to the post-nonlinear causal model
K. Zhang, Z. Wang, J. Zhang, and B. Schölkopf · 2016
Cited alongside, same era.
Mimic-iii, a freely accessible critical care database
A. EW. Johnson, T. J. Pollard, L. Shen, H. L. Li-wei, M. Feng, M. Ghassemi, B. Moody, P. Szolovits, L. A. Celi, and R. G. Mark · 2016
Cited alongside, same era.
Convergence of q-learning: A simple proof
F. S. Melo · 2016
Cited alongside, same era.
Bidirectional conditional generative adversarial networks
A. Jaiswal, W. AbdAlmageed, Y. Wu, and P. Natarajan · 2017
Cited alongside, same era.
Deep reinforcement learning for sepsis treatment
A. Raghu, M. Komorowski, I. Ahmed, L. Celi, P. Szolovits, and M. Ghassemi · 2017
Cited alongside, same era.
J. Buckman, D. Hafner, G. Tucker, E. Brevdo, and H. Lee · 2018
Later among the works it cites.
Woulda, coulda, shoulda: Counterfactually-guided policy search
L. Buesing, T. Weber, Y. Zwols, S. Racaniere, A. Guez, J. B. Lespiau, and N. Heess · 2018
Later among the works it cites.
A simple neural attentive meta-learner
N. Mishra, M. Rohaninejad, X. Chen, and P. Abbeel · 2018
Later among the works it cites.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
K. Chua, R. Calandra, R. McAllister, and S. Levine · 2018
Later among the works it cites.
Off-policy deep reinforcement learning without exploration
Scott Fujimoto, David Meger, and Doina Precup · 2018
Later among the works it cites.
Recall traces: Backtracking models for efficient reinforcement learning
Anirudh Goyal, Philemon Brakel, William Fedus, Soumye Singhal, Timothy Lillicrap, Sergey Levine, Hugo Larochelle, and Yoshua Bengio · 2018
Later among the works it cites.
Efficient off-policy meta-reinforcement learning via probabilistic context variables
K. Rakelly, A. Zhou, D. Quillen, C. Finn, and S. Levine · 2019
Later among the works it cites.
Counterfactual off-policy evaluation with gumbel-max structural causal models
M. Oberst and D. Sontag · 2019
Later among the works it cites.