Fetching the paper…
Reading the bibliography…
Semi-markovian decision processes
R. Howard · 1964
Earlier work this paper cites.
Linear optimal control systems , volume 1
H. Kwakernaak and R. Sivan · 1972
Earlier work this paper cites.
Reinforcement learning in continuous time: Advantage updating
L. C. Baird · 1994
Earlier work this paper cites.
Reinforcement learning methods for continuous-time markov decision problems
S. J. Bradtke and M. O. Duff · 1995
Earlier work this paper cites.
Reinforcement learning: A survey
L. P. Kaelbling, M. L. Littman, and A. W. Moore · 1996
Earlier work this paper cites.
Generalization in reinforcement learning: Successful examples using sparse coarse coding
R. S. Sutton · 1996
Earlier work this paper cites.
Hierarchical control and learning for Markov decision processes
R. E. Parr and S. Russell · 1998
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
R. S. Sutton, D. Precup, and S. Singh · 1999
Earlier work this paper cites.
Reinforcement learning in continuous time and space
K. Doya · 2000
Earlier work this paper cites.
Constrained model predictive control: Stability and optimality
D. Q. Mayne, J. B. Rawlings, C. V. Rao, and P. O. Scokaert · 2000
Earlier work this paper cites.
Learning precise timing with lstm recurrent networks
F. A. Gers, N. N. Schraudolph, and J. Schmidhuber · 2002
Earlier work this paper cites.
Recent advances in hierarchical reinforcement learning
A. G. Barto and S. Mahadevan · 2003
Earlier work this paper cites.
Dynamic multidrug therapies for hiv: Optimal and sti control approaches
B. M. Adams, H. T. Banks, H.-D. Kwon, and H. T. Tran · 2004
Earlier work this paper cites.
Clinical data based optimal sti strategies for hiv: a reinforcement learning approach
D. Ernst, G.-B. Stan, J. Goncalves, and L. Wehenkel · 2006
Earlier work this paper cites.
Policy gradient in continuous time
R. Munos · 2006
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Earlier work this paper cites.
Auto-encoding variational bayes
D. P. Kingma and M. Welling · 2013
Earlier work this paper cites.
Model-based hierarchical reinforcement learning and human action control
M. Botvinick and A. Weinstein · 2014
Earlier work this paper cites.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
K. Cho, B. Van Merriënboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio · 2014
Earlier work this paper cites.
Stochastic backpropagation and approximate inference in deep generative models
D. J. Rezende, S. Mohamed, and D. Wierstra · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
I. Sutskever, O. Vinyals, and Q. V. Le · 2014
Cited alongside, same era.
Scheduled sampling for sequence prediction with recurrent neural networks
S. Bengio, O. Vinyals, N. Jaitly, and N. Shazeer · 2015
Cited alongside, same era.
Continuous control with deep reinforcement learning
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Cited alongside, same era.
Action-conditional video prediction using deep networks in atari games
J. Oh, X. Guo, H. Lee, R. L. Lewis, and S. Singh · 2015
Cited alongside, same era.
Deep variational reinforcement learning for pomdps
M. Igl, L. Zintgraf, T. A. Le, F. Wood, and S. Whiteson · 2018
Later among the works it cites.
Model-ensemble trust-region policy optimization
T. Kurutach, I. Clavera, Y. Duan, A. Tamar, and P. Abbeel · 2018
Later among the works it cites.
Plan online, learn offline: Efficient learning and exploration via model-based control
K. Lowrey, A. Rajeswaran, S. Kakade, E. Todorov, and I. Mordatch · 2018
Later among the works it cites.
Adaptive skip intervals: Temporal abstraction for recurrent dynamical models
A. Neitz, G. Parascandolo, S. Bauer, and B. Schölkopf · 2018
Later among the works it cites.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Schaul, J. Quan, I. Antonoglou, and D. Silver · 2015
Cited alongside, same era.
J. Schmidhuber · 2015
Cited alongside, same era.
Openai gym, 2016
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Cited alongside, same era.
Recurrent environment simulators
S. Chiappa, S. Racaniere, D. Wierstra, and S. Mohamed · 2017
Cited alongside, same era.
Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates
S. Gu, E. Holly, T. Lillicrap, and S. Levine · 2017
Cited alongside, same era.
Robust and efficient transfer learning with hidden parameter markov decision processes
T. W. Killian, S. Daulton, G. Konidaris, and F. Doshi-Velez · 2017
Cited alongside, same era.
An efficient approach to model-based hierarchical reinforcement learning
Z. Li, A. Narayan, and T.-Y. Leong · 2017
Cited alongside, same era.
Gru-ode-bayes: Continuous modeling of sporadically-observed time series
E. De Brouwer, J. Simm, A. Arany, and Y. Moreau · 2019
Later among the works it cites.
Model-based lookahead reinforcement learning
Z.-W. Hong, J. Pajarinen, and J. Peters · 2019
Later among the works it cites.
When to trust your model: Model-based policy optimization
M. Janner, J. Fu, M. Zhang, and S. Levine · 2019
Later among the works it cites.
Latent odes for irregularly-sampled time series
Y. Rubanova, R. T. Chen, and D. Duvenaud · 2019
Later among the works it cites.
Making deep q-learning methods robust to time discretization
C. Tallec, L. Blier, and Y. Ollivier · 2019
Later among the works it cites.
Exploring model-based planning with policy networks
T. Wang and J. Ba · 2019
Later among the works it cites.
Benchmarking model-based reinforcement learning
T. Wang, X. Bao, I. Clavera, J. Hoang, Y. Wen, E. Langlois, S. Zhang, G. Zhang, P. Abbeel, and J. Ba · 2019
Later among the works it cites.
Model-augmented actor-critic: Backpropagating through paths
I. Clavera, V. Fu, and P. Abbeel · 2020
Closest in time.
C. Finlay, J.-H. Jacobsen, L. Nurbekyan, and A. M. Oberman · 2020
Closest in time.
Interpretable off-policy evaluation in reinforcement learning by highlighting influential transitions, 2020
O. Gottesman, J. Futoma, Y. Liu, S. Parbhoo, L. A. Celi, E. Brunskill, and F. Doshi-Velez · 2020
Closest in time.
Learning differential equations that are easy to solve
J. Kelly, J. Bettencourt, M. J. Johnson, and D. Duvenaud · 2020
Closest in time.
Morel: Model-based offline reinforcement learning
R. Kidambi, A. Rajeswaran, P. Netrapalli, and T. Joachims · 2020
Closest in time.
Neural Controlled Differential Equations for Irregular Time Series
P. Kidger, J. Morrill, J. Foster, and T. Lyons · 2020
Closest in time.
Mopo: Model-based offline policy optimization
T. Yu, G. Thomas, L. Yu, S. Ermon, J. Zou, S. Levine, C. Finn, and T. Ma · 2020
Closest in time.