Fetching the paper…
Reading the bibliography…
Classical value iteration approaches are not applicable to environments with continuous states and actions.
Dynamic Programming
Bellman, R · 1957
Earlier work this paper cites.
Optimal control theory: an introduction
Kirk, D. E · 1970
Earlier work this paper cites.
Practical issues in temporal difference learning
Tesauro, G · 1992
Earlier work this paper cites.
Generalization in reinforcement learning: Safely approximating the value function
Boyan, J. and Moore, A · 1994
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Puterman, M. L · 1994
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
Baird, L · 1995
Earlier work this paper cites.
Learning to act using real-time dynamic programming
Barto, A. G., Bradtke, S. J., and Singh, S. P · 1995
Earlier work this paper cites.
Feature-based methods for large scale dynamic programming
Tsitsiklis, J. N. and Van Roy, B · 1996
Earlier work this paper cites.
Optimal control of nonlinear continuous-time systems: design of bounded controllers via generalized nonquadratic functionals
Lyshevski, S. E · 1998
Earlier work this paper cites.
Introduction to reinforcement learning , volume 135
Sutton, R. S., Barto, A. G., et al · 1998
Earlier work this paper cites.
Differential games: a mathematical theory with applications to warfare and pursuit, control and optimization
Isaacs, R · 1999
Earlier work this paper cites.
Reinforcement learning in continuous time and space
Doya, K · 2000
Earlier work this paper cites.
Nonlinear systems , volume 3
Khalil, H. K. and Grizzle, J. W · 2002
Earlier work this paper cites.
Tree-based batch mode reinforcement learning
Ernst, D., Geurts, P., and Wehenkel, L · 2005
Earlier work this paper cites.
Linear theory for control of nonlinear stochastic systems
Kappen, H. J · 2005
Earlier work this paper cites.
Neural fitted q iteration–first experiences with a data efficient neural reinforcement learning method
Riedmiller, M · 2005
Earlier work this paper cites.
Least squares solutions of the hjb equation with neural network value-function approximators
Tassa, Y. and Erez, T · 2007
Earlier work this paper cites.
Linearly-solvable markov decision problems
Todorov, E · 2007
Earlier work this paper cites.
Event based control
Aström, K. J · 2008
Earlier work this paper cites.
Finite-time bounds for fitted value iteration
Munos, R. and Szepesvári, C · 2008
Earlier work this paper cites.
Regularized fitted q-iteration for planning in continuous-space markovian decision problems
Farahmand, A. M., Ghavamzadeh, M., Szepesvári, C., and Mannor, S · 2009
Earlier work this paper cites.
Reinforcement learning of motor skills in high dimensions: A path integral approach
Theodorou, E., Buchli, J., and Schaal, S · 2010
Cited alongside, same era.
Optimal reinforcement learning for gaussian systems
Hennig, P · 2011
Cited alongside, same era.
Neural-network-based online hjb solution for optimal robust guaranteed cost control of continuous-time uncertain nonlinear systems
Liu, D., Wang, D., Wang, F.-Y., Li, H., and Yang, X · 2014
Cited alongside, same era.
Model-based path integral stochastic control: A bayesian nonparametric approach
Pan, Y., Theodorou, E. A., and Kontitsis, M · 2014
Cited alongside, same era.
Reinforcement learning for adaptive optimal control of unknown continuous-time nonlinear systems with input constraints
Yang, X., Liu, D., and Wang, D · 2014
Cited alongside, same era.
The lyapunov neural network: Adaptive stability certification for safe learning of dynamical systems
Richards, S. M., Berkenkamp, F., and Krause, A · 2018
Later among the works it cites.
Neural lyapunov control
Chang, Y.-C., Roohi, N., and Gao, S · 2019
Later among the works it cites.
Closing the sim-to-real loop: Adapting simulation randomization with real world experience, 2019
Chebotar, Y., Handa, A., Makoviychuk, V., Macklin, M., Issac, J., Ratliff, N., and Fox, D · 2019
Later among the works it cites.
Learning stable deep dynamics models
Kolter, J. Z. and Manek, G · 2019
Later among the works it cites.
HJB optimal feedback control with deep differential value functions and action constraints
Lutter, M., Belousov, B., Listmann, K., Clever, D., and Peters, J · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Cited alongside, same era.
High-dimensional continuous control using generalized advantage estimation
Schulman, J., Moritz, P., Levine, S., Jordan, M., and Abbeel, P · 2015
Cited alongside, same era.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Cited alongside, same era.
Neural network-based solutions for stochastic optimal control using path integrals
Rajagopal, K., Balakrishnan, S. N., and Busemeyer, J. R · 2016
Cited alongside, same era.
Hamilton-Jacobi reachability: A brief overview and recent advances
Bansal, S., Chen, M., Herbert, S., and Tomlin, C. J · 2017
Cited alongside, same era.
Safe model-based reinforcement learning with stability guarantees
Berkenkamp, F., Turchetta, M., Schoellig, A., and Krause, A · 2017
Cited alongside, same era.
Scaling robot supervision to hundreds of hours with roboturk: Robotic manipulation dataset through human reasoning and dexterity
Mandlekar, A., Booher, J., Spero, M., Tung, A., Gupta, A., Zhu, Y., Garg, A., Savarese, S., and Fei-Fei, L · 2019
Later among the works it cites.
Learning dexterous in-hand manipulation, 2019
OpenAI, Andrychowicz, M., Baker, B., Chociej, M., Jozefowicz, R., McGrew, B., Pachocki, J., Petron, A., Plappert, M., Powell, G., Ray, A., Schneider, J., Sidor, S., Tobin, J., Welinder, P., Weng, L., and Zaremba, W · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S · 2019
Later among the works it cites.
Learning deep stochastic optimal control policies using forward-backward sdes
Pereira, M., Wang, Z., Exarchos, I., and Theodorou, E · 2019
Later among the works it cites.
Bayessim: adaptive domain randomization via probabilistic inference for robotics simulators, 2019
Ramos, F., Possas, R. C., and Fox, D · 2019
Later among the works it cites.
Learning a Contact-Adaptive Controller for Robust, Efficient Legged Locomotion
Da, X., Xie, Z., Hoeller, D., Boots, B., Anandkumar, A., Zhu, Y., Babich, B., and Garg, A · 2020
Later among the works it cites.
Mushroomrl: Simplifying reinforcement learning research
D’Eramo, C., Tateo, D., Bonarini, A., Restelli, M., and Peters, J · 2020
Later among the works it cites.
Rl unplugged: Benchmarks for offline reinforcement learning
Gulcehre, C., Wang, Z., Novikov, A., Paine, T. L., Colmenarejo, S. G., Zolna, K., Agarwal, R., Merel, J., Mankowitz, D., Paduraru, C., et al · 2020
Later among the works it cites.
Array programming with NumPy
Harris, C. R., Millman, K. J., van der Walt, S. J., Gommers, R., Virtanen, P., Cournapeau, D., Wieser, E., Taylor, J., Berg, S., Smith, N. J., Kern, R., Picus, M., Hoyer, S., van Kerkwijk, M. H., Brett, M., Haldane, A., del Río, J. F., Wiebe, M., Peterson, P., Gérard-Marchant, P., Sheppard, K., Reddy, T., Weckesser, W., Abbasi, H., Gohlke, C., and Oliphant, T. E · 2020
Later among the works it cites.
Kim, J., Shin, J., and Yang, I · 2020
Later among the works it cites.
Simurlacra - a framework for reinforcement learning from randomized simulations
Muratore, F · 2020
Later among the works it cites.
Safe optimal control using stochastic barrier functions and deep forward-backward sdes
Pereira, M. A., Wang, Z., Exarchos, I., and Theodorou, E. A · 2020
Later among the works it cites.
Hidden fluid mechanics: Learning velocity and pressure fields from flow visualizations
Raissi, M., Yazdani, A., and Karniadakis, G. E · 2020
Later among the works it cites.
Conservative safety critics for exploration
Bharadhwaj, H., Kumar, A., Rhinehart, N., Levine, S., Shkurti, F., and Garg, A · 2021
Closest in time.
Dynamics Randomization Revisited: A Case Study for Quadrupedal Locomotion
Xie, Z., Da, X., van de Panne, M., Babich, B., and Garg, A · 2021
Closest in time.