Fetching the paper…
Reading the bibliography…
In this paper, we study the learning of safe policies in the setting of reinforcement learning problems.
CA: Stanford University Press, 1958
K. J. Arrow and L. Hurwicz, Studies in Linear and Nonlinear Programming · 1958
Earlier work this paper cites.
Wiley for The Massachusetts Institute of Technology, 1964
R. A. Howard, Dynamic programming and Markov processes · 1964
Earlier work this paper cites.
R. A. Howard and J. E. Matheson, “Risk-sensitive Markov decision processes,” Management Science
1972
Earlier work this paper cites.
O. Khatib, “Real-time obstacle avoidance for manipulators and mobile robots,” in Proceedings. 1985 IEEE International Conference on Robotics and Automation
1985
Earlier work this paper cites.
V. S. Borkar, “A convex analytic approach to Markov decision processes,” Probability Theory and Related Fields
1988
Earlier work this paper cites.
K.-I. Funahashi, “On the approximate realization of continuous mappings by neural networks,” Neural Networks
1989
Earlier work this paper cites.
G. Cybenko, “Approximation by superpositions of a sigmoidal function,” Mathematics of Control, Signals and Systems
1989
Earlier work this paper cites.
K. Hornik, M. Stinchcombe, and H. White, “Multilayer feedforward networks are universal approximators,” Neural Networks
1989
Earlier work this paper cites.
D. E. Koditschek and E. Rimon, “Robot navigation functions on manifolds with boundary,” Advances in applied mathematics
1990
Earlier work this paper cites.
D. E. Koditschek, “The control of natural motion in mechanical systems,” 1991
1991
Earlier work this paper cites.
J. Park and I. W. Sandberg, “Universal approximation using radial-basis-function networks,” Neural Computation
1991
Earlier work this paper cites.
C. J. Watkins and P. Dayan, “Q-learning,” Machine learning
1992
Earlier work this paper cites.
R. J. Williams, “Simple statistical gradient-following algorithms for connectionist reinforcement learning,” Machine Learning
1992
Earlier work this paper cites.
M. Heger, “Consideration of risk in reinforcement learning,” in Machine Learning Proceedings 1994
1994
Earlier work this paper cites.
M. Deistler, K. Peternell, and W. Scherrer, “Consistency and relative efficiency of subspace methods,” Automatica
1995
Earlier work this paper cites.
Athena scientific Belmont, MA, 1995
D. P. Bertsekas, Dynamic programming and optimal control · 1995
Earlier work this paper cites.
Athena Scientific Belmont, MA, 1996
D. P. Bertsekas and J. N. Tsitsiklis, Neuro-Dynamic Programming · 1996
Earlier work this paper cites.
S. P. Coraluppi and S. I. Marcus, “Risk-sensitive and minimax control of discrete-time, finite-state Markov decision processes,” Automatica
1999
Earlier work this paper cites.
Athena Scientific, Belmont, 1999
D. P. Bertsekas, Nonlinear Programming · 1999
Earlier work this paper cites.
R. S. Sutton, D. A. McAllester, S. P. Singh, and Y. Mansour, “Policy gradient methods for reinforcement learning with function approximation,” in Advances in Neural Information Processing Systems
2000
Earlier work this paper cites.
D. Q. Mayne, J. B. Rawlings, C. V. Rao, and P. O. Scokaert, “Constrained model predictive control: Stability and optimality,” Automatica
2000
Earlier work this paper cites.
Springer, 2000
V. Vapnik, The Nature of Statistical Learning Theory · 2000
Cited alongside, same era.
V. R. Konda and J. N. Tsitsiklis, “Actor-critic algorithms,” in Advances in Neural Information Processing Systems
2000
Cited alongside, same era.
O. Mihatsch and R. Neuneier, “Risk-sensitive reinforcement learning,” Machine Learning
2002
Cited alongside, same era.
M. Hutter, “Self-optimizing and Pareto-optimal policies in general environments based on Bayes-mixtures,” in International Conference on Computational Learning Theory
2002
Cited alongside, same era.
Prentice hall Upper Saddle River, NJ, 2002
H. K. Khalil and J. W. Grizzle, Nonlinear systems · 2002
Cited alongside, same era.
Cambridge University Press, 2004
S. Boyd and L. Vandenberghe, Convex Optimization · 2004
S. Paternain and A. Ribeiro, “Online learning of feasible strategies in unknown environments,” IEEE Transactions on Automatic Control
2016
Later among the works it cites.
Y. Chow, M. Ghavamzadeh, L. Janson, and M. Pavone, “Risk-constrained reinforcement learning with percentile risk criteria.,” Journal of Machine Learning Research
2017
Later among the works it cites.
J. Achiam, D. Held, A. Tamar, and P. Abbeel, “Constrained policy optimization,” in Proceedings of the 34th International Conference on Machine Learning-Volume 70
2017
Later among the works it cites.
2017
Later among the works it cites.
T. Chen, Q. Ling, and G. B. Giannakis, “An online convex optimization approach to proactive network resource allocation,” IEEE Transactions on Signal Processing
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
P. Geibel and F. Wysotzki, “Risk-sensitive reinforcement learning applied to control under constraints,” Journal of Artificial Intelligence Research
2005
Cited alongside, same era.
P. Geibel, “Reinforcement learning for MDPs with constraints,” in European Conference on Machine Learning
2006
Cited alongside, same era.
Y. Kadota, M. Kurano, and M. Yasuda, “Discounted Markov decision processes with utility constraints,” Computers & Mathematics with Applications
2006
Cited alongside, same era.
Springer Science & Business Media, 2007
T. Van Cutsem and C. Vournas, Voltage stability of electric power systems · 2007
Cited alongside, same era.
P. Wieland and F. Allgöwer, “Constructive safety using control barrier functions,” IFAC Proceedings Volumes
2007
Cited alongside, same era.
A. Nedić and A. Ozdaglar, “Subgradient methods for saddle-point problems,” Journal of optimization theory and applications
2009
Cited alongside, same era.
2017
Later among the works it cites.
Z. Lu, H. Pu, F. Wang, Z. Hu, and L. Wang, “The expressive power of neural networks: A view from the width,” in Advances in Neural Information Processing Systems
2017
Later among the works it cites.
MIT press, 2018
R. S. Sutton and A. G. Barto, Reinforcement Learning: An Introduction · 2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
T. Koller, F. Berkenkamp, M. Turchetta, and A. Krause, “Learning-based model predictive control for safe exploration,” in 2018 IEEE Conference on Decision and Control (CDC)
2018
Later among the works it cites.
2018
Later among the works it cites.
H. Lin and S. Jegelka, “Resnet with one-neuron hidden layers is a universal approximator,” in Advances in Neural Information Processing Systems
2018
Later among the works it cites.
M. Fazlyab, S. Paternain, A. Ribeiro, and V. M. Preciado, “Distributed smooth and strongly convex optimization with inexact dual methods,” in 2018 Annual American Control Conference (ACC)
2018
Later among the works it cites.
H. Robbins and S. Monro, “A stochastic approximation method,” The Annals of Mathematical Statistics
2018
Later among the works it cites.
S. Paternain, M. Calvo-Fullana, L. F. Chamon, and A. Ribeiro, “Learning safe policies via primal–dual methods,” in Proceedings of the 58th IEEE Conference on Decision and Control
2019
Closest in time.
S. Paternain, M. Calvo-Fullana, L. Chamon, and A. Ribeiro, “Constrained reinforcement learning has zero duality gap,” in Proc. 33rd Conference on Neural Information Processing Systems
2019
Closest in time.
A. Tsiamis and G. J. Pappas, “Finite sample analysis of stochastic system identification,” in 2019 IEEE 58th Conference on Decision and Control (CDC)
2019
Closest in time.
S. Dean, H. Mania, N. Matni, B. Recht, and S. Tu, “On the sample complexity of the linear quadratic regulator,” Foundations of Computational Mathematics
2019
Closest in time.
D. Ding, K. Zhang, T. Basar, and M. Jovanovic, “Natural policy gradient primal-dual method for constrained markov decision processes,” Advances in Neural Information Processing Systems
2020
Closest in time.
A. Agarwal, S. M. Kakade, J. D. Lee, and G. Mahajan, “Optimality and approximation with policy gradient methods in markov decision processes,” in Conference on Learning Theory
2020
Closest in time.