Fetching the paper…
Reading the bibliography…
We consider online learning for episodic stochastically constrained Markov decision processes (CMDPs), which plays a central role in ensuring the safety of reinforcement learning.
Q-learning with ucb exploration is sample efficient for infinite-horizon mdp
Dong, K · 1901
Earlier work this paper cites.
Batch policy learning under constraints
Le, H. M · 1903
Earlier work this paper cites.
Online convex optimization in adversarial markov decision processes
Rosenberg, A · 1905
Earlier work this paper cites.
Global optimality guarantees for policy gradient methods
Bhandari, J · 1906
Earlier work this paper cites.
Neural proximal/trust region policy optimization attains globally optimal policy
Liu, B · 1906
Earlier work this paper cites.
Abbasi-Yadkori, Y · 1908
Earlier work this paper cites.
Optimality and approximation with policy gradient methods in markov decision processes
Agarwal, A · 1908
Earlier work this paper cites.
Online primal-dual mirror descent under stochastic constraints
Wei, X · 1908
Earlier work this paper cites.
Neural policy gradient methods: Global optimality and rates of convergence
Wang, L · 1909
Earlier work this paper cites.
Model-free reinforcement learning in infinite-horizon average-reward markov decision processes
Wei, C.-Y · 1910
Earlier work this paper cites.
Provably efficient exploration in policy optimization
Cai, Q · 1912
Earlier work this paper cites.
Learning adversarial mdps with bandit feedback and unknown transition
Jin, C · 1912
Earlier work this paper cites.
Markov renewal programming by linear fractional programming
Fox, B · 1966
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
Constrained Markov decision processes
Altman, E · 1999
Earlier work this paper cites.
Direct gradient-based reinforcement learning
Baxter, J · 2000
Earlier work this paper cites.
Actor-critic algorithms
Konda, V. R · 2000
Earlier work this paper cites.
Constrained upper confidence reinforcement learning
Zheng, L · 2001
Cited alongside, same era.
Finite-time analysis of the multiarmed bandit problem
Auer, P · 2002
Cited alongside, same era.
Optimistic policy optimization with bandit feedback
Efroni, Y · 2002
Cited alongside, same era.
A natural policy gradient
Kakade, S. M · 2002
Cited alongside, same era.
Exploration-exploitation in constrained mdps
Efroni, Y · 2003
Cited alongside, same era.
The on-line shortest path problem under partial monitoring
Trust region policy optimization
Schulman, J · 2015
Later among the works it cites.
Dynamic service migration and workload scheduling in edge-clouds
Urgaonkar, R · 2015
Later among the works it cites.
Dynamic service migration in mobile edge-clouds
Wang, S · 2015
Later among the works it cites.
Constrained policy optimization
Achiam, J · 2017
Later among the works it cites.
Minimax regret bounds for reinforcement learning
Azar, M. G · 2017
Later among the works it cites.
Risk-constrained reinforcement learning with percentile risk criteria
Chow, Y · 2017
Later among the works it cites.
Learning unknown markov decision processes: A thompson sampling approach
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
György, A · 2007
Cited alongside, same era.
Online markov decision processes
Even-Dar, E · 2009
Cited alongside, same era.
Online learning with sample path constraints
Mannor, S · 2009
Cited alongside, same era.
Approximate primal solutions and rate analysis for dual subgradient methods
Nedić, A · 2009
Cited alongside, same era.
Markov decision processes with arbitrary reward processes
Yu, J. Y · 2009
Cited alongside, same era.
Near-optimal regret bounds for reinforcement learning
Jaksch, T · 2010
Cited alongside, same era.
The online loop-free stochastic shortest-path problem
Neu, G · 2010
Cited alongside, same era.
Ouyang, Y · 2017
Later among the works it cites.
Proximal policy optimization algorithms
Schulman, J · 2017
Later among the works it cites.
Online convex optimization with stochastic constraints
Yu, H · 2017
Later among the works it cites.
Global convergence of policy gradient methods for the linear quadratic regulator
Fazel, M · 2018
Later among the works it cites.
Efficient bias-span-constrained exploration-exploitation in reinforcement learning
Fruit, R · 2018
Later among the works it cites.
Is q-learning provably efficient?
Jin, C · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Sutton, R. S · 2018
Later among the works it cites.
Online learning in weakly coupled markov decision processes: A convergence time study
Wei, X · 2018
Later among the works it cites.
Regret minimization for reinforcement learning by evaluating the optimal bias function
Zhang, Z · 2019
Later among the works it cites.
Provably efficient safe exploration via primal-dual policy optimization
Ding, D · 2020
Closest in time.