Fetching the paper…
Reading the bibliography…
We propose an algorithm for tabular episodic reinforcement learning with constraints.
Active exploration in markov decision processes
Tarbouriech, J. and Lazaric, A. (2019) · 1902
Earlier work this paper cites.
Batch policy learning under constraints
Le, H. M., Voloshin, C., and Yue, Y. (2019) · 1903
Earlier work this paper cites.
Introduction to multi-armed bandits
Slivkins, A. (2019) · 1904
Earlier work this paper cites.
Efficient model-free reinforcement learning in metric spaces
Song, Z. and Sun, W. (2019) · 1905
Earlier work this paper cites.
Provably efficient imitation learning from observation alone
Sun, W., Vemula, A., Boots, B., and Bagnell, J. A. (2019) · 1905
Earlier work this paper cites.
Provably efficient reinforcement learning with linear function approximation
Jin, C., Yang, Z., Wang, Z., and Jordan, M. I. (2019) · 1907
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K. (2016) · 1937
Earlier work this paper cites.
A markovian decision process
Bellman, R. (1957) · 1957
Earlier work this paper cites.
Dyna, an integrated architecture for learning, planning, and reacting
Sutton, R. S. (1991) · 1991
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Sutton, R. S. and Barto, A. G. (1998) · 1998
Earlier work this paper cites.
Constrained Markov Decision Processes
Altman, E. (1999) · 1999
Earlier work this paper cites.
Constrained upper confidence reinforcement learning
Zheng, L. and Ratliff, L. J. (2020) · 2001
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Kearns, M. and Singh, S. (2002) · 2002
Earlier work this paper cites.
Learning in markov decision processes under constraints
Singh, R., Gupta, A., and Shroff, N. B. (2020) · 2002
Earlier work this paper cites.
Provably efficient safe exploration via primal-dual policy optimization
Ding, D., Wei, X., Yang, Z., Wang, Z., and Jovanović, M. R. (2020) · 2003
Cited alongside, same era.
Exploration-exploitation in constrained mdps
Efroni, Y., Mannor, S., and Pirotta, M. (2020) · 2003
Cited alongside, same era.
Qiu, S., Wei, X., Yang, Z., Ye, J., and Wang, Z. (2020) · 2003
Cited alongside, same era.
Online convex programming and generalized infinitesimal gradient ascent
Zinkevich, M. (2003) · 2003
Cited alongside, same era.
Prediction, learning, and games
Cesa-Bianchi, N. and Lugosi, G. (2006) · 2006
Close the gaps: A learning-while-doing algorithm for single-product revenue management problems
Wang, Z., Deng, S., and Ye, Y. (2014) · 2014
Later among the works it cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al. (2015) · 2015
Later among the works it cites.
Trust region policy optimization
Schulman, J., Levine, S., Moritz, P., Jordan, M. I., and Abbeel, P. (2015) · 2015
Later among the works it cites.
Resource management with deep reinforcement learning
Mao, H., Alizadeh, M., Menache, I., and Kandula, S. (2016) · 2016
Later among the works it cites.
Constrained policy optimization
Achiam, J., Held, D., Tamar, A., and Abbeel, P. (2017) · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A game-theoretic approach to apprenticeship learning
Syed, U. and Schapire, R. E. (2007) · 2007
Cited alongside, same era.
Maximum entropy inverse reinforcement learning
Ziebart, B. D., Maas, A. L., Bagnell, J. A., and Dey, A. K. (2008) · 2008
Cited alongside, same era.
Dynamic pricing without knowing the demand function: Risk bounds and near-optimal algorithms
Besbes, O. and Zeevi, A. (2009) · 2009
Cited alongside, same era.
Near-optimal regret bounds for reinforcement learning
Jaksch, T., Ortner, R., and Auer, P. (2010) · 2010
Cited alongside, same era.
On the minimax complexity of pricing in a changing environment
Besbes, O. and Zeevi, A. (2011) · 2011
Cited alongside, same era.
Dynamic pricing with limited supply
Babaioff, M., Dughmi, S., Kleinberg, R. D., and Slivkins, A. (2015) · 2012
Cited alongside, same era.
Bandits with knapsacks
Badanidiyuru, A., Kleinberg, R., and Slivkins, A. (2018) · 2013
Cited alongside, same era.
Azar, M. G., Osband, I., and Munos, R. (2017) · 2017
Later among the works it cites.
Unifying pac and regret: Uniform pac bounds for episodic reinforcement learning
Dann, C., Lattimore, T., and Brunskill, E. (2017) · 2017
Later among the works it cites.
Leike, J., Martic, M., Krakovna, V., Ortega, P. A., Everitt, T., Lefrancq, A., Orseau, L., and Legg, S. (2017) · 2017
Later among the works it cites.
Reinforcement Learning: An Introduction
Sutton, R. S. and Barto, A. G. (2018) · 2018
Later among the works it cites.
Regret minimization for reinforcement learning with vectorial feedback and complex objectives
Cheung, W. C. (2019) · 2019
Later among the works it cites.
Reinforcement learning with convex constraints
Miryoosefi, S., Brantley, K., Daume III, H., Dudik, M., and Schapire, R. E. (2019) · 2019
Later among the works it cites.
Online convex optimization in adversarial markov decision processes
Rosenberg, A. and Mansour, Y. (2019) · 2019
Later among the works it cites.
Reward constrained policy optimization
Tessler, C., Mankowitz, D. J., and Mannor, S. (2019) · 2019
Later among the works it cites.
Benchmarking safe exploration in deep reinforcement learning
Ray, A., Achiam, J., and Amodei, D. (2020) · 2020
Closest in time.