Fetching the paper…
Reading the bibliography…
In constrained reinforcement learning (RL), a learning agent seeks to not only optimize the overall reward but also satisfy the additional safety, diversity, or budget constraints.
A short note on concentration inequalities for random vectors with subgaussian norm
Jin, C., Netrapalli, P., Ge, R., Kakade, S. M., and Jordan, M. I. (2019) · 1902
Earlier work this paper cites.
Batch policy learning under constraints
Le, H. M., Voloshin, C., and Yue, Y. (2019) · 1903
Earlier work this paper cites.
Provably efficient imitation learning from observation alone
Sun, W., Vemula, A., Boots, B., and Bagnell, J. A. (2019) · 1905
Earlier work this paper cites.
Zur theorie der gesellschaftsspiele
Neumann, J. v. (1928) · 1928
Earlier work this paper cites.
An analog of the minimax theorem for vector payoffs
Blackwell, D. et al. (1956) · 1956
Earlier work this paper cites.
Constrained Markov decision processes
Altman, E. (1999) · 1999
Earlier work this paper cites.
Adaptive game playing using multiplicative weights
Freund, Y. and Schapire, R. E. (1999) · 1999
Earlier work this paper cites.
Learning in markov decision processes under constraints
Singh, R., Gupta, A., and Shroff, N. B. (2020) · 2002
Earlier work this paper cites.
Exploration-exploitation in constrained mdps
Efroni, Y., Mannor, S., and Pirotta, M. (2020) · 2003
Earlier work this paper cites.
Qiu, S., Wei, X., Yang, Z., Ye, J., and Wang, Z. (2020) · 2003
Earlier work this paper cites.
Online convex programming and generalized infinitesimal gradient ascent
Zinkevich, M. (2003) · 2003
Earlier work this paper cites.
Provably efficient reinforcement learning with general value function approximation
Wang, R., Salakhutdinov, R., and Yang, L. F. (2020b) · 2005
Earlier work this paper cites.
On reward-free reinforcement learning with linear function approximation
Wang, R., Du, S. S., Yang, L. F., and Salakhutdinov, R. (2020a) · 2006
Cited alongside, same era.
Task-agnostic exploration in reinforcement learning
Zhang, X., Singla, A., et al. (2020) · 2006
Cited alongside, same era.
A game-theoretic approach to apprenticeship learning
Syed, U. and Schapire, R. E. (2007) · 2007
Cited alongside, same era.
Provably efficient reward-agnostic navigation with linear value iteration
Zanette, A., Lazaric, A., Kochenderfer, M. J., and Brunskill, E. (2020) · 2008
Cited alongside, same era.
Maximum entropy inverse reinforcement learning
Ziebart, B. D., Maas, A. L., Bagnell, J. A., and Dey, A. K. (2008) · 2008
Cited alongside, same era.
Markov decision processes: discrete stochastic dynamic programming
Puterman, M. L. (2014) · 2014
Later among the works it cites.
Convex analysis
Rockafellar, R. T. (2015) · 2015
Later among the works it cites.
Resource management with deep reinforcement learning
Mao, H., Alizadeh, M., Menache, I., and Kandula, S. (2016) · 2016
Later among the works it cites.
Minimax regret bounds for reinforcement learning
Azar, M. G., Osband, I., and Munos, R. (2017) · 2017
Later among the works it cites.
Fairness in reinforcement learning
Jabbari, S., Joseph, M., Kearns, M., Morgenstern, J., and Roth, A. (2017) · 2017
Later among the works it cites.
Reinforcement learning with convex constraints
Miryoosefi, S., Brantley, K., Daume III, H., Dudik, M., and Schapire, R. E. (2019) · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dynamic pricing without knowing the demand function: Risk bounds and near-optimal algorithms
Besbes, O. and Zeevi, A. (2009) · 2009
Cited alongside, same era.
A sharp analysis of model-based reinforcement learning with self-play
Liu, Q., Yu, T., Bai, Y., and Jin, C. (2020) · 2010
Cited alongside, same era.
Blackwell approachability and no-regret learning are equivalent
Abernethy, J., Bartlett, P. L., and Hazan, E. (2011) · 2011
Cited alongside, same era.
Wu, J., Braverman, V., and Yang, L. F. (2020) · 2011
Cited alongside, same era.
Dynamic pricing with limited supply
Babaioff, M., Dughmi, S., Kleinberg, R. D., and Slivkins, A. (2015) · 2012
Cited alongside, same era.
Competitive Markov decision processes
Filar, J. and Vrieze, K. (2012) · 2012
Cited alongside, same era.
Reward-free exploration for reinforcement learning
Jin, C., Krishnamurthy, A., Simchowitz, M., and Yu, T. (2020a)
Cited in the paper.
Reward constrained policy optimization
Tessler, C., Mankowitz, D. J., and Mannor, S. (2019) · 2019
Later among the works it cites.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., Choi, D. H., Powell, R., Ewalds, T., Georgiev, P., et al. (2019) · 2019
Later among the works it cites.
Constrained episodic reinforcement learning in concave-convex and knapsack settings
Brantley, K., Dudik, M., Lykouris, T., Miryoosefi, S., Simchowitz, M., Slivkins, A., and Sun, W. (2020) · 2020
Later among the works it cites.
Provably efficient safe exploration via primal-dual policy optimization
Ding, D., Wei, X., Yang, Z., Wang, Z., and Jovanovic, M. (2021) · 2021
Closest in time.
Provably efficient algorithms for multi-objective competitive rl
Yu, T., Tian, Y., Zhang, J., and Sra, S. (2021) · 2021
Closest in time.