Fetching the paper…
Reading the bibliography…
In standard reinforcement learning (RL), a learning agent seeks to optimize the overall reward.
Batch policy learning under constraints
Le, H. M., Voloshin, C., and Yue, Y. (2019) · 1903
Earlier work this paper cites.
Zur theorie der gesellschaftsspiele
von Neumann, J. (1928) · 1928
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K. (2016) · 1937
Earlier work this paper cites.
An analog of the minimax theorem for vector payoffs
Blackwell, D. (1956) · 1956
Earlier work this paper cites.
On general minimax theorems
Sion, M. (1958) · 1958
Earlier work this paper cites.
A method for finding projections onto the intersection of convex sets in hilbert spaces
Boyle, J. P. and Dykstra, R. L. (1986) · 1986
Cited alongside, same era.
Projections onto convex cones in hilbert space
Ingram, J. M. and Marsh, M. (1991) · 1991
Cited alongside, same era.
Constrained Markov decision processes
Altman, E. (1999) · 1999
Cited alongside, same era.
Adaptive game playing using multiplicative weights
Freund, Y. and Schapire, R. E. (1999) · 1999
Cited alongside, same era.
Online convex programming and generalized infinitesimal gradient ascent
Zinkevich, M. (2003) · 2003
Cited alongside, same era.
A game-theoretic approach to apprenticeship learning
Syed, U. and Schapire, R. E. (2008) · 2008
Later among the works it cites.
Blackwell approachability and no-regret learning are equivalent
Abernethy, J., Bartlett, P. L., and Hazan, E. (2011) · 2011
Later among the works it cites.
Constrained policy optimization
Achiam, J., Held, D., Tamar, A., and Abbeel, P. (2017) · 2017
Later among the works it cites.
Provably efficient maximum entropy exploration
Hazan, E., Kakade, S. M., Singh, K., and Van Soest, A. (2018) · 2018
Later among the works it cites.
Reward constrained policy optimization
Tessler, C., Mankowitz, D. J., and Mannor, S. (2019) · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…