Fetching the paper…
Reading the bibliography…
Maximising a cumulative reward function that is Markov and stationary, i.e., defined over state-action pairs and independent of time, is sufficient to capture many kinds of goals in a Markov decision process (MDP).
Zur theorie der gesellschaftsspiele
J. Von Neumann · 1928
Earlier work this paper cites.
An algorithm for quadratic programming
M. Frank and P. Wolfe · 1956
Earlier work this paper cites.
The relaxation method of finding the common point of convex sets and its application to the solution of problems in convex programming
L. M. Bregman · 1967
Earlier work this paper cites.
Convex analysis
R. T. Rockafellar · 1970
Earlier work this paper cites.
Problem complexity and method efficiency in optimization
A. S. Nemirovskij and D. B. Yudin · 1983
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
M. L. Puterman · 1984
Earlier work this paper cites.
Q-learning
C. J. Watkins and P. Dayan · 1992
Earlier work this paper cites.
A course in game theory
M. J. Osborne and A. Rubinstein · 1994
Earlier work this paper cites.
A decision-theoretic generalization of on-line learning and an application to boosting
Y. Freund and R. E. Schapire · 1997
Earlier work this paper cites.
Information theory and statistics
S. Kullback · 1997
Earlier work this paper cites.
Constrained Markov decision processes , volume 7
E. Altman · 1999
Earlier work this paper cites.
Mirror descent and nonlinear projected subgradient methods for convex optimization
A. Beck and M. Teboulle · 2003
Earlier work this paper cites.
Online convex programming and generalized infinitesimal gradient ascent
M. Zinkevich · 2003
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
P. Abbeel and A. Y. Ng · 2004
Earlier work this paper cites.
Convex optimization
S. Boyd and L. Vandenberghe · 2004
Earlier work this paper cites.
An actor-critic algorithm for constrained markov decision processes
V. S. Borkar · 2005
Earlier work this paper cites.
Stochastic matrix games with bandit feedback
B. O’Donoghue, T. Lattimore, and I. Osband · 2006
Earlier work this paper cites.
Logarithmic regret algorithms for online convex optimization
E. Hazan, A. Agarwal, and S. Kale · 2007
Earlier work this paper cites.
Variational policy gradient method for reinforcement learning with general utilities
J. Zhang, A. Koppel, A. S. Bedi, C. Szepesvari, and M. Wang · 2007
Earlier work this paper cites.
An analysis of model-based interval estimation for markov decision processes
A. L. Strehl and M. L. Littman · 2008
Earlier work this paper cites.
A game-theoretic approach to apprenticeship learning
U. Syed and R. E. Schapire · 2008
Earlier work this paper cites.
Apprenticeship learning using linear programming
U. Syed, M. Bowling, and R. E. Schapire · 2008
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
T. Jaksch, R. Ortner, and P. Auer · 2010
Earlier work this paper cites.
Follow-the-regularized-leader and mirror descent: Equivalence theorems and l1 regularization
B. McMahan · 2011
Earlier work this paper cites.
An online actor–critic algorithm with function approximation for constrained markov decision processes
S. Bhatnagar and K. Lakshmanan · 2012
Earlier work this paper cites.
Pac bounds for discounted mdps
T. Lattimore and M. Hutter · 2012
Earlier work this paper cites.
Revisiting frank-wolfe: Projection-free sparse convex optimization
M. Jaggi · 2013
Cited alongside, same era.
(More) efficient reinforcement learning via posterior sampling
I. Osband, D. Russo, and B. Van Roy · 2013
Cited alongside, same era.
Fast algorithms for online stochastic convex programming
S. Agrawal and N. R. Devanur · 2014
Cited alongside, same era.
Sample complexity of episodic fixed-horizon reinforcement learning
C. Dann and E. Brunskill · 2015
Cited alongside, same era.
On the global linear convergence of frank-wolfe optimization variants
M. Jaggi and S. Lacoste-Julien · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Cited alongside, same era.
A theory of regularized markov decision processes
M. Geist, B. Scherrer, and O. Pietquin · 2019
Later among the works it cites.
Provably efficient maximum entropy exploration
E. Hazan, S. Kakade, K. Singh, and A. Van Soest · 2019
Later among the works it cites.
Efficient exploration via state marginal matching
L. Lee, B. Eysenbach, E. Parisotto, E. Xing, S. Levine, and R. Salakhutdinov · 2019
Later among the works it cites.
Reinforcement learning with convex constraints
S. Miryoosefi, K. Brantley, H. Daumé III, M. Dudík, and R. Schapire · 2019
Later among the works it cites.
Dualdice: Behavior-agnostic estimation of discounted stationary distribution corrections
O. Nachum, Y. Chow, B. Dai, and L. Li · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Trust region policy optimization
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz · 2015
Cited alongside, same era.
Stochastic neural networks for hierarchical reinforcement learning
C. Florensa, Y. Duan, and P. Abbeel · 2016
Cited alongside, same era.
Introduction to online convex optimization
E. Hazan · 2016
Cited alongside, same era.
Generative adversarial imitation learning
J. Ho and S. Ermon · 2016
Cited alongside, same era.
Following the leader and fast rates in linear prediction: Curved constraint sets and other regularities
R. Huang, T. Lattimore, A. György, and C. Szepesvári · 2016
Cited alongside, same era.
Deep exploration via bootstrapped dqn
I. Osband, C. Blundell, A. Pritzel, and B. V. Roy · 2016
Cited alongside, same era.
Behaviour suite for reinforcement learning
I. Osband, Y. Doron, M. Hessel, J. Aslanides, E. Sezener, A. Saraiva, K. McKinney, T. Lattimore, C. Szepesvari, S. Singh, et al · 2019
Later among the works it cites.
Online convex optimization in adversarial markov decision processes
A. Rosenberg and Y. Mansour · 2019
Later among the works it cites.
Reward constrained policy optimization
C. Tessler, D. J. Mankowitz, and S. Mannor · 2019
Later among the works it cites.
Wasserstein adversarial imitation learning
H. Xiao, M. Herman, J. Wagner, S. Ziesche, J. Etesami, and T. H. Linh · 2019
Later among the works it cites.
Optimality and approximation with policy gradient methods in markov decision processes
A. Agarwal, S. M. Kakade, J. D. Lee, and G. Mahajan · 2020
Later among the works it cites.
Provably efficient exploration in policy optimization
Q. Cai, Z. Yang, C. Jin, and Z. Wang · 2020
Later among the works it cites.
Exploration-exploitation in constrained mdps
Y. Efroni, S. Mannor, and M. Pirotta · 2020
Later among the works it cites.
Efficiently solving mdps with stochastic mirror descent
Y. Jin and A. Sidford · 2020
Later among the works it cites.
Constrained mdps and the reward hypothesis, 2020
C. Szepesvári · 2020
Later among the works it cites.
Behavior priors for efficient reinforcement learning
D. Tirumala, A. Galashov, H. Noh, L. Hasenclever, R. Pascanu, J. Schwarz, G. Desjardins, W. M. Czarnecki, A. Ahuja, Y. W. Teh, et al · 2020
Later among the works it cites.
Mirror descent policy optimization
M. Tomar, L. Shani, Y. Efroni, and M. Ghavamzadeh · 2020
Later among the works it cites.
Off-policy evaluation via the regularized lagrangian
M. Yang, O. Nachum, B. Dai, L. Li, and D. Schuurmans · 2020
Later among the works it cites.
Apprenticeship learning via frank-wolfe
T. Zahavy, A. Cohen, H. Kaplan, and Y. Mansour · 2020
Later among the works it cites.
Wasserstein distance guided adversarial imitation learning with reward shape exploration
M. Zhang, Y. Wang, X. Ma, L. Xia, J. Yang, Z. Li, and X. Li · 2020
Later among the works it cites.
Inverse reinforcement learning in contextual mdps
S. Belogolovsky, P. Korsunsky, S. Mannor, C. Tessler, and T. Zahavy · 2021
Closest in time.
Balancing constraints and rewards with meta-gradient d4{pg}
D. A. Calian, D. J. Mankowitz, T. Zahavy, Z. Xu, J. Oh, N. Levine, and T. Mann · 2021
Closest in time.
Concave utility reinforcement learning: the mean-field game viewpoint
M. Geist, J. Pérolat, M. Laurière, R. Elie, S. Perrin, O. Bachem, R. Munos, and O. Pietquin · 2021
Closest in time.
Towards tight bounds on the sample complexity of average-reward mdps
Y. Jin and A. Sidford · 2021
Closest in time.
Adaptive reward-free exploration
E. Kaufmann, P. Ménard, O. D. Domingues, A. Jonsson, E. Leurent, and M. Valko · 2021
Closest in time.
Online apprenticeship learning
L. Shani, T. Zahavy, and S. Mannor · 2021
Closest in time.