Fetching the paper…
Reading the bibliography…
We consider an agent who is involved in a Markov decision process and receives a vector of outcomes every round.
Convex optimization: Algorithms and complexity
Sébastien Bubeck · 1935
Earlier work this paper cites.
An algorithm for quadratic programming
Marguerite Frank and Philip Wolfe · 1956
Earlier work this paper cites.
On general minimax theorems
Maurice Sion · 1958
Earlier work this paper cites.
Probability inequalities for sums of bounded random variables
Wassily Hoeffding · 1963
Earlier work this paper cites.
Problem complexity and method efficiency in optimization
Arkadi Nemirovski and David Berkovich Yudin · 1983
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Martin L. Puterman · 1994
Earlier work this paper cites.
Multi-criteria reinforcement learning
Zoltán Gábor, Zsolt Kalmár, and Csaba Szepesvári · 1998
Earlier work this paper cites.
Markov chains
James R. Norris · 1998
Earlier work this paper cites.
Constrained Markov Decision Processes
E. Altman · 1999
Earlier work this paper cites.
Online convex programming and generalized infinitesimal gradient ascent
Martin Zinkevich · 2003
Earlier work this paper cites.
A geometric approach to multi-criterion reinforcement learning
Shie Mannor and Nahum Shimkin · 2004
Earlier work this paper cites.
Dynamic preferences in multi-criteria reinforcement learning
Sriraam Natarajan and Prasad Tadepalli · 2005
Earlier work this paper cites.
Logarithmic online regret bounds for undiscounted reinforcement learning
Peter Auer and Ronald Ortner · 2006
Earlier work this paper cites.
Online learning: Theory, algorithms, and applications
Shai Shalev-Shwartz · 2007
Cited alongside, same era.
Learning all optimal policies with multiple criteria
Leon Barrett and Srini Narayanan · 2008
Cited alongside, same era.
Exploration-exploitation tradeoff using variance estimates in multi-armed bandits
Jean-Yves Audibert, Rémi Munos, and Csaba Szepesvári · 2009
Cited alongside, same era.
REGAL: A regularization based algorithm for reinforcement learning in weakly communicating mdps
Peter L. Bartlett and Ambuj Tewari · 2009
Cited alongside, same era.
Online learning with global cost functions
Eyal Even-Dar, Robert Kleinberg, Shie Mannor, and Yishay Mansour · 2009
Cited alongside, same era.
Online learning with sample path constraints
Shie Mannor, John N. Tsitsiklis, and Jia Yuan Yu · 2009
Cited alongside, same era.
Resourceful contextual bandits
Ashwinkumar Badanidiyuru, John Langford, and Aleksandrs Slivkins · 2014
Later among the works it cites.
Multi-objective reinforcement learning using sets of pareto dominating policies
Kristof Van Moffaert and Ann Nowé · 2014
Later among the works it cites.
Multiobjective reinforcement learning: A comprehensive overview
C. Liu, X. Xu, and D. Hu · 2015
Later among the works it cites.
Linear contextual bandits with knapsacks
Shipra Agrawal and Nikhil R. Devanur · 2016
Later among the works it cites.
An efficient algorithm for contextual bandits with knapsacks, and an extension to concave objectives
Shipra Agrawal, Nikhil R. Devanur, and Lihong Li · 2016
Later among the works it cites.
Optimistic posterior sampling for reinforcement learning: worst-case regret bounds
Shipra Agrawal and Randy Jia · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Cited alongside, same era.
Linear fitted-q iteration with multiple reward functions
Daniel J. Lizotte, Michael Bowling, and Susan A. Murphy · 2012
Cited alongside, same era.
Online learning and online convex optimization
Shai Shalev-Shwartz · 2012
Cited alongside, same era.
Bandits with knapsacks
Ashwinkumar Badanidiyuru, Robert Kleinberg, and Aleksandrs Slivkins · 2013
Cited alongside, same era.
A survey of multi-objective sequential decision-making
Diederik M. Roijers, Peter Vamplew, Shimon Whiteson, and Richard Dazeley · 2013
Cited alongside, same era.
Bandits with concave rewards and convex knapsacks
Shipra Agrawal and Nikhil R Devanur · 2014
Cited alongside, same era.
Later among the works it cites.
Fast rates for bandit optimization with upper-confidence frank-wolfe
Quentin Berthet and Vianney Perchet · 2017
Later among the works it cites.
Multi-objective bandits: Optimizing the generalized Gini index
Róbert Busa-Fekete, Balázs Szörényi, Paul Weng, and Shie Mannor · 2017
Later among the works it cites.
Bandits with knapsacks
Ashwinkumar Badanidiyuru, Robert Kleinberg, and Aleksandrs Slivkins · 2018
Later among the works it cites.
Provably efficient maximum entropy exploration
Elad Hazan, Sham M. Kakade, Karan Singh, and Abby Van Soest · 2018
Later among the works it cites.
Adversarial bandits with knapsacks
Nicole Immorlica, Karthik Abinav Sankararaman, Robert E. Schapire, and Aleksandrs Slivkins · 2018
Later among the works it cites.
Regret bounds for reinforcement learning via markov chain concentration
Ronald Ortner · 2018
Later among the works it cites.
Active exploration in markov decision processes
Jean Tarbouriech and Alessandro Lazaric · 2019
Closest in time.