Fetching the paper…
Reading the bibliography…
In many sequential decision-making problems, the goal is to optimize a utility function while satisfying a set of constraints on different utilities.
Alternative theoretical frameworks for finite horizon discrete-time stochastic optimal control
Steven E Shreve and Dimitri P Bertsekas · 1978
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Martin L Puterman · 1994
Earlier work this paper cites.
Constrained Markov decision processes , volume 7
Eitan Altman · 1999
Earlier work this paper cites.
Improved algorithms for conservative exploration in bandits
Evrard Garcelon, Mohammad Ghavamzadeh, Alessandro Lazaric, and Matteo Pirotta · 2002
Earlier work this paper cites.
Conservative exploration in reinforcement learning
Evrard Garcelon, Mohammad Ghavamzadeh, Alessandro Lazaric, and Matteo Pirotta · 2002
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Sham Kakade and John Langford · 2002
Earlier work this paper cites.
Mirror descent and nonlinear projected subgradient methods for convex optimization
Amir Beck and Marc Teboulle · 2003
Earlier work this paper cites.
On the empirical state-action frequencies in markov decision processes under general policies
Shie Mannor and John N. Tsitsiklis · 2005
Earlier work this paper cites.
Tuning bandit algorithms in stochastic environments
Jean-Yves Audibert, Rémi Munos, and Csaba Szepesvári · 2007
Earlier work this paper cites.
Empirical bernstein bounds and sample variance penalization
Andreas Maurer and Massimiliano Pontil · 2009
Earlier work this paper cites.
Subgradient methods for saddle-point problems
Angelia Nedić and Asuman Ozdaglar · 2009
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Earlier work this paper cites.
An online actor–critic algorithm with function approximation for constrained markov decision processes
Shalabh Bhatnagar and K Lakshmanan · 2012
Earlier work this paper cites.
Trading regret for efficiency: online convex optimization with long term constraints
Mehrdad Mahdavi, Rong Jin, and Tianbao Yang · 2012
Earlier work this paper cites.
Bandits with knapsacks
Ashwinkumar Badanidiyuru, Robert Kleinberg, and Aleksandrs Slivkins · 2013
Earlier work this paper cites.
Multi-armed bandit with budget constraint and variable costs
Wenkui Ding, Tao Qin, Xu-Dong Zhang, and Tie-Yan Liu · 2013
Earlier work this paper cites.
Online learning in episodic markovian decision processes by relative entropy policy search
Alexander Zimin and Gergely Neu · 2013
Earlier work this paper cites.
Bandits with concave rewards and convex knapsacks
Shipra Agrawal and Nikhil R. Devanur · 2014
Cited alongside, same era.
Bandits with budgets: Regret lower bounds and optimal algorithms
Richard Combes, Chong Jiang, and Rayadurgam Srikant · 2015
Cited alongside, same era.
Sample complexity of episodic fixed-horizon reinforcement learning
Christoph Dann and Emma Brunskill · 2015
Cited alongside, same era.
A comprehensive survey on safe reinforcement learning
Javier Garcıa and Fernando Fernández · 2015
Cited alongside, same era.
Fairness in learning: Classic and contextual bandits
Matthew Joseph, Michael J. Kearns, Jamie H. Morgenstern, and Aaron Roth · 2016
Cited alongside, same era.
Conservative bandits
Yifan Wu, Roshan Shariff, Tor Lattimore, and Csaba Szepesvári · 2016
Cited alongside, same era.
A lyapunov-based approach to safe reinforcement learning
Yinlam Chow, Ofir Nachum, Edgar A. Duéñez-Guzmán, and Mohammad Ghavamzadeh · 2018
Later among the works it cites.
Learning-based model predictive control for safe exploration
Torsten Koller, Felix Berkenkamp, Matteo Turchetta, and Andreas Krause · 2018
Later among the works it cites.
Safe exploration and optimization of constrained mdps using gaussian processes
Akifumi Wachi, Yanan Sui, Yisong Yue, and Masahiro Ono · 2018
Later among the works it cites.
Bandits with global convex constraints and objective
Shipra Agrawal and Nikhil R Devanur · 2019
Later among the works it cites.
Provably efficient exploration in policy optimization
Qi Cai, Zhuoran Yang, Chi Jin, and Zhaoran Wang · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Constrained policy optimization
Joshua Achiam, David Held, Aviv Tamar, and Pieter Abbeel · 2017
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Cited alongside, same era.
First-order methods in optimization , volume 25
Amir Beck · 2017
Cited alongside, same era.
Safe model-based reinforcement learning with stability guarantees
Felix Berkenkamp, Matteo Turchetta, Angela P. Schoellig, and Andreas Krause · 2017
Cited alongside, same era.
Linear programming formulation for non-stationary, finite-horizon markov decision process models
Arnab Bhattacharya and Jeffrey P Kharoufeh · 2017
Cited alongside, same era.
Risk-constrained reinforcement learning with percentile risk criteria
Yinlam Chow, Mohammad Ghavamzadeh, Lucas Janson, and Marco Pavone · 2017
Cited alongside, same era.
Richard Cheng, Gábor Orosz, Richard M. Murray, and Joel W. Burdick · 2019
Later among the works it cites.
Lyapunov-based safe policy optimization for continuous control
Yinlam Chow, Ofir Nachum, Aleksandra Faust, Mohammad Ghavamzadeh, and Edgar A. Duéñez-Guzmán · 2019
Later among the works it cites.
Tight regret bounds for model-based reinforcement learning with greedy policies
Yonathan Efroni, Nadav Merlis, Mohammad Ghavamzadeh, and Shie Mannor · 2019
Later among the works it cites.
Learning adversarial mdps with bandit feedback and unknown transition
Chi Jin, Tiancheng Jin, Haipeng Luo, Suvrit Sra, and Tiancheng Yu · 2019
Later among the works it cites.
A modern introduction to online learning
Francesco Orabona · 2019
Later among the works it cites.
Constrained reinforcement learning has zero duality gap
Santiago Paternain, Luiz F. O. Chamon, Miguel Calvo-Fullana, and Alejandro Ribeiro · 2019
Later among the works it cites.
Online convex optimization in adversarial markov decision processes
Aviv Rosenberg and Yishay Mansour · 2019
Later among the works it cites.
Reward constrained policy optimization
Chen Tessler, Daniel J. Mankowitz, and Shie Mannor · 2019
Later among the works it cites.
Tighter problem-dependent regret bounds in reinforcement learning without domain knowledge using value function bounds
Andrea Zanette and Emma Brunskill · 2019
Later among the works it cites.
Optimistic policy optimization with bandit feedback
Yonathan Efroni, Lior Shani, Aviv Rosenberg, and Shie Mannor · 2020
Closest in time.
Constrained upper confidence reinforcement learning
Liyuan Zheng and Lillian J. Ratliff · 2020
Closest in time.