Fetching the paper…
Reading the bibliography…
Autonomous agents must often deal with conflicting requirements, such as completing tasks using the least amount of time/energy, learning multiple tasks, or dealing with multiple opponents.
Studies in linear and nonlinear programming
Kenneth J. Arrow and Leonard Hurwicz, · 1958
Earlier work this paper cites.
Convex analysis
R. T. Rockafellar, · 1970
Earlier work this paper cites.
“A convex analytic approach to markov decision processes,”
Vivek S Borkar, · 1988
Earlier work this paper cites.
“Multilayer feedforward networks are universal approximators,”
Kurt Hornik, Maxwell Stinchcombe, and Halbert White, · 1989
Earlier work this paper cites.
“On the approximate realization of continuous mappings by neural networks,”
Ken-Ichi Funahashi, · 1989
Earlier work this paper cites.
“Approximation by superpositions of a sigmoidal function,”
George Cybenko, · 1989
Earlier work this paper cites.
“Universal approximation using radial-basis-function networks,”
Jooyoung Park and Irwin W Sandberg, · 1991
Earlier work this paper cites.
Constrained Markov decision processes
Eitan Altman, · 1999
Earlier work this paper cites.
“Policy gradient methods for reinforcement learning with function approximation,”
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour, · 2000
Earlier work this paper cites.
“Actor-critic algorithms,”
Vijay R Konda and John N Tsitsiklis, · 2000
Earlier work this paper cites.
“Portfolio optimization with conditional value-at-risk objective and constraints,”
Pavlo Krokhmal, Jonas Palmquist, and Stanislav Uryasev, · 2002
Earlier work this paper cites.
“A geometric approach to multi-criterion reinforcement learning,”
Shie Mannor and Nahum Shimkin, · 2004
Earlier work this paper cites.
Convex optimization
S. Boyd and L. Vandenberghe, · 2004
Cited alongside, same era.
“An actor-critic algorithm for constrained markov decision processes,”
Vivek S Borkar, · 2005
Cited alongside, same era.
“On the relation between universality, characteristic kernels and rkhs embedding of measures,”
Bharath Sriperumbudur, Kenji Fukumizu, and Gert Lanckriet, · 2010
Cited alongside, same era.
“Control and optimization meet the smart power grid: Scheduling of power demands for optimal energy management,”
Iordanis Koutsopoulos and Leandros Tassiulas, · 2011
Cited alongside, same era.
“An online actor–critic algorithm with function approximation for constrained markov decision processes,”
Shalabh Bhatnagar and K Lakshmanan, · 2012
Cited alongside, same era.
Jan Leike, Miljan Martic, Victoria Krakovna, Pedro A Ortega, Tom Everitt, Andrew Lefrancq, Laurent Orseau, and Shane Legg, · 2017
Later among the works it cites.
“Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates,”
Shixiang Gu, Ethan Holly, Timothy Lillicrap, and Sergey Levine, · 2017
Later among the works it cites.
“The expressive power of neural networks: A view from the width,”
Zhou Lu, Hongming Pu, Feicheng Wang, Zhiqiang Hu, and Liwei Wang, · 2017
Later among the works it cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto, · 2018
Later among the works it cites.
“Deepmimic: Example-guided deep reinforcement learning of physics-based character skills,”
Xue Bin Peng, Pieter Abbeel, Sergey Levine, and Michiel van de Panne, · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dotan Di Castro, Aviv Tamar, and Shie Mannor, · 2012
Cited alongside, same era.
“Variance adjusted actor critic algorithms,”
Aviv Tamar and Shie Mannor, · 2013
Cited alongside, same era.
“Multi-objective reinforcement learning using sets of Pareto dominating policies,”
Kristof Van Moffaert and Ann Nowé, · 2014
Cited alongside, same era.
“Risk-sensitive and robust decision-making: a cvar optimization approach,”
Yinlam Chow, Aviv Tamar, Shie Mannor, and Marco Pavone, · 2015
Cited alongside, same era.
Convex optimization algorithms
Dimitri P Bertsekas and Athena Scientific, · 2015
Cited alongside, same era.
“Trust region policy optimization,”
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz, · 2015
Cited alongside, same era.
“Constrained policy optimization,”
Joshua Achiam, David Held, Aviv Tamar, and Pieter Abbeel, · 2017
Cited alongside, same era.
Chen Tessler, Daniel J Mankowitz, and Shie Mannor, · 2018
Later among the works it cites.
“Simple random search provides a competitive approach to reinforcement learning,”
Horia Mania, Aurelia Guy, and Benjamin Recht, · 2018
Later among the works it cites.
“Optimization of web service-based control system for balance between network traffic and delay,”
Chen Hou and Qianchuan Zhao, · 2018
Later among the works it cites.
“Safe exploration in continuous action spaces,”
Gal Dalal, Krishnamurthy Dvijotham, Matej Vecerik, Todd Hester, Cosmin Paduraru, and Yuval Tassa, · 2018
Later among the works it cites.
“Resnet with one-neuron hidden layers is a universal approximator,”
Hongzhou Lin and Stefanie Jegelka, · 2018
Later among the works it cites.
“Ray interference: a source of plateaus in deep reinforcement learning,”
Tom Schaul, Diana Borsa, Joseph Modayil, and Razvan Pascanu, · 2019
Closest in time.