Fetching the paper…
Reading the bibliography…
Finding optimal policies which maximize long term rewards of Markov Decision Processes requires the use of dynamic programming and backward induction to solve the Bellman optimality equation.
Risk aversion in the small and in the large
John W. Pratt · 1964
Earlier work this paper cites.
Fairness in processor scheduling in time sharing systems
Sibsankar Haldar and DK Subramanian · 1991
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Multi-agent reinforcement learning: Independent vs. cooperative agents
Ming Tan · 1993
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
Michael L Littman · 1994
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Martin L. Puterman · 1994
Earlier work this paper cites.
Dynamic programming and optimal control , volume 1
Dimitri P Bertsekas · 1995
Earlier work this paper cites.
Instance-based utile distinctions for reinforcement learning with hidden state
R Andrew McCallum · 1995
Earlier work this paper cites.
Multi-agent reinforcement learning: A modular approach
Norihiko Ono and Kenji Fukumoto · 1996
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
Opportunistic beamforming using dumb antennas
P. Viswanath, D. N. C. Tse, and R. Laroia · 2002
Earlier work this paper cites.
Introductory lectures on convex optimization: A basic course , volume 87
Yurii Nesterov · 2003
Earlier work this paper cites.
Multi-agent reinforcement learning: a critical survey
Yoav Shoham, Rob Powers, and Trond Grenager · 2003
Earlier work this paper cites.
Inequalities for the l1 deviation of the empirical distribution
Tsachy Weissman, Erik Ordentlich, Gadiel Seroussi, Sergio Verdu, and Marcelo J Weinberger · 2003
Earlier work this paper cites.
WCDMA for UMTS.: Radio Access for Third Generation Mobile Communications
Harri Holma and Antti Toskala · 2005
Earlier work this paper cites.
A comparative study of random waypoint and gauss-markov mobility models in the performance evaluation of manet
Jinthana Ariyakhajorn, Pattana Wannawilai, and Chanboon Sathitwiriyawong · 2006
Earlier work this paper cites.
Generalized proportional fair scheduling in third generation wireless data networks
T Bu, L Li, and R Ramjee · 2006
Earlier work this paper cites.
Towards a unified theory of state abstraction for mdps
Lihong Li, Thomas J Walsh, and Michael L Littman · 2006
Earlier work this paper cites.
Decision-theoretic planning with non-markovian rewards
Sylvie Thiébaux, Charles Gretton, John Slaney, David Price, and Froduald Kabanza · 2006
Earlier work this paper cites.
Generalized
Eitan Altman, Konstantin Avrachenkov, and Andrey Garnaev · 2008
Earlier work this paper cites.
Proportional fair multiuser scheduling in lte
Raymond Kwan, Cyril Leung, and Jie Zhang · 2009
Earlier work this paper cites.
Responsive elastic computing
Julien Perez, Cécile Germain-Renaud, Balázs Kégl, and Charles Loomis · 2009
Cited alongside, same era.
Multi-agent reinforcement learning: An overview
Lucian Buşoniu, Robert Babuška, and Bart De Schutter · 2010
Cited alongside, same era.
Leen: Locality/fairness-aware key partitioning for mapreduce in the cloud
Shadi Ibrahim, Hai Jin, Lu Lu, Song Wu, Bingsheng He, and Li Qi · 2010
Cited alongside, same era.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Cited alongside, same era.
An axiomatic theory of fairness in network resource allocation
Tian Lan, David Kao, Mung Chiang, and Ashutosh Sabharwal · 2010
Cited alongside, same era.
Characterizing fairness for 3g wireless networks
Vaneet Aggarwal, Rittwik Jana, Jeffrey Pang, KK Ramakrishnan, and NK Shankaranarayanan · 2011
Cited alongside, same era.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Later among the works it cites.
Dueling network architectures for deep reinforcement learning
Ziyu Wang, Tom Schaul, Matteo Hessel, Hado Van Hasselt, Marc Lanctot, and Nando De Freitas · 2015
Later among the works it cites.
Joint latency and cost optimization for erasure-coded data center storage
Yu Xiang, Tian Lan, Vaneet Aggarwal, and Yih-Farn R Chen · 2015
Later among the works it cites.
On fairness in decision-making under uncertainty: Definitions, computation, and comparison
Chongjie Zhang and Julie A Shah · 2015
Later among the works it cites.
Cvxpy: A python-embedded modeling language for convex optimization
Steven Diamond and Stephen Boyd · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Saves: A sustainable multiagent application to conserve building energy considering occupants
Jun-young Kwak, Pradeep Varakantham, Rajiv Maheswaran, Milind Tambe, Farrokh Jazizadeh, Geoffrey Kavulya, Laura Klein, Burcin Becerik-Gerber, Timothy Hayes, and Wendy Wood · 2012
Cited alongside, same era.
A multiobjective reinforcement learning approach to water resources systems operation: Pareto frontier approximation in a single run
A Castelletti, Francesca Pianosi, and Marcello Restelli · 2013
Cited alongside, same era.
(more) efficient reinforcement learning via posterior sampling
Ian Osband, Daniel Russo, and Benjamin Van Roy · 2013
Cited alongside, same era.
A survey of multi-objective sequential decision-making
Diederik M. Roijers, Peter Vamplew, Shimon Whiteson, and Richard Dazeley · 2013
Cited alongside, same era.
Extreme state aggregation beyond mdps
Marcus Hutter · 2014
Cited alongside, same era.
Multiobjective reinforcement learning: A comprehensive overview
Chunming Liu, Xin Xu, and Dewen Hu · 2014
Cited alongside, same era.
Robert Margolies, Ashwin Sridharan, Vaneet Aggarwal, Rittwik Jana, NK Shankaranarayanan, Vinay A Vaishampayan, and Gil Zussman · 2016
Later among the works it cites.
Safe, multi-agent, reinforcement learning for autonomous driving
Shai Shalev-Shwartz, Shaked Shammah, and Amnon Shashua · 2016
Later among the works it cites.
Optimistic posterior sampling for reinforcement learning: worst-case regret bounds
Shipra Agrawal and Randy Jia · 2017
Later among the works it cites.
Markov chains and mixing times , volume 107
David A Levin and Yuval Peres · 2017
Later among the works it cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Later among the works it cites.
Lbp: Robust rate adaptation algorithm for svc video streaming
Anis Elgabli, Vaneet Aggarwal, Shuai Hao, Feng Qian, and Subhabrata Sen · 2018
Later among the works it cites.
Is q-learning provably efficient?
Chi Jin, Zeyuan Allen-Zhu, Sebastien Bubeck, and Michael I Jordan · 2018
Later among the works it cites.
Introduction to stochastic processes
Gregory F Lawler · 2018
Later among the works it cites.
Resource allocation for underlay d2d communication with proportional fairness
Xiaoshuai Li, Rajan Shankaran, Mehmet A Orgun, Gengfa Fang, and Yubin Xu · 2018
Later among the works it cites.
On q-learning convergence for non-markov decision processes
Sultan Javed Majeed and Marcus Hutter · 2018
Later among the works it cites.
Multi-task learning as multi-objective optimization
Ozan Sener and Vladlen Koltun · 2018
Later among the works it cites.
Source Code for Non-Linear Reinforcement Learning
Mridul Agarwal and Vaneet Aggarwal · 2019
Closest in time.
Learning fairness in multi-agent systems
Jiechuan Jiang and Zongqing Lu · 2019
Closest in time.
E2e: Embracing user heterogeneity to improve quality of experience on the web
Xu Zhang, Siddhartha Sen, Daniar Kurniawan, Haryadi Gunawi, and Junchen Jiang · 2019
Closest in time.
A multi-objective deep reinforcement learning framework
Thanh Thi Nguyen, Ngoc Duy Nguyen, Peter Vamplew, Saeid Nahavandi, Richard Dazeley, and Chee Peng Lim · 2020
Closest in time.