Fetching the paper…
Reading the bibliography…
We consider reinforcement learning (RL) in Markov Decision Processes in which an agent repeatedly interacts with an environment that is modeled by a controlled Markov process.
K. Azuma, “Weighted sums of certain dependent random variables,” Tohoku Mathematical Journal, Second Series , vol. 19, no. 3, pp. 357–367, 1967
1967
Earlier work this paper cites.
A. Lazar, “Optimal flow control of a class of queueing networks in equilibrium,” IEEE transactions on Automatic Control , vol. 28, no. 11, pp. 1001–1007, 1983
1983
Earlier work this paper cites.
T. L. Lai and H. Robbins, “Asymptotically efficient adaptive allocation rules,” Advances in applied mathematics , vol. 6, no. 1, pp. 4–22, 1985
1985
Earlier work this paper cites.
P. Nain and K. Ross, “Optimal priority assignment with hard constraint,” IEEE transactions on Automatic Control , vol. 31, no. 10, pp. 883–888, 1986
1986
Earlier work this paper cites.
M.-T. T. Hsiao and A. A. Lazar, “Optimal decentralized flow control of markovian queueing networks with multiple controllers,” Performance evaluation , vol. 13, no. 3, pp. 181–204, 1991
1991
Earlier work this paper cites.
E. Altman and A. Schwartz, “Adaptive control of constrained Markov chains,” IEEE Transactions on Automatic Control , vol. 36, no. 4, pp. 454–462, 1991
1991
Earlier work this paper cites.
R. Agrawal, “Sample mean based index policies by o (log n) regret for the multi-armed bandit problem,” Advances in Applied Probability , vol. 27, no. 4, pp. 1054–1078, 1995
1995
Earlier work this paper cites.
——, “Stochastic approximation with two time scales,” Systems & Control Letters , vol. 29, no. 5, pp. 291–294, 1997
1997
Earlier work this paper cites.
D. P. Bertsekas, Nonlinear programming . Taylor & Francis, 1997, vol. 48, no. 3
1997
Earlier work this paper cites.
R. S. Sutton and A. G. Barto, Reinforcement learning - an introduction , ser. Adaptive computation and machine learning. MIT Press, 1998. [Online]. Available: http://www.worldcat.org/oclc/37293240
1998
Earlier work this paper cites.
E. Altman, Constrained Markov Decision Processes . Chapman and Hall/CRC, March 1999
1999
Earlier work this paper cites.
V. R. Konda and V. S. Borkar, “Actor-critic–type learning algorithms for markov decision processes,” SIAM Journal on control and Optimization , vol. 38, no. 1, pp. 94–123, 1999
1999
Earlier work this paper cites.
V. R. Konda and J. N. Tsitsiklis, “Actor-critic algorithms,” in Advances in neural information processing systems , 2000, pp. 1008–1014
2000
Cited alongside, same era.
R. I. Brafman and M. Tennenholtz, “R-max-a general polynomial time algorithm for near-optimal reinforcement learning,” Journal of Machine Learning Research , vol. 3, no. Oct, pp. 213–231, 2002
2002
Cited alongside, same era.
H. Kushner and G. G. Yin, Stochastic approximation and recursive algorithms and applications . Springer Science & Business Media, 2003, vol. 35
2003
Cited alongside, same era.
S. Boyd and L. Vandenberghe, Convex optimization . Cambridge university press, 2004
2004
Cited alongside, same era.
V. S. Borkar, “An actor-critic algorithm for constrained Markov decision processes,” Systems & control letters , vol. 54, no. 3, pp. 207–213, 2005
2005
V. S. Borkar, Stochastic approximation: a dynamical systems viewpoint . Springer, 2009, vol. 48
2009
Later among the works it cites.
T. Jaksch, R. Ortner, and P. Auer, “Near-optimal regret bounds for reinforcement learning,” Journal of Machine Learning Research , vol. 11, no. Apr, pp. 1563–1600, 2010
2010
Later among the works it cites.
M. L. Puterman, Markov Decision Processes.: Discrete Stochastic Dynamic Programming . John Wiley & Sons, 2014
2014
Later among the works it cites.
P. R. Kumar and P. Varaiya, Stochastic systems: Estimation, identification, and adaptive control . SIAM, 2015
2015
Later among the works it cites.
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
2005
Cited alongside, same era.
P. Auer and R. Ortner, “Logarithmic online regret bounds for undiscounted reinforcement learning,” in Advances in Neural Information Processing Systems , 2007, pp. 49–56
2007
Cited alongside, same era.
E. Uchibe and K. Doya, “Constrained reinforcement learning from intrinsic and extrinsic rewards,” in 2007 IEEE 6th International Conference on Development and Learning . IEEE, 2007, pp. 163–168
2007
Cited alongside, same era.
C. Villani, Optimal transport: old and new . Springer Science & Business Media, 2008, vol. 338
2008
Cited alongside, same era.
J. Peters and S. Schaal, “Natural actor-critic,” Neurocomputing , vol. 71, no. 7-9, pp. 1180–1190, 2008
2008
Cited alongside, same era.
L. I. Sennott, Stochastic dynamic programming and the control of queueing systems . John Wiley & Sons, 2009, vol. 504
2009
Cited alongside, same era.
P. L. Bartlett and A. Tewari, “REGAL: A regularization based algorithm for reinforcement learning in weakly communicating mdps,” in Proceedings of the Twenty-Fifth Conference on Uncertainty in Artificial Intelligence, Montreal, QC, Canada, June 18-21, 2009 . AUAI Press, 2009, pp. 35–42
2009
Cited alongside, same era.
R. Singh and P. Kumar, “Throughput optimal decentralized scheduling of multihop networks with end-to-end deadline constraints: Unreliable links,” IEEE Transactions on Automatic Control , vol. 64, no. 1, pp. 127–142, 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
2019
Later among the works it cites.
S. Resnick, A probability path . Springer, 2019
2019
Later among the works it cites.
2020
Closest in time.
2020
Closest in time.
——, “Adaptive csma for decentralized scheduling of multi-hop networks with end-to-end deadline constraints,” IEEE/ACM Transactions on Networking , 2021
2021
Closest in time.