Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) algorithms find applications in inventory control, recommender systems, vehicular traffic management, cloud computing and robotics.
1905
Earlier work this paper cites.
T. Banerjee, G. Whipps, P. Gurram, and V. Tarokh, “Sequential Event Detection Using Multimodal Data in Nonstationary Environments,” in 2018 21st International Conference on Information Fusion (FUSION) , 2018, pp. 1940–1947
1947
Earlier work this paper cites.
E. S. Page, “Continuous Inspection Schemes,” Biometrika , vol. 41, no. 1/2, pp. 100–115, 1954
1954
Earlier work this paper cites.
A. Shiryaev, “On Optimum Methods in Quickest Detection Problems,” Theory of Probability and Its Applications , vol. 8, no. 1, pp. 22–46, 1963
1963
Earlier work this paper cites.
C. J. Watkins and P. Dayan, “Q-learning,” Machine learning , vol. 8, no. 3-4, pp. 279–292, 1992
1992
Earlier work this paper cites.
D. P. Bertsekas and J. N. Tsitsiklis, Neuro-Dynamic Programming . Athena Scientific, 1996
1996
Earlier work this paper cites.
M. B. Ring, “CHILD: A first step towards continual learning,” in Learning to learn . Springer, 1998, pp. 261–292
1998
Earlier work this paper cites.
R. S. Sutton, D. McAllester, S. Singh, and Y. Mansour, “Policy Gradient Methods for Reinforcement Learning with Function Approximation,” in Proceedings of the 12th International Conference on Neural Information Processing Systems , ser. NIPS’99. MIT Press, 1999, p. 1057–1063
1999
Earlier work this paper cites.
S. P. Choi, D.-Y. Yeung, and N. L. Zhang, “An Environment Model for Nonstationary Reinforcement Learning,” in Advances in neural information processing systems , 2000, pp. 987–993
2000
Earlier work this paper cites.
——, “Hidden-Mode Markov Decision Processes for Nonstationary Sequential Decision Making,” in Sequence Learning . Springer, 2000, pp. 264–287
2000
Earlier work this paper cites.
V. R. Konda and J. N. Tsitsiklis, “On Actor-Critic Algorithms,” SIAM J. Control Optim. , vol. 42, no. 4, p. 1143–1166, April 2003
2003
Earlier work this paper cites.
M. L. Puterman, Markov Decision Processes: Discrete Stochastic Dynamic Programming , 2nd ed. New York, NY, USA: John Wiley & Sons, Inc., 2005
2005
Earlier work this paper cites.
B. C. da Silva et al. , “Dealing with Non-stationary Environments Using Context Detection,” in Proceedings of the 23rd International Conference on Machine Learning , 2006, pp. 217–224
2006
Earlier work this paper cites.
L. Busoniu, R. Babuska, and B. De Schutter, “A Comprehensive Survey of Multiagent Reinforcement Learning,” IEEE Transactions on Systems, Man, and Cybernetics, Part C (Applications and Reviews) , vol. 38, no. 2, pp. 156–172, 2008
2008
Earlier work this paper cites.
B. C. Csáji and L. Monostori, “Value Function Based Reinforcement Learning in Changing Markovian Environments,” Journal of Machine Learning Research , vol. 9, pp. 1679–1709, jun 2008
2008
Earlier work this paper cites.
V. S. Borkar, Stochastic Approximation: A Dynamical Systems Viewpoint . Springer, 2009, vol. 48
2009
Earlier work this paper cites.
J. Y. Yu and S. Mannor, “Arbitrarily modulated Markov decision processes,” in Proceedings of the 48h IEEE Conference on Decision and Control (CDC) Conference , 2009, pp. 2946–2953
2009
Earlier work this paper cites.
T. Sun, Q. Zhao, and P. B. Luh, “A Rollout Algorithm for Multichain Markov Decision Processes with Average Cost,” in Positive Systems . Springer, 2009, pp. 151–162
2009
Earlier work this paper cites.
T. Jaksch, R. Ortner, and P. Auer, “Near-optimal regret bounds for reinforcement learning,” Journal of Machine Learning Research , vol. 11, no. Apr, pp. 1563–1600, 2010
2010
Earlier work this paper cites.
A. Salkham and V. Cahill, “Soilse: A decentralized approach to optimization of fluctuating urban traffic using reinforcement learning,” in 13th International IEEE Conference on Intelligent Transportation Systems , Sept 2010, pp. 531–538
2010
Earlier work this paper cites.
C.-Z. Xu, J. Rao, and X. Bu, “Url: A unified reinforcement learning approach for autonomic cloud management,” Journal of Parallel and Distributed Computing , vol. 72, no. 2, pp. 95 – 105, 2012
2012
Earlier work this paper cites.
P. Wang and B. Goertzel, Theoretical Foundations of Artificial General Intelligence . Springer, 2012, vol. 4
2012
Earlier work this paper cites.
S. Shalev-Shwartz, “Online Learning and Online Convex Optimization,” Foundations and Trends® in Machine Learning , vol. 4, no. 2, pp. 107–194, 2012
2012
Earlier work this paper cites.
D. Bertsekas, Dynamic Programming and Optimal Control , 4th ed. Belmont,MA: Athena Scientific, 2013, vol. II
2013
Earlier work this paper cites.
T. Dick, A. Gyorgy, and C. Szepesvari, “Online learning in Markov decision processes with changing cost sequences,” in International Conference on Machine Learning , 2014, pp. 512–520
2014
Earlier work this paper cites.
E. Hadoux, A. Beynier, and P. Weng, “Sequential Decision-Making under Non-stationary Environments via Sequential Change-point Detection,” in Learning over Multiple Contexts (LMCE) , Nancy, France, Sep 2014
2014
Cited alongside, same era.
J. Kober and J. Peters, Reinforcement Learning in Robotics: A Survey . Springer International Publishing, 2014, pp. 9–67
2014
Cited alongside, same era.
Prabuchandran K.J., Hemanth Kumar A.N, and S. Bhatnagar, “Multi-agent reinforcement learning for traffic signal control,” in 17th International IEEE Conference on Intelligent Transportation Systems (ITSC) , 2014, pp. 2529–2534
2014
Cited alongside, same era.
R. Allamaraju, H. Kingravi, A. Axelrod, G. Chowdhary, R. Grande, J. P. How, C. Crick, and W. Sheng, “Human aware UAS path planning in urban environments using nonstationary MDPs,” in 2014 IEEE International Conference on Robotics and Automation (ICRA) , 2014, pp. 1161–1167
2014
Cited alongside, same era.
E. Liang, R. Liaw, R. Nishihara, P. Moritz, R. Fox, K. Goldberg, J. Gonzalez, M. Jordan, and I. Stoica, “RLlib: Abstractions for Distributed Reinforcement Learning,” in Proceedings of the 35th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, vol. 80. PMLR, 2018, pp. 3053–3062
2018
Later among the works it cites.
A. Coluccia and A. Fascista, “An alternative procedure to cumulative sum for cyber-physical attack detection,” Internet Technology Letters , vol. 1, no. 3, p. e2, 2018
2018
Later among the works it cites.
E. Liebman, E. Zavesky, and P. Stone, “A Stitch in Time - Autonomous Model Management via Reinforcement Learning,” in Proceedings of the 17th International Conference on Autonomous Agents and Multiagent Systems , ser. AAMAS ’18. International Foundation for Autonomous Agents and Multiagent Systems, 2018, p. 990–998
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
C. O’Reilly, A. Gluhak, M. A. Imran, and S. Rajasegarar, “Anomaly Detection in Wireless Sensor Networks in a Non-Stationary Environment,” IEEE Communications Surveys Tutorials , vol. 16, no. 3, pp. 1413–1432, 2014
2014
Cited alongside, same era.
R. Rana and F. S. Oliveira, “Real-time dynamic pricing in a non-stationary environment using model-free reinforcement learning,” Omega , vol. 47, pp. 116 – 126, 2014
2014
Cited alongside, same era.
A. Hallak, D. D. Castro, and S. Mannor, “Contextual Markov Decision Processes,” in Proceedings of the 12th European Workshop on Reinforcement Learning (EWRL 2015) , 2015
2015
Cited alongside, same era.
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz, “Trust Region Policy Optimization,” in International conference on machine learning , 2015, pp. 1889–1897
2015
Cited alongside, same era.
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski et al. , “Human-level control through deep reinforcement learning,” Nature , vol. 518, no. 7540, p. 529, 2015
2015
Cited alongside, same era.
2015
Cited alongside, same era.
C. A. Gomez-Uribe and N. Hunt, “The Netflix Recommender System: Algorithms, Business Value, and Innovation,” ACM Trans. Manage. Inf. Syst. , vol. 6, no. 4, Dec 2016
2016
Cited alongside, same era.
S. Abdallah and M. Kaisers, “Addressing Environment Non-Stationarity by Repeating Q-learning Updates,” Journal of Machine Learning Research , vol. 17, no. 46, pp. 1–31, 2016
2016
Cited alongside, same era.
2018
Later among the works it cites.
R. Fruit, M. Pirotta, and A. Lazaric, “Near Optimal Exploration-Exploitation in Non-Communicating Markov Decision Processes,” in Proceedings of the 32nd International Conference on Neural Information Processing Systems , ser. NIPS’18. Curran Associates Inc., 2018, p. 2998–3008
2018
Later among the works it cites.
O. N. Granichin and V. A. Erofeeva, “Cyclic Stochastic Approximation with Disturbance on Input in the Parameter Tracking Problem based on a Multiagent Algorithm,” Automation and Remote Control , vol. 79, no. 6, pp. 1013–1028, 2018
2018
Later among the works it cites.
Y. Qian, J. Wu, R. Wang, F. Zhu, and W. Zhang, “Survey on Reinforcement Learning Applications in Communication Networks,” Journal of Communications and Information Networks , vol. 4, no. 2, pp. 30–39, June 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
G. I. Parisi, R. Kemker, J. L. Part, C. Kanan, and S. Wermter, “Continual lifelong learning with neural networks: A review,” Neural Networks , vol. 113, pp. 54 – 71, 2019
2019
Later among the works it cites.
J. Vanschoren, Meta-Learning . Springer International Publishing, 2019, pp. 35–61
2019
Later among the works it cites.
R. Ortner, P. Gajane, and P. Auer, “Variational Regret Bounds for Reinforcement Learning,” in Proceedings of the 35th Conference on Uncertainty in Artificial Intelligence , 2019
2019
Later among the works it cites.
Y. Li and N. Li, “Online Learning for Markov Decision Processes in Nonstationary Environments: A Dynamic Regret Analysis,” in 2019 American Control Conference (ACC) , July 2019, pp. 1232–1237
2019
Later among the works it cites.
E. Lecarpentier and E. Rachelson, “Non-Stationary Markov Decision Processes, a Worst-Case Approach using Model-Based Reinforcement Learning,” in Advances in Neural Information Processing Systems , 2019, pp. 7214–7223
2019
Later among the works it cites.
K. J. Prabuchandran, N. Singh, P. Dayama, and V. Pandit, “Change Point Detection for Compositional Multivariate Data,” arXiv , 2019
2019
Later among the works it cites.
C. Kaplanis et al. , “Policy Consolidation for Continual Reinforcement Learning,” in Proceedings of the 36th International Conference on Machine Learning , vol. 97. PMLR, 09–15 Jun 2019, pp. 3242–3251
2019
Later among the works it cites.
H. Liu, R. Socher, and C. Xiong, “Taming MAML: Efficient Unbiased Meta-Reinforcement Learning,” in Proceedings of the 36th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, vol. 97. PMLR, 09–15 Jun 2019, pp. 4061–4071
2019
Later among the works it cites.
M. Schneckenreither and S. Haeussler, “Reinforcement learning methods for operations research applications: The order release problem,” in Machine Learning, Optimization, and Data Science , G. Nicosia, P. Pardalos, G. Giuffrida, R. Umeton, and V. Sciacca, Eds. Springer International Publishing, 2019, pp. 545–559
2019
Later among the works it cites.
K. Shao, Z. Tang, Y. Zhu, N. Li, and D. Zhao, “A Survey of Deep Reinforcement Learning in Video Games,” 2019
2019
Later among the works it cites.
Y. Jaafra, A. Deruyver, J. L. Laurent, and M. S. Naceur, “Context-Aware Autonomous Driving Using Meta-Reinforcement Learning,” in 2019 18th IEEE International Conference On Machine Learning And Applications (ICMLA) , Dec 2019, pp. 450–455
2019
Later among the works it cites.
G. Mesesan, J. Englsberger, G. Garofalo, C. Ott, and A. Albu-Schäffer, “Dynamic Walking on Compliant and Uneven Terrain using DCM and Passivity-based Whole-body Control,” in 2019 IEEE-RAS 19th International Conference on Humanoid Robots (Humanoids) , 2019, pp. 25–32
2019
Later among the works it cites.
X. Lin, P. Guo, C. Florensa, and D. Held, “Adaptive Variance for Changing Sparse-Reward Environments,” in 2019 International Conference on Robotics and Automation (ICRA) , 2019, pp. 3210–3216
2019
Later among the works it cites.
M. Turchetta, A. Krause, and S. Trimpe, “Robust Model-free Reinforcement Learning with Multi-objective Bayesian Optimization,” 2019
2019
Later among the works it cites.
S. Banerjee, R. Bhattacharjee, and A. Sinha, “Fundamental Limits of Age-of-Information in Stationary and Non-stationary Environments,” 2020
2020
Closest in time.
T. Azayev and K. Zimmerman, “Blind Hexapod Locomotion in Complex Terrain with Gait Adaptation Using Deep Reinforcement Learning and Classification,” Journal of Intelligent & Robotic Systems , pp. 1–13, 2020
2020
Closest in time.