Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) methods learn optimal decisions in the presence of a stationary environment.
Prabuchandran KJ, Singh N, Dayama P, Pandit V (2019) Change Point Detection for Compositional Multivariate Data. arXiv 1901.04935
1901
Earlier work this paper cites.
Tatbul N, Lee TJ, Zdonik S, Alam M, Gottschlich J (2018) Precision and Recall for Time Series. In: Advances in Neural Information Processing Systems, pp 1920–1930
1930
Earlier work this paper cites.
Page ES (1954) Continuous Inspection Schemes. Biometrika 41(1/2):100–115
1954
Earlier work this paper cites.
Shiryaev A (1963) On Optimum Methods in Quickest Detection Problems. Theory of Probability & Its Applications 8(1):22–46
1963
Earlier work this paper cites.
Watkins CJ, Dayan P (1992) Q-learning. Machine learning 8(3-4):279–292
1992
Earlier work this paper cites.
Sutton RS, McAllester D, Singh S, Mansour Y (1999) Policy Gradient Methods for Reinforcement Learning with Function Approximation. In: Proceedings of the 12th International Conference on Neural Information Processing Systems, pp 1057–1063
1999
Earlier work this paper cites.
Minka T (2000) Estimating a Dirichlet distribution
2000
Earlier work this paper cites.
Abounadi J, Bertsekas D, Borkar V (2001) Learning Algorithms for Markov Decision Processes with Average Cost. SIAM Journal on Control and Optimization 40(3):681–698, DOI 10.1137/S0363012999361974
2001
Earlier work this paper cites.
Konda VR, Tsitsiklis JN (2003) On Actor-Critic Algorithms. SIAM Journal on Control and Optimization 42(4):1143–1166
2003
Earlier work this paper cites.
Puterman ML (2005) Markov Decision Processes: Discrete Stochastic Dynamic Programming, 2nd edn. John Wiley & Sons, Inc., New York, NY, USA
2005
Earlier work this paper cites.
Levin DA, Peres Y, Wilmer EL (2006) Markov Chains and Mixing Times. American Mathematical Society
2006
Earlier work this paper cites.
da Silva BC, Basso EW, Bazzan ALC, Engel PM (2006) Dealing with Non-Stationary Environments Using Context Detection. In: Proceedings of the 23rd International Conference on Machine Learning, Association for Computing Machinery, ICML ’06, p 217–224, DOI 10.1145/1143844.1143872
2006
Earlier work this paper cites.
Csáji BC, Monostori L (2008) Value Function Based Reinforcement Learning in Changing Markovian Environments. J Mach Learn Res 9:1679–1709
2008
Earlier work this paper cites.
Yu JY, Mannor S (2009) Online learning in markov decision processes with arbitrarily changing rewards and transitions. In: 2009 International Conference on Game Theory for Networks, pp 314–322, DOI 10.1109/GAMENETS.2009.5137416
2009
Earlier work this paper cites.
Jaksch T, Ortner R, Auer P (2010) Near-optimal regret bounds for reinforcement learning. Journal of Machine Learning Research 11:1563–1600
2010
Earlier work this paper cites.
Salkham A, Cahill V (2010) Soilse: A decentralized approach to optimization of fluctuating urban traffic using Reinforcement Learning. In: 13th International IEEE Conference on Intelligent Transportation Systems, pp 531–538, DOI 10.1109/ITSC.2010.5625145
2010
Cited alongside, same era.
Prashanth LA, Bhatnagar S (2011) Reinforcement learning with average cost for adaptive control of traffic lights at intersections. In: 2011 14th International IEEE Conference on Intelligent Transportation Systems (ITSC), pp 1640–1645, DOI 10.1109/ITSC.2011.6082823
2011
Cited alongside, same era.
Prabuchandran KJ, Meena SK, Bhatnagar S (2013) Q-learning based energy management policies for a single sensor node with finite buffer. Wireless Communications Letters, IEEE 2(1):82–85, DOI 10.1109/WCL.2012.112012.120754
2012
Cited alongside, same era.
Bertsekas D (2013) Dynamic Programming and Optimal Control, vol II, 4th edn. Athena Scientific, Belmont,MA
2013
Cited alongside, same era.
Everett R, Roberts S (2018) Learning against non-stationary agents with opponent modelling and deep reinforcement learning. In: 2018 AAAI Spring Symposium Series
2018
Later among the works it cites.
Iwashita AS, Papa JP (2019) An Overview on Concept Drift Learning. IEEE Access 7:1532–1547, DOI 10.1109/ACCESS.2018.2886026
2018
Later among the works it cites.
Kemker R, et al. (2018) Measuring catastrophic forgetting in neural networks. In: Thirty-second AAAI conference on artificial intelligence
2018
Later among the works it cites.
Liebman E, Zavesky E, Stone P (2018) A Stitch in Time - Autonomous Model Management via Reinforcement Learning. In: Proceedings of the 17th International Conference on Autonomous Agents and MultiAgent Systems, International Foundation for Autonomous Agents and Multiagent Systems, AAMAS ’18, p 990–998
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dick T, György A, Szepesvári C (2014) Online Learning in Markov Decision Processes with Changing Cost Sequences. In: Proceedings of the 31st International Conference on International Conference on Machine Learning - Volume 32, JMLR.org, ICML’14, p I–512–I–520
2014
Cited alongside, same era.
Hadoux E, Beynier A, Weng P (2014) Sequential Decision-Making under Non-stationary Environments via Sequential Change-point Detection. In: Learning over Multiple Contexts (LMCE), Nancy, France
2014
Cited alongside, same era.
Harel M, Mannor S, El-Yaniv R, Crammer K (2014) Concept Drift Detection Through Resampling. In: International Conference on Machine Learning, pp 1009–1017
2014
Cited alongside, same era.
Matteson DS, James NA (2014) A Nonparametric Approach for Multiple Change Point Analysis of Multivariate Data. Journal of the American Statistical Association 109(505):334–345
2014
Cited alongside, same era.
Hallak A, Castro DD, Mannor S (2015) Contextual Markov Decision Processes. In: Proceedings of the 12th European Workshop on Reinforcement Learning (EWRL 2015)
2015
Cited alongside, same era.
Abdallah S, Kaisers M (2016) Addressing Environment Non-Stationarity by Repeating Q-learning Updates. Journal of Machine Learning Research 17(46):1–31
2016
Cited alongside, same era.
Tijsma AD, Drugan MM, Wiering MA (2016) Comparing exploration strategies for q-learning in random stochastic mazes. In: 2016 IEEE Symposium Series on Computational Intelligence (SSCI), pp 1–8, DOI 10.1109/SSCI.2016.7849366
2016
Cited alongside, same era.
Banerjee T, Miao Liu, How JP (2017) Quickest change detection approach to optimal control in markov decision processes with model changes. In: 2017 American Control Conference (ACC), pp 399–405, DOI 10.23919/ACC.2017.7962986
2017
Cited alongside, same era.
Mohammadi M, Al-Fuqaha A (2018) Enabling cognitive smart cities using big data and machine learning: Approaches and challenges. IEEE Communications Magazine 56(2):94–101, DOI 10.1109/MCOM.2018.1700298
2018
Later among the works it cites.
2018
Later among the works it cites.
Roveri M (2019) Learning Discrete-Time Markov Chains Under Concept Drift. IEEE Transactions on Neural Networks and Learning Systems 30(9):2570–2582, DOI 10.1109/TNNLS.2018.2886956
2018
Later among the works it cites.
Sutton RS, Barto AG (2018) Reinforcement Learning: An Introduction, 2nd edn. MIT Press, Cambridge, MA, USA
2018
Later among the works it cites.
Andrychowicz M, et al. (2019) Learning dexterous in-hand manipulation. The International Journal of Robotics Research DOI 10.1177/0278364919887447
2019
Closest in time.
Ding S, Du W, Zhao X, Wang L, Jia W (2019) A new asynchronous reinforcement learning algorithm based on improved parallel PSO. Appl Intell 49(12):4211–4222, DOI 10.1007/s10489-019-01487-4
2019
Closest in time.
Kaplanis C, et al. (2019) Policy consolidation for continual reinforcement learning. In: Proceedings of the 36th International Conference on Machine Learning, PMLR, vol 97, pp 3242–3251
2019
Closest in time.
Niroui F, Zhang K, Kashino Z, Nejat G (2019) Deep Reinforcement Learning Robot for Search and Rescue Applications: Exploration in Unknown Cluttered Environments. IEEE Robotics and Automation Letters 4(2):610–617, DOI 10.1109/LRA.2019.2891991
2019
Closest in time.
Ortner R, Gajane P, Auer P (2019) Variational Regret Bounds for Reinforcement Learning. In: Proceedings of the 35th Conference on Uncertainty in Artificial Intelligence
2019
Closest in time.
Zhao X, et al. (2019) Applications of Asynchronous Deep Reinforcement Learning Based on Dynamic Updating Weights. Applied Intelligence 49(2):581–591, DOI 10.1007/s10489-018-1296-x
2019
Closest in time.