Fetching the paper…
Reading the bibliography…
Most reinforcement learning methods are based upon the key assumption that the transition dynamics and reward functions are fixed, that is, the underlying Markov decision process is stationary.
Optimistic adaptive acceleration for optimization
Wang, J.-K., Li, X., and Li, P · 1903
Earlier work this paper cites.
Be aware of non-stationarity: Nearly optimal algorithms for piecewise-stationary cascading bandits
Wang, L., Zhou, H., Li, B., Varshney, L. R., and Zhao, Z · 1909
Earlier work this paper cites.
The central role of the propensity score in observational studies for causal effects
Rosenbaum, P. R. and Rubin, D. B · 1983
Earlier work this paper cites.
A new optimality criterion for nonhomogeneous Markov decision processes
Hopp, W. J., Bean, J. C., and Smith, R. L · 1987
Earlier work this paper cites.
Continual learning in reinforcement environments
Ring, M. B · 1994
Earlier work this paper cites.
Lifelong learning algorithms
Thrun, S · 1998
Earlier work this paper cites.
A general method for incremental self-improvement and multi-agent learning
Schmidhuber, J · 1999
Earlier work this paper cites.
An environment model for nonstationary reinforcement learning
Choi, S. P., Yeung, D.-Y., and Zhang, N. L · 2000
Earlier work this paper cites.
Solving nonstationary infinite horizon dynamic optimization problems
Garcia, A. and Smith, R. L · 2000
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Precup, D · 2000
Earlier work this paper cites.
Nash convergence of gradient dynamics in general-sum games
Singh, S., Kearns, M., and Mansour, Y · 2000
Earlier work this paper cites.
Econometric analysis
Greene, W. H · 2003
Earlier work this paper cites.
Variance reduction techniques for gradient estimates in reinforcement learning
Greensmith, E., Bartlett, P. L., and Baxter, J · 2004
Earlier work this paper cites.
Convergence and no-regret in multiagent learning
Bowling, M · 2005
Earlier work this paper cites.
Experts in a Markov decision process
Even-Dar, E., Kakade, S. M., and Mansour, Y · 2005
Earlier work this paper cites.
Bayesian models of nonstationary Markov decision processes
Jong, N. K. and Stone, P · 2005
Earlier work this paper cites.
Pattern recognition and machine learning
Bishop, C. M · 2006
Earlier work this paper cites.
Solution and forecast horizons for infinite-horizon nonhomogeneous Markov decision processes
Cheevaprawatdomrong, T., Schochetman, I. E., Smith, R. L., and Garcia, A · 2007
Earlier work this paper cites.
Awesome: A general multiagent learning algorithm that converges in self-play and learns a best response against stationary opponents
Conitzer, V. and Sandholm, T · 2007
Earlier work this paper cites.
On upper-confidence bound policies for non-stationary bandit problems
Moulines, E · 2008
Earlier work this paper cites.
Reinforcement learning in non-stationary continuous time and space scenarios
Basso, E. W. and Engel, P. M · 2009
Earlier work this paper cites.
Online learning in Markov decision processes with arbitrarily changing rewards and transitions
Yu, J. Y. and Mannor, S · 2009
Earlier work this paper cites.
Multi-agent learning with policy prediction
Zhang, C. and Lesser, V · 2010
Earlier work this paper cites.
Importance-weighted least-squares probabilistic classifier for covariate shift adaptation with application to human activity recognition
Hachiya, H., Sugiyama, M., and Ueda, N · 2012
Cited alongside, same era.
Online learning and online convex optimization
Shalev-Shwartz, S. et al · 2012
Cited alongside, same era.
Online learning in Markov decision processes with adversarially chosen transition probability distributions
Abbasi, Y., Bartlett, P. L., Kanade, V., Seldin, Y., and Szepesvári, C · 2013
Cited alongside, same era.
A linear programming approach to nonstationary infinite-horizon Markov decision processes
Ghate, A. and Smith, R. L · 2013
Cited alongside, same era.
Learning in non-stationary mdps as transfer learning
Mahmud, M. and Ramamoorthy, S · 2013
Cited alongside, same era.
Opponent modelling by sequence prediction and lookahead in two-player games
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Later among the works it cites.
Predictive off-policy policy evaluation for nonstationary decision problems, with applications to digital marketing
Thomas, P. S., Theocharous, G., Ghavamzadeh, M., Durugkar, I., and Brunskill, E · 2017
Later among the works it cites.
Unifying task specification in reinforcement learning
White, M · 2017
Later among the works it cites.
Learning with opponent-learning awareness
Foerster, J., Chen, R. Y., Al-Shedivat, M., Whiteson, S., Abbeel, P., and Mordatch, I · 2018
Later among the works it cites.
Gajane, P., Ortner, R., and Auer, P · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mealing, R. and Shapiro, J. L · 2013
Cited alongside, same era.
Online learning with predictable sequences
Rakhlin, A. and Sridharan, K · 2013
Cited alongside, same era.
Model-free intelligent diabetes management using machine learning
Bastani, M · 2014
Cited alongside, same era.
Stochastic multi-armed-bandit problem with non-stationary rewards
Besbes, O., Gur, Y., and Zeevi, A · 2014
Cited alongside, same era.
Weighted importance sampling for off-policy learning with linear function approximation
Mahmood, A. R., van Hasselt, H. P., and Sutton, R. S · 2014
Cited alongside, same era.
The UVA/PADOVA type 1 diabetes simulator: New features
Man, C. D., Micheletto, F., Lv, D., Breton, M., Kovatchev, B., and Cobelli, C · 2014
Cited alongside, same era.
Reinforcement learning for closed-loop propofol anesthesia: A study in human volunteers
Moore, B. L., Pyeatt, L. D., Kulkarni, V., Panousis, P., Padrez, K., and Doufas, A. G · 2014
Cited alongside, same era.
Soft actor-critic algorithms and applications
Haarnoja, T., Zhou, A., Hartikainen, K., Tucker, G., Ha, S., Tan, J., Kumar, V., Zhu, H., Gupta, A., Abbeel, P., et al · 2018
Later among the works it cites.
TIDBD: Adapting temporal-difference step-sizes through stochastic meta-descent
Kearney, A., Veeriah, V., Travnik, J. B., Sutton, R. S., and Pilarski, P. M · 2018
Later among the works it cites.
Rotting bandits are no harder than stochastic ones
Seznec, J., Locatelli, A., Carpentier, A., Lazaric, A., and Valko, M · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Later among the works it cites.
Reinforcement learning under drift
Cheung, W. C., Simchi-Levi, D., and Zhu, R · 2019
Later among the works it cites.
Finn, C., Rajeswaran, A., Kakade, S., and Levine, S · 2019
Later among the works it cites.
Meta-descent for online, continual prediction
Jacobsen, A., Schlegel, M., Linke, C., Degris, T., White, A., and White, M · 2019
Later among the works it cites.
When people change their mind: Off-policy evaluation in non-stationary recommendation environments
Jagerman, R., Markov, I., and de Rijke, M · 2019
Later among the works it cites.
Lecarpentier, E. and Rachelson, E · 2019
Later among the works it cites.
Cascading non-stationary bandits: Online learning to rank in the non-stationary cascade model
Li, C. and de Rijke, M · 2019
Later among the works it cites.
Online Markov decision processes with time-varying transition probabilities and rewards
Li, Y., Zhong, A., Qu, G., and Li, N · 2019
Later among the works it cites.
Adaptive online planning for continual lifelong learning
Lu, K., Mordatch, I., and Abbeel, P · 2019
Later among the works it cites.
Learning and planning for time-varying mdps using maximum likelihood estimation
Ornik, M. and Topcu, U · 2019
Later among the works it cites.
Reinforcement learning in non-stationary environments
Padakandla, S., J., P. K., and Bhatnagar, S · 2019
Later among the works it cites.
An online learning approach to model predictive control
Wagener, N., Cheng, C.-A., Sacks, J., and Boots, B · 2019
Later among the works it cites.
Simglucose v0.2.1 (2018) , 2019
Xie, J · 2019
Later among the works it cites.
A survey of reinforcement learning algorithms for dynamically varying environments
Padakandla, S · 2020
Closest in time.
Deep reinforcement learning amidst lifelong non-stationarity
Xie, A., Harrison, J., and Finn, C · 2020
Closest in time.