Regret minimization experience replay in off-policy reinforcement learning
Liu, X.-H., Xue, Z., Pang, J., Jiang, S., Xu, F., and Yu, Y. (2021) · 2021
Later among the works it cites.
Pac-bayes control: learning policies that provably generalize to novel environments
Majumdar, A., Farid, A., and Sonar, A. (2021) · 2021
Later among the works it cites.
Reinforcement learning with sparse-executing actions via sparsity regularization
Original
Pang, J.-C., Xu, T., Jiang, S., Liu, Y.-R., and Yu, Y. (2021) · 2021
Later among the works it cites.
Syndicated bandits: A framework for auto tuning hyper-parameters in contextual bandit algorithms
Ding, Q., Kang, Y., Liu, Y.-W., Lee, T. C. M., Hsieh, C.-J., and Sharpnack, J. (2022) · 2022
Later among the works it cites.
Pac-bayesian lifelong learning for multi-armed bandits
Flynn, H., Reeb, D., Kandemir, M., and Peters, J. (2022) · 2022
Later among the works it cites.
Model-based lifelong reinforcement learning with bayesian exploration
Fu, H., Yu, S., Littman, M., and Konidaris, G. (2022) · 2022
Later among the works it cites.
What can transformers learn in-context? a case study of simple function classes
Garg, S., Tsipras, D., Liang, P. S., and Valiant, G. (2022) · 2022
Later among the works it cites.
Efficient frameworks for generalized low-rank matrix bandit problems
Kang, Y., Hsieh, C.-J., and Lee, T. C. M. (2022) · 2022
Later among the works it cites.
Towards continual reinforcement learning: A review and perspectives
Khetarpal, K., Riemer, M., Rish, I., and Precup, D. (2022) · 2022
Later among the works it cites.
Reactive exploration to cope with non-stationarity in lifelong reinforcement learning
Steinparz, C. A., Schmied, T., Paischer, F., Dinu, M.-C., Patil, V. P., Bitto-Nemling, A., Eghbal-zadeh, H., and Hochreiter, S. (2022) · 2022
Later among the works it cites.
A general sample complexity analysis of vanilla policy gradient
Yuan, R., Gower, R. M., and Lazaric, A. (2022) · 2022
Later among the works it cites.
Prediction and control in continual reinforcement learning
Anand, N. and Precup, D. (2023) · 2023
Later among the works it cites.
Robust lipschitz bandits to adversarial corruptions
Kang, Y., Hsieh, C.-J., and Lee, T. C. M. (2023) · 2023
Later among the works it cites.
The effectiveness of world models for continual reinforcement learning
Kessler, S., Ostaszewski, M., Bortkiewicz, M., Żarski, M., Wolczyk, M., Parker-Holder, J., Roberts, S. J., Mi, P., et al. (2023) · 2023
Later among the works it cites.
Statistical guarantees for variational autoencoders using pac-bayesian theory
Mbacke, S. D., Clerc, F., and Germain, P. (2023) · 2023
Later among the works it cites.
A definition of continual reinforcement learning
Abel, D., Barreto, A., Van Roy, B., Precup, D., van Hasselt, H. P., and Singh, S. (2024) · 2024
Closest in time.
Online continuous hyperparameter optimization for generalized linear contextual bandits
Kang, Y., Hsieh, C.-J., and Lee, T. (2024) · 2024
Closest in time.