Fetching the paper…
Reading the bibliography…
Efficient exploration in bandits is a fundamental online learning problem.
Meta dynamic pricing: Transfer learning across experiments
Bastani, H., Simchi-Levi, D., and Zhu, R · 1902
Earlier work this paper cites.
Empirical Bayes regret minimization
Hsu, C.-W., Kveton, B., Meshi, O., Mladenov, M., and Szepesvari, C · 1904
Earlier work this paper cites.
Meta-learning of sequential strategies
Ortega, P., Wang, J., Rowland, M., Genewein, T., Kurth-Nelson, Z., Pascanu, R., Heess, N., Veness, J., Pritzel, A., Sprechmann, P., Jayakumar, S., McGrath, T., Miller, K., Azar, M. G., Osband, I., Rabinowitz, N., Gyorgy, A., Chiappa, S., Osindero, S., Teh, Y. W., van Hasselt, H., de Freitas, N., Botvinick, M., and Legg, S · 1905
Earlier work this paper cites.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
Thompson, W. R · 1933
Earlier work this paper cites.
Mesures dans les espaces produits
Tulcea, C. I · 1949
Earlier work this paper cites.
Bayes estimates for the linear model
Lindley, D. and Smith, A · 1972
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
Lai, T. L. and Robbins, H · 1985
Earlier work this paper cites.
Explanation-Based Neural Network Learning - A Lifelong Learning Approach
Thrun, S · 1996
Earlier work this paper cites.
Theoretical models of learning to learn
Baxter, J · 1998
Earlier work this paper cites.
Lifelong learning algorithms
Thrun, S · 1998
Earlier work this paper cites.
A model of inductive bias learning
Baxter, J · 2000
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Auer, P., Cesa-Bianchi, N., and Fischer, P · 2002
Earlier work this paper cites.
Multi-armed bandit algorithms and empirical evaluation
Vermorel, J. and Mohri, M · 2005
Cited alongside, same era.
Differentiable meta-learning in contextual bandits
Kveton, B., Mladenov, M., Hsu, C.-W., Zaheer, M., Szepesvari, C., and Boutilier, C · 2006
Cited alongside, same era.
Policy gradient optimization of Thompson sampling policies
Min, S., Moallemi, C., and Russo, D · 2006
Cited alongside, same era.
Differentiable linear bandit algorithm
Yang, K. and Toni, L · 2006
Cited alongside, same era.
Data Analysis Using Regression and Multilevel/Hierarchical Models
Gelman, A. and Hill, J · 2007
Cited alongside, same era.
Online clustering of bandits
Gentile, C., Li, S., and Zappella, G · 2014
Later among the works it cites.
Algorithms for multi-armed bandit problems
Kuleshov, V. and Precup, D · 2014
Later among the works it cites.
Learning to optimize via posterior sampling
Russo, D. and Van Roy, B · 2014
Later among the works it cites.
RL 2 : Fast reinforcement learning via slow reinforcement learning
Duan, Y., Schulman, J., Chen, X., Bartlett, P., Sutskever, I., and Abbeel, P · 2016
Later among the works it cites.
Multi-task learning for contextual bandits
Deshmukh, A. A., Dogan, U., and Scott, C · 2017
Later among the works it cites.
A tutorial on Thompson sampling
Russo, D., Van Roy, B., Kazerouni, A., Osband, I., and Wen, Z · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yang, J., Hu, W., Lee, J., and Du, S · 2010
Cited alongside, same era.
Analysis of Thompson sampling for the multi-armed bandit problem
Agrawal, S. and Goyal, N · 2012
Cited alongside, same era.
An empirical evaluation of Thompson sampling
Chapelle, O. and Li, L · 2012
Cited alongside, same era.
Meta-learning of exploration/exploitation strategies: The multi-armed bandit case
Maes, F., Wehenkel, L., and Ernst, D · 2012
Cited alongside, same era.
Sequential transfer in multi-armed bandit with finite set of models
Azar, M. G., Lazaric, A., and Brunskill, E · 2013
Cited alongside, same era.
Bayesian Data Analysis
Gelman, A., Carlin, J., Stern, H., Dunson, D., Vehtari, A., and Rubin, D · 2013
Cited alongside, same era.
Later among the works it cites.
Bandit Algorithms
Lattimore, T. and Szepesvari, C · 2019
Later among the works it cites.
Information-theoretic confidence bounds for reinforcement learning
Lu, X. and Van Roy, B · 2019
Later among the works it cites.
Differentiable meta-learning of bandit policies
Boutilier, C., Hsu, C.-W., Kveton, B., Mladenov, M., Szepesvari, C., and Zaheer, M · 2020
Later among the works it cites.
Meta-learning with stochastic linear bandits
Cella, L., Lazaric, A., and Pontil, M · 2020
Later among the works it cites.
Latent bandits revisited
Hong, J., Kveton, B., Zaheer, M., Chow, Y., Ahmed, A., and Boutilier, C · 2020
Later among the works it cites.