Fetching the paper…
Reading the bibliography…
We describe MELEE, a meta-learning algorithm for learning a good exploration policy in the interactive contextual bandit setting.
Associative reinforcement learning: Functions ink-dnf
Kaelbling, L. P · 1994
Earlier work this paper cites.
Generalization in reinforcement learning: Successful examples using sparse coarse coding
Sutton, R. S · 1996
Earlier work this paper cites.
Introduction to Reinforcement Learning
Sutton, R. S. and Barto, A. G · 1998
Earlier work this paper cites.
Beating the hold-out: Bounds for k-fold and progressive cross-validation
Blum, A., Kalai, A., and Langford, J · 1999
Earlier work this paper cites.
Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods
Platt, J. C · 1999
Earlier work this paper cites.
Random forests
Breiman, L · 2001
Earlier work this paper cites.
Using confidence bounds for exploitation-exploration trade-offs
Auer, P · 2003
Earlier work this paper cites.
Relating reinforcement learning performance to classification performance
Langford, J. and Zadrozny, B · 2005
Earlier work this paper cites.
A note on platt’s probabilistic outputs for support vector machines
Lin, H.-T., Lin, C.-J., and Weng, R. C · 2007
Earlier work this paper cites.
Efficient bandit algorithms for online multiclass prediction
Kakade, S. M., Shalev-Shwart, S., and Tewari, A · 2008
Earlier work this paper cites.
The epoch-greedy algorithm for multi-armed bandits with side information
Langford, J. and Zhang, T · 2008
Earlier work this paper cites.
The offset tree for learning with partial labels
Beygelzimer, A. and Langford, J · 2009
Earlier work this paper cites.
Search-based structured prediction
Daumé, III, H., Langford, J., and Marcu, D · 2009
Cited alongside, same era.
Efficient optimal learning for contextual bandits
Dudik, M., Hsu, D., Kale, S., Karampatziakis, N., Langford, J., Reyzin, L., and Zhang, T · 2011
Cited alongside, same era.
Scikit-learn: Machine learning in Python
Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., and Duchesnay, E · 2011
Cited alongside, same era.
A reduction of imitation learning and structured prediction to no-regret online learning
Ross, S., Gordon, G., and Bagnell, J. A · 2011
Cited alongside, same era.
Meta-learning of exploration/exploitation strategies: The multi-armed bandit case
Maes, F., Wehenkel, L., and Ernst, D · 2012
Cited alongside, same era.
Multi-armed bandits: Competing with optimal sequences
Karnin, Z. S. and Anava, O · 2016
Later among the works it cites.
Li, K. and Malik, J · 2016
Later among the works it cites.
Deep exploration via bootstrapped dqn
Osband, I., Blundell, C., Pritzel, A., and Van Roy, B · 2016
Later among the works it cites.
Neural architecture search with reinforcement learning
Zoph, B. and Le, Q. V · 2016
Later among the works it cites.
Learning algorithms for active learning
Bachman, P., Sordoni, A., and Trischler, A · 2017
Later among the works it cites.
Learning how to active learn: A deep reinforcement learning approach
Fang, M., Li, Y., and Cohn, T · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Taming the monster: A fast and simple algorithm for contextual bandits
Agarwal, A., Hsu, D., Kale, S., Langford, J., Li, L., and Schapire, R. E · 2014
Cited alongside, same era.
Causal discovery with continuous additive noise models
Peters, J., Mooij, J. M., Janzing, D., and Schölkopf, B · 2014
Cited alongside, same era.
Reinforcement and imitation learning via interactive no-regret learning
Ross, S. and Bagnell, J. A · 2014
Cited alongside, same era.
Learning to search better than your teacher
Chang, K.-W., Krishnamurthy, A., Agarwal, A., Daumé, III, H., and Langford, J · 2015
Cited alongside, same era.
Learning to learn by gradient descent by gradient descent
Andrychowicz, M., Denil, M., Gomez, S., Hoffman, M. W., Pfau, D., Schaul, T., and de Freitas, N · 2016
Cited alongside, same era.
Random forest for the contextual bandit problem
Féraud, R., Allesiardo, R., Urvoy, T., and Clérot, F · 2016
Cited alongside, same era.
Introduction to online convex optimization
Hazan, E. et al · 2016
Cited alongside, same era.
Later among the works it cites.
Learning active learning from data
Konyushkova, K., Sznitman, R., and Fua, P · 2017
Later among the works it cites.
Woodward, M. and Finn, C · 2017
Later among the works it cites.
A Contextual Bandit Bake-off
Bietti, A., Agarwal, A., and Langford, J · 2018
Later among the works it cites.
Meta-reinforcement learning of structured exploration strategies
Gupta, A., Mendonca, R., Liu, Y., Abbeel, P., and Levine, S · 2018
Later among the works it cites.
Learning to explore with meta-policy gradient
Xu, T., Liu, Q., Zhao, L., Xu, W., and Peng, J · 2018
Later among the works it cites.