Fetching the paper…
Reading the bibliography…
This work addresses the efficiency concern on inferring a nonlinear contextual bandit when the number of arms $n$ is very large.
1901
Earlier work this paper cites.
1908
Earlier work this paper cites.
H. Robbins, S. Monro, A stochastic approximation method, The annals of mathematical statistics (1951) 400–407
1951
Earlier work this paper cites.
J. R. Blum, Multidimensional stochastic approximation methods, The Annals of Mathematical Statistics (1954) 737–744
1954
Earlier work this paper cites.
P. Auer, N. Cesa-Bianchi, Y. Freund, R. E. Schapire, The nonstochastic multiarmed bandit problem, SIAM journal on computing 32 (1) (2002) 48–77
2002
Earlier work this paper cites.
C. Andrieu, N. De Freitas, A. Doucet, M. I. Jordan, An introduction to mcmc for machine learning, Machine learning 50 (1-2) (2003) 5–43
2003
Earlier work this paper cites.
R. Martí, Multi-start methods, in: Handbook of metaheuristics, Springer, 2003, pp. 355–368
2003
Earlier work this paper cites.
Y. Wang, J.-Y. Audibert, R. Munos, Infinitely many-armed bandits, in: Advances in Neural Information Processing Systems, 2008
2008
Earlier work this paper cites.
G. Tsoumakas, I. Katakis, I. Vlahavas, Effective and efficient multilabel classification in domains with large number of labels, in: Proc. ECML/PKDD 2008 Workshop on Mining Multidimensional Data (MMD’08), Vol. 21, 2008, pp. 53–59
2008
Earlier work this paper cites.
2008
Earlier work this paper cites.
S. Filippi, O. Cappe, A. Garivier, C. Szepesvári, Parametric bandits: The generalized linear case., in: NIPS, Vol. 23, 2010, pp. 586–594
2010
Earlier work this paper cites.
W. Chu, L. Li, L. Reyzin, R. Schapire, Contextual bandits with linear payoff functions, in: Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics, JMLR Workshop and Conference Proceedings, 2011, pp. 208–214
2011
Earlier work this paper cites.
Y. Abbasi-Yadkori, D. Pál, C. Szepesvári, Improved algorithms for linear stochastic bandits., in: NIPS, Vol. 11, 2011, pp. 2312–2320
2011
Earlier work this paper cites.
O. Chapelle, L. Li, An empirical evaluation of thompson sampling, in: Advances in neural information processing systems, 2011, pp. 2249–2257
2011
Earlier work this paper cites.
Y. Gai, B. Krishnamachari, R. Jain, Combinatorial network optimization with unknown variables: Multi-armed bandits with linear rewards and individual observations, IEEE/ACM Transactions on Networking 20 (5) (2012) 1466–1478
2012
Cited alongside, same era.
Y. Malkov, A. Ponomarenko, A. Logvinov, V. Krylov, Scalable distributed algorithm for approximate nearest neighbor search problem in high dimensional general metric spaces, in: International Conference on Similarity Search and Applications, Springer, 2012, pp. 132–147
2012
Cited alongside, same era.
S. Agrawal, N. Goyal, Further optimal regret bounds for thompson sampling, in: Artificial intelligence and statistics, PMLR, 2013, pp. 99–107
2013
Cited alongside, same era.
R. Martí, M. G. Resende, C. C. Ribeiro, Multi-start methods for combinatorial optimization, European Journal of Operational Research 226 (1) (2013) 1–8
2013
Cited alongside, same era.
S. Mandt, M. D. Hoffman, D. M. Blei, Stochastic gradient descent as approximate bayesian inference, The Journal of Machine Learning Research 18 (1) (2017) 4873–4907
2017
Later among the works it cites.
D. M. Blei, A. Kucukelbir, J. D. McAuliffe, Variational inference: A review for statisticians, Journal of the American statistical Association 112 (518) (2017) 859–877
2017
Later among the works it cites.
Y. Gal, J. Hron, A. Kendall, Concrete dropout, in: Advances in neural information processing systems, 2017, pp. 3581–3590
2017
Later among the works it cites.
S. R. Chowdhury, A. Gopalan, On kernelized multi-armed bandits, in: International Conference on Machine Learning, PMLR, 2017, pp. 844–853
2017
Later among the works it cites.
I. Urteaga, C. Wiggins, Variational inference for the multi-armed contextual bandit, in: International Conference on Artificial Intelligence and Statistics, PMLR, 2018, pp. 698–706
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
R. Allesiardo, R. Féraud, D. Bouneffouf, A neural networks committee for the contextual bandit problem, in: International Conference on Neural Information Processing, Springer, 2014, pp. 374–381
2014
Cited alongside, same era.
B. Kveton, C. Szepesvari, Z. Wen, A. Ashkan, Cascading bandits: Learning to rank in the cascade model, in: International Conference on Machine Learning, PMLR, 2015, pp. 767–776
2015
Cited alongside, same era.
R. Combes, S. Magureanu, A. Proutiere, C. Laroche, Learning to rank: Regret lower bounds and efficient algorithms, in: Proceedings of the 2015 ACM SIGMETRICS International Conference on Measurement and Modeling of Computer Systems, 2015, pp. 231–244
2015
Cited alongside, same era.
Z. Liu, P. Luo, X. Wang, X. Tang, Deep learning face attributes in the wild, in: Proceedings of International Conference on Computer Vision (ICCV), 2015
2015
Cited alongside, same era.
F. M. Harper, J. A. Konstan, The movielens datasets: History and context, Acm transactions on interactive intelligent systems (tiis) 5 (4) (2015) 1–19
2015
Cited alongside, same era.
A. Carpentier, M. Valko, Revealing graph bandits for maximizing local influence, in: Artificial Intelligence and Statistics, PMLR, 2016, pp. 10–18
2016
Cited alongside, same era.
E. Hazan, Z. Karnin, Volumetric spanners: an efficient exploration basis for learning, The Journal of Machine Learning Research 17 (1) (2016) 4062–4095
2016
Cited alongside, same era.
D. Russo, B. Van Roy, An information-theoretic analysis of thompson sampling, The Journal of Machine Learning Research 17 (1) (2016) 2442–2471
2016
Cited alongside, same era.
2018
Later among the works it cites.
Z. Lipton, X. Li, J. Gao, L. Li, F. Ahmed, L. Deng, Bbq-networks: Efficient exploration in deep reinforcement learning for task-oriented dialogue systems, in: Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 32, 2018
2018
Later among the works it cites.
K. Azizzadenesheli, E. Brunskill, A. Anandkumar, Efficient exploration through bayesian deep q-networks, in: 2018 Information Theory and Applications Workshop (ITA), IEEE, 2018, pp. 1–9
2018
Later among the works it cites.
Y. A. Malkov, D. A. Yashunin, Efficient and robust approximate nearest neighbor search using hierarchical navigable small world graphs, IEEE transactions on pattern analysis and machine intelligence 42 (4) (2018) 824–836
2018
Later among the works it cites.
E. Fouché, J. Komiyama, K. Böhm, Scaling multi-armed bandit algorithms, in: Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, 2019, pp. 1449–1459
2019
Later among the works it cites.
B. Tóth, S. Sachidanandan, E. S. Jørgensen, Balancing relevance and discovery to inspire customers in the ikea app, in: Fourteenth ACM Conference on Recommender Systems, 2020, pp. 563–563
2020
Later among the works it cites.
D. Guo, S. I. Ktena, P. K. Myana, F. Huszar, W. Shi, A. Tejani, M. Kneier, S. Das, Deep bayesian bandits: Exploring in online personalized recommendations, in: Fourteenth ACM Conference on Recommender Systems, 2020, pp. 456–461
2020
Later among the works it cites.
D. Zhou, L. Li, Q. Gu, Neural contextual bandits with ucb-based exploration, in: International Conference on Machine Learning, PMLR, 2020, pp. 11492–11502
2020
Later among the works it cites.
R. Guo, P. Sun, E. Lindgren, Q. Geng, D. Simcha, F. Chern, S. Kumar, Accelerating large-scale inference with anisotropic vector quantization, in: International Conference on Machine Learning, PMLR, 2020, pp. 3887–3896
2020
Later among the works it cites.