Fetching the paper…
Reading the bibliography…
Motivated by modern applications, such as online advertisement and recommender systems, we study the top-$k$ extreme contextual bandits problem, where the total number of arms can be enormous, and the learner is allowed to select $k$ arms and observe all or some of the rewards for the chosen arms.
Associative reinforcement learning using linear probabilistic concepts
Naoki Abe and Philip M Long · 1999
Earlier work this paper cites.
Co-clustering documents and words using bipartite spectral graph partitioning
Inderjit S Dhillon · 2001
Earlier work this paper cites.
The nonstochastic multiarmed bandit problem
Peter Auer, Nicolo Cesa-Bianchi, Yoav Freund, and Robert E Schapire · 2002
Earlier work this paper cites.
The epoch-greedy algorithm for contextual multi-armed bandits
John Langford and Tong Zhang · 2007
Earlier work this paper cites.
High-probability regret bounds for bandit online linear optimization
Peter Bartlett, Varsha Dani, Thomas Hayes, Sham Kakade, Alexander Rakhlin, and Ambuj Tewari · 2008
Earlier work this paper cites.
Stochastic linear optimization under bandit feedback
Varsha Dani, Thomas P Hayes, and Sham M Kakade · 2008
Earlier work this paper cites.
Liblinear: A library for large linear classification
Rong-En Fan, Kai-Wei Chang, Cho-Jui Hsieh, Xiang-Rui Wang, and Chih-Jen Lin · 2008
Earlier work this paper cites.
Tighter bounds for multi-armed bandits with expert advice
H Brendan McMahan and Matthew Streeter · 2009
Earlier work this paper cites.
Parametric bandits: The generalized linear case
Sarah Filippi, Olivier Cappe, Aurélien Garivier, and Csaba Szepesvári · 2010
Earlier work this paper cites.
Dylan J Foster, Alexander Rakhlin, David Simchi-Levi, and Yunzong Xu · 2010
Earlier work this paper cites.
Eigen v3
Gaël Guennebaud, Benoît Jacob, et al · 2010
Earlier work this paper cites.
Contextual bandit algorithms with supervised learning guarantees
Alina Beygelzimer, John Langford, Lihong Li, Lev Reyzin, and Robert Schapire · 2011
Earlier work this paper cites.
Contextual bandits with linear payoff functions
Wei Chu, Lihong Li, Lev Reyzin, and Robert Schapire · 2011
Earlier work this paper cites.
Contextual gaussian process bandit optimization
Andreas Krause and Cheng Ong · 2011
Earlier work this paper cites.
Linear submodular bandits and their application to diversified retrieval
Yisong Yue and Carlos Guestrin · 2011
Earlier work this paper cites.
Contextual bandit learning with predictable rewards
Alekh Agarwal, Miroslav Dudík, Satyen Kale, John Langford, and Robert Schapire · 2012
Earlier work this paper cites.
Combinatorial bandits
Nicolo Cesa-Bianchi and Gábor Lugosi · 2012
Cited alongside, same era.
Thompson sampling for contextual bandits with linear payoffs
Shipra Agrawal and Navin Goyal · 2013
Cited alongside, same era.
Taming the monster: A fast and simple algorithm for contextual bandits
Alekh Agarwal, Daniel Hsu, Satyen Kale, John Langford, Lihong Li, and Robert Schapire · 2014
Cited alongside, same era.
Combinatorial partial monitoring game with linear feedback and its applications
Tian Lin, Bruno Abrahao, Robert Kleinberg, John Lui, and Wei Chen · 2014
Cited alongside, same era.
Contextual combinatorial bandit and its application on diversified online recommendation
Lijing Qin, Shouyuan Chen, and Xiaoyan Zhu · 2014
Cited alongside, same era.
Combinatorial bandits revisited
Richard Combes, Mohammad Sadegh Talebi Mazraeh Shahi, Alexandre Proutiere, et al · 2015
Alberto Bietti, Alekh Agarwal, and John Langford · 2018
Later among the works it cites.
Contextual bandits for adapting treatment in a mouse model of de novo carcinogenesis
Audrey Durand, Charis Achilleos, Demetris Iacovides, Katerina Strati, Georgios D Mitsis, and Joelle Pineau · 2018
Later among the works it cites.
Practical contextual bandits with regression oracles
Dylan J Foster, Alekh Agarwal, Miroslav Dudík, Haipeng Luo, and Robert E Schapire · 2018
Later among the works it cites.
Parabel: Partitioned label trees for extreme classification with application to dynamic search advertising
Yashoteja Prabhu, Anil Kag, Shrutendra Harsola, Rahul Agrawal, and Manik Varma · 2018
Later among the works it cites.
A no-regret generalization of hierarchical softmax to extreme multi-label classification
Marek Wydmuch, Kalina Jasinska, Mikhail Kuznetsov, Róbert Busa-Fekete, and Krzysztof Dembczynski · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Tight regret bounds for stochastic combinatorial semi-bandits
Branislav Kveton, Zheng Wen, Azin Ashkan, and Csaba Szepesvari · 2015
Cited alongside, same era.
Multi-armed bandit models for the optimal design of clinical trials: benefits and challenges
Sofía S Villar, Jack Bowden, and James Wason · 2015
Cited alongside, same era.
The extreme classification repository: Multi-label datasets and code, 2016
K. Bhatia, K. Dahiya, H. Jain, A. Mittal, Y. Prabhu, and M. Varma · 2016
Cited alongside, same era.
Combinatorial multi-armed bandit with general reward functions
Wei Chen, Wei Hu, Fu Li, Jian Li, Yu Liu, and Pinyan Lu · 2016
Cited alongside, same era.
Extreme f-measure maximization using sparse probability estimates
Kalina Jasinska, Krzysztof Dembczynski, Róbert Busa-Fekete, Karlson Pfannschmidt, Timo Klerx, and Eyke Hullermeier · 2016
Cited alongside, same era.
Collaborative filtering bandits
Shuai Li, Alexandros Karatzoglou, and Claudio Gentile · 2016
Cited alongside, same era.
Later among the works it cites.
Batch-size independent regret bounds for the combinatorial multi-armed bandit problem
Nadav Merlis and Shie Mannor · 2019
Later among the works it cites.
Efficient counterfactual learning from bandit feedback
Yusuke Narita, Shota Yasui, and Kohei Yata · 2019
Later among the works it cites.
Attentionxml: Label tree-based attention-aware deep model for high-performance extreme multi-label text classification
Ronghui You, Zihan Zhang, Ziye Wang, Suyang Dai, Hiroshi Mamitsuka, and Shanfeng Zhu · 2019
Later among the works it cites.
Beyond ucb: Optimal and efficient contextual bandits with regression oracles
Dylan J Foster and Alexander Rakhlin · 2020
Later among the works it cites.
Bonsai: diverse and shallow trees for extreme multi-label classification
Sujay Khandagale, Han Xiao, and Rohit Babbar · 2020
Later among the works it cites.
Learning from extreme bandit feedback
Romain Lopez, Inderjit Dhillon, and Michael I Jordan · 2020
Later among the works it cites.
Efficient contextual bandits with continuous actions
Maryam Majzoubi, Chicheng Zhang, Rajan Chari, Akshay Krishnamurthy, John Langford, and Aleksandrs Slivkins · 2020
Later among the works it cites.
Top- k k combinatorial bandits with full-bandit feedback
Idan Rejwan and Yishay Mansour · 2020
Later among the works it cites.
Bypassing the monster: A faster and simpler optimal algorithm for contextual bandits under realizability
David Simchi-Levi and Yunzong Xu · 2020
Later among the works it cites.
Pecos: Prediction for enormous and correlated output spaces
Hsiang-Fu Yu, Kai Zhong, and Inderjit S Dhillon · 2020
Later among the works it cites.