Fetching the paper…
Reading the bibliography…
We create a computationally tractable algorithm for contextual bandits with continuous actions having unknown structure.
A generalization of sampling without replacement from a finite universe
Daniel G Horvitz and Donovan J Thompson · 1952
Earlier work this paper cites.
On tail probabilities for martingales
David A Freedman · 1975
Earlier work this paper cites.
Learning decision lists
Ronald L Rivest · 1987
Earlier work this paper cites.
Learning decision trees from random examples
Andrzej Ehrenfeucht and David Haussler · 1989
Earlier work this paper cites.
Rank-r decision trees are a subclass of r-decision lists
Avrim Blum · 1992
Earlier work this paper cites.
Learning decision trees using the fourier spectrum
Eyal Kushilevitz and Yishay Mansour · 1993
Earlier work this paper cites.
The continuum-armed bandit problem
Rajeev Agrawal · 1995
Earlier work this paper cites.
The nature of statistical learning theory
Vladimir Vapnik · 1995
Earlier work this paper cites.
Lower bounds on learning decision lists and trees
Thomas Hancock, Tao Jiang, Ming Li, and John Tromp · 1996
Earlier work this paper cites.
Beating the hold-out: Bounds for k-fold and progressive cross-validation
Avrim Blum, Adam Kalai, and John Langford · 1999
Earlier work this paper cites.
The nonstochastic multiarmed bandit problem
Peter Auer, Nicolo Cesa-Bianchi, Yoav Freund, and Robert E Schapire · 2002
Earlier work this paper cites.
Reinforcement learning with immediate rewards and linear hypotheses
Naoki Abe, Alan W Biermann, and Philip M Long · 2003
Earlier work this paper cites.
Nearly tight bounds for the continuum-armed bandit problem
Robert Kleinberg · 2004
Earlier work this paper cites.
Weighted one-against-all
Alina Beygelzimer, John Langford, and Bianca Zadrozny · 2005
Earlier work this paper cites.
Improved rates for the stochastic continuum-armed bandit problem
Peter Auer, Ronald Ortner, and Csaba Szepesvári · 2007
Earlier work this paper cites.
The epoch-greedy algorithm for contextual multi-armed bandits
John Langford and Tong Zhang · 2007
Earlier work this paper cites.
High-probability regret bounds for bandit online linear optimization
Peter L Bartlett, Varsha Dani, Thomas Hayes, Sham Kakade, Alexander Rakhlin, and Ambuj Tewari · 2008
Earlier work this paper cites.
Agnostically learning decision trees
Parikshit Gopalan, Adam Tauman Kalai, and Adam R Klivans · 2008
Earlier work this paper cites.
Multi-armed bandits in metric spaces
Robert Kleinberg, Aleksandrs Slivkins, and Eli Upfal · 2008
Earlier work this paper cites.
The offset tree for learning with partial labels
Alina Beygelzimer and John Langford · 2009
Cited alongside, same era.
Error-correcting tournaments
Alina Beygelzimer, John Langford, and Pradeep Ravikumar · 2009
Cited alongside, same era.
Estimation of the warfarin dose with clinical and pharmacogenetic data
TE Klein, RB Altman, Niclas Eriksson, BF Gage, SE Kimmel, MT Lee, NA Limdi, D Page, DM Roden, MJ Wagner, et al · 2009
Cited alongside, same era.
Bandits and experts in metric spaces
Robert Kleinberg, Aleksandrs Slivkins, and Eli Upfal · 2010
Cited alongside, same era.
A contextual-bandit approach to personalized news article recommendation
Lihong Li, Wei Chu, John Langford, and Robert E Schapire · 2010
Cited alongside, same era.
X-armed bandits
Sébastien Bubeck, Rémi Munos, Gilles Stoltz, and Csaba Szepesvári · 2011
Efficient algorithms for adversarial contextual learning
Vasilis Syrgkanis, Akshay Krishnamurthy, and Robert Schapire · 2016
Later among the works it cites.
Making contextual decisions with low technical debt
Alekh Agarwal, Sarah Bird, Markus Cozowicz, Luong Hoang, John Langford, Stephen Lee, Jiaji Li, Dan Melamed, Gal Oshri, Oswaldo Ribas, Siddhartha Sen, and Alex Slivkins · 2017
Later among the works it cites.
Corralling a band of bandit algorithms
Alekh Agarwal, Haipeng Luo, Behnam Neyshabur, and Robert E Schapire · 2017
Later among the works it cites.
Algorithmic chaining and the role of partial feedback in online nonparametric learning
Nicolò Cesa-Bianchi, Pierre Gaillard, Claudio Gentile, and Sébastien Gerchinovitz · 2017
Later among the works it cites.
On context-dependent clustering of bandits
Claudio Gentile, Shuai Li, Purushottam Kar, Alexandros Karatzoglou, Giovanni Zappella, and Evans Etrue · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Efficient optimal learning for contextual bandits
Miroslav Dudik, Daniel Hsu, Satyen Kale, Nikos Karampatziakis, John Langford, Lev Reyzin, and Tong Zhang · 2011
Cited alongside, same era.
Doubly robust policy evaluation and learning
Miroslav Dudík, John Langford, and Lihong Li · 2011
Cited alongside, same era.
Contextual gaussian process bandit optimization
Andreas Krause and Cheng S. Ong · 2011
Cited alongside, same era.
Multi-armed bandits on implicit metric spaces
Aleksandrs Slivkins · 2011
Cited alongside, same era.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
Sébastien Bubeck, Nicolo Cesa-Bianchi, et al · 2012
Cited alongside, same era.
Estimation of extreme values and associated level sets of a regression function via selective sampling
Stanislav Minsker · 2013
Cited alongside, same era.
From ads to interventions: Contextual bandits in mobile health
Ambuj Tewari and Susan A Murphy · 2017
Later among the works it cites.
Alberto Bietti, Alekh Agarwal, and John Langford · 2018
Later among the works it cites.
Policy evaluation and optimization with continuous treatments
Nathan Kallus and Angela Zhou · 2018
Later among the works it cites.
Bandit algorithms
Tor Lattimore and Csaba Szepesvári · 2018
Later among the works it cites.
Data center cooling using model-predictive control
Nevena Lazic, Craig Boutilier, Tyler Lu, Eehern Wong, Binz Roy, MK Ryu, and Greg Imwalle · 2018
Later among the works it cites.
Top-down induction of decision trees: rigorous guarantees and inherent limitations
Guy Blanc, Jane Lange, and Li-Yang Tan · 2019
Later among the works it cites.
On the optimality of trees generated by id3
Alon Brutzkus, Amit Daniely, and Eran Malach · 2019
Later among the works it cites.
A deep reinforcement learning perspective on internet congestion control
Nathan Jay, Noga Rotman, Brighten Godfrey, Michael Schapira, and Aviv Tamar · 2019
Later among the works it cites.
Contextual bandits with continuous actions: smoothing, zooming, and adapting
Akshay Krishnamurthy, John Langford, Aleksandrs Slivkins, and Chicheng Zhang · 2019
Later among the works it cites.
Improved algorithm on online clustering of bandits
Shuai Li, Wei Chen, Shuai Li, and Kwong-Sak Leung · 2019
Later among the works it cites.
Introduction to multi-armed bandits
Aleksandrs Slivkins · 2019
Later among the works it cites.
Towards practical lipschitz stochastic bandits
Tianyu Wang, Weicheng Ye, Dawei Geng, and Cynthia Rudin · 2019
Later among the works it cites.
Fast distributed bandits for online recommendation systems
Kanak Mahadik, Qingyun Wu, Shuai Li, and Amit Sabne · 2020
Closest in time.