Fetching the paper…
Reading the bibliography…
We provide a sound and consistent foundation for the use of \emph{nonrandom} exploration data in "contextual bandit" or "partially labeled" settings where only the value of a chosen action is learned.
Probability inequalities for sums of bounded random variables
Hoeffding, Wassily · 1963
Earlier work this paper cites.
Safe and effective importance sampling
Owen, Art and Zhou, Yi · 1998
Earlier work this paper cites.
Approximate planning in large pomdps via reusable trajectories
Kearns, Michael, Mansour, Yishay, and Ng, Andrew Y · 2000
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Precup, Doina, Sutton, Rich, and Singh, Satinder · 2000
Cited alongside, same era.
The nonstochastic multiarmed bandit problem
Auer, Peter, Bianchi, Nicolò C., Freund, Yoav, and Schapire, Robert E · 2002
Cited alongside, same era.
The epoch-greedy algorithm for multi-armed bandits with side information
Langford, John and Zhang, Tong · 2008
Later among the works it cites.
Exploration scavenging
Langford, John, Strehl, Alexander L., and Wortman, Jenn · 2008
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…