Fetching the paper…
Reading the bibliography…
We study the problem of offline policy optimization in stochastic contextual bandit problems, where the goal is to learn a near-optimal policy based on a dataset of decision data collected by a suboptimal behavior policy.
A generalization of sampling without replacement from a finite universe
Daniel G Horvitz and Donovan J Thompson · 1952
Earlier work this paper cites.
Some PAC-Bayesian theorems
David A. McAllester · 1998
Earlier work this paper cites.
Smoothed analysis of algorithms: why the simplex algorithm usually takes polynomial time
Daniel A. Spielman and Shang-Hua Teng · 2001
Earlier work this paper cites.
The nonstochastic multiarmed bandit problem
Peter Auer, Nicolò Cesa-Bianchi, Yoav Freund, and Robert E. Schapire · 2002
Earlier work this paper cites.
Optimal dynamic treatment regimes
Susan A Murphy · 2003
Earlier work this paper cites.
PAC-Bayesian statistical learning theory
Jean-Yves Audibert · 2004
Earlier work this paper cites.
PAC-Bayesian supervised classification
Olivier Catoni · 2007
Earlier work this paper cites.
The epoch-greedy algorithm for multi-armed bandits with side information
John Langford and Tong Zhang · 2007
Earlier work this paper cites.
Truncated importance sampling
Edward L Ionides · 2008
Earlier work this paper cites.
Efficient optimal learning for contextual bandits
Miroslav Dudík, Daniel J. Hsu, Satyen Kale, Nikos Karampatziakis, John Langford, Lev Reyzin, and Tong Zhang · 2011
Earlier work this paper cites.
Cancer discovery , 1(1):44–53, 2011
Edward S. Kim, Roy S. Herbst, Ignacio I. Wistuba, J. Jack Lee, Jr. Blumenschein, George R., Anne Tsao, David J. Stewart, Marshall E. Hicks, Jr. Erasmus, Jeremy, Sanjay Gupta, Christine M. Alden, Suyu Liu, Ximing Tang, Fadlo R. Khuri, Hai T. Tran, Bruce E. Johnson, John V. Heymach, Li Mao, Frank Fossella, Merrill S. Kies, Vassiliki Papadimitrakopoulou, Suzanne E. Davis, Scott M. Lippman, and Waun K. Hong · 2011
Earlier work this paper cites.
Unbiased offline evaluation of contextual-bandit-based news article recommendation algorithms
Lihong Li, Wei Chu, John Langford, and Xuanhui Wang · 2011
Earlier work this paper cites.
Counterfactual reasoning and learning systems: the example of computational advertising
Léon Bottou, Jonas Peters, Joaquin Quiñonero Candela, Denis Xavier Charles, Max Chickering, Elon Portugaly, Dipankar Ray, Patrice Y. Simard, and Ed Snelson · 2013
Cited alongside, same era.
Concentration Inequalities - A Nonasymptotic Theory of Independence
Stéphane Boucheron, Gábor Lugosi, and Pascal Massart · 2013
Cited alongside, same era.
Taming the monster: A fast and simple algorithm for contextual bandits
Alekh Agarwal, Daniel J. Hsu, Satyen Kale, John Langford, Lihong Li, and Robert E. Schapire · 2014
Cited alongside, same era.
Efficient learning by implicit exploration in bandit problems with side observations
Tomáš Kocák, Gergely Neu, Michal Valko, and Rémi Munos · 2014
Cited alongside, same era.
Toward minimax off-policy value estimation
Lihong Li, Rémi Munos, and Csaba Szepesvári · 2015
Cited alongside, same era.
Bayesian counterfactual risk minimization
Ben London and Ted Sandler · 2019
Later among the works it cites.
User-friendly introduction to PAC-Bayes bounds
Pierre Alquier · 2021
Later among the works it cites.
PAC-Bayes, MAC-Bayes and conditional mutual information: Fast rate bounds that handle general VC classes
Peter Grünwald, Thomas Steinke, and Lydia Zakynthinou · 2021
Later among the works it cites.
Is pessimism provably efficient for offline rl?
Ying Jin, Zhuoran Yang, and Zhaoran Wang · 2021
Later among the works it cites.
Bridging offline reinforcement learning and imitation learning: A tale of pessimism
Paria Rashidinejad, Banghua Zhu, Cong Ma, Jiantao Jiao, and Stuart Russell · 2021
Later among the works it cites.
On the optimality of batch policy optimization algorithms
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gergely Neu · 2015
Cited alongside, same era.
Batch learning from logged bandit feedback through counterfactual risk minimization
Adith Swaminathan and Thorsten Joachims · 2015
Cited alongside, same era.
Recommendations as treatments: Debiasing learning and evaluation
Tobias Schnabel, Adith Swaminathan, Ashudeep Singh, Navin Chandak, and Thorsten Joachims · 2016
Cited alongside, same era.
Personalized diabetes management using electronic medical records
Dimitris Bertsimas, Nathan Kallus, Alexander M Weinstein, and Ying Daisy Zhuo · 2017
Cited alongside, same era.
Mobile health
James M Rehg, Susan A Murphy, and Santosh Kumar · 2017
Cited alongside, same era.
Learning preferences with side information
Vivek F. Farias and Andrew A. Li · 2019
Cited alongside, same era.
Contextual bandits with continuous actions: Smoothing, zooming, and adapting
Akshay Krishnamurthy, John Langford, Aleksandrs Slivkins, and Chicheng Zhang · 2019
Cited alongside, same era.
Chenjun Xiao, Yifan Wu, Jincheng Mei, Bo Dai, Tor Lattimore, Lihong Li, Csaba Szepesvári, and Dale Schuurmans · 2021
Later among the works it cites.
Policy learning “without” overlap: Pessimism and generalized empirical Bernstein’s inequality
Ying Jin, Zhimei Ren, Zhuoran Yang, and Zhaoran Wang · 2022
Later among the works it cites.
Pessimism for offline linear contextual bandits using ℓ p \ell_{p} confidence sets
Gene Li, Cong Ma, and Nati Srebro · 2022
Later among the works it cites.
PAC-Bayes bounds for bandit problems: A survey and experimental comparison
Hamish Flynn, David Reeb, Melih Kandemir, and Jan Peters · 2023
Closest in time.
Online learning with off-policy feedback
Germano Gabbianelli, Gergely Neu, and Matteo Papini · 2023
Closest in time.
PAC-Bayesian offline contextual bandits with guarantees
Otmane Sakhi, Pierre Alquier, and Nicolas Chopin · 2023
Closest in time.
Oracle-efficient pessimism: Offline policy optimization in contextual bandits
Lequn Wang, Akshay Krishnamurthy, and Aleksandrs Slivkins · 2023
Closest in time.