Fetching the paper…
Reading the bibliography…
We consider offline policy optimization (OPO) in contextual bandits, where one is given a fixed dataset of logged interactions.
A generalization of sampling without replacement from a finite universe
D. G. Horvitz and D. J. Thompson · 1952
Earlier work this paper cites.
Probability inequalities for the sum of independent random variables
G. Bennett · 1962
Earlier work this paper cites.
The nonstochastic multiarmed bandit problem
P. Auer, N. Cesa-Bianchi, Y. Freund, and R. E. Schapire · 1995
Earlier work this paper cites.
A general agnostic active learning algorithm
S. Dasgupta, D. J. Hsu, and C. Monteleoni · 2007
Earlier work this paper cites.
The Epoch-Greedy Algorithm for Contextual Multi-armed Bandits
J. Langford and T. Zhang · 2007
Earlier work this paper cites.
Exploration scavenging
J. Langford, A. Strehl, and J. Wortman · 2008
Earlier work this paper cites.
The offset tree for learning with partial labels
A. Beygelzimer and J. Langford · 2009
Earlier work this paper cites.
Search-based structured prediction
H. Daumé, J. Langford, and D. Marcu · 2009
Earlier work this paper cites.
Empirical bernstein bounds and sample variance penalization
A. Maurer and M. Pontil · 2009
Earlier work this paper cites.
Showing Relevant Ads via Lipschitz Context Multi-Armed Bandits
T. Lu, D. Pál, and M. Pál · 2010
Earlier work this paper cites.
Learning from logged implicit exploration data
A. Strehl, J. Langford, L. Li, and S. M. Kakade · 2010
Earlier work this paper cites.
Efficient optimal learning for contextual bandits
M. Dudik, D. Hsu, S. Kale, N. Karampatziakis, J. Langford, L. Reyzin, and T. Zhang · 2011
Earlier work this paper cites.
Doubly robust policy evaluation and optimization
M. Dudík, D. Erhan, J. Langford, and L. Li · 2011
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
S. Ross, G. Gordon, and D. Bagnell · 2011
Earlier work this paper cites.
Contextual bandits with similarity information
A. Slivkins · 2011
Earlier work this paper cites.
Counterfactual reasoning and learning systems: The example of computational advertising
L. Bottou, J. Peters, J. Quiñonero-Candela, D. X. Charles, D. M. Chickering, E. Portugaly, D. Ray, P. Simard, and E. Snelson · 2013
Earlier work this paper cites.
Openml: networked science in machine learning
J. Vanschoren, J. N. van Rijn, B. Bischl, and L. Torgo · 2013
Earlier work this paper cites.
Taming the monster: A fast and simple algorithm for contextual bandits
A. Agarwal, D. Hsu, S. Kale, J. Langford, L. Li, and R. Schapire · 2014
Earlier work this paper cites.
Batch learning from logged bandit feedback through counterfactual risk minimization
A. Swaminathan and T. Joachims · 2015
Earlier work this paper cites.
Recursive partitioning for heterogeneous causal effects
S. Athey and G. Imbens · 2016
Cited alongside, same era.
Optimal and adaptive off-policy evaluation in contextual bandits
Y.-X. Wang, A. Agarwal, and M. Dudık · 2017
Cited alongside, same era.
More robust doubly robust off-policy evaluation
M. Farajtabar, Y. Chow, and M. Ghavamzadeh · 2018
Cited alongside, same era.
Practical contextual bandits with regression oracles
D. Foster, A. Agarwal, M. Dudík, H. Luo, and R. Schapire · 2018
Cited alongside, same era.
Deep learning with logged bandit feedback
T. Joachims, A. Swaminathan, and M. De Rijke · 2018
Cited alongside, same era.
Policy evaluation and optimization with continuous treatments
N. Kallus and A. Zhou · 2018
Cited alongside, same era.
A contextual bandit bake-off
A. Bietti, A. Agarwal, and J. Langford · 2021
Later among the works it cites.
Is pessimism provably efficient for offline rl?
Y. Jin, Z. Yang, and Z. Wang · 2021
Later among the works it cites.
Bridging offline reinforcement learning and imitation learning: A tale of pessimism
P. Rashidinejad, B. Zhu, C. Ma, J. Jiao, and S. Russell · 2021
Later among the works it cites.
Conservative objective models for effective offline model-based optimization
B. Trabucco, A. Kumar, X. Geng, and S. Levine · 2021
Later among the works it cites.
Instabilities of offline rl with pre-trained neural representation
R. Wang, Y. Wu, R. Salakhutdinov, and S. Kakade · 2021
Later among the works it cites.
Bellman-consistent pessimism for offline reinforcement learning
T. Xie, C.-A. Cheng, N. Jiang, P. Mineiro, and A. Agarwal · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Semi-parametric efficient policy learning with continuous actions
V. Chernozhukov, M. Demirer, G. Lewis, and V. Syrgkanis · 2019
Cited alongside, same era.
Off-policy deep reinforcement learning without exploration
S. Fujimoto, D. Meger, and D. Precup · 2019
Cited alongside, same era.
Bayesian counterfactual risk minimization
B. London and T. Sandler · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, Albanand Köpf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala · 2019
Cited alongside, same era.
Cab: Continuous adaptive blending for policy evaluation and learning
Y. Su, L. Wang, M. Santacatterina, and T. Joachims · 2019
Cited alongside, same era.
Morel: Model-based offline reinforcement learning
R. Kidambi, A. Rajeswaran, P. Netrapalli, and T. Joachims · 2020
Cited alongside, same era.
Contextual bandits in large action spaces: Made practical
Y. Zhu, D. J. Foster, P. Mineiro, and J. Langford · 2021
Later among the works it cites.
Smoothed online learning is as easy as statistical learning
A. Block, Y. Dagan, N. Golowich, and A. Rakhlin · 2022
Later among the works it cites.
Oracle-efficient online learning for beyond worst-case adversaries
N. Haghtalab, Y. Han, A. Shetty, and K. Yang · 2022
Later among the works it cites.
Policy learning “without” overlap: Pessimism and generalized empirical bernstein’s inequality
Y. Jin, Z. Ren, Z. Yang, and Z. Wang · 2022
Later among the works it cites.
Pessimism for offline linear contextual bandits using ℓ p \ell_{p} confidence sets
G. Li, C. Ma, and N. Srebro · 2022
Later among the works it cites.
Offline neural contextual bandits: Pessimism, optimization and generalization
T. Nguyen-Tang, S. Gupta, A. T. Nguyen, and S. Venkatesh · 2022
Later among the works it cites.
Optimal conservative offline rl with general function approximation via augmented lagrangian
P. Rashidinejad, H. Zhu, K. Yang, S. Russell, and J. Jiao · 2022
Later among the works it cites.
Introduction to multi-armed bandits
A. Slivkins · 2022
Later among the works it cites.
Pessimistic model-based offline reinforcement learning under partial coverage
M. Uehara and W. Sun · 2022
Later among the works it cites.
Contextual bandits with smooth regret: Efficient learning in continuous action spaces
Y. Zhu and P. Mineiro · 2022
Later among the works it cites.
Exponential smoothing for off-policy learning
I. Aouali, V.-E. Brunel, D. Rohde, and A. Korba · 2023
Closest in time.
Boosted off-policy learning
B. London, L. Lu, T. Sandler, and T. Joachims · 2023
Closest in time.
Pac-bayesian offline contextual bandits with guarantees
O. Sakhi, P. Alquier, and N. Chopin · 2023
Closest in time.