Fetching the paper…
Reading the bibliography…
Contextual bandit algorithms are ubiquitous tools for active sequential experimentation in healthcare and the tech industry.
Design of experiments
Ronald Aylmer Fisher · 1936
Earlier work this paper cites.
Etude critique de la notion de collectif
Jean Ville · 1939
Earlier work this paper cites.
Asymptotic minimax character of the sample distribution function and of the classical multinomial estimator
Aryeh Dvoretzky, Jack Kiefer, and Jacob Wolfowitz · 1956
Earlier work this paper cites.
Probability Inequalities for Sums of Bounded Random Variables
Wassily Hoeffding · 1963
Earlier work this paper cites.
Randomized response: A survey technique for eliminating evasive answer bias
Stanley L Warner · 1965
Earlier work this paper cites.
Confidence sequences for mean, variance, and median
DA Darling and Herbert Robbins · 1967
Earlier work this paper cites.
Statistical methods related to the law of the iterated logarithm
Herbert Robbins · 1970
Earlier work this paper cites.
Forcing a sequential experiment to be balanced
Bradley Efron · 1971
Earlier work this paper cites.
Estimating causal effects of treatments in randomized and nonrandomized studies
Donald B Rubin · 1974
Earlier work this paper cites.
On confidence sequences
Tze Leung Lai · 1976
Earlier work this paper cites.
A new approach to causal inference in mortality studies with a sustained exposure period—application to control of the healthy worker survivor effect
James Robins · 1986
Earlier work this paper cites.
The tight constant in the Dvoretzky-Kiefer-Wolfowitz inequality
Pascal Massart · 1990
Earlier work this paper cites.
On the application of probability theory to agricultural experiments, essay on principles, section 9
Jerzy Neyman · 1990
Earlier work this paper cites.
Estimation of regression coefficients when some regressors are not always observed
James M Robins, Andrea Rotnitzky, and Lue Ping Zhao · 1994
Earlier work this paper cites.
Controlling the false discovery rate: a practical and powerful approach to multiple testing
Yoav Benjamini and Yosef Hochberg · 1995
Earlier work this paper cites.
False discovery rate–adjusted multiple confidence intervals for selected parameters
Yoav Benjamini and Daniel Yekutieli · 2005
Earlier work this paper cites.
The epoch-greedy algorithm for contextual multi-armed bandits
John Langford and Tong Zhang · 2007
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
Lihong Li, Wei Chu, John Langford, and Robert E Schapire · 2010
Earlier work this paper cites.
Doubly robust policy evaluation and learning
Miroslav Dudík, John Langford, and Lihong Li · 2011
Earlier work this paper cites.
Test martingales, Bayes factors and p-values
Glenn Shafer, Alexander Shen, Nikolai Vereshchagin, and Vladimir Vovk · 2011
Earlier work this paper cites.
Targeted learning: causal inference for observational and experimental data
Mark J van der Laan and Sherri Rose · 2011
Earlier work this paper cites.
Estimation of the effect of interventions that modify the received treatment
Sebastian Haneuse and Andrea Rotnitzky · 2013
Earlier work this paper cites.
Doubly robust policy evaluation and optimization
Miroslav Dudík, Dumitru Erhan, John Langford, and Lihong Li · 2014
Cited alongside, same era.
lil’ UCB: An optimal exploration algorithm for multi-armed bandits
Kevin Jamieson, Matthew Malloy, Robert Nowak, and Sébastien Bubeck · 2014
Cited alongside, same era.
Exponential inequalities for martingales with applications
Xiequan Fan, Ion Grama, and Quansheng Liu · 2015
Cited alongside, same era.
Causal inference in statistics, social, and biomedical sciences
Guido W Imbens and Donald B Rubin · 2015
Cited alongside, same era.
High-confidence off-policy evaluation
Philip Thomas, Georgios Theocharous, and Mohammad Ghavamzadeh · 2015
Cited alongside, same era.
On the complexity of best-arm identification in multi-armed bandit models
Emilie Kaufmann, Olivier Cappé, and Aurélien Garivier · 2016
Cited alongside, same era.
Off-policy risk assessment in contextual bandits
Audrey Huang, Liu Leqi, Zachary Lipton, and Kamyar Azizzadenesheli · 2021
Later among the works it cites.
Off-policy confidence sequences
Nikos Karampatziakis, Paul Mineiro, and Aaditya Ramdas · 2021
Later among the works it cites.
Near-optimal inference in adaptive linear regression
Koulik Khamaru, Yash Deshpande, Lester Mackey, and Martin J Wainwright · 2021
Later among the works it cites.
Testing by betting: A strategy for statistical and scientific communication
Glenn Shafer · 2021
Later among the works it cites.
Doubly robust interval estimation for optimal policy evaluation in online learning
Ye Shen, Hengrui Cai, and Rui Song · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Double/debiased machine learning for treatment and structural parameters
Victor Chernozhukov, Denis Chetverikov, Mert Demirer, Esther Duflo, Christian Hansen, Whitney Newey, and James Robins · 2018
Cited alongside, same era.
Peter Grünwald, Rianne de Heide, and Wouter M Koolen · 2019
Cited alongside, same era.
Nonparametric causal effects based on incremental propensity score interventions
Edward H Kennedy · 2019
Cited alongside, same era.
Connections between mirror descent, Thompson sampling and the information ratio
Julian Zimmert and Tor Lattimore · 2019
Cited alongside, same era.
Sampling-based versus design-based uncertainty in regression analysis
Alberto Abadie, Susan Athey, Guido W Imbens, and Jeffrey M Wooldridge · 2020
Cited alongside, same era.
Time-uniform Chernoff bounds via nonnegative supermartingales
Steven R. Howard, Aaditya Ramdas, Jon McAuliffe, and Jasjeet Sekhon · 2020
Cited alongside, same era.
Judith ter Schure and Peter Grünwald · 2021
Later among the works it cites.
E-values: Calibration, combination and applications
Vladimir Vovk and Ruodu Wang · 2021
Later among the works it cites.
Off-policy evaluation via adaptive weighting with data from contextual bandits
Ruohan Zhan, Vitor Hadad, David A Hirshberg, and Susan Athey · 2021
Later among the works it cites.
Statistical inference with m-estimators on adaptively collected data
Kelly Zhang, Lucas Janson, and Susan Murphy · 2021
Later among the works it cites.
Tsallis-INF: An optimal algorithm for stochastic and adversarial bandits
Julian Zimmert and Yevgeny Seldin · 2021
Later among the works it cites.
Design-based confidence sequences for anytime-valid causal inference
Dae Woong Ham, Iavor Bojinov, Michael Lindon, and Martin Tingley · 2022
Closest in time.
Sequential estimation of quantiles with applications to A/B testing and best-arm identification
Steven R Howard and Aaditya Ramdas · 2022
Closest in time.
Semiparametric doubly robust targeted double machine learning: a review
Edward H Kennedy · 2022
Closest in time.
A review of off-policy evaluation in reinforcement learning
Masatoshi Uehara, Chengchun Shi, and Nathan Kallus · 2022
Closest in time.
False discovery rate control with e-values
Ruodu Wang and Aaditya Ramdas · 2022
Closest in time.
Post-selection inference for e-value based confidence intervals
Ziyu Xu, Ruodu Wang, and Aaditya Ramdas · 2022
Closest in time.
Comparing sequential forecasters
Yo Joong Choe and Aaditya Ramdas · 2023
Closest in time.
Game-theoretic statistics and safe anytime-valid inference
Aaditya Ramdas, Peter Grünwald, Vladimir Vovk, and Glenn Shafer · 2023
Closest in time.
Online bootstrap inference for policy evaluation in reinforcement learning
Pratik Ramprasad, Yuantong Li, Zhuoran Yang, Zhaoran Wang, Will Wei Sun, and Guang Cheng · 2023
Closest in time.
Nonparametric extensions of randomized response for private confidence sets
Ian Waudby-Smith, Zhiwei Steven Wu, and Aaditya Ramdas · 2023
Closest in time.
Estimating means of bounded random variables by betting
Ian Waudby-Smith and Aaditya Ramdas · 2024
Closest in time.