Fetching the paper…
Reading the bibliography…
Contextual bandit and reinforcement learning algorithms have been successfully used in various interactive learning systems such as online advertising, recommender systems, and dynamic pricing.
Use of ranks in one-criterion variance analysis
W. H. Kruskal and W. A. Wallis · 1952
Earlier work this paper cites.
Multiple comparisons among means
O. J. Dunn · 1961
Earlier work this paper cites.
Independence properties of directed Markov fields
S. L. Lauritzen, A. P. Dawid, B. N. Larsen, and H.-G. Leimer · 1990
Earlier work this paper cites.
Chapter 36 large sample estimation and hypothesis testing
W. K. Newey and D. McFadden · 1994
Earlier work this paper cites.
Strong completeness and faithfulness in bayesian networks
C. Meek · 1995
Earlier work this paper cites.
Causal diagrams for empirical research
J. Pearl · 1995
Earlier work this paper cites.
Regression shrinkage and selection via the lasso
R. Tibshirani · 1996
Earlier work this paper cites.
Finding minimal d-separators
J. Tian, A. Paz, and J. Pearl · 1998
Earlier work this paper cites.
Efficient reinforcement learning in factored MDPs
M. Kearns and D. Koller · 1999
Earlier work this paper cites.
Causation, Prediction, and Search
P. Spirtes, C. Glymour, and R. Scheines · 2000
Earlier work this paper cites.
Random forests
L. Breiman · 2001
Earlier work this paper cites.
Influence diagrams for causal modelling and inference
A. P. Dawid · 2002
Earlier work this paper cites.
Multiagent planning with factored MDPs
C. Guestrin, D. Koller, and R. Parr · 2002
Earlier work this paper cites.
Efficient solution algorithms for factored MDPs
C. Guestrin, D. Koller, R. Parr, and S. Venkataraman · 2003
Earlier work this paper cites.
Solving large scale linear prediction problems using stochastic gradient descent algorithms
T. Zhang · 2004
Earlier work this paper cites.
Causal graph based decomposition of factored MDPs
A. Jonsson and A. Barto · 2006
Earlier work this paper cites.
The epoch-greedy algorithm for multi-armed bandits with side information
J. Langford and T. Zhang · 2008
Earlier work this paper cites.
The offset tree for learning with partial labels
A. Beygelzimer and J. Langford · 2009
Earlier work this paper cites.
Estimation of the warfarin dose with clinical and pharmacogenetic data
I. W. P. Consortium · 2009
Earlier work this paper cites.
Causality
J. Pearl · 2009
Earlier work this paper cites.
Learning from logged implicit exploration data
A. Strehl, J. Langford, L. Li, and S. M. Kakade · 2010
Earlier work this paper cites.
Doubly robust policy evaluation and learning
M. Dudik, J. Langford, and L. Li · 2011
Earlier work this paper cites.
Transportability of causal and statistical relations: A formal approach
J. Pearl and E. Bareinboim · 2011
Earlier work this paper cites.
A kernel two-sample test
A. Gretton, K. M. Borgwardt, M. J. Rasch, B. Schölkopf, and A. Smola · 2012
Earlier work this paper cites.
On causal and anticausal learning
B. Schölkopf, D. Janzing, J. Peters, E. Sgouritsa, K. Zhang, and J. M. Mooij · 2012
Cited alongside, same era.
Machine learning in non-stationary environments: Introduction to covariate shift adaptation
M. Sugiyama and M. Kawanabe · 2012
Cited alongside, same era.
Counterfactual reasoning and learning systems: The example of computational advertising
L. Bottou, J. Peters, J. Quiñonero-Candela, D. X. Charles, D. M. Chickering, E. Portugaly, D. Ray, P. Simard, and E. Snelson · 2013
Cited alongside, same era.
Domain generalization via invariant feature representation
K. Muandet, D. Balduzzi, and B. Schölkopf · 2013
Cited alongside, same era.
Transportability from multiple environments with limited experiments: Completeness results
E. Bareinboim and J. Pearl · 2014
Cited alongside, same era.
Bandits with unobserved confounders: A causal approach
M. Arjovsky, L. Bottou, I. Gulrajani, and D. Lopez-Paz · 2019
Later among the works it cites.
Review of causal discovery methods based on graphical models
C. Glymour, K. Zhang, and P. Spirtes · 2019
Later among the works it cites.
Learning stable and predictive structures in kinetic systems
N. Pfister, S. Bauer, and J. Peters · 2019
Later among the works it cites.
Preventing failures due to dataset shift: Learning predictive models that transport
A. Subbaswamy, P. Schulam, and S. Saria · 2019
Later among the works it cites.
General transportability of soft interventions: Completeness results
J. Correa and E. Bareinboim · 2020
Later among the works it cites.
Causal discovery for causal bandits utilizing separating sets
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
E. Bareinboim, A. Forney, and J. Pearl · 2015
Cited alongside, same era.
A comprehensive survey on safe reinforcement learning
J. Garcıa and F. Fernández · 2015
Cited alongside, same era.
Concrete problems in ai safety
D. Amodei, C. Olah, J. Steinhardt, P. Christiano, J. Schulman, and D. Mané · 2016
Cited alongside, same era.
Causal inference and the data-fusion problem
E. Bareinboim and J. Pearl · 2016
Cited alongside, same era.
Foundations of structural causal models with cycles and latent variables
S. Bongers, P. Forré, J. Peters, and J. M. Mooij · 2016
Cited alongside, same era.
Causal bandits: Learning good interventions via causal inference
F. Lattimore, T. Lattimore, and M. D. Reid · 2016
Cited alongside, same era.
Causal inference in statistics : a primer
J. Pearl · 2016
Cited alongside, same era.
A. A. de Kroon, D. Belgrave, and J. M. Mooij · 2020
Later among the works it cites.
Distributional robustness of k-class estimators and the pulse
M. E. Jakobsen and J. Peters · 2020
Later among the works it cites.
Confounding-robust policy evaluation in infinite-horizon reinforcement learning
N. Kallus and A. Zhou · 2020
Later among the works it cites.
Generalized transportability: Synthesis of experiments from heterogeneous domains
S. Lee, J. D. Correa, and E. Bareinboim · 2020
Later among the works it cites.
Off-policy evaluation in partially observable environments
G. Tennenholtz, U. Shalit, and S. Mannor · 2020
Later among the works it cites.
Counterfactual learning of continuous stochastic policies
H. Zenati, A. Bietti, M. Martin, E. Diemert, and J. Mairal · 2020
Later among the works it cites.
Invariant causal prediction for block MDPs
A. Zhang, C. Lyle, S. Sodhani, A. Filos, M. Kwiatkowska, J. Pineau, Y. Gal, and D. Precup · 2020
Later among the works it cites.
Policy learning with observational data
S. Athey and S. Wager · 2021
Closest in time.
Recent advances in adversarial training for adversarial robustness
T. Bai, J. Luo, J. Zhao, B. Wen, and Q. Wang · 2021
Closest in time.
A causal framework for distribution generalization
R. Christiansen, N. Pfister, M. E. Jakobsen, N. Gnecco, and J. Peters · 2021
Closest in time.
Regularizing towards causal invariance: Linear models with proxies
M. Oberst, N. Thams, J. Peters, and D. Sontag · 2021
Closest in time.
Stabilizing variable selection and regression
N. Pfister, E. G. Williams, J. Peters, R. Aebersold, and P. Bühlmann · 2021
Closest in time.
Anchor regression: Heterogeneous data meet causality
D. Rothenhäusler, N. Meinshausen, P. Bühlmann, and J. Peters · 2021
Closest in time.
Invariant policy optimization: Towards stronger generalization in reinforcement learning
A. Sonar, V. Pacelli, and A. Majumdar · 2021
Closest in time.
Bandits with partially observable confounded data
G. Tennenholtz, U. Shalit, S. Mannor, and Y. Efroni · 2021
Closest in time.
Statistical testing under distributional shifts
N. Thams, S. Saengkyongam, N. Pfister, and J. Peters · 2021
Closest in time.
Towards a theoretical framework of out-of-distribution generalization
H. Ye, C. Xie, T. Cai, R. Li, Z. Li, and L. Wang · 2021
Closest in time.
Exploiting independent instruments: Identification and distribution generalization
S. Saengkyongam, L. Henckel, N. Pfister, and J. Peters · 2022
Closest in time.
Offline multi-action policy learning: Generalization and optimization
Z. Zhou, S. Athey, and S. Wager · 2022
Closest in time.