Fetching the paper…
Reading the bibliography…
We develop a learning principle and an efficient algorithm for batch learning from logged bandit feedback.
The central role of the propensity score in observational studies for causal effects
Rosenbaum, Paul R. and Rubin, Donald B · 1983
Earlier work this paper cites.
Statistical Learning Theory
Vapnik, V · 1998
Earlier work this paper cites.
Beating the hold-out: Bounds for k-fold and progressive cross-validation
Blum, Avrim, Kalai, Adam, and Langford, John · 1999
Earlier work this paper cites.
Conditional random fields: Probabilistic models for segmenting and labeling sequence data
Lafferty, John D., McCallum, Andrew, and Pereira, Fernando C. N · 2001
Earlier work this paper cites.
Cost-sensitive learning by cost-proportionate example weighting
Zadrozny, Bianca, Langford, John, and Abe, Naoki · 2003
Earlier work this paper cites.
Support vector machine learning for interdependent and structured output spaces
Tsochantaridis, Ioannis, Hofmann, Thomas, Joachims, Thorsten, and Altun, Yasemin · 2004
Earlier work this paper cites.
Truncated importance sampling
Ionides, Edward L · 2008
Earlier work this paper cites.
The epoch-greedy algorithm for multi-armed bandits with side information
Langford, John and Zhang, Tong · 2008
Earlier work this paper cites.
Exploration scavenging
Langford, John, Strehl, Alexander, and Wortman, Jennifer · 2008
Earlier work this paper cites.
Neural Network Learning: Theoretical Foundations
Anthony, Martin and Bartlett, Peter L · 2009
Earlier work this paper cites.
The offset tree for learning with partial labels
Beygelzimer, Alina and Langford, John · 2009
Earlier work this paper cites.
Empirical bernstein bounds and sample-variance penalization
Maurer, Andreas and Pontil, Massimiliano · 2009
Cited alongside, same era.
A contextual-bandit approach to personalized news article recommendation
Li, Lihong, Chu, Wei, Langford, John, and Schapire, Robert E · 2010
Cited alongside, same era.
Learning from logged implicit exploration data
Strehl, Alexander L., Langford, John, Li, Lihong, and Kakade, Sham · 2010
Cited alongside, same era.
A quasi-Newton approach to nonsmooth convex optimization problems in machine learning
Yu, Jin, Vishwanathan, S. V. N., Günter, Simon, and Schraudolph, Nicol N · 2010
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, John, Hazan, Elad, and Singer, Yoram · 2011
Cited alongside, same era.
Doubly robust policy evaluation and learning
Langford, John, Li, Lihong, and Dudík, Miroslav · 2011
Exploration vs exploitation vs safety: Risk-aware multi-armed bandits
Galichet, Nicolas, Sebag, Michèle, and Teytaud, Olivier · 2013
Later among the works it cites.
Reusing historical interaction data for faster online learning to rank for IR
Hofmann, Katja, Schuth, Anne, Whiteson, Shimon, and de Rijke, Maarten · 2013
Later among the works it cites.
Nonsmooth optimization via quasi-newton methods
Lewis, Adrian S. and Overton, Michael L · 2013
Later among the works it cites.
Monte Carlo theory, methods and examples
Owen, Art B · 2013
Later among the works it cites.
Taming the monster: A fast and simple algorithm for contextual bandits
Agarwal, Alekh, Hsu, Daniel, Kale, Satyen, Langford, John, Li, Lihong, and Schapire, Robert · 2014
Later among the works it cites.
Counterfactual estimation and optimization of click metrics for search engines
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Unbiased offline evaluation of contextual-bandit-based news article recommendation algorithms
Li, Lihong, Chu, Wei, Langford, John, and Wang, Xuanhui · 2011
Cited alongside, same era.
Scikit-learn: Machine learning in Python
Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., and Duchesnay, E · 2011
Cited alongside, same era.
Safe exploration of state and action spaces in reinforcement learning
Garcia, J. and Fernandez, F · 2012
Cited alongside, same era.
Multi-armed bandit problems with history
Shivaswamy, Pannagadatta K. and Joachims, Thorsten · 2012
Cited alongside, same era.
Counterfactual reasoning and learning systems: the example of computational advertising
Bottou, Léon, Peters, Jonas, Candela, Joaquin Q., Charles, Denis X., Chickering, Max, Portugaly, Elon, Ray, Dipankar, Simard, Patrice Y., and Snelson, Ed · 2013
Cited alongside, same era.
Li, Lihong, Chen, Shunbao, Kleban, Jim, and Gupta, Ankur · 2014
Later among the works it cites.
Improving offline evaluation of contextual bandit algorithms via bootstrapping techniques
Mary, Jérémie, Preux, Philippe, and Nicol, Olivier · 2014
Later among the works it cites.
GenSVM: A Generalized Multiclass Support Vector Machine
van den Burg, G.J.J. and Groenen, P.J.F · 2014
Later among the works it cites.
Toward minimax off-policy value estimation
Li, Lihong, Munos, Remi, and Szepesvari, Csaba · 2015
Closest in time.
High-confidence off-policy evaluation
Thomas, Philip S., Theocharous, Georgios, and Ghavamzadeh, Mohammad · 2015
Closest in time.