Fetching the paper…
Reading the bibliography…
Recent work has demonstrated that problems-- particularly imitation learning and structured prediction-- where a learner's predictions influence the input-distribution it is tested on can be naturally addressed by an interactive approach and analyzed using no-regret online learning.
A generalization of sampling without replacement from a finite universe
D. G. Horvitz and D. J. Thompson · 1952
Earlier work this paper cites.
The right way to do reinforcement learning with function approximation
R. Sutton · 2000
Earlier work this paper cites.
The nonstochastic multiarmed bandit problem
P. Auer, N. Cesa-Bianchi, Y. Freund, and R.E. Schapire · 2002
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
S. Kakade and J. Langford · 2002
Earlier work this paper cites.
Policy search by dynamic programming
J. A. Bagnell, A. Y. Ng, S. Kakade, and J. Schneider · 2003
Earlier work this paper cites.
On the sample complexity of reinforcement learning
S. M. Kakade · 2003
Earlier work this paper cites.
Error limiting reductions between classification tasks
A. Beygelzimer, V. Dani, T. Hayes, J. Langford, and B. Zadrozny · 2005
Earlier work this paper cites.
Sensitive error correcting output codes
J. Langford and A. Beygelzimer · 2005
Earlier work this paper cites.
Prediction, learning, and games
N. Cesa-Bianchi and G. Lugosi · 2006
Earlier work this paper cites.
Machine learning techniques—reductions between prediction quality metrics
A. Beygelzimer, J. Langford, and B. Zadrozny · 2008
Earlier work this paper cites.
A survey of robot learning from demonstration
B. D. Argall, S. Chernova, M. Veloso, and B. Browning · 2009
Cited alongside, same era.
Error-correcting tournaments
A. Beygelzimer, J. Langford, and P. Ravikumar · 2009
Cited alongside, same era.
Search-based structured prediction
H. Daumé III, J. Langford, and D. Marcu · 2009
Cited alongside, same era.
A contextual-bandit approach to personalized news article recommendation
L. Li, W. Chu, J. Langford, and R. E. Schapire · 2010
Cited alongside, same era.
Error and regret bounds for cost-sensitive multiclass classification reduction to regression, 2010
P. Mineiro · 2010
Cited alongside, same era.
Stacked hierarchical labeling
D. Munoz, J. A. Bagnell, and M. Hebert · 2010
Cited alongside, same era.
Learning message-passing inference machines for structured prediction
S. Ross, D. Munoz, M. Hebert, and J. A. Bagnell · 2011
Later among the works it cites.
Decoupling exploration and exploitation in multi-armed bandits
O. Avner, S. Mannor, and O. Shamir · 2012
Later among the works it cites.
Projection-free online learning
E. Hazan and S. Kale · 2012
Later among the works it cites.
Agnostic system identification for model-based reinforcement learning
S. Ross and J. A. Bagnell · 2012
Later among the works it cites.
Stability conditions for online learnability
S. Ross and J. A. Bagnell · 2012
Later among the works it cites.
The interplay between stability and regret in online learning
Ankan Saha, Prateek Jain, and Ambuj Tewari · 2012
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Efficient reductions for imitation learning
S. Ross and J. A. Bagnell · 2010
Cited alongside, same era.
Contextual bandit algorithms with supervised learning guarantees
A. Beygelzimer, J. Langford, L. Li, L. Reyzin, and R. E. Schapire · 2011
Cited alongside, same era.
Doubly robust policy evaluation and learning
M. Dudik, J. Langford, and L. Li · 2011
Cited alongside, same era.
A reduction of imitation learning and structured prediction to no-regret online learning
S. Ross, G. J. Gordon, and J. A. Bagnell · 2011
Cited alongside, same era.
Later among the works it cites.
Imitation learning for natural language direction following through unknown environments
F. Duvallet, T. Kollar, and A. Stentz · 2013
Later among the works it cites.
Learning policies for contextual submodular prediction
S. Ross, J. Zhou, Y. Yue, D. Dey, and J. A. Bagnell · 2013
Later among the works it cites.
Approximate policy iteration schemes: A comparison
Bruno Scherrer · 2014
Closest in time.