Selecting the state-representation in reinforcement learning
Maillard, Odalric-Ambrym, Ryabko, Daniil, and Munos, Rémi · 2011
Later among the works it cites.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
Bubeck, Sébastien · 2012
Later among the works it cites.
Pattern classification
Duda, Richard O, Hart, Peter E, and Stork, David G · 2012
Later among the works it cites.
Reinforcement learning to adjust parametrized motor primitives to new situations
Kober, Jens, Wilhelm, Andreas, Oztop, Erhan, and Peters, Jan · 2012
Later among the works it cites.
Predicting current user intent with contextual markov models
Kiseleva, Julia, Lam, Hoang Thanh, Pechenizkiy, Mykola, and Calders, Toon · 2013
Later among the works it cites.
Competing with an infinite set of models in reinforcement learning
Nguyen, Phuong, Maillard, Odalric-Ambrym, Ryabko, Daniil, and Ortner, Ronald · 2013
Later among the works it cites.
Concurrent reinforcement learning from customer interactions
Silver, David, Newnham, Leonard, Barker, David, Weller, Suzanne, and McFall, Jason · 2013
Later among the works it cites.
Learning multiple models via regularized weighting
Vainsencher, Daniel, Mannor, Shie, and Xu, Huan · 2013
Later among the works it cites.
Robust markov decision processes
Wiesemann, Wolfram, Kuhn, Daniel, and Rustem, Berç · 2013
Later among the works it cites.
Latent bandits
Maillard, Odalric-Ambrym and Mannor, Shie · 2014
Later among the works it cites.
Handling signal variability with contextual markovian models
Radenen, Mathieu and Artières, Thierry · 2014
Later among the works it cites.