Fetching the paper…
Reading the bibliography…
This paper formalises the problem of online algorithm selection in the context of Reinforcement Learning.
The algorithm selection problem
John R. Rice · 1976
Earlier work this paper cites.
Learning from Delayed Rewards
C.J.C.H. Watkins · 1989
Earlier work this paper cites.
Computationally feasible bounds for partially observed markov decision processes
William S. Lovejoy · 1991
Earlier work this paper cites.
Reinforcement Learning: An Introduction (Adaptive Computation and Machine Learning)
Richard S. Sutton and Andrew G. Barto · 1998
Earlier work this paper cites.
Ensemble learning
Thomas G. Dietterich · 2002
Earlier work this paper cites.
Pac bounds for multi-armed bandit and markov decision processes
Eyal Even-Dar, Shie Mannor, and Yishay Mansour · 2002
Earlier work this paper cites.
Meta-learning in reinforcement learning
Nicolas Schweighofer and Kenji Doya · 2003
Earlier work this paper cites.
Tree-based batch mode reinforcement learning
Damien Ernst, Pierre Geurts, and Louis Wehenkel · 2005
Earlier work this paper cites.
Prediction, learning, and games
Nicolo Cesa-Bianchi and Gábor Lugosi · 2006
Earlier work this paper cites.
Learning dynamic algorithm portfolios
Matteo Gagliolo and Jürgen Schmidhuber · 2006
Earlier work this paper cites.
Ensemble algorithms in reinforcement learning
Marco A Wiering and Hado Van Hasselt · 2008
Earlier work this paper cites.
Cross-disciplinary perspectives on meta-learning for algorithm selection
Kate A. Smith-Miles · 2009
Earlier work this paper cites.
Best Arm Identification in Multi-Armed Bandits
Jean-Yves Audibert and Sébastien Bubeck · 2010
Cited alongside, same era.
Analyzing bandit-based adaptive operator selection mechanisms
Álvaro Fialho, Luis Da Costa, Marc Schoenauer, and Michele Sebag · 2010
Cited alongside, same era.
Algorithm selection as a bandit problem with unbounded losses
Matteo Gagliolo and Jürgen Schmidhuber · 2010
Cited alongside, same era.
On Upper-Confidence Bound Policies for Switching Bandit Problems , pp. 174–188
Aurélien Garivier and Eric Moulines · 2011
Cited alongside, same era.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
Sébastien Bubeck and Nicolò Cesa-Bianchi · 2012
Cited alongside, same era.
Algorithm selection for combinatorial search problems: A survey
Meta online learning: experiments on a unit commitment problem
Jialin Liu and Olivier Teytaud · 2014
Later among the works it cites.
Exp3 with drift detection for the switching bandit problem
Robin Allesiardo and Raphaël Féraud · 2015
Later among the works it cites.
Algorithm Portfolios for Noisy Optimization
Marie-Liesse Cauwet, Jialin Liu, Baptiste Rozière, and Olivier Teytaud · 2015
Later among the works it cites.
Optimising turn-taking strategies with reinforcement learning
Hatim Khouzaimi, Romain Laroche, and Fabrice Lefevre · 2015
Later among the works it cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Later among the works it cites.
Reinforcement learning for turn-taking management in incremental spoken dialogue systems
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lars Kotthoff · 2012
Cited alongside, same era.
Batch reinforcement learning
Sascha Lange, Thomas Gabel, and Martin Riedmiller · 2012
Cited alongside, same era.
The arcade learning environment: An evaluation platform for general agents
Marc G. Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2013
Cited alongside, same era.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Cited alongside, same era.
Stochastic multi-armed-bandit problem with non-stationary rewards
Omar Besbes, Yonatan Gur, and Assaf Zeevi · 2014
Cited alongside, same era.
Algorithm portfolios for noisy optimization: Compare solvers early
Marie-Liesse Cauwet, Jialin Liu, and Olivier Teytaud · 2014
Cited alongside, same era.
Finite-time analysis of the multiarmed bandit problem
Peter Auer, Nicolò Cesa-Bianchi, and Paul Fischer
Cited in the paper.
Hatim Khouzaimi, Romain Laroche, and Fabrice Lefèvre · 2016
Later among the works it cites.
A negotiation dialogue game
Romain Laroche and Aude Genevay · 2016
Later among the works it cites.
Safe and efficient off-policy reinforcement learning
Rémi Munos, Tom Stepleton, Anna Harutyunyan, and Marc G. Bellemare · 2016
Later among the works it cites.
Deep exploration via bootstrapped dqn
Ian Osband, Charles Blundell, Alexander Pritzel, and Benjamin Van Roy · 2016
Later among the works it cites.
Learning to reinforcement learn
Jane X. Wang, Zeb Kurth-Nelson, Dhruva Tirumala, Hubert Soyer, Joel Z. Leibo, Rémi Munos, Charles Blundell, Dharshan Kumaran, and Matt Botvinick · 2016
Later among the works it cites.
Adapting the trace parameter in reinforcement learning
Martha White and Adam White · 2016
Later among the works it cites.