Algorithms for inverse reinforcement learning
Andrew Y Ng, Stuart J Russell, et al · 2000
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Peter Auer, Nicolo Cesa-Bianchi, and Paul Fischer · 2002
Earlier work this paper cites.
Optimizing dialogue management with reinforcement learning: Experiments with the njfun system
Satinder Singh, Diane Litman, Michael Kearns, and Marilyn Walker · 2002
Earlier work this paper cites.
Regret bounds for kernel-based reinforcement learning
Original
Omar Darwiche Domingues, Pierre Ménard, Matteo Pirotta, Emilie Kaufmann, and Michal Valko · 2004
Earlier work this paper cites.
Autonomous inverted helicopter flight via reinforcement learning
Andrew Y Ng, Adam Coates, Mark Diel, Varun Ganapathi, Jamie Schulte, Ben Tse, Eric Berger, and Eric Liang · 2006
Earlier work this paper cites.
High-probability regret bounds for bandit online linear optimization
Peter Bartlett, Varsha Dani, Thomas Hayes, Sham Kakade, Alexander Rakhlin, and Ambuj Tewari · 2008
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Peter Auer, Thomas Jaksch, and Ronald Ortner · 2009
Earlier work this paper cites.
Interactively optimizing information retrieval systems as a dueling bandits problem
Yisong Yue and Thorsten Joachims · 2009
Earlier work this paper cites.
Episodic reinforcement learning in finite mdps: Minimax lower bounds revisited
Original
Omar Darwiche Domingues, Pierre Ménard, Emilie Kaufmann, and Michal Valko · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Yasin Abbasi-Yadkori, Dávid Pál, and Csaba Szepesvári · 2011
Earlier work this paper cites.
Relative entropy inverse reinforcement learning
Abdeslam Boularias, Jens Kober, and Jan Peters · 2011
Earlier work this paper cites.
Beat the mean bandit
Yisong Yue and Thorsten Joachims · 2011
Earlier work this paper cites.
Pac bounds for discounted mdps
Tor Lattimore and Marcus Hutter · 2012
Earlier work this paper cites.
Apprenticeship learning using inverse reinforcement learning and gradient methods
Original
Gergely Neu and Csaba Szepesvári · 2012
Earlier work this paper cites.
The k k -armed dueling bandits problem
Yisong Yue, Josef Broder, Robert Kleinberg, and Thorsten Joachims · 2012
Earlier work this paper cites.
Minimax pac bounds on the sample complexity of reinforcement learning with a generative model
Mohammad Gheshlaghi Azar, Rémi Munos, and Hilbert J Kappen · 2013
Earlier work this paper cites.
Learning trajectory preferences for manipulators via iterative improvement
Ashesh Jain, Brian Wojcik, Thorsten Joachims, and Ashutosh Saxena · 2013
Earlier work this paper cites.
Preference-based reinforcement learning: A preliminary survey
Christian Wirth and Johannes Fürnkranz · 2013
Earlier work this paper cites.
Reducing dueling bandits to cardinal bandits
Nir Ailon, Zohar Shay Karnin, and Thorsten Joachims · 2014
Earlier work this paper cites.
Preference-based reinforcement learning: evolutionary direct policy search using a preference-based racing algorithm
Róbert Busa-Fekete, Balázs Szörényi, Paul Weng, Weiwei Cheng, and Eyke Hüllermeier · 2014
Earlier work this paper cites.