Fetching the paper…
Reading the bibliography…
Active Reinforcement Learning (ARL) is a twist on RL where the agent observes reward information only if it pays a cost.
Bayesian q-learning
R. Dearden, N. Friedman, and S. Russell · 1998
Earlier work this paper cites.
Reinforcement learning: An introduction , volume 1
R. S. Sutton and A. G. Barto · 1998
Earlier work this paper cites.
A bayesian framework for reinforcement learning
M. Strens · 2000
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
P. Auer, N. Cesa-Bianchi, and P. Fischer · 2002
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
P. Abbeel and A. Y. Ng · 2004
Earlier work this paper cites.
Active reinforcement learning
A. Epshteyn, A. Vogel, and G. DeJong · 2008
Earlier work this paper cites.
Near-bayesian exploration in polynomial time
J. Z. Kolter and A. Y Ng · 2009
Earlier work this paper cites.
Gaussian process optimization in the bandit setting: No regret and experimental design
N. Srinivas, A. Krause, S. M. Kakade, and M. Seeger · 2009
Earlier work this paper cites.
Near-optimal brl using optimistic local transitions
M. Araya, O. Buffet, and V. Thomas · 2012
Earlier work this paper cites.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
S. Bubeck, N. Cesa-Bianchi, et al · 2012
Earlier work this paper cites.
Efficient bayes-adaptive reinforcement learning using sample-based search
A. Guez, D. Silver, and P. Dayan · 2012
Earlier work this paper cites.
Active imitation learning via reduction to iid active learning
K. Judah, A. Fern, and T. G. Dietterich · 2012
Earlier work this paper cites.
Active learning
B. Settles · 2012
Cited alongside, same era.
(more) efficient reinforcement learning via posterior sampling
I. Osband, D. Russo, and B. Van Roy · 2013
Cited alongside, same era.
Active reward learning
C. Daniel, M. Viering, J. Metz, O. Kroemer, and J. Peters · 2014
Cited alongside, same era.
Algorithms for multi-armed bandit problems
V. Kuleshov and D. Precup · 2014
Cited alongside, same era.
Bayesian reinforcement learning: A survey
M. Ghavamzadeh, S. Mannor, J. Pineau, A. Tamar, et al · 2015
Cited alongside, same era.
Sample-Based Search Methods For Bayes-Adaptive Planning
A. Guez · 2015
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, et al · 2016
Later among the works it cites.
Exploration from demonstration for interactive reinforcement learning
K. Subramanian, C. L. Isbell Jr, and Andrea L. Thomaz · 2016
Later among the works it cites.
Thinking fast and slow with deep learning and tree search
T. Anthony, Z. Tian, and D. Barber · 2017
Later among the works it cites.
Minimax regret bounds for reinforcement learning
M. G. Azar, I. Osband, and R. Munos · 2017
Later among the works it cites.
Deep reinforcement learning from human preferences
P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei · 2017
Later among the works it cites.
Active preference-based learning of reward functions
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Multi-armed bandit experiments in the online service economy
S. L. Scott · 2015
Cited alongside, same era.
Benchmarking for bayesian reinforcement learning
M. Castronovo, D. Ernst, A. Couëtoux, and R. Fonteneau · 2016
Cited alongside, same era.
Learning the preferences of ignorant, inconsistent agents
O. Evans, A. Stuhlmüller, and N. D. Goodman · 2016
Cited alongside, same era.
Generative adversarial imitation learning
J. Ho and S. Ermon · 2016
Cited alongside, same era.
Active reinforcement learning: Observing rewards at a cost
D. Krueger, J. Leike, O. Evans, and J. Salvatier · 2016
Cited alongside, same era.
A. D. Dragan D. Sadigh, Shankar S., and Sanjit A S · 2017
Later among the works it cites.
Robot planning with mathematical models of human state and action
A. D. Dragan · 2017
Later among the works it cites.
Trial without error: Towards safe reinforcement learning via human intervention
W. Saunders, G. Sastry, A. Stuhlmueller, and O. Evans · 2017
Later among the works it cites.
Mastering the game of go without human knowledge
D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, et al · 2017
Later among the works it cites.
Deep tamer: Interactive agent shaping in high-dimensional state spaces
G. Warnell, N. Waytowich, V. Lawhern, and P. Stone · 2017
Later among the works it cites.
A survey of preference-based reinforcement learning methods
C. Wirth, R. Akrour, G. Neumann, J. Fürnkranz, et al · 2017
Later among the works it cites.