Fetching the paper…
Reading the bibliography…
In preference-based reinforcement learning (RL), an agent interacts with the environment while receiving preferences instead of absolute feedback.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
W. R. Thompson · 1933
Earlier work this paper cites.
Convergence of probability measures
P. Billingsley · 1968
Earlier work this paper cites.
Support vector learning for ordinal regression
R. Herbrich, T. Graepel, and K. Obermayer · 1999
Earlier work this paper cites.
Learning to rank using gradient descent
C. Burges, T. Shaked, E. Renshaw, A. Lazier, M. Deeds, N. Hamilton, and G. N. Hullender · 2005
Earlier work this paper cites.
Preference learning with Gaussian processes
W. Chu and Z. Ghahramani · 2005
Earlier work this paper cites.
A support vector method for multivariate performance measures
T. Joachims · 2005
Earlier work this paper cites.
Query chains: Learning to rank from implicit feedback
F. Radlinski and T. Joachims · 2005
Earlier work this paper cites.
Gaussian processes for machine learning
C. E. Rasmussen and C. K. Williams · 2006
Earlier work this paper cites.
Learning to rank with nonsmooth cost functions
C. J. Burges, R. Ragno, and Q. V. Le · 2007
Earlier work this paper cites.
A support vector method for optimizing average precision
Y. Yue, T. Finley, F. Radlinski, and T. Joachims · 2007
Earlier work this paper cites.
Active preference learning with discrete choice data
B. Eric, N. D. Freitas, and A. Ghosh · 2008
Earlier work this paper cites.
A survey of robot learning from demonstration
B. D. Argall, S. Chernova, M. Veloso, and B. Browning · 2009
Earlier work this paper cites.
Learning to rank for information retrieval
T.-Y. Liu · 2009
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
T. Jaksch, R. Ortner, and P. Auer · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Y. Abbasi-Yadkori, D. Pál, and C. Szepesvári · 2011
Earlier work this paper cites.
An empirical evaluation of Thompson sampling
O. Chapelle and L. Li · 2011
Earlier work this paper cites.
Bayesian active learning for classification and preference learning
N. Houlsby, F. Huszár, Z. Ghahramani, and M. Lengyel · 2011
Earlier work this paper cites.
APRIL: Active preference learning-based reinforcement learning
R. Akrour, M. Schoenauer, and M. Sebag · 2012
Earlier work this paper cites.
Elements of information theory
T. M. Cover and J. A. Thomas · 2012
Cited alongside, same era.
Preference-based reinforcement learning: A formal framework and a policy iteration algorithm
J. Fürnkranz, E. Hüllermeier, W. Cheng, and S.-H. Park · 2012
Cited alongside, same era.
Online structured prediction via coactive learning
P. Shivaswamy and T. Joachims · 2012
Cited alongside, same era.
A Bayesian approach for policy learning from trajectory preference queries
A. Wilson, A. Fern, and P. Tadepalli · 2012
Cited alongside, same era.
The k-armed dueling bandits problem
Y. Yue, J. Broder, R. Kleinberg, and T. Joachims · 2012
Cited alongside, same era.
Thompson sampling for contextual bandits with linear payoffs
S. Agrawal and N. Goyal · 2013
Cited alongside, same era.
Dueling bandits: Beyond Condorcet winners to general tournament solutions
S. Y. Ramamohan, A. Rajkumar, and S. Agarwal · 2016
Later among the works it cites.
An information-theoretic analysis of Thompson sampling
D. Russo and B. Van Roy · 2016
Later among the works it cites.
Model-free preference-based reinforcement learning
C. Wirth, J. Fürnkranz, and G. Neumann · 2016
Later among the works it cites.
Double Thompson sampling for dueling bandits
H. Wu and X. Liu · 2016
Later among the works it cites.
Linear Thompson sampling revisited
M. Abeille and A. Lazaric · 2017
Later among the works it cites.
Optimistic posterior sampling for reinforcement learning: Worst-case regret bounds
S. Agrawal and R. Jia · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Preference-based evolutionary direct policy search
R. Busa-Fekete, B. Szörényi, P. Weng, W. Cheng, and E. Hüllermeier · 2013
Cited alongside, same era.
(More) efficient reinforcement learning via posterior sampling
I. Osband, D. Russo, and B. Van Roy · 2013
Cited alongside, same era.
Stable coactive learning via perturbation
K. Raman, T. Joachims, P. Shivaswamy, and T. Schnabel · 2013
Cited alongside, same era.
Reducing dueling bandits to cardinal bandits
N. Ailon, Z. Karnin, and T. Joachims · 2014
Cited alongside, same era.
Programming by feedback
R. Akrour, M. Schoenauer, M. Sebag, and J.-C. Souplet · 2014
Cited alongside, same era.
Relative upper confidence bound for the k-armed dueling bandit problem
M. Zoghi, S. Whiteson, R. Munos, and M. De Rijke · 2014
Cited alongside, same era.
Do you want your autonomous car to drive like you?
C. Basu, Q. Yang, D. Hungerman, M. Sinahal, and A. D. Dragan · 2017
Later among the works it cites.
Deep reinforcement learning from human preferences
P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei · 2017
Later among the works it cites.
Why is posterior sampling better than optimism for reinforcement learning?
I. Osband and B. Van Roy · 2017
Later among the works it cites.
Active preference-based learning of reward functions
D. Sadigh, A. D. Dragan, S. Sastry, and S. A. Seshia · 2017
Later among the works it cites.
Multi-dueling bandits with dependent arms
Y. Sui, V. Zhuang, J. W. Burdick, and Y. Yue · 2017
Later among the works it cites.
Efficient Preference-based Reinforcement Learning
C. Wirth · 2017
Later among the works it cites.
A survey of preference-based reinforcement learning methods
C. Wirth, R. Akrour, G. Neumann, and J. Fürnkranz · 2017
Later among the works it cites.
Information directed reinforcement learning
A. Zanette and R. Sarkar · 2017
Later among the works it cites.
An information-theoretic analysis for Thompson sampling with many actions
S. Dong and B. Van Roy · 2018
Later among the works it cites.
Learning dynamic robot-to-human object handover from human feedback
A. Kupcsik, D. Hsu, and W. S. Lee · 2018
Later among the works it cites.
Information-directed exploration for deep reinforcement learning
N. Nikolov, J. Kirschner, F. Berkenkamp, and A. Krause · 2018
Later among the works it cites.
On the performance of Thompson sampling on logistic bandits
S. Dong, T. Ma, and B. Van Roy · 2019
Closest in time.