Fetching the paper…
Reading the bibliography…
Preference-based feedback is important for many applications where direct evaluation of a reward function is not feasible.
The method of paired comparisons for social values
Thurstone, L. L · 1927
Earlier work this paper cites.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Bradley, R. A. and Terry, M. E · 1952
Earlier work this paper cites.
Gaussian processes for machine learning , volume 1
Rasmussen, C. E., Williams, C. K., et al · 2006
Earlier work this paper cites.
Random features for large-scale kernel machines
Rahimi, A. and Recht, B · 2007
Earlier work this paper cites.
Gaussian process optimization in the bandit setting: No regret and experimental design
Srinivas, N., Krause, A., Kakade, S. M., and Seeger, M · 2010
Earlier work this paper cites.
Preference-based reinforcement learning: a formal framework and a policy iteration algorithm
Fürnkranz, J., Hüllermeier, E., Cheng, W., and Park, S.-H · 2012
Earlier work this paper cites.
The k-armed dueling bandits problem
Yue, Y., Broder, J., Kleinberg, R., and Joachims, T · 2012
Earlier work this paper cites.
Robust preference learning-based reinforcement learning
Akour, R · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Best-arm identification in linear bandits
Soare, M., Lazaric, A., and Munos, R · 2014
Cited alongside, same era.
Contextual dueling bandits, 2015
Dudík, M., Hofmann, K., Schapire, R. E., Slivkins, A., and Zoghi, M · 2015
Cited alongside, same era.
On kernelized multi-armed bandits
Chowdhury, S. R. and Gopalan, A · 2017
Cited alongside, same era.
Deep reinformcement learning from human preferences
Christiano, P., Leike, J., Brown, T. B., Martic, M., Legg, S., and Amodei, D · 2017
Cited alongside, same era.
Towards coherent and cohesive long-form text generation
Cho, W. S., Zhang, P., Zhang, Y., Li, X., Galley, M., Brockett, C., Wang, M., and Gao, J · 2018
Cited alongside, same era.
Gpytorch: Blackbox matrix-matrix gaussian process inference with gpu acceleration
Gardner, J. R., Pleiss, G., Bindel, D., Weinberger, K. Q., and Wilson, A. G · 2018
Cited alongside, same era.
Multi-fidelity gaussian process bandit optimisation
Kandasamy, K., Dasarathy, G., Oliva, J., Schneider, J., and Poczos, B · 2019
Later among the works it cites.
Finding generalizable evidence by learning to convince q&a models
Perez, E., Karamcheti, S., Fergus, R., Weston, J., Kiela, D., and Cho, K · 2019
Later among the works it cites.
Fine-tuning language models from human preferences
Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G · 2019
Later among the works it cites.
Learning to summarize from human feedback
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D. M., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P · 2020
Later among the works it cites.
Zeroth order non-convex optimization with dueling-choice bandits
Xu, Y., Joshi, A., Singh, A., and Dubrawski, A · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Can neural machine translation be improved with user feedback?
Kreutzer, J., Khadivi, S., Matusov, E., and Riezler, S · 2018
Cited alongside, same era.
Improving a neural semantic parser by counterfactual learning from human bandit feedback
Lawrence, C. and Riezler, S · 2018
Cited alongside, same era.
Offline contextual bayesian optimization
Char, I., Chung, Y., Neiswanger, W., Kandasamy, K., Nelson, A. O., Boyer, M., Kolemen, E., and Schneider, J · 2019
Cited alongside, same era.
Bengs, V., Busa-Fekete, R., Mesaoudi-Paul, A. E., and Hüllermeier, E · 2021
Later among the works it cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Later among the works it cites.
Near-optimal policy identification in active reinforcement learning
Li, X., Mehta, V., Kirschner, J., Char, I., Neiswanger, W., Schneider, J., Krause, A., and Bogunovic, I · 2023
Closest in time.
Principled reinforcement learning with human feedback from pairwise or k k -wise comparisons, 2023
Zhu, B., Jiao, J., and Jordan, M. I · 2023
Closest in time.