Fetching the paper…
Reading the bibliography…
Preference-based feedback is important for many applications in machine learning where evaluation of a reward function is not feasible.
The method of paired comparisons for social values
Thurstone, L. L · 1927
Earlier work this paper cites.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Bradley, R. A. and Terry, M. E · 1952
Earlier work this paper cites.
Gaussian processes for machine learning , volume 1
Rasmussen, C. E., Williams, C. K., et al · 2006
Earlier work this paper cites.
Random features for large-scale kernel machines
Rahimi, A. and Recht, B · 2007
Earlier work this paper cites.
Gaussian process optimization in the bandit setting: No regret and experimental design
Srinivas, N., Krause, A., Kakade, S. M., and Seeger, M · 2010
Earlier work this paper cites.
Preference-based reinforcement learning: a formal framework and a policy iteration algorithm
Fürnkranz, J., Hüllermeier, E., Cheng, W., and Park, S.-H · 2012
Earlier work this paper cites.
The k-armed dueling bandits problem
Yue, Y., Broder, J., Kleinberg, R., and Joachims, T · 2012
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Mnih, V · 2013
Earlier work this paper cites.
Robust preference learning-based reinforcement learning
Akour, R · 2014
Earlier work this paper cites.
Contextual dueling bandits, 2015
Dudík, M., Hofmann, K., Schapire, R. E., Slivkins, A., and Zoghi, M · 2015
Earlier work this paper cites.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Gal, Y. and Ghahramani, Z · 2016
Earlier work this paper cites.
On kernelized multi-armed bandits
Chowdhury, S. R. and Gopalan, A · 2017
Earlier work this paper cites.
Deep reinformcement learning from human preferences
Christiano, P., Leike, J., Brown, T. B., Martic, M., Legg, S., and Amodei, D · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Towards coherent and cohesive long-form text generation
Cho, W. S., Zhang, P., Zhang, Y., Li, X., Galley, M., Brockett, C., Wang, M., and Gao, J · 2018
Earlier work this paper cites.
Gpytorch: Blackbox matrix-matrix gaussian process inference with gpu acceleration
Gardner, J. R., Pleiss, G., Bindel, D., Weinberger, K. Q., and Wilson, A. G · 2018
Earlier work this paper cites.
Can neural machine translation be improved with user feedback?
Kreutzer, J., Khadivi, S., Matusov, E., and Riezler, S · 2018
Cited alongside, same era.
Improving a neural semantic parser by counterfactual learning from human bandit feedback
Lawrence, C. and Riezler, S · 2018
Cited alongside, same era.
Offline contextual bayesian optimization
Char, I., Chung, Y., Neiswanger, W., Kandasamy, K., Nelson, A. O., Boyer, M., Kolemen, E., and Schneider, J · 2019
Cited alongside, same era.
Multi-fidelity gaussian process bandit optimisation
Kandasamy, K., Dasarathy, G., Oliva, J., Schneider, J., and Poczos, B · 2019
Cited alongside, same era.
Finding generalizable evidence by learning to convince q&a models
Perez, E., Karamcheti, S., Fergus, R., Weston, J., Kiela, D., and Cho, K · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Pythia: A suite for analyzing large language models across training and scaling
Biderman, S., Schoelkopf, H., Anthony, Q. G., Bradley, H., O’Brien, K., Hallahan, E., Khan, M. A., Purohit, S., Prashanth, U. S., Raff, E., et al · 2023
Closest in time.
Symbolic discovery of optimization algorithms
Chen, X., Liang, C., Huang, D., Real, E., Wang, K., Liu, Y., Pham, H., Dong, X., Luong, T., Hsieh, C.-J., et al · 2023
Closest in time.
Qlora: Efficient finetuning of quantized llms, 2023
Dettmers, T., Pagnoni, A., Holtzman, A., and Zettlemoyer, L · 2023
Closest in time.
Near-optimal policy identification in active reinforcement learning
Li, X., Mehta, V., Kirschner, J., Char, I., Neiswanger, W., Schneider, J., Krause, A., and Bogunovic, I · 2023
Closest in time.
Direct preference optimization: Your language model is secretly a reward model, 2023
Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., and Finn, C · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Fine-tuning language models from human preferences
Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G · 2019
Cited alongside, same era.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Cited alongside, same era.
Learning to summarize from human feedback
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D. M., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P · 2020
Cited alongside, same era.
Zeroth order non-convex optimization with dueling-choice bandits
Xu, Y., Joshi, A., Singh, A., and Dubrawski, A · 2020
Cited alongside, same era.
Preference-based online learning with dueling bandits: A survey, 2021
Bengs, V., Busa-Fekete, R., Mesaoudi-Paul, A. E., and Hüllermeier, E · 2021
Cited alongside, same era.
Lora: Low-rank adaptation of large language models, 2021
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Cited alongside, same era.
Llama: Open and efficient foundation language models, 2023
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., and Lample, G · 2023
Closest in time.
Jeopardy dataset, 2023
Wolf, T., Tunstall, L., and von Platen, P · 2023
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena, 2023
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E. P., Zhang, H., Gonzalez, J. E., and Stoica, I · 2023
Closest in time.
Llama 3 model card
AI@Meta · 2024
Closest in time.
Provably sample efficient rlhf via active preference optimization
Das, N., Chakraborty, S., Pacchiano, A., and Chowdhury, S. R · 2024
Closest in time.
Efficiently learning at test-time: Active fine-tuning of llms
Hübotter, J., Bongni, S., Hakimi, I., and Krause, A · 2024
Closest in time.
Reinforcement learning from human feedback with active queries
Ji, K., He, J., and Gu, Q · 2024
Closest in time.
Active preference learning for large language models
Muldrew, W., Hayes, P., Zhang, M., and Barber, D · 2024
Closest in time.
Exploratory preference optimization: Harnessing implicit q*-approximation for sample-efficient rlhf
Xie, T., Foster, D. J., Krishnamurthy, A., Rosset, C., Awadallah, A., and Rakhlin, A · 2024
Closest in time.
Iterative preference learning from human feedback: Bridging theory and practice for rlhf under kl-constraint
Xiong, W., Dong, H., Ye, C., Wang, Z., Zhong, H., Ji, H., Jiang, N., and Zhang, T · 2024
Closest in time.
Self-exploring language models: Active preference elicitation for online alignment
Zhang, S., Yu, D., Sharma, H., Yang, Z., Wang, S., Hassan, H., and Wang, Z · 2024
Closest in time.