Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have achieved remarkable success, yet aligning their generations with human preferences remains a critical challenge.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Bradley, R. A. and Terry, M. E · 1952
Earlier work this paper cites.
Visualizing data using t-sne
Van der Maaten, L. and Hinton, G · 2008
Earlier work this paper cites.
Overview of supervised learning
Hastie, T., Tibshirani, R., Friedman, J., Hastie, T., Tibshirani, R., and Friedman, J · 2009
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P · 2013
Earlier work this paper cites.
Generating sentences from a continuous space
Bowman, S., Vilnis, L., Vinyals, O., Dai, A., Jozefowicz, R., and Bengio, S · 2016
Earlier work this paper cites.
Imitation learning: A survey of learning methods
Hussein, A., Gaber, M. M., Elyan, E., and Jayne, C · 2017
Earlier work this paper cites.
Categorical reparametrization with gumble-softmax
Jang, E., Gu, S., and Poole, B · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Neural discrete representation learning
Van Den Oord, A., Vinyals, O., et al · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. u., and Polosukhin, I · 2017
Earlier work this paper cites.
A survey of preference-based reinforcement learning methods
Wirth, C., Akrour, R., Neumann, G., and Fürnkranz, J · 2017
Earlier work this paper cites.
Learning discourse-level diversity for neural dialog models using conditional variational autoencoders
Zhao, T., Zhao, R., and Eskenazi, M · 2017
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O · 2018
Earlier work this paper cites.
Generating informative responses with controlled sentence function
Ke, P., Guan, J., Huang, M., and Zhu, X · 2018
Earlier work this paper cites.
Unsupervised discrete sentence representation learning for interpretable neural dialog generation
Zhao, T., Lee, K., and Eskenazi, M · 2018
Earlier work this paper cites.
On the weaknesses of reinforcement learning for neural machine translation
Choshen, L., Fox, L., Aizenbud, Z., and Abend, O · 2020
Earlier work this paper cites.
A general language assistant as a laboratory for alignment
Askell, A., Bai, Y., Chen, A., Drain, D., Ganguli, D., Henighan, T., Jones, A., Joseph, N., Mann, B., DasSarma, N., et al · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., Hesse, C., and Schulman, J · 2021
Earlier work this paper cites.
Webgpt: Browser-assisted question-answering with human feedback
Nakano, R., Hilton, J., Balaji, S., Wu, J., Ouyang, L., Kim, C., Hesse, C., Jain, S., Kosaraju, V., Saunders, W., et al · 2021
Earlier work this paper cites.
Reward is enough
Silver, D., Singh, S., Precup, D., and Sutton, R. S · 2021
Cited alongside, same era.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., et al · 2022
Cited alongside, same era.
Discrete latent variable models
Bartolucci, F., Pandolfi, S., and Pennoni, F · 2022
Cited alongside, same era.
Truthfulqa: Measuring how models mimic human falsehoods
Lin, S., Hilton, J., and Evans, O · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Cited alongside, same era.
Open problems and fundamental limitations of reinforcement learning from human feedback
Human alignment of large language models through online preference optimisation
Calandriello, D., Guo, D., Munos, R., Rowland, M., Tang, Y., Pires, B. A., Richemond, P. H., Lan, C. L., Valko, M., Liu, T., et al · 2024
Later among the works it cites.
Learning to plan for language modeling from unlabeled data
Cornille, N., Moens, M.-F., and Mai, F · 2024
Later among the works it cites.
Safe RLHF: Safe reinforcement learning from human feedback
Dai, J., Pan, X., Sun, R., Ji, J., Xu, X., Liu, M., Wang, Y., and Yang, Y · 2024
Later among the works it cites.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Later among the works it cites.
Model alignment as prospect theoretic optimization
Ethayarajh, K., Xu, W., Muennighoff, N., Jurafsky, D., and Kiela, D · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Casper, S., Davies, X., Shi, C., Gilbert, T. K., Scheurer, J., Rando, J., Freedman, R., Korbak, T., Lindner, D., Freire, P., et al · 2023
Cited alongside, same era.
Ultrafeedback: Boosting language models with high-quality feedback, 2023
Cui, G., Yuan, L., Ding, N., Yao, G., Zhu, W., Ni, Y., Xie, G., Liu, Z., and Sun, M · 2023
Cited alongside, same era.
SteerLM: Attribute conditioned SFT as an (user-steerable) alternative to RLHF
Dong, Y., Wang, Z., Sreedhar, M., Wu, X., and Kuchaiev, O · 2023
Cited alongside, same era.
Aligning language models with preferences through f-divergence minimization
Go, D., Korbak, T., Kruszewski, G., Rozen, J., Ryu, N., and Dymetman, M · 2023
Cited alongside, same era.
Generating coherent narratives by learning dynamic and discrete entity states with a contrastive framework
Guan, J., Yang, Z., Zhang, R., Hu, Z., and Huang, M · 2023
Cited alongside, same era.
Reinforced self-training (rest) for language modeling
Gulcehre, C., Paine, T. L., Srinivasan, S., Konyushkova, K., Weerts, L., Sharma, A., Siddhant, A., Ahern, A., Wang, M., Gu, C., et al · 2023
Cited alongside, same era.
Personalized soups: Personalized large language model alignment via post-hoc parameter merging
Jang, J., Kim, S., Lin, B. Y., Wang, Y., Hessel, J., Zettlemoyer, L., Hajishirzi, H., Choi, Y., and Ammanabrolu, P · 2023
Cited alongside, same era.
Later among the works it cites.
A framework for few-shot language model evaluation, 07 2024
Gao, L., Tow, J., Abbasi, B., Biderman, S., Black, S., DiPofi, A., Foster, C., Golding, L., Hsu, J., Le Noac’h, A., Li, H., McDonell, K., Muennighoff, N., Ociepa, C., Phang, J., Reynolds, L., Schoelkopf, H., Skowron, A., Sutawika, L., Tang, E., Thite, A., Wang, B., Wang, K., and Zou, A · 2024
Later among the works it cites.
Openrlhf: An easy-to-use, scalable and high-performance rlhf framework
Hu, J., Wu, X., Wang, W., Xianyu, Zhang, D., and Cao, Y · 2024
Later among the works it cites.
Simpo: Simple preference optimization with a reference-free reward
Meng, Y., Xia, M., and Chen, D · 2024
Later among the works it cites.
Rule based rewards for fine-grained llm safety
Mu, T., Helyar, A., Heidecke, J., Achiam, J., Vallone, A., Kivlichan, I. D., Lin, M., Beutel, A., Schulman, J., and Weng, L · 2024
Later among the works it cites.
Personalizing reinforcement learning from human feedback with variational preference learning
Poddar, S., Wan, Y., Ivison, H., Gupta, A., and Jaques, N · 2024
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., and Finn, C · 2024
Later among the works it cites.
Rewarded soups: towards pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards
Rame, A., Couairon, G., Dancette, C., Gaya, J.-B., Shukor, M., Soulier, L., and Cord, M · 2024
Later among the works it cites.
Offline regularised reinforcement learning for large language models alignment
Richemond, P. H., Tang, Y., Guo, D., Calandriello, D., Azar, M. G., Rafailov, R., Pires, B. A., Tarassov, E., Spangher, L., Ellsworth, W., et al · 2024
Later among the works it cites.
Generalized preference optimization: A unified approach to offline alignment
Tang, Y., Guo, Z. D., Zheng, Z., Calandriello, D., Munos, R., Rowland, M., Richemond, P. H., Valko, M., Avila Pires, B., and Piot, B · 2024
Later among the works it cites.
Fine-tuning language models for factuality
Tian, K., Mitchell, E., Yao, H., Manning, C. D., and Finn, C · 2024
Later among the works it cites.
Arithmetic control of LLMs for diverse user preferences: Directional preference alignment with multi-objective rewards
Wang, H., Lin, Y., Xiong, W., Yang, R., Diao, S., Qiu, S., Zhao, H., and Zhang, T · 2024
Later among the works it cites.
No preference left behind: Group distributional preference optimization
Yao, B., Cai, Z., Chuang, Y.-S., Yang, S., Jiang, M., Yang, D., and Hu, J · 2024
Later among the works it cites.
Quiet-star: Language models can teach themselves to think before speaking
Zelikman, E., Harik, G., Shao, Y., Jayasiri, V., Haber, N., and Goodman, N. D · 2024
Later among the works it cites.
Beyond one-preference-fits-all alignment: Multi-objective direct preference optimization
Zhou, Z., Liu, J., Shao, J., Yue, X., Yang, C., Ouyang, W., and Qiao, Y · 2024
Later among the works it cites.