Fetching the paper…
Reading the bibliography…
Alignment of Large Language models (LLMs) is crucial for safe and trustworthy deployment in applications.
On the limitations of scalarisation for multi-objective reinforcement learning of pareto fronts
Peter Vamplew, John Yearwood, Richard Dazeley, and Adam Berry · 2008
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Human-aligned artificial intelligence is a multiobjective problem
Peter Vamplew, Richard Dazeley, Cameron Foale, Sally Firmin, and Jane Mummery · 2018
Earlier work this paper cites.
Fine-tuning language models from human preferences, 2020
Daniel M. Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B. Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving · 2020
Earlier work this paper cites.
Webgpt: Browser-assisted question-answering with human feedback
Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, et al · 2021
Earlier work this paper cites.
FUDGE: Controlled text generation with future discriminators
Kevin Yang and Dan Klein · 2021
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al · 2022
Earlier work this paper cites.
Scaling laws for reward model overoptimization, 2022
Leo Gao, John Schulman, and Jacob Hilton · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback, 2022
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe · 2022
Earlier work this paper cites.
Learning to summarize from human feedback, 2022
Nisan Stiennon, Long Ouyang, Jeff Wu, Daniel M. Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul Christiano · 2022
Earlier work this paper cites.
A contrastive framework for neural text generation
Yixuan Su, Tian Lan, Yan Wang, Dani Yogatama, Lingpeng Kong, and Nigel Collier · 2022
Earlier work this paper cites.
Scaling laws for reward model overoptimization
Leo Gao, John Schulman, and Jacob Hilton · 2023
Cited alongside, same era.
Aligning language models with preferences through f-divergence minimization, 2023
Dongyoung Go, Tomasz Korbak, Germán Kruszewski, Jos Rozen, Nahyeon Ryu, and Marc Dymetman · 2023
Cited alongside, same era.
Personalized soups: Personalized large language model alignment via post-hoc parameter merging
Joel Jang, Seungone Kim, Bill Yuchen Lin, Yizhong Wang, Jack Hessel, Luke Zettlemoyer, Hannaneh Hajishirzi, Yejin Choi, and Prithviraj Ammanabrolu · 2023
Cited alongside, same era.
Augmented language models: a survey, 2023
Grégoire Mialon, Roberto Dessì, Maria Lomeli, Christoforos Nalmpantis, Ram Pasunuru, Roberta Raileanu, Baptiste Rozière, Timo Schick, Jane Dwivedi-Yu, Asli Celikyilmaz, Edouard Grave, Yann LeCun, and Thomas Scialom · 2023
Cited alongside, same era.
Direct preference optimization: Your language model is secretly a reward model, 2023
Collaborative decoding of critical tokens for boosting factuality of large language models, 2024
Lifeng Jin, Baolin Peng, Linfeng Song, Haitao Mi, Ye Tian, and Dong Yu · 2024
Later among the works it cites.
Args: Alignment as reward-guided search, 2024
Maxim Khanov, Jirayu Burapacheep, and Yixuan Li · 2024
Later among the works it cites.
Rewardbench: Evaluating reward models for language modeling, 2024
Nathan Lambert, Valentina Pyatkin, Jacob Morrison, LJ Miranda, Bill Yuchen Lin, Khyathi Chandu, Nouha Dziri, Sachin Kumar, Tom Zick, Yejin Choi, Noah A. Smith, and Hannaneh Hajishirzi · 2024
Later among the works it cites.
Controllable text generation for large language models: A survey, 2024
Xun Liang, Hanyu Wang, Yezhaohui Wang, Shichao Song, Jiawei Yang, Simin Niu, Jie Hu, Dan Liu, Shunyu Yao, Feiyu Xiong, and Zhiyu Li · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning, and Chelsea Finn · 2023
Cited alongside, same era.
Rrhf: Rank responses to align language models with human feedback without tears, 2023
Zheng Yuan, Hongyi Yuan, Chuanqi Tan, Wei Wang, Songfang Huang, and Fei Huang · 2023
Cited alongside, same era.
Starling-7b: Improving llm helpfulness and harmlessness with rlaif, 2023
Banghua Zhu, Evan Frick, Tianhao Wu, Hanlin Zhu, and Jiantao Jiao · 2023
Cited alongside, same era.
Self-play fine-tuning converts weak language models to strong language models
Zixiang Chen, Yihe Deng, Huizhuo Yuan, Kaixuan Ji, and Quanquan Gu · 2024
Cited alongside, same era.
Attacks, defenses and evaluations for llm conversation safety: A survey, 2024
Zhichen Dong, Zhanhui Zhou, Chao Yang, Jing Shao, and Yu Qiao · 2024
Cited alongside, same era.
Deal: Decoding-time alignment for large language models
James Y Huang, Sailik Sengupta, Daniele Bonadiman, Yi-an Lai, Arshit Gupta, Nikolaos Pappas, Saab Mansour, Katrin Kirchoff, and Dan Roth · 2024
Cited alongside, same era.
Domain-specialized llm: Financial fine-tuning and utilization method using mistral 7b
Cheonsu Jeong · 2024
Cited alongside, same era.
Parl: A unified framework for policy alignment in reinforcement learning
Souradip Chakraborty, Amrit Bedi, Alec Koppel, Huazheng Wang, Dinesh Manocha, Mengdi Wang, and Furong Huang
Cited in the paper.
Alisa Liu, Xiaochuang Han, Yizhong Wang, Yulia Tsvetkov, Yejin Choi, and Noah A. Smith · 2024
Later among the works it cites.
Controlled decoding from language models, 2024
Sidharth Mudgal, Jong Lee, Harish Ganapathy, YaGuang Li, Tao Wang, Yanping Huang, Zhifeng Chen, Heng-Tze Cheng, Michael Collins, Trevor Strohman, Jilin Chen, Alex Beutel, and Ahmad Beirami · 2024
Later among the works it cites.
Learning to decode collaboratively with multiple language models
Zejiang Shen, Hunter Lang, Bailin Wang, Yoon Kim, and David Sontag · 2024
Later among the works it cites.
Factuality of large language models in the year 2024, 2024
Yuxia Wang, Minghan Wang, Muhammad Arslan Manzoor, Fei Liu, Georgi Georgiev, Rocktim Jyoti Das, and Preslav Nakov · 2024
Later among the works it cites.
Personalized large language models, 2024
Stanisław Woźniak, Bartłomiej Koptyra, Arkadiusz Janz, Przemysław Kazienko, and Jan Kocoń · 2024
Later among the works it cites.
Sorry-bench: Systematically evaluating large language model safety refusal behaviors
Tinghao Xie, Xiangyu Qi, Yi Zeng, Yangsibo Huang, Udari Madhushani Sehwag, Kaixuan Huang, Luxi He, Boyi Wei, Dacheng Li, Ying Sheng, et al · 2024
Later among the works it cites.
Iterative data smoothing: Mitigating reward overfitting and overoptimization in rlhf, 2024
Banghua Zhu, Michael I. Jordan, and Jiantao Jiao · 2024
Later among the works it cites.