Fetching the paper…
Reading the bibliography…
While existing alignment paradigms have been integral in developing large language models (LLMs), LLMs often learn an averaged human preference and struggle to model diverse preferences across cultures, demographics, and communities.
Two concepts of liberty
Isaiah Berlin. 1969 · 1969
Earlier work this paper cites.
The fragmentation of value
Thomas Nagel. 1979 · 1979
Earlier work this paper cites.
Truth and Objectivity
Crispin Wright. 1992 · 1992
Earlier work this paper cites.
Value-focused thinking: A path to creative decisionmaking
Ralph L Keeney and Ralph L Keeney. 2009 · 2009
Earlier work this paper cites.
Language and Gender , 2 edition
Penelope Eckert and Sally McConnell-Ginet. 2013 · 2013
Earlier work this paper cites.
Mixture of experts: a literature survey
Saeed Masoudnia and Reza Ebrahimpour. 2014 · 2014
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017 · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017 · 2017
Earlier work this paper cites.
Scalable agent alignment via reward modeling: a research direction
Jan Leike, David Krueger, Tom Everitt, Miljan Martic, Vishal Maini, and Shane Legg. 2018 · 2018
Earlier work this paper cites.
Artificial intelligence, values, and alignment
Iason Gabriel. 2020 · 2020
Earlier work this paper cites.
Learning to summarize with human feedback
Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano. 2020 · 2020
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Edward J Hu, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et al. 2021 · 2021
Earlier work this paper cites.
Dexperts: Decoding-time controlled text generation with experts and anti-experts
Alisa Liu, Maarten Sap, Ximing Lu, Swabha Swayamdipta, Chandra Bhagavatula, Noah A Smith, and Yejin Choi. 2021 · 2021
Earlier work this paper cites.
Get your vitamin c! robust fact verification with contrastive evidence
Tal Schuster, Adam Fisch, and Regina Barzilay. 2021 · 2021
Earlier work this paper cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. 2024 · 2021
Earlier work this paper cites.
Constitutional ai: Harmlessness from ai feedback
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al. 2022 · 2022
Earlier work this paper cites.
Communitylm: Probing partisan worldviews from language models
Hang Jiang, Doug Beeferman, Brandon Roy, and Deb Roy. 2022 · 2022
Earlier work this paper cites.
Identifying the human values behind arguments
Johannes Kiesel, Milad Alshomary, Nicolas Handke, Xiaoni Cai, Henning Wachsmuth, and Benno Stein. 2022 · 2022
Earlier work this paper cites.
Wanli: Worker and ai collaboration for natural language inference dataset creation
Alisa Liu, Swabha Swayamdipta, Noah A Smith, and Yejin Choi. 2022 · 2022
Earlier work this paper cites.
Who is GPT-3? an exploration of personality, values and demographics
Marilù Miotto, Nicola Rossberg, and Bennett Kleinberg. 2022 · 2022
Earlier work this paper cites.
ArtELingo: A million emotion annotations of WikiArt with emphasis on diversity over language and culture
Youssef Mohamed, Mohamed Abdelfattah, Shyma Alhuwaider, Feifan Li, Xiangliang Zhang, Kenneth Church, and Mohamed Elhoseiny. 2022 · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022 · 2022
Earlier work this paper cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023 · 2023
Earlier work this paper cites.
Confidence-based ensembling of perspective-aware models
Silvia Casola, Soda Lo, Valerio Basile, Simona Frenda, Alessandra Cignarella, Viviana Patti, and Cristina Bosco. 2023 · 2023
Earlier work this paper cites.
Open problems and fundamental limitations of reinforcement learning from human feedback
Stephen Casper, Xander Davies, Claudia Shi, Thomas Krendl Gilbert, Jérémy Scheurer, Javier Rando, Rachel Freedman, Tomasz Korbak, David Lindner, Pedro Freire, et al. 2023 · 2023
Earlier work this paper cites.
Sociocultural norm similarities and differences via situational alignment and explainable textual entailment
Sky CH-Wang, Arkadiy Saakyan, Oliver Li, Zhou Yu, and Smaranda Muresan. 2023 · 2023
Cited alongside, same era.
SOUL: Towards sentiment and opinion understanding of language
Yue Deng, Wenxuan Zhang, Sinno Pan, and Lidong Bing. 2023 · 2023
Cited alongside, same era.
Towards measuring the representation of subjective global opinions in language models
Esin Durmus, Karina Nyugen, Thomas I Liao, Nicholas Schiefer, Amanda Askell, Anton Bakhtin, Carol Chen, Zac Hatfield-Dodds, Danny Hernandez, Nicholas Joseph, et al. 2023 · 2023
Cited alongside, same era.
From pretraining data to language models to downstream tasks: Tracking the trails of political biases leading to unfair nlp models
Shangbin Feng, Chan Young Park, Yuhan Liu, and Yulia Tsvetkov. 2023 · 2023
Cited alongside, same era.
NORMSAGE: Multi-lingual multi-cultural norm discovery from conversations on-the-fly
Yi Fung, Tuhin Chakrabarty, Hao Guo, Owen Rambow, Smaranda Muresan, and Heng Ji. 2023 · 2023
Dices dataset: Diversity in conversational ai evaluation for safety
Lora Aroyo, Alex Taylor, Mark Diaz, Christopher Homan, Alicia Parrish, Gregory Serapio-García, Vinodkumar Prabhakaran, and Ding Wang. 2024 · 2024
Closest in time.
Maxmin-rlhf: Towards equitable alignment of large language models with diverse human preferences
Souradip Chakraborty, Jiahao Qiu, Hui Yuan, Alec Koppel, Furong Huang, Dinesh Manocha, Amrit Singh Bedi, and Mengdi Wang. 2024 · 2024
Closest in time.
IterAlign: Iterative constitutional alignment of large language models
Xiusi Chen, Hongzhi Wen, Sreyashi Nag, Chen Luo, Qingyu Yin, Ruirui Li, Zheng Li, and Wei Wang. 2024a · 2024
Closest in time.
Knowledge card: Filling llms’ knowledge gaps with plug-in specialized language models
Shangbin Feng, Weijia Shi, Yuyang Bai, Vidhisha Balachandran, Tianxing He, and Yulia Tsvetkov. 2024 · 2024
Closest in time.
Rebel: Reinforcement learning via regressing relative rewards
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Reinforced self-training (rest) for language modeling
Caglar Gulcehre, Tom Le Paine, Srivatsan Srinivasan, Ksenia Konyushkova, Lotte Weerts, Abhishek Sharma, Aditya Siddhant, Alex Ahern, Miaosen Wang, Chenjie Gu, et al. 2023 · 2023
Cited alongside, same era.
Culturally aware natural language inference
Jing Huang and Diyi Yang. 2023 · 2023
Cited alongside, same era.
Aligning language models to user opinions
EunJeong Hwang, Bodhisattwa Majumder, and Niket Tandon. 2023 · 2023
Cited alongside, same era.
Personalized soups: Personalized large language model alignment via post-hoc parameter merging
Joel Jang, Seungone Kim, Bill Yuchen Lin, Yizhong Wang, Jack Hessel, Luke Zettlemoyer, Hannaneh Hajishirzi, Yejin Choi, and Prithviraj Ammanabrolu. 2023 · 2023
Cited alongside, same era.
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. 2023 · 2023
Cited alongside, same era.
From values to opinions: Predicting human behaviors and stances using value-injected large language models
Dongjun Kang, Joonsuk Park, Yohan Jo, and JinYeong Bak. 2023 · 2023
Cited alongside, same era.
DLAMA: A framework for curating culturally diverse facts for probing the knowledge of pretrained language models
Amr Keleg and Walid Magdy. 2023 · 2023
Cited alongside, same era.
Zhaolin Gao, Jonathan D Chang, Wenhao Zhan, Owen Oertell, Gokul Swamy, Kianté Brantley, Thorsten Joachims, J Andrew Bagnell, Jason D Lee, and Wen Sun. 2024 · 2024
Closest in time.
Building knowledge-guided lexica to model cultural variation
Shreya Havaldar, Salvatore Giorgi, Sunny Rai, Thomas Talhelm, Sharath Chandra Guntuku, and Lyle Ungar. 2024 · 2024
Closest in time.
Flames: Benchmarking value alignment of LLMs in Chinese
Kexin Huang, Xiangyang Liu, Qianyu Guo, Tianxiang Sun, Jiawei Sun, Yaru Wang, Zeyang Zhou, Yixu Wang, Yan Teng, Xipeng Qiu, Yingchun Wang, and Dahua Lin. 2024 · 2024
Closest in time.
Hannah Rose Kirk, Alexander Whitefield, Paul Röttger, Andrew Bean, Katerina Margatina, Juan Ciro, Rafael Mosquera, Max Bartolo, Adina Williams, He He, et al. 2024 · 2024
Closest in time.
Are multilingual LLMs culturally-diverse reasoners? an investigation into multicultural proverbs and sayings
Chen Liu, Fajri Koto, Timothy Baldwin, and Iryna Gurevych. 2024a · 2024
Closest in time.
P 3 Sum: Preserving author’s perspective in news summarization with diffusion language models
Yuhan Liu, Shangbin Feng, Xiaochuang Han, Vidhisha Balachandran, Chan Young Park, Sachin Kumar, and Yulia Tsvetkov. 2024b · 2024
Closest in time.
Self-alignment of large language models via monopolylogue-based social scene simulation
Xianghe Pang, Shuo Tang, Rui Ye, Yuxin Xiong, Bolun Zhang, Yanfeng Wang, and Siheng Chen. 2024 · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. 2024 · 2024
Closest in time.
Normad: A benchmark for measuring the cultural adaptability of large language models
Abhinav Rao, Akhila Yerukola, Vishwa Shah, Katharina Reinecke, and Maarten Sap. 2024 · 2024
Closest in time.
Evaluating the moral beliefs encoded in llms
Nino Scherrer, Claudia Shi, Amir Feder, and David Blei. 2024 · 2024
Closest in time.
Understanding the capabilities and limitations of large language models for cultural commonsense
Siqi Shen, Lajanugen Logeswaran, Moontae Lee, Honglak Lee, Soujanya Poria, and Rada Mihalcea. 2024 · 2024
Closest in time.
Weiyan Shi, Ryan Li, Yutong Zhang, Caleb Ziems, Raya Horesh, Rogério Abreu de Paula, Diyi Yang, et al. 2024 · 2024
Closest in time.
Principle-driven self-alignment of language models from scratch with minimal human supervision
Zhiqing Sun, Yikang Shen, Qinhong Zhou, Hongxin Zhang, Zhenfang Chen, David Cox, Yiming Yang, and Chuang Gan. 2024 · 2024
Closest in time.
Gemma: Open models based on gemini research and technology
Gemma Team, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivière, Mihir Sanjay Kale, Juliette Love, et al. 2024 · 2024
Closest in time.
SeaEval for multilingual foundation models: From cross-lingual alignment to cultural reasoning
Bin Wang, Zhengyuan Liu, Xin Huang, Fangkai Jiao, Yang Ding, AiTi Aw, and Nancy Chen. 2024a · 2024
Closest in time.
Fake alignment: Are LLMs really aligned well?
Yixu Wang, Yan Teng, Kexin Huang, Chengqi Lyu, Songyang Zhang, Wenwei Zhang, Xingjun Ma, Yu-Gang Jiang, Yu Qiao, and Yingchun Wang. 2024b · 2024
Closest in time.
Self-play preference optimization for language model alignment
Yue Wu, Zhiqing Sun, Huizhuo Yuan, Kaixuan Ji, Yiming Yang, and Quanquan Gu. 2024 · 2024
Closest in time.
Aligning as debiasing: Causality-aware alignment via reinforcement learning with interventional feedback
Yu Xia, Tong Yu, Zhankui He, Handong Zhao, Julian McAuley, and Shuai Li. 2024 · 2024
Closest in time.
Value FULCRA: Mapping large language models to the multidimensional spectrum of basic human value
Jing Yao, Xiaoyuan Yi, Yifan Gong, Xiting Wang, and Xing Xie. 2024 · 2024
Closest in time.
Self-rewarding language models
Weizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho, Sainbayar Sukhbaatar, Jing Xu, and Jason Weston. 2024 · 2024
Closest in time.
Fair abstractive summarization of diverse perspectives
Yusen Zhang, Nan Zhang, Yixin Liu, Alexander Fabbri, Junru Liu, Ryo Kamoi, Xiaoxin Lu, Caiming Xiong, Jieyu Zhao, Dragomir Radev, Kathleen McKeown, and Rui Zhang. 2024 · 2024
Closest in time.