Fetching the paper…
Reading the bibliography…
The rapid advancement of language models (LMs) necessitates robust alignment with diverse user values.
The next big five inventory (bfi-2): Developing and assessing a hierarchical model with 15 facets to enhance bandwidth, fidelity, and predictive power
Christopher J Soto and Oliver P John · 2017
Earlier work this paper cites.
The potential and issues in data-driven development of web personas
Tea Mijač, Mario Jadrić, and Maja Ćukušić · 2018
Earlier work this paper cites.
Prolific. ac—a subject pool for online experiments
Stefan Palan and Christian Schitter · 2018
Earlier work this paper cites.
Language (technology) is power: A critical survey of" bias" in nlp
Su Lin Blodgett, Solon Barocas, Hal Daumé III, and Hanna Wallach · 2020
Earlier work this paper cites.
Learning to summarize with human feedback
Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano · 2020
Earlier work this paper cites.
Constitutional ai: Harmlessness from ai systems using constitutional principles
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, et al · 2022
Earlier work this paper cites.
Challenges and strategies in cross-cultural nlp
Daniel Hershcovich, Stella Frank, Heather Lent, Miryam de Lhoneux, Mostafa Abdou, Stephanie Brandl, Emanuele Bugliarello, Laura Cabello Piqueras, Ilias Chalkidis, Ruixiang Cui, et al · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Earlier work this paper cites.
Ultrafeedback: Boosting language models with high-quality feedback, 2023
Ganqu Cui, Lifan Yuan, Ning Ding, Guanming Yao, Wei Zhu, Yuan Ni, Guotong Xie, Zhiyuan Liu, and Maosong Sun · 2023
Earlier work this paper cites.
Alpacafarm: A simulation framework for methods that learn from human feedback
Y. Dubois, N. Du, K. Zhang, Y. Zhang, S. Agrawal, Y. Cao, S. Salehi, J. Kim, S. Li, S. Zhang, et al · 2023
Earlier work this paper cites.
Towards measuring the representation of subjective global opinions in language models
Esin Durmus, Karina Nyugen, Thomas I Liao, Nicholas Schiefer, Amanda Askell, Anton Bakhtin, Carol Chen, Zac Hatfield-Dodds, Danny Hernandez, Nicholas Joseph, et al · 2023
Earlier work this paper cites.
Social contract ai: Aligning ai assistants with implicit group norms
Jan-Philipp Fränken, Sam Kwok, Peixuan Ye, Kanishk Gandhi, Dilip Arumugam, Jared Moore, Alex Tamkin, Tobias Gerstenberg, and Noah D Goodman · 2023
Earlier work this paper cites.
Aligning language models to user opinions
EunJeong Hwang, Bodhisattwa Prasad Majumder, and Niket Tandon · 2023
Earlier work this paper cites.
Personalized soups: Personalized large language model alignment via post-hoc parameter merging, 2023
Joel Jang, Seungone Kim, Bill Yuchen Lin, Yizhong Wang, Jack Hessel, Luke Zettlemoyer, Hannaneh Hajishirzi, Yejin Choi, and Prithviraj Ammanabrolu · 2023
Earlier work this paper cites.
Can large language models capture dissenting human voices?, 2023
Noah Lee, Na Min An, and James Thorne · 2023
Earlier work this paper cites.
Eliciting human preferences with language models
Belinda Z Li, Alex Tamkin, Noah Goodman, and Jacob Andreas · 2023
Earlier work this paper cites.
Chatharuhi: Reviving anime character in reality via large language model
Cheng Li, Ziang Leng, Chenxi Yan, Junyi Shen, Hao Wang, Weishi Mi, Yaying Fei, Xiaoyang Feng, Song Yan, HaoSheng Wang, et al · 2023
Earlier work this paper cites.
Lost in the middle: How language models use long contexts, 2023
Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang · 2023
Earlier work this paper cites.
Sample efficient reinforcement learning from human feedback via active exploration, 2023
Viraj Mehta, Vikramjeet Das, Ojash Neopane, Yijia Dai, Ilija Bogunovic, Jeff Schneider, and Willie Neiswanger · 2023
Earlier work this paper cites.
Generative agents: Interactive simulacra of human behavior
Joon Sung Park, Joseph O’Brien, Carrie Jun Cai, Meredith Ringel Morris, Percy Liang, and Michael S Bernstein · 2023
Cited alongside, same era.
Active preference inference using language models and probabilistic reasoning, 2023
Top Piriyakulkij, Volodymyr Kuleshov, and Kevin Ellis · 2023
Cited alongside, same era.
Whose opinions do language models reflect?
Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto · 2023
Cited alongside, same era.
Whose opinions do language models reflect?, 2023
Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto · 2023
Cited alongside, same era.
Character-llm: A trainable agent for role-playing
Yunfan Shao, Linyang Li, Junqi Dai, and Xipeng Qiu · 2023
Cited alongside, same era.
Personas as a way to model truthfulness in language models, 2024
Nitish Joshi, Javier Rando, Abulhair Saparov, Najoung Kim, and He He · 2024
Closest in time.
Hannah Rose Kirk, Alexander Whitefield, Paul Röttger, Andrew Bean, Katerina Margatina, Juan Ciro, Rafael Mosquera, Max Bartolo, Adina Williams, He He, et al · 2024
Closest in time.
Personalized language modeling from personalized human feedback, 2024
Xinyu Li, Zachary C. Lipton, and Liu Leqi · 2024
Closest in time.
The generation gap:exploring age bias in the underlying value systems of large language models, 2024
Siyang Liu, Trish Maturi, Bowen Yi, Siqi Shen, and Rada Mihalcea · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Not all countries celebrate thanksgiving: On the cultural dominance in large language models
Wenxuan Wang, Wenxiang Jiao, Jingyuan Huang, Ruyi Dai, Jen-tse Huang, Zhaopeng Tu, and Michael R Lyu · 2023
Cited alongside, same era.
Are language models worse than humans at following prompts? it’s complicated
Albert Webson, Alyssa Marie Loo, Qinan Yu, and Ellie Pavlick · 2023
Cited alongside, same era.
Group preference optimization: Few-shot alignment of large language models, 2023
Siyan Zhao, John Dang, and Aditya Grover · 2023
Cited alongside, same era.
Judging llm-as-a-judge with mt-bench and chatbot arena, 2023
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica · 2023
Cited alongside, same era.
Star-gate: Teaching language models to ask clarifying questions
Chinmaya Andukuri, Jan-Philipp Fränken, Tobias Gerstenberg, and Noah D Goodman · 2024
Cited alongside, same era.
Star-gate: Teaching language models to ask clarifying questions, 2024
Chinmaya Andukuri, Jan-Philipp Fränken, Tobias Gerstenberg, and Noah D. Goodman · 2024
Cited alongside, same era.
Measuring political bias in large language models: What is said and how it is said, 2024
Yejin Bang, Delong Chen, Nayeon Lee, and Pascale Fung · 2024
Cited alongside, same era.
Keming Lu, Bowen Yu, Chang Zhou, and Jingren Zhou · 2024
Closest in time.
Introducing llama3
Meta · 2024
Closest in time.
Active preference learning for large language models, 2024
William Muldrew, Peter Hayes, Mingtian Zhang, and David Barber · 2024
Closest in time.
Scaling laws for reward model overoptimization in direct alignment algorithms, 2024
Rafael Rafailov, Yaswanth Chittepu, Ryan Park, Harshit Sikchi, Joey Hejna, Bradley Knox, Chelsea Finn, and Scott Niekum · 2024
Closest in time.
Group robust preference optimization in reward-free rlhf, 2024
Shyam Sundhar Ramesh, Yifan Hu, Iason Chaimalas, Viraj Mehta, Pier Giuseppe Sessa, Haitham Bou Ammar, and Ilija Bogunovic · 2024
Closest in time.
Assessing political bias in large language models, 2024
Luca Rettenberger, Markus Reischl, and Mark Schutera · 2024
Closest in time.
Show, don’t tell: Aligning language models with demonstrated feedback
Omar Shaikh, Michelle Lam, Joey Hejna, Yijia Shao, Michael Bernstein, and Diyi Yang · 2024
Closest in time.
Distributional preference learning: Understanding and accounting for hidden context in rlhf, 2024
Anand Siththaranjan, Cassidy Laidlaw, and Dylan Hadfield-Menell · 2024
Closest in time.
A roadmap to pluralistic alignment, 2024
Taylor Sorensen, Jared Moore, Jillian Fisher, Mitchell Gordon, Niloofar Mireshghallah, Christopher Michael Rytting, Andre Ye, Liwei Jiang, Ximing Lu, Nouha Dziri, Tim Althoff, and Yejin Choi · 2024
Closest in time.
Persona-db: Efficient large language model personalization for response prediction with collaborative data refinement, 2024
Chenkai Sun, Ke Yang, Revanth Gangi Reddy, Yi R. Fung, Hou Pong Chan, ChengXiang Zhai, and Heng Ji · 2024
Closest in time.
Gemini: A family of highly capable multimodal models, 2024
Gemini Team · 2024
Closest in time.
American community survey (acs) public use microdata sample (pums), 2024
United States Census Bureau · 2024
Closest in time.
Character is destiny: Can large language models simulate persona-driven decisions in role-playing?, 2024
Rui Xu, Xintao Wang, Jiangjie Chen, Siyu Yuan, Xinfeng Yuan, Jiaqing Liang, Zulong Chen, Xiaoqing Dong, and Yanghua Xiao · 2024
Closest in time.
Self-exploring language models: Active preference elicitation for online alignment, 2024
Shenao Zhang, Donghan Yu, Hiteshi Sharma, Ziyi Yang, Shuohang Wang, Hany Hassan, and Zhaoran Wang · 2024
Closest in time.