Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF, December 2023
Original
Anand Siththaranjan, Cassidy Laidlaw, and Dylan Hadfield-Menell · 2023
Later among the works it cites.
Towards Measuring the Representation of Subjective Global Opinions in Language Models, June 2023
Original
Esin Durmus, Karina Nyugen, Thomas I. Liao, Nicholas Schiefer, Amanda Askell, Anton Bakhtin, Carol Chen, Zac Hatfield-Dodds, Danny Hernandez, Nicholas Joseph, Liane Lovitt, Sam McCandlish, Orowa Sikder, Alex Tamkin, Janel Thamkul, Jared Kaplan, Jack Clark, and Deep Ganguli · 2023
Later among the works it cites.
Llama 2: Open Foundation and Fine-Tuned Chat Models, July 2023
Original
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom · 2023
Later among the works it cites.
Mistral 7B, October 2023
Original
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed · 2023
Later among the works it cites.
A Long Way to Go: Investigating Length Correlations in RLHF, October 2023
Original
Prasann Singhal, Tanya Goyal, Jiacheng Xu, and Greg Durrett · 2023
Later among the works it cites.
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica · 2023
Later among the works it cites.
Taking the Person Seriously: Ethically Aware IS Research in the Era of Reinforcement Learning-Based Personalization
Travis Greene, Copenhagen Business School, Galit Shmueli, National Tsing Hua University, Soumya Ray, and National Tsing Hua University · 2023
Later among the works it cites.
Inside the AI Factory, June 2023
Josh Dzieza · 2023
Later among the works it cites.
DICES Dataset: Diversity in Conversational AI Evaluation for Safety, June 2023
Original
Lora Aroyo, Alex S. Taylor, Mark Diaz, Christopher M. Homan, Alicia Parrish, Greg Serapio-Garcia, Vinodkumar Prabhakaran, and Ding Wang · 2023
Later among the works it cites.
LIMA: Less Is More for Alignment, May 2023
Original
Chunting Zhou, Pengfei Liu, Puxin Xu, Srini Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, Lili Yu, Susan Zhang, Gargi Ghosh, Mike Lewis, Luke Zettlemoyer, and Omer Levy · 2023
Later among the works it cites.
Stanford Human Preferences Dataset, September 2023
StanfordNLP · 2023
Later among the works it cites.
(InThe)WildChat: 570K ChatGPT Interaction Logs In The Wild
Wenting Zhao, Xiang Ren, Jack Hessel, Claire Cardie, Yejin Choi, and Yuntian Deng · 2023
Later among the works it cites.
OpenAssistant Conversations – Democratizing Large Language Model Alignment, October 2023
Original
Andreas Köpf, Yannic Kilcher, Dimitri von Rütte, Sotiris Anagnostidis, Zhi-Rui Tam, Keith Stevens, Abdullah Barhoum, Nguyen Minh Duc, Oliver Stanley, Richárd Nagyfi, Shahul ES, Sameer Suri, David Glushkov, Arnav Dantuluri, Andrew Maguire, Christoph Schuhmann, Huu Nguyen, and Alexander Mattick · 2023
Later among the works it cites.
How do people feel about AI? A nationally representative survey of public attitudes to artificial intelligence in Britain
The Alan Turing Institute and The Ada Lovelace Institute · 2023
Later among the works it cites.
Mirages. On Anthropomorphism in Dialogue Systems
Gavin Abercrombie, Amanda Cercas Curry, Tanvi Dinkar, Verena Rieser, and Zeerak Talat · 2023
Later among the works it cites.
The Impact of Preference Agreement in Reinforcement Learning from Human Feedback: A Case Study in Summarization, November 2023
Original
Sian Gooding and Hassan Mansoor · 2023
Later among the works it cites.
Gemini: A Family of Highly Capable Multimodal Models, December 2023
Original
Gemini Team · 2023
Later among the works it cites.
Whose Opinions Do Language Models Reflect?, March 2023
Original
Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto · 2023
Later among the works it cites.
Evaluating Agents using Social Choice Theory, December 2023
Original
Marc Lanctot, Kate Larson, Yoram Bachrach, Luke Marris, Zun Li, Avishkar Bhoopchand, Thomas Anthony, Brian Tanner, and Anna Koop · 2023
Later among the works it cites.
Elo Uncovered: Robustness and Best Practices in Language Model Evaluation, November 2023
Original
Meriem Boubdir, Edward Kim, Beyza Ermis, Sara Hooker, and Marzieh Fadaee · 2023
Later among the works it cites.
The benefits, risks and bounds of personalizing the alignment of large language models to individuals
Hannah Rose Kirk, Bertie Vidgen, Paul Röttger, and Scott A. Hale · 2024
Closest in time.
MaxMin-RLHF: Towards Equitable Alignment of Large Language Models with Diverse Human Preferences, February 2024
Original
Souradip Chakraborty, Jiahao Qiu, Hui Yuan, Alec Koppel, Furong Huang, Dinesh Manocha, Amrit Singh Bedi, and Mengdi Wang · 2024
Closest in time.
Social Choice for AI Alignment: Dealing with Diverse Human Feedback, April 2024
Original
Vincent Conitzer, Rachel Freedman, Jobst Heitzig, Wesley H. Holliday, Bob M. Jacobs, Nathan Lambert, Milan Mossé, Eric Pacuit, Stuart Russell, Hailey Schoelkopf, Emanuel Tewolde, and William S. Zwicker · 2024
Closest in time.
Investigating Cultural Alignment of Large Language Models, February 2024
Original
Badr AlKhamissi, Muhammad ElNokrashy, Mai AlKhamissi, and Mona Diab · 2024
Closest in time.
Unintended Impacts of LLM Alignment on Global Representation, February 2024
Original
Michael J. Ryan, William Held, and Diyi Yang · 2024
Closest in time.
A Roadmap to Pluralistic Alignment, February 2024
Original
Taylor Sorensen, Jared Moore, Jillian Fisher, Mitchell Gordon, Niloofar Mireshghallah, Christopher Michael Rytting, Andre Ye, Liwei Jiang, Ximing Lu, Nouha Dziri, Tim Althoff, and Yejin Choi · 2024
Closest in time.
Not All Countries Celebrate Thanksgiving: On the Cultural Dominance in Large Language Models, February 2024
Original
Wenxuan Wang, Wenxiang Jiao, Jingyuan Huang, Ruyi Dai, Jen-tse Huang, Zhaopeng Tu, and Michael R. Lyu · 2024
Closest in time.
Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs, February 2024
Original
Arash Ahmadian, Chris Cremer, Matthias Gallé, Marzieh Fadaee, Julia Kreutzer, Olivier Pietquin, Ahmet Üstün, and Sara Hooker · 2024
Closest in time.
RewardBench: Evaluating Reward Models for Language Modeling, March 2024
Original
Nathan Lambert, Valentina Pyatkin, Jacob Morrison, L. J. Miranda, Bill Yuchen Lin, Khyathi Chandu, Nouha Dziri, Sachin Kumar, Tom Zick, Yejin Choi, Noah A. Smith, and Hannaneh Hajishirzi · 2024
Closest in time.
KTO: Model Alignment as Prospect Theoretic Optimization, February 2024
Original
Kawin Ethayarajh, Winnie Xu, Niklas Muennighoff, Dan Jurafsky, and Douwe Kiela · 2024
Closest in time.
The illusion of artificial inclusion, February 2024
Original
William Agnew, A. Stevie Bergman, Jennifer Chien, Mark Díaz, Seliem El-Sayed, Jaylen Pittman, Shakir Mohamed, and Kevin R. McKee · 2024
Closest in time.
AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback, January 2024
Original
Yann Dubois, Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto · 2024
Closest in time.
Controllable Preference Optimization: Toward Controllable Multi-Objective Alignment, February 2024
Original
Yiju Guo, Ganqu Cui, Lifan Yuan, Ning Ding, Jiexin Wang, Huimin Chen, Bowen Sun, Ruobing Xie, Jie Zhou, Yankai Lin, Zhiyuan Liu, and Maosong Sun · 2024
Closest in time.
LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset, March 2024
Original
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Tianle Li, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zhuohan Li, Zi Lin, Eric P. Xing, Joseph E. Gonzalez, Ion Stoica, and Hao Zhang · 2024
Closest in time.
Aya Dataset: An Open-Access Collection for Multilingual Instruction Tuning, February 2024
Original
Shivalika Singh, Freddie Vargus, Daniel Dsouza, Börje F. Karlsson, Abinaya Mahendiran, Wei-Yin Ko, Herumb Shandilya, Jay Patel, Deividas Mataciunas, Laura OMahony, Mike Zhang, Ramith Hettiarachchi, Joseph Wilson, Marina Machado, Luisa Souza Moura, Dominik Krzemiński, Hakimeh Fadaei, Irem Ergün, Ifeoma Okoh, Aisha Alaagib, Oshan Mudannayake, Zaid Alyafeai, Vu Minh Chien, Sebastian Ruder, Surya Guthikonda, Emad A. Alghamdi, Sebastian Gehrmann, Niklas Muennighoff, Max Bartolo, Julia Kreutzer, Ahmet Üstün, Marzieh Fadaee, and Sara Hooker · 2024
Closest in time.
Particip-AI: A Democratic Surveying Framework for Anticipating Future AI Use Cases, Harms and Benefits, March 2024
Original
Jimin Mun, Liwei Jiang, Jenny Liang, Inyoung Cheong, Nicole DeCario, Yejin Choi, Tadayoshi Kohno, and Maarten Sap · 2024
Closest in time.
Meta Community Forum: Results Analysis
Samuel Chang, Estelle Ciesla, Michael Finch, James Fishkin, Lodewijk Gelauff, Ashish Goel, Ricky Hernandez Marquez, Shoaib Mohammed, and Alice Siu · 2024
Closest in time.
AnthroScore: A Computational Linguistic Measure of Anthropomorphism, February 2024
Original
Myra Cheng, Kristina Gligoric, Tiziano Piccardi, and Dan Jurafsky · 2024
Closest in time.
Can LLM be a Personalized Judge?, June 2024
Original
Yijiang River Dong, Tiancheng Hu, and Nigel Collier · 2024
Closest in time.
Inside OpenAI’s Plan to Make AI More ’Democratic’, February 2024
Billy Perrigo · 2024
Closest in time.
Human Feedback is not Gold Standard, January 2024
Original
Tom Hosking, Phil Blunsom, and Max Bartolo · 2024
Closest in time.
Mixtral of Experts, January 2024
Original
Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, Lélio Renard Lavaud, Lucile Saulnier, Marie-Anne Lachaux, Pierre Stock, Sandeep Subramanian, Sophia Yang, Szymon Antoniak, Teven Le Scao, Théophile Gervet, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed · 2024
Closest in time.
Introducing the next generation of Claude, April 2024
Anthropic · 2024
Closest in time.
Command R, April 2024
Cohere · 2024
Closest in time.
Introducing Meta Llama 3: The most capable openly available LLM to date, April 2024
MetaAI · 2024
Closest in time.
Representative samples, February 2024
Prolific · 2024
Closest in time.
How social relationships shape moral wrongness judgments
Brian D. Earp, Killian L. McLoughlin, Joshua T. Monrad, Margaret S. Clark, and Molly J. Crockett · 2041
Closest in time.
The misuse of colour in science communication
Fabio Crameri, Grace E. Shephard, and Philip J. Heron · 2041
Closest in time.
STELA: a community-centred approach to norm elicitation for AI alignment
Stevie Bergman, Nahema Marchal, John Mellor, Shakir Mohamed, Iason Gabriel, and William Isaac · 2045
Closest in time.
Comparing attentional disengagement between Prolific and MTurk samples
Derek A. Albert and Daniel Smilek · 2045
Closest in time.
Controversies, contradiction, and “participation” in AI
Mona Sloane · 2053
Closest in time.