Fetching the paper…
Reading the bibliography…
Building pluralistic AI requires designing models that are able to be shaped to represent a wide range of value systems and cultures.
Language models are few-shot learners
Tom B Brown. 2020 · 2005
Earlier work this paper cites.
Moral foundations questionnaire
Jesse Graham, Brian A Nosek, Jonathan Haidt, Ravi Iyer, Koleva Spassena, and Peter H Ditto. 2008 · 2008
Earlier work this paper cites.
An overview of the Schwartz theory of basic values
Shalom H Schwartz. 2012 · 2012
Earlier work this paper cites.
The power of scale for parameter-efficient prompt tuning
Brian Lester, Rami Al-Rfou, and Noah Constant. 2021 · 2021
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al. 2022 · 2022
Earlier work this paper cites.
Discovering language model behaviors with model-written evaluations
Ethan Perez, Sam Ringer, Kamilė Lukošiūtė, Karina Nguyen, Edwin Chen, Scott Heiner, Craig Pettit, Catherine Olsson, Sandipan Kundu, Saurav Kadavath, et al. 2022 · 2022
Earlier work this paper cites.
Steering large language models using APE
Yongchao Zhou, Andrei Ioan Muresanu, Ziwen Han, Keiran Paster, Silviu Pitis, Harris Chan, and Jimmy Ba. 2022 · 2022
Earlier work this paper cites.
Moral foundations of large language models
Marwa Abdulhai, Gregory Serapio-Garcia, Clément Crepy, Daria Valter, John Canny, and Natasha Jaques. 2023 · 2023
Earlier work this paper cites.
Steering large language models for machine translation with finetuning and in-context learning
Duarte M Alves, Nuno M Guerreiro, João Alves, José Pombal, Ricardo Rei, José GC de Souza, Pierre Colombo, and André FT Martins. 2023 · 2023
Earlier work this paper cites.
What’s the magic word? A control theory of LLM prompting
Aman Bhargava, Cameron Witkowski, Manav Shah, and Matt Thomson. 2023 · 2023
Earlier work this paper cites.
Marked personas: Using natural language prompts to measure stereotypes in language models
Myra Cheng, Esin Durmus, and Dan Jurafsky. 2023 · 2023
Earlier work this paper cites.
Large language models as superpositions of cultural perspectives
Grgur Kovač, Masataka Sawayama, Rémy Portelas, Cédric Colas, Peter Ford Dominey, and Pierre-Yves Oudeyer. 2023 · 2023
Earlier work this paper cites.
On the steerability of large language models toward data-driven personas
Junyi Li, Ninareh Mehrabi, Charith Peris, Palash Goyal, Kai-Wei Chang, Aram Galstyan, Richard Zemel, and Rahul Gupta. 2023 · 2023
Earlier work this paper cites.
Steering Llama 2 via contrastive activation addition
Nina Rimsky, Nick Gabrieli, Julian Schulz, Meg Tong, Evan Hubinger, and Alexander Matt Turner. 2023 · 2023
Earlier work this paper cites.
Whose opinions do language models reflect?
Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto. 2023 · 2023
Cited alongside, same era.
Activation addition: Steering language models without optimization
Alexander Matt Turner, Lisa Thiergart, Gavin Leech, David Udell, Juan J Vazquez, Ulisse Mini, and Monte MacDiarmid. 2023 · 2023
Cited alongside, same era.
Larger language models do in-context learning differently
Jerry Wei, Jason Wei, Yi Tay, Dustin Tran, Albert Webson, Yifeng Lu, Xinyun Chen, Hanxiao Liu, Da Huang, Denny Zhou, et al. 2023 · 2023
Cited alongside, same era.
Fundamental limitations of alignment in large language models
Yotam Wolf, Noam Wies, Oshri Avnery, Yoav Levine, and Amnon Shashua. 2023 · 2023
Cited alongside, same era.
Value FULCRA: Mapping large language models to the multidimensional spectrum of basic human values
What are human values, and how do we align AI to them?
Oliver Klingefjord, Ryan Lowe, and Joe Edelman. 2024 · 2024
Closest in time.
Propulsion: Steering LLM with tiny fine-tuning
Md Kowsher, Nusrat Jahan Prottasha, and Prakash Bhat. 2024 · 2024
Closest in time.
Programming refusal with conditional activation steering
Bruce W. Lee, Inkit Padhi, Karthikeyan Natesan Ramamurthy, Erik Miehling, Pierre Dognin, Manish Nagireddy, and Amit Dhurandhar. 2024 · 2024
Closest in time.
Guiding large language models via directional stimulus prompting
Zekun Li, Baolin Peng, Pengcheng He, Michel Galley, Jianfeng Gao, and Xifeng Yan. 2024 · 2024
Closest in time.
Evaluating large language model biases in persona-steered generation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jing Yao, Xiaoyuan Yi, Xiting Wang, Yifan Gong, and Xing Xie. 2023 · 2023
Cited alongside, same era.
Many-shot jailbreaking
Cem Anil, Esin Durmus, Mrinank Sharma, Joe Benton, Sandipan Kundu, Joshua Batson, Nina Rimsky, Meg Tong, Jesse Mu, Daniel Ford, et al. 2024 · 2024
Cited alongside, same era.
PAD: Personalized alignment at decoding-time
Ruizhe Chen, Xiaotian Zhang, Meng Luo, Wenhao Chai, and Zuozhu Liu. 2024 · 2024
Cited alongside, same era.
Social choice for AI alignment: Dealing with diverse human feedback
Vincent Conitzer, Rachel Freedman, Jobst Heitzig, Wesley H Holliday, Bob M Jacobs, Nathan Lambert, Milan Mossé, Eric Pacuit, Stuart Russell, Hailey Schoelkopf, et al. 2024 · 2024
Cited alongside, same era.
Modular pluralism: Pluralistic alignment via multi-LLM collaboration
Shangbin Feng, Taylor Sorensen, Yuhan Liu, Jillian Fisher, Chan Young Park, Yejin Choi, and Yulia Tsvetkov. 2024 · 2024
Cited alongside, same era.
CharED: Character-wise ensemble decoding for large language models
Kevin Gu, Eva Tuecke, Dmitriy Katz, Raya Horesh, David Alvarez-Melis, and Mikhail Yurochkin. 2024 · 2024
Cited alongside, same era.
Word embeddings are steers for language models
Chi Han, Jialiang Xu, Manling Li, Yi Fung, Chenkai Sun, Nan Jiang, Tarek Abdelzaher, and Heng Ji. 2024 · 2024
Cited alongside, same era.
CoS: Enhancing personalization and mitigating bias with context steering
Jerry Zhi-Yang He, Sashrika Pandey, Mariah L Schrum, and Anca Dragan. 2024 · 2024
Cited alongside, same era.
Andy Liu, Mona Diab, and Daniel Fried. 2024 · 2024
Closest in time.
Language models in dialogue: Conversational maxims for human-AI interactions
Erik Miehling, Manish Nagireddy, Prasanna Sattigeri, Elizabeth M Daly, David Piorkowski, and John T Richards. 2024 · 2024
Closest in time.
PersonaGym: Evaluating persona agents and LLMs
Vinay Samuel, Henry Peng Zou, Yue Zhou, Shreyas Chaudhari, Ashwin Kalyan, Tanmay Rajpurohit, Ameet Deshpande, Karthik Narasimhan, and Vishvak Murahari. 2024 · 2024
Closest in time.
The transient nature of emergent in-context learning in transformers
Aaditya Singh, Stephanie Chan, Ted Moskovitz, Erin Grant, Andrew Saxe, and Felix Hill. 2024 · 2024
Closest in time.
Steering without side effects: Improving post-deployment control of language models
Asa Cooper Stickland, Alexander Lyzhov, Jacob Pfau, Salsabila Mahdi, and Samuel R Bowman. 2024 · 2024
Closest in time.
Exploring and steering the moral compass of large language models
Alejandro Tlaie. 2024 · 2024
Closest in time.
Xinpeng Wang, Bolei Ma, Chengzhi Hu, Leon Weber-Genzel, Paul Röttger, Frauke Kreuter, Dirk Hovy, and Barbara Plank. 2024 · 2024
Closest in time.
The learnability of in-context learning
Noam Wies, Yoav Levine, and Amnon Shashua. 2024 · 2024
Closest in time.
Aligning LLMs with individual preferences via interaction
Shujin Wu, May Fung, Cheng Qian, Jeonghwan Kim, Dilek Hakkani-Tur, and Heng Ji. 2024 · 2024
Closest in time.