Fetching the paper…
Reading the bibliography…
The rapid advancement of Large Language Models (LLMs) has attracted much attention to value alignment for their responsible development.
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer. 2019 · 1910
Earlier work this paper cites.
Values, expectations, and the prediction of social action: An expectancy-valence analysis
Norman T Feather and James W Newton. 1982 · 1982
Earlier work this paper cites.
Measuring and predicting managerial success
Gabriel Jaskolka, Janice M Beyer, and Harrison M Trice. 1985 · 1985
Earlier work this paper cites.
Universals in the content and structure of values: Theoretical advances and empirical tests in 20 countries
Shalom H Schwartz. 1992 · 1992
Earlier work this paper cites.
Extending the cross-cultural validity of the theory of basic human values with a different method of measurement
Shalom H Schwartz, Gila Melech, Arielle Lehmann, Steven Burgess, Mari Harris, and Vicki Owens. 2001 · 2001
Earlier work this paper cites.
Common morality: Deciding what to do
Bernard Gert. 2004 · 2004
Earlier work this paper cites.
Deberta: Decoding-enhanced bert with disentangled attention
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2020 · 2006
Earlier work this paper cites.
Random k-labelsets: An ensemble method for multilabel classification
Grigorios Tsoumakas and Ioannis Vlahavas. 2007 · 2007
Earlier work this paper cites.
Aligning ai with shared human values
Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt. 2020 · 2008
Earlier work this paper cites.
Visualizing data using t-sne
Laurens Van der Maaten and Geoffrey Hinton. 2008 · 2008
Earlier work this paper cites.
Random k-labelsets for multilabel classification
Grigorios Tsoumakas, Ioannis Katakis, and Ioannis Vlahavas. 2010 · 2010
Earlier work this paper cites.
Dimensionalizing cultures: The hofstede model in context
Geert Hofstede. 2011 · 2011
Earlier work this paper cites.
Measuring social value orientation
Ryan O Murphy, Kurt A Ackermann, and Michel JJ Handgraaf. 2011 · 2011
Earlier work this paper cites.
An overview of the schwartz theory of basic values
Shalom H Schwartz. 2012 · 2012
Earlier work this paper cites.
Moral foundations theory: The pragmatic validity of moral pluralism
Jesse Graham, Jonathan Haidt, Sena Koleva, Matt Motyl, Ravi Iyer, Sean P Wojcik, and Peter H Ditto. 2013 · 2013
Earlier work this paper cites.
Predicting ideological prejudice
Mark J Brandt. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Some moral and technical consequences of automation
Norbert Wiener. 2017 · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Earlier work this paper cites.
Artificial intelligence, values, and alignment
Iason Gabriel. 2020 · 2020
Earlier work this paper cites.
A general language assistant as a laboratory for alignment
Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain, Deep Ganguli, Tom Henighan, Andy Jones, Nicholas Joseph, Ben Mann, Nova DasSarma, et al. 2021 · 2021
Earlier work this paper cites.
On the opportunities and risks of foundation models
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al. 2021 · 2021
Cited alongside, same era.
Zachary Kenton, Tom Everitt, Laura Weidinger, Iason Gabriel, Vladimir Mikulik, and Geoffrey Irving. 2021 · 2021
Cited alongside, same era.
Webgpt: Browser-assisted question-answering with human feedback
Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, et al. 2021 · 2021
Cited alongside, same era.
The effects of reward misspecification: Mapping and mitigating misaligned models
Alexander Pan, Kush Bhatia, and Jacob Steinhardt. 2021 · 2021
Cited alongside, same era.
Critic: Large language models can self-correct with tool-interactive critiquing
Zhibin Gou, Zhihong Shao, Yeyun Gong, Yelong Shen, Yujiu Yang, Nan Duan, and Weizhu Chen. 2023 · 2023
Closest in time.
Evaluating the robustness of discrete prompts
Yoichi Ishibashi, Danushka Bollegala, Katsuhito Sudoh, and Satoshi Nakamura. 2023 · 2023
Closest in time.
Beavertails: Towards improved safety alignment of llm via a human-preference dataset
Jiaming Ji, Mickel Liu, Juntao Dai, Xuehai Pan, Chi Zhang, Ce Bian, Ruiyang Sun, Yizhou Wang, and Yaodong Yang. 2023 · 2023
Closest in time.
Better zero-shot reasoning with role-play prompting
Aobo Kong, Shiwan Zhao, Hao Chen, Qicheng Li, Yong Qin, Ruiqi Sun, and Xin Zhou. 2023 · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jeff Wu, Long Ouyang, Daniel M Ziegler, Nisan Stiennon, Ryan Lowe, Jan Leike, and Paul Christiano. 2021 · 2021
Cited alongside, same era.
Probing pre-trained language models for cross-cultural differences in values
Arnav Arora, Lucie-Aimée Kaffee, and Isabelle Augenstein. 2022 · 2022
Cited alongside, same era.
Measuring progress on scalable oversight for large language models
Samuel R Bowman, Jeeyoon Hyun, Ethan Perez, Edwin Chen, Craig Pettit, Scott Heiner, Kamile Lukosuite, Amanda Askell, Andy Jones, Anna Chen, et al. 2022 · 2022
Cited alongside, same era.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
Deep Ganguli, Liane Lovitt, Jackson Kernion, Amanda Askell, Yuntao Bai, Saurav Kadavath, Ben Mann, Ethan Perez, Nicholas Schiefer, Kamal Ndousse, et al. 2022 · 2022
Cited alongside, same era.
The ghost in the machine has an american accent: value conflict in gpt-3
Rebecca L Johnson, Giada Pistilli, Natalia Menédez-González, Leslye Denisse Dias Duran, Enrico Panai, Julija Kalpokiene, and Donald Jay Bertulfo. 2022 · 2022
Cited alongside, same era.
Safetext: A benchmark for exploring physical safety in language models
Sharon Levy, Emily Allaway, Melanie Subbiah, Lydia Chilton, Desmond Patton, Kathleen McKeown, and William Yang Wang. 2022 · 2022
Cited alongside, same era.
Who is gpt-3? an exploration of personality, values and demographics
Marilù Miotto, Nicola Rossberg, and Bennett Kleinberg. 2022 · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022 · 2022
Cited alongside, same era.
Hao Liu, Carmelo Sferrazza, and Pieter Abbeel. 2023 · 2023
Closest in time.
Inverse scaling: When bigger isn’t better
Ian R McKenzie, Alexander Lyzhov, Michael Pieler, Alicia Parrish, Aaron Mueller, Ameya Prabhu, Euan McLean, Aaron Kirtland, Alexis Ross, Alisa Liu, et al. 2023 · 2023
Closest in time.
OpenAI. 2023 · 2023
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D Manning, and Chelsea Finn. 2023 · 2023
Closest in time.
Towards coding social science datasets with language models
Christopher Michael Rytting, Taylor Sorensen, Lisa Argyle, Ethan Busby, Nancy Fulda, Joshua Gubler, and David Wingate. 2023 · 2023
Closest in time.
Transparent value alignment
Lindsay Sanneman and Julie Shah. 2023 · 2023
Closest in time.
Preference ranking optimization for human alignment
Feifan Song, Bowen Yu, Minghao Li, Haiyang Yu, Fei Huang, Yongbin Li, and Houfeng Wang. 2023 · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023 · 2023
Closest in time.
Decolonial ai alignment: Visesadharma, argument, and artistic expression
Kush R Varshney. 2023 · 2023
Closest in time.
Simple synthetic data reduces sycophancy in large language models
Jerry Wei, Da Huang, Yifeng Lu, Denny Zhou, and Quoc V Le. 2023 · 2023
Closest in time.
Cvalues: Measuring the values of chinese large language models from safety to responsibility
Guohai Xu, Jiayi Liu, Ming Yan, Haotian Xu, Jinghui Si, Zhuoran Zhou, Peng Yi, Xing Gao, Jitao Sang, Rong Zhang, et al. 2023 · 2023
Closest in time.
From instructions to intrinsic human values–a survey of alignment goals for big models
Jing Yao, Xiaoyuan Yi, Xiting Wang, Jindong Wang, and Xing Xie. 2023 · 2023
Closest in time.
Rrhf: Rank responses to align language models with human feedback without tears
Zheng Yuan, Hongyi Yuan, Chuanqi Tan, Wei Wang, Songfang Huang, and Fei Huang. 2023 · 2023
Closest in time.
Heterogeneous value evaluation for large language models
Zhaowei Zhang, Nian Liu, Siyuan Qi, Ceyao Zhang, Ziqi Rong, Yaodong Yang, and Shuguang Cui. 2023 · 2023
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. 2023 · 2023
Closest in time.
Red teaming chatgpt via jailbreaking: Bias, robustness, reliability and toxicity
Terry Yue Zhuo, Yujin Huang, Chunyang Chen, and Zhenchang Xing. 2023 · 2023
Closest in time.
Can large language models transform computational social science?
Caleb Ziems, William Held, Omar Shaikh, Jiaao Chen, Zhehao Zhang, and Diyi Yang. 2023 · 2023
Closest in time.