Fetching the paper…
Reading the bibliography…
Much recent work seeks to evaluate values and opinions in large language models (LLMs) using multiple-choice surveys and questionnaires.
Question wording as an independent variable in survey analysis
Howard Schuman and Stanley Presser. 1977 · 1977
Earlier work this paper cites.
The effect of the question on survey responses: A review
Graham Kalton and Howard Schuman. 1982 · 1982
Earlier work this paper cites.
Framing theory
Dennis Chong and James N Druckman. 2007 · 2007
Earlier work this paper cites.
Studying framing effects on political preferences
Ethan Busby, D Flynn, James N Druckman, and P D’Angelo. 2018 · 2018
Earlier work this paper cites.
Aligning ai with shared human values
Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt. 2020 · 2020
Earlier work this paper cites.
Measuring and improving consistency in pretrained language models
Yanai Elazar, Nora Kassner, Shauli Ravfogel, Abhilasha Ravichander, Eduard Hovy, Hinrich Schütze, and Yoav Goldberg. 2021 · 2021
Earlier work this paper cites.
Adversarial glue: A multi-task benchmark for robustness evaluation of language models
Boxin Wang, Chejian Xu, Shuohang Wang, Zhe Gan, Yu Cheng, Jianfeng Gao, Ahmed Hassan Awadallah, and Bo Li. 2021 · 2021
Earlier work this paper cites.
Simulators. LessWrong online forum, 2nd September
Janus. 2022 · 2022
Earlier work this paper cites.
Who is GPT-3? an exploration of personality, values and demographics
Marilù Miotto, Nicola Rossberg, and Bennett Kleinberg. 2022 · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Gray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. 2022 · 2022
Earlier work this paper cites.
Mirages. on anthropomorphism in dialogue systems
Gavin Abercrombie, Amanda Cercas Curry, Tanvi Dinkar, Verena Rieser, and Zeerak Talat. 2023 · 2023
Earlier work this paper cites.
Simple Generation
Giuseppe Attanasio. 2023 · 2023
Earlier work this paper cites.
Using cognitive psychology to understand gpt-3
Marcel Binz and Eric Schulz. 2023 · 2023
Earlier work this paper cites.
Towards measuring the representation of subjective global opinions in language models
Esin Durmus, Karina Nyugen, Thomas I Liao, Nicholas Schiefer, Amanda Askell, Anton Bakhtin, Carol Chen, Zac Hatfield-Dodds, Danny Hernandez, Nicholas Joseph, et al. 2023 · 2023
Earlier work this paper cites.
Multilingual coarse political stance classification of media. the editorial line of a ChatGPT and bard newspaper
Cristina España-Bonet. 2023 · 2023
Cited alongside, same era.
From pretraining data to language models to downstream tasks: Tracking the trails of political biases leading to unfair NLP models
Shangbin Feng, Chan Young Park, Yuhan Liu, and Yulia Tsvetkov. 2023 · 2023
Cited alongside, same era.
Revisiting the political biases of chatgpt
Sasuke Fujimoto and Takemoto Kazuhiro. 2023 · 2023
Cited alongside, same era.
Ai in the gray: Exploring moderation policies in dialogic large language models vs. human answers in controversial topics
Vahid Ghafouri, Vibhor Agarwal, Yong Zhang, Nishanth Sastry, Jose Such, and Guillermo Suarez-Tangil. 2023 · 2023
Cited alongside, same era.
Jochen Hartmann, Jasper Schwenzow, and Maximilian Witte. 2023 · 2023
Evaluating the moral beliefs encoded in llms
Nino Scherrer, Claudia Shi, Amir Feder, and David Blei. 2023 · 2023
Later among the works it cites.
Role play with large language models
Murray Shanahan, Kyle McDonell, and Laria Reynolds. 2023 · 2023
Later among the works it cites.
Bangzhao Shu, Lechen Zhang, Minje Choi, Lavinia Dunagan, Dallas Card, and David Jurgens. 2023 · 2023
Later among the works it cites.
Assessing political inclination of Bangla language models
Surendrabikram Thapa, Ashwarya Maratha, Khan Md Hasib, Mehwish Nasim, and Usman Naseem. 2023 · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. 2023 · 2023
Cited alongside, same era.
The past, present and better future of feedback learning in large language models for subjective human preferences and values
Hannah Kirk, Andrew Bean, Bertie Vidgen, Paul Rottger, and Scott Hale. 2023 · 2023
Cited alongside, same era.
R Thomas McCoy, Shunyu Yao, Dan Friedman, Matthew Hardy, and Thomas L Griffiths. 2023 · 2023
Cited alongside, same era.
More human than human: Measuring chatgpt political bias
Fabio Motoki, Valdemar Pinho Neto, and Victor Rodrigues. 2023 · 2023
Cited alongside, same era.
Does chatgpt have a liberal bias?
Arvind Narayanan and Sayash Kapoor. 2023 · 2023
Cited alongside, same era.
The shifted and the overlooked: A task-oriented investigation of user-GPT interactions
Siru Ouyang, Shuohang Wang, Yang Liu, Ming Zhong, Yizhu Jiao, Dan Iter, Reid Pryzant, Chenguang Zhu, Heng Ji, and Jiawei Han. 2023 · 2023
Cited alongside, same era.
Xstest: A test suite for identifying exaggerated safety behaviours in large language models
Paul Röttger, Hannah Rose Kirk, Bertie Vidgen, Giuseppe Attanasio, Federico Bianchi, and Dirk Hovy. 2023 · 2023
Cited alongside, same era.
Later among the works it cites.
Zephyr: Direct distillation of lm alignment
Lewis Tunstall, Edward Beeching, Nathan Lambert, Nazneen Rajani, Kashif Rasul, Younes Belkada, Shengyi Huang, Leandro von Werra, Clémentine Fourrier, Nathan Habib, et al. 2023 · 2023
Later among the works it cites.
Chatgpt’s left-leaning liberal bias
Merel van den Broek. 2023 · 2023
Later among the works it cites.
On the robustness of chatgpt: An adversarial and out-of-distribution perspective
Jindong Wang, HU Xixu, Wenxin Hou, Hao Chen, Runkai Zheng, Yidong Wang, Linyi Yang, Wei Ye, Haojun Huang, Xiubo Geng, et al. 2023 · 2023
Later among the works it cites.
Jailbroken: How does llm safety training fail?
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt. 2023 · 2023
Later among the works it cites.
Cvalues: Measuring the values of chinese large language models from safety to responsibility
Guohai Xu, Jiayi Liu, Ming Yan, Haotian Xu, Jinghui Si, Zhuoran Zhou, Peng Yi, Xing Gao, Jitao Sang, Rong Zhang, et al. 2023 · 2023
Later among the works it cites.
The political preferences of llms
David Rozado. 2024 · 2024
Closest in time.
Quantifying language models’ sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt formatting
Melanie Sclar, Yejin Choi, Yulia Tsvetkov, and Alane Suhr. 2024 · 2024
Closest in time.
(inthe)wildchat: 570k chatGPT interaction logs in the wild
Wenting Zhao, Xiang Ren, Jack Hessel, Claire Cardie, Yejin Choi, and Yuntian Deng. 2024 · 2024
Closest in time.