Fetching the paper…
Reading the bibliography…
Recent advances in Large Language Models (LLMs) have sparked wide interest in validating and comprehending the human-like cognitive-behavioral traits LLMs may capture and convey.
The functional approach to the study of attitudes
Daniel Katz. 1960 · 1960
Earlier work this paper cites.
Beliefs, Attitudes and Values: A Theory of Organization and Change
Milton Rokeach. 1968 · 1968
Earlier work this paper cites.
Development in Judging Moral Issues
James R. Rest. 1979 · 1979
Earlier work this paper cites.
Attitudes, Personality, and Behavior , 1st edition
Icek Ajzen. 1988 · 1988
Earlier work this paper cites.
A theoretical note on the differences between attitudes, opinions, and values
Manfred Max Bergman. 1998 · 1998
Earlier work this paper cites.
The Psychology of Survey Response
Roger Tourangeau, Lance J. Rips, and Kenneth Rasinski. 2000 · 2000
Earlier work this paper cites.
Survey Methodology
Robert M. Groves, Floyd J. Fowler, Mick P. Couper, James M. Lepkowski, Eleanor Singer, and Roger Tourangeau. 2004 · 2004
Earlier work this paper cites.
Culture’s recent consequences
Geert Hofstede. 2005 · 2005
Earlier work this paper cites.
Aligning ai with shared human values
Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt. 2023 · 2008
Earlier work this paper cites.
Mapping the moral domain
Jesse Graham, Brian A Nosek, Jonathan Haidt, Ravi Iyer, Spassena Koleva, and Peter H Ditto. 2011 · 2011
Earlier work this paper cites.
Training needs in survey research methods: An overview
Barbara O’Hare, Matt Jans, and Stanislav Kolenikov. 2015 · 2015
Earlier work this paper cites.
Behavior in models: A framework for representing human behavior
Andrew Greasley and Chris Owen. 2016 · 2016
Earlier work this paper cites.
Post-election cross section (gles 2017)
GLES. 2019 · 2017
Earlier work this paper cites.
Ethical considerations for data collection using surveys
Marilyn J. Hammer. 2017 · 2017
Earlier work this paper cites.
Large language models sensitivity to the order of options in multiple-choice questions
Pouya Pezeshkpour and Estevam Hruschka. 2024 · 2017
Earlier work this paper cites.
The survey response process from a cognitive viewpoint
Roger Tourangeau. 2018 · 2018
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Earlier work this paper cites.
Machine behaviour
Iyad Rahwan, Manuel Cebrian, Nick Obradovich, Josh Bongard, Jean-François Bonnefon, Cynthia Breazeal, Jacob W. Crandall, Nicholas A. Christakis, Iain D. Couzin, Matthew O. Jackson, Nicholas R. Jennings, Ece Kamar, Isabel M. Kloumann, Hugo Larochelle, David Lazer, Richard McElreath, Alan Mislove, David C. Parkes, Alex ‘Sandy’ Pentland, Margaret E. Roberts, Azim Shariff, Joshua B. Tenenbaum, and Michael Wellman. 2019 · 2019
Earlier work this paper cites.
"I Call Alexa to the Stand": The Privacy Implications of Anthropomorphizing Virtual Assistants Accompanying Smart-Home Technology
Christopher B. Burkett. 2020 · 2020
Earlier work this paper cites.
Social chemistry 101: Learning to reason about social and moral norms
Maxwell Forbes, Jena D. Hwang, Vered Shwartz, Maarten Sap, and Yejin Choi. 2020 · 2020
Earlier work this paper cites.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mohammad Bavarian, Clemens Winter, Philippe Tillet, Felipe Petroski Such, Dave Cummings, Matthias Plappert, Fotios Chantzis, Elizabeth Barnes, Ariel Herbert-Voss, William Hebgen Guss, Alex Nichol, Alex Paino, Nikolas Tezak, Jie Tang, Igor Babuschkin, Suchir Balaji, Shantanu Jain, William Saunders, Christopher Hesse, Andrew N. Carr, Jan Leike, Josh Achiam, Vedant Misra, Evan Morikawa, Alec Radford, Matthew Knight, Miles Brundage, Mira Murati, Katie Mayer, Peter Welinder, Bob McGrew, Dario Amodei, Sam McCandlish, Ilya Sutskever, and Wojciech Zaremba. 2021 · 2021
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021 · 2021
Earlier work this paper cites.
Anthropomorphism, privacy and security concerns: preliminary work
Eloïse Zehnder, Jérôme Dinet, and François Charpillet. 2021 · 2021
Earlier work this paper cites.
Cultural value resonance in folktales: A transformer-based analysis with the world value corpus
Noam Benkler, Scott Friedman, Sonja Schmer-Galunder, Drisana Mosaphir, Vasanth Sarathy, Pavan Kantharaju, Matthew D. McLure, and Robert P. Goldman. 2022 · 2022
Earlier work this paper cites.
Does moral code have a moral code? probing delphi’s moral philosophy
Kathleen C. Fraser, Svetlana Kiritchenko, and Esma Balkir. 2022 · 2022
Earlier work this paper cites.
World Values Survey: Round Seven - Country-Pooled Datafile Version 5.0
C. Haerpfer, R. Inglehart, A. Moreno, C. Welzel, K. Kizilova, J. Diez-Medrano, M. Lagos, P. Norris, E. Ponarin, and B. Puranen, editors. 2022 · 2022
Earlier work this paper cites.
Challenges and strategies in cross-cultural NLP
Daniel Hershcovich, Stella Frank, Heather Lent, Miryam de Lhoneux, Mostafa Abdou, Stephanie Brandl, Emanuele Bugliarello, Laura Cabello Piqueras, Ilias Chalkidis, Ruixiang Cui, Constanza Fierro, Katerina Margatina, Phillip Rust, and Anders Søgaard. 2022 · 2022
Earlier work this paper cites.
LoRA: Low-rank adaptation of large language models
Edward J Hu, yelong shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022 · 2022
Earlier work this paper cites.
When to make exceptions: Exploring language models as accounts of human moral judgment
Zhijing Jin, Sydney Levine, Fernando Gonzalez Adauto, Ojasv Kamal, Maarten Sap, Mrinmaya Sachan, Rada Mihalcea, Josh Tenenbaum, and Bernhard Schölkopf. 2022 · 2022
Earlier work this paper cites.
The ghost in the machine has an american accent: value conflict in gpt-3
Rebecca L Johnson, Giada Pistilli, Natalia Menédez-González, Leslye Denisse Dias Duran, Enrico Panai, Julija Kalpokiene, and Donald Jay Bertulfo. 2022 · 2022
Earlier work this paper cites.
TruthfulQA: Measuring how models mimic human falsehoods
Stephanie Lin, Jacob Hilton, and Owain Evans. 2022 · 2022
Earlier work this paper cites.
Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity
Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp. 2022 · 2022
Earlier work this paper cites.
Who is GPT-3? an exploration of personality, values and demographics
Marilù Miotto, Nicola Rossberg, and Bennett Kleinberg. 2022 · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F Christiano, Jan Leike, and Ryan Lowe. 2022 · 2022
Earlier work this paper cites.
Neural theory-of-mind? on the limits of social intelligence in large LMs
Maarten Sap, Ronan Le Bras, Daniel Fried, and Yejin Choi. 2022 · 2022
Earlier work this paper cites.
On the machine learning of ethical judgments from natural language
Zeerak Talat, Hagen Blix, Josef Valvoda, Maya Indira Ganesh, Ryan Cotterell, and Adina Williams. 2022 · 2022
Earlier work this paper cites.
Finding skill neurons in pre-trained transformer-based language models
Xiaozhi Wang, Kaiyue Wen, Zhengyan Zhang, Lei Hou, Zhiyuan Liu, and Juanzi Li. 2022 · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed Chi, Quoc V Le, and Denny Zhou. 2022 · 2022
Earlier work this paper cites.
Taxonomy of risks posed by language models
Laura Weidinger, Jonathan Uesato, Maribeth Rauh, Conor Griffin, Po-Sen Huang, John Mellor, Amelia Glaese, Myra Cheng, Borja Balle, Atoosa Kasirzadeh, Courtney Biles, Sasha Brown, Zac Kenton, Will Hawkins, Tom Stepleton, Abeba Birhane, Lisa Anne Hendricks, Laura Rimell, William Isaac, Julia Haas, Sean Legassick, Geoffrey Irving, and Iason Gabriel. 2022 · 2022
Earlier work this paper cites.
Moral foundations of large language models
Marwa Abdulhai, Gregory Serapio-Garcia, Clément Crepy, Daria Valter, John Canny, and Natasha Jaques. 2023 · 2023
Earlier work this paper cites.
Out of one, many: Using language models to simulate human samples
Lisa P. Argyle, Ethan C. Busby, Nancy Fulda, Joshua R. Gubler, Christopher Rytting, and David Wingate. 2023 · 2023
Earlier work this paper cites.
Probing pre-trained language models for cross-cultural differences in values
Arnav Arora, Lucie-aimée Kaffee, and Isabelle Augenstein. 2023 · 2023
Earlier work this paper cites.
Assessing llms for moral value pluralism
Noam Benkler, Drisana Mosaphir, Scott Friedman, Andrew Smart, and Sonja Schmer-Galunder. 2023 · 2023
Earlier work this paper cites.
Using cognitive psychology to understand gpt-3
Marcel Binz and Eric Schulz. 2023 · 2023
Earlier work this paper cites.
Synthetic replacements for human survey data? the perils of large language models
James Bisbee, Joshua Clinton, Cassy Dorff, Brenton Kenkel, and Jennifer Larson. 2023 · 2023
Earlier work this paper cites.
Assessing cross-cultural alignment between ChatGPT and human societies: An empirical study
Yong Cao, Li Zhou, Seolhwa Lee, Laura Cabello, Min Chen, and Daniel Hershcovich. 2023 · 2023
Earlier work this paper cites.
Manipulating the perceived personality traits of language models
Graham Caron and Shashank Srivastava. 2023 · 2023
Earlier work this paper cites.
Open problems and fundamental limitations of reinforcement learning from human feedback
Stephen Casper, Xander Davies, Claudia Shi, Thomas Krendl Gilbert, Jérémy Scheurer, Javier Rando, Rachel Freedman, Tomasz Korbak, David Lindner, Pedro Freire, et al. 2023 · 2023
Earlier work this paper cites.
CoMPosT: Characterizing and evaluating caricature in LLM simulations
Myra Cheng, Tiziano Piccardi, and Diyi Yang. 2023b · 2023
Cited alongside, same era.
Toxicity in chatgpt: Analyzing persona-assigned language models
Ameet Deshpande, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, and Karthik Narasimhan. 2023 · 2023
Cited alongside, same era.
Questioning the survey responses of large language models
Ricardo Dominguez-Olmedo, Moritz Hardt, and Celestine Mendler-Dünner. 2023 · 2023
Cited alongside, same era.
Opinion paper: “so what if chatgpt wrote it?” multidisciplinary perspectives on opportunities, challenges and implications of generative conversational ai for research, practice and policy
Yogesh K. Dwivedi, Nir Kshetri, Laurie Hughes, Emma Louise Slade, Anand Jeyaraj, Arpan Kumar Kar, Abdullah M. Baabdullah, Alex Koohang, Vishnupriya Raghavan, Manju Ahuja, Hanaa Albanna, Mousa Ahmad Albashrawi, Adil S. Al-Busaidi, Janarthanan Balakrishnan, Yves Barlette, Sriparna Basu, Indranil Bose, Laurence Brooks, Dimitrios Buhalis, Lemuria Carter, Soumyadeb Chowdhury, Tom Crick, Scott W. Cunningham, Gareth H. Davies, Robert M. Davison, Rahul Dé, Denis Dennehy, Yanqing Duan, Rameshwar Dubey, Rohita Dwivedi, John S. Edwards, Carlos Flavián, Robin Gauld, Varun Grover, Mei Chih Hu, Marijn Janssen, Paul Jones, Iris Junglas, Sangeeta Khorana, Sascha Kraus, Kai R. Larsen, Paul Latreille, Sven Laumer, F. Tegwen Malik, Abbas Mardani, Marcello Mariani, Sunil Mithas, Emmanuel Mogaji, Jeretta Horn Nord, Siobhan O’Connor, Fevzi Okumus, Margherita Pagani, Neeraj Pandey, Savvas Papagiannidis, Ilias O. Pappas, Nishith Pathak, Jan Pries-Heje, Ramakrishnan Raman, Nripendra P. Rana, Sven Volker Rehm, Samuel Ribeiro-Navarrete, Alexander Richter, Frantz Rowe, Suprateek Sarker, Bernd Carsten Stahl, Manoj Kumar Tiwari, Wil van der Aalst, Viswanath Venkatesh, Giampaolo Viglia, Michael Wade, Paul Walton, Jochen Wirtz, and Ryan Wright. 2023 · 2023
Defending against unforeseen failure modes with latent adversarial training
Stephen Casper, Lennart Schulze, Oam Patel, and Dylan Hadfield-Menell. 2024 · 2024
Closest in time.
Tanise Ceron, Neele Falk, Ana Barić, Dmitry Nikolaev, and Sebastian Padó. 2024 · 2024
Closest in time.
Llama meets EU: Investigating the European political spectrum through the lens of LLMs
Ilias Chalkidis and Stephanie Brandl. 2024 · 2024
Closest in time.
Rlrf:reinforcement learning from reflection through debates as feedback for bias mitigation in llms
Ruoxi Cheng, Haoxuan Ma, Shuirong Cao, and Tianyu Shi. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
From pretraining data to language models to downstream tasks: Tracking the trails of political biases leading to unfair NLP models
Shangbin Feng, Chan Young Park, Yuhan Liu, and Yulia Tsvetkov. 2023 · 2023
Cited alongside, same era.
Moral foundations theory: On the advantages of moral pluralism over moral monism
Jesse Graham, Jonathan Haidt, Matt Motyl, Peter Meindl, Carol Iskiwitch, and Marlon Mooijman. 2018 · 2023
Cited alongside, same era.
Speaking multiple languages affects the moral bias of language models
Katharina Haemmerl, Bjoern Deiseroth, Patrick Schramowski, Jindřich Libovický, Constantin Rothkopf, Alexander Fraser, and Kristian Kersting. 2023 · 2023
Cited alongside, same era.
Evaluating large language models in generating synthetic hci research data: a case study
Perttu Hämäläinen, Mikke Tavast, and Anton Kunnari. 2023 · 2023
Cited alongside, same era.
Jochen Hartmann, Jasper Schwenzow, and Maximilian Witte. 2023 · 2023
Cited alongside, same era.
Aligning language models to user opinions
EunJeong Hwang, Bodhisattwa Majumder, and Niket Tandon. 2023 · 2023
Cited alongside, same era.
Co-writing with opinionated language models affects users’ views
Maurice Jakesch, Advait Bhat, Daniel Buschek, Lior Zalmanson, and Mor Naaman. 2023 · 2023
Cited alongside, same era.
Consistency analysis of ChatGPT
Myeongjun Jang and Thomas Lukasiewicz. 2023 · 2023
Cited alongside, same era.
Hyeong Kyu Choi and Yixuan Li. 2024 · 2024
Closest in time.
Social choice for ai alignment: Dealing with diverse human feedback
Vincent Conitzer, Rachel Freedman, Jobst Heitzig, Wesley H Holliday, Bob M Jacobs, Nathan Lambert, Milan Mossé, Eric Pacuit, Stuart Russell, Hailey Schoelkopf, et al. 2024 · 2024
Closest in time.
Towards measuring the representation of subjective global opinions in language models
Esin Durmus, Karina Nguyen, Thomas Liao, Nicholas Schiefer, Amanda Askell, Anton Bakhtin, Carol Chen, Zac Hatfield-Dodds, Danny Hernandez, Nicholas Joseph, Liane Lovitt, Sam McCandlish, Orowa Sikder, Alex Tamkin, Janel Thamkul, Jared Kaplan, Jack Clark, and Deep Ganguli. 2024 · 2024
Closest in time.
Position: Insights from survey methodology can improve training data
Stephanie Eckman, Barbara Plank, and Frauke Kreuter. 2024 · 2024
Closest in time.
Determinants of llm-assisted decision-making
Eva Eigner and Thorsten Händler. 2024 · 2024
Closest in time.
Are large language models chameleons?
Mingmeng Geng, Sihong He, and Roberto Trotta. 2024 · 2024
Closest in time.
Playwriting with large language models: Perceived features, interaction strategies and outcomes
Paolo Grigis and Antonella De Angeli. 2024 · 2024
Closest in time.
OpinionGPT: Modelling explicit biases in instruction-tuned LLMs
Patrick Haller, Ansar Aynetdinov, and Alan Akbik. 2024 · 2024
Closest in time.
Interpreting black-box models: A review on explainable artificial intelligence
Vikas Hassija, Vinay Chamola, Atmesh Mahapatra, Abhinandan Singal, Divyansh Goel, Kaizhu Huang, Simone Scardapane, Indro Spinelli, Mufti Mahmud, and Amir Hussain. 2024 · 2024
Closest in time.
Quantifying the persona effect in LLM simulations
Tiancheng Hu and Nigel Collier. 2024 · 2024
Closest in time.
Collective constitutional ai: Aligning a language model with public input
Saffron Huang, Divya Siddarth, Liane Lovitt, Thomas I. Liao, Esin Durmus, Alex Tamkin, and Deep Ganguli. 2024 · 2024
Closest in time.
PersonaLLM: Investigating the ability of large language models to express personality traits
Hang Jiang, Xiajie Zhang, Xubo Cao, Cynthia Breazeal, Deb Roy, and Jad Kabbara. 2024 · 2024
Closest in time.
Does cross-cultural alignment change the commonsense morality of language models?
Yuu Jinnai. 2024 · 2024
Closest in time.
Personas as a way to model truthfulness in language models
Nitish Joshi, Javier Rando, Abulhair Saparov, Najoung Kim, and He He. 2024 · 2024
Closest in time.
Ai-augmented surveys: Leveraging large language models and surveys for opinion prediction
Junsol Kim and Byungkyu Lee. 2024 · 2024
Closest in time.
What are human values, and how do we align ai to them?
Oliver Klingefjord, Ryan Lowe, and Joe Edelman. 2024 · 2024
Closest in time.
Evaluating large language models in theory of mind tasks
Michal Kosinski. 2024 · 2024
Closest in time.
Evaluating psychological safety of large language models
Xingxuan Li, Yutong Li, Lin Qiu, Shafiq Joty, and Lidong Bing. 2024 · 2024
Closest in time.
Culturally aware and adapted nlp: A taxonomy and a survey of the state of the art
Chen Cecilia Liu, Iryna Gurevych, and Anna Korhonen. 2024 · 2024
Closest in time.
Using llms to model the beliefs and preferences of targeted populations
Keiichi Namikoshi, Alex Filipowicz, David A. Shamma, Rumen Iliev, Candice L. Hogan, and Nikos Arechiga. 2024 · 2024
Closest in time.
Performance and biases of large language models in public opinion simulation
Yao Qu and Jue Wang. 2024 · 2024
Closest in time.
Political compass or spinning arrow? towards more meaningful evaluations for values and opinions in large language models
Paul Röttger, Valentin Hofmann, Valentina Pyatkin, Musashi Hinck, Hannah Kirk, Hinrich Schuetze, and Dirk Hovy. 2024 · 2024
Closest in time.
The political preferences of llms
David Rozado. 2024 · 2024
Closest in time.
Generative echo chamber? effect of llm-powered search systems on diverse information seeking
Nikhil Sharma, Q. Vera Liao, and Ziang Xiao. 2024 · 2024
Closest in time.
Hua Shen, Tiffany Knearem, Reshmi Ghosh, Kenan Alkiek, Kundan Krishna, Yachuan Liu, Ziqiao Ma, Savvas Petridis, Yi-Hao Peng, Li Qiwei, Sushrita Rakshit, Chenglei Si, Yutong Xie, Jeffrey P. Bigham, Frank Bentley, Joyce Chai, Zachary Lipton, Qiaozhu Mei, Rada Mihalcea, Michael Terry, Diyi Yang, Meredith Ringel Morris, Paul Resnick, and David Jurgens. 2024 · 2024
Closest in time.
You don’t need a personality test to know these models are unreliable: Assessing the reliability of large language models on psychometric instruments
Bangzhao Shu, Lechen Zhang, Minje Choi, Lavinia Dunagan, Lajanugen Logeswaran, Moontae Lee, Dallas Card, and David Jurgens. 2024 · 2024
Closest in time.
Large human language models: A need and the challenges
Nikita Soni, H. Schwartz, João Sedoc, and Niranjan Balasubramanian. 2024 · 2024
Closest in time.
Value kaleidoscope: Engaging ai with pluralistic human values, rights, and duties
Taylor Sorensen, Liwei Jiang, Jena D. Hwang, Sydney Levine, Valentina Pyatkin, Peter West, Nouha Dziri, Ximing Lu, Kavel Rao, Chandra Bhagavatula, Maarten Sap, John Tasioulas, and Yejin Choi. 2024 · 2024
Closest in time.
Testing theory of mind in large language models and humans
James W. A. Strachan, Dalila Albergo, Giulia Borghini, Oriana Pansardi, Eugenio Scaliti, Saurabh Gupta, Krati Saxena, Alessandro Rufo, Stefano Panzeri, Guido Manzi, Michael S. A. Graziano, and Cristina Becchio. 2024 · 2024
Closest in time.
Seungjong Sun, Eungu Lee, Dongyan Nan, Xiangying Zhao, Wonbyung Lee, Bernard J. Jansen, and Jang Hyun Kim. 2024 · 2024
Closest in time.
Gemma: Open models based on gemini research and technology
Gemma Team, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivière, Mihir Sanjay Kale, Juliette Love, Pouya Tafti, Léonard Hussenot, Pier Giuseppe Sessa, Aakanksha Chowdhery, Adam Roberts, Aditya Barua, Alex Botev, Alex Castro-Ros, Ambrose Slone, Amélie Héliou, Andrea Tacchetti, Anna Bulanova, Antonia Paterson, Beth Tsai, Bobak Shahriari, Charline Le Lan, Christopher A. Choquette-Choo, Clément Crepy, Daniel Cer, Daphne Ippolito, David Reid, Elena Buchatskaya, Eric Ni, Eric Noland, Geng Yan, George Tucker, George-Christian Muraru, Grigory Rozhdestvenskiy, Henryk Michalewski, Ian Tenney, Ivan Grishchenko, Jacob Austin, James Keeling, Jane Labanowski, Jean-Baptiste Lespiau, Jeff Stanway, Jenny Brennan, Jeremy Chen, Johan Ferret, Justin Chiu, Justin Mao-Jones, Katherine Lee, Kathy Yu, Katie Millican, Lars Lowe Sjoesund, Lisa Lee, Lucas Dixon, Machel Reid, Maciej Mikuła, Mateo Wirth, Michael Sharman, Nikolai Chinaev, Nithum Thain, Olivier Bachem, Oscar Chang, Oscar Wahltinez, Paige Bailey, Paul Michel, Petko Yotov, Rahma Chaabouni, Ramona Comanescu, Reena Jana, Rohan Anil, Ross McIlroy, Ruibo Liu, Ryan Mullins, Samuel L Smith, Sebastian Borgeaud, Sertan Girgin, Sholto Douglas, Shree Pandya, Siamak Shakeri, Soham De, Ted Klimenko, Tom Hennigan, Vlad Feinberg, Wojciech Stokowiec, Yu hui Chen, Zafarali Ahmed, Zhitao Gong, Tris Warkentin, Ludovic Peran, Minh Giang, Clément Farabet, Oriol Vinyals, Jeff Dean, Koray Kavukcuoglu, Demis Hassabis, Zoubin Ghahramani, Douglas Eck, Joelle Barral, Fernando Pereira, Eli Collins, Armand Joulin, Noah Fiedel, Evan Senter, Alek Andreev, and Kathleen Kenealy. 2024 · 2024
Closest in time.
Do LLMs Exhibit Human-like Response Biases? A Case Study in Survey Design
Lindia Tjuatja, Valerie Chen, Tongshuang Wu, Ameet Talwalkwar, and Graham Neubig. 2024 · 2024
Closest in time.
Neurons in large language models: Dead, n-gram, positional
Elena Voita, Javier Ferrando, and Christoforos Nalmpantis. 2024 · 2024
Closest in time.
“my answer is C”: First-token probabilities do not match text answers in instruction-tuned language models
Xinpeng Wang, Bolei Ma, Chengzhi Hu, Leon Weber-Genzel, Paul Röttger, Frauke Kreuter, Dirk Hovy, and Barbara Plank. 2024b · 2024
Closest in time.
Unveiling selection biases: Exploring order and token sensitivity in large language models
Sheng-Lun Wei, Cheng-Kuang Wu, Hen-Hsen Huang, and Hsin-Hsi Chen. 2024 · 2024
Closest in time.
"this chatbot would never…": Perceived moral agency of mental health chatbots
Joel Wester, Henning Pohl, Simo Hosio, and Niels van Berkel. 2024 · 2024
Closest in time.
Revealing fine-grained values and opinions in large language models
Dustin Wright, Arnav Arora, Nadav Borenstein, Srishti Yadav, Serge Belongie, and Isabelle Augenstein. 2024 · 2024
Closest in time.
The call for socially aware language technologies
Diyi Yang, Dirk Hovy, David Jurgens, and Barbara Plank. 2024 · 2024
Closest in time.
CMoralEval: A moral evaluation benchmark for Chinese large language models
Linhao Yu, Yongqi Leng, Yufei Huang, Shang Wu, Haixin Liu, Xinmeng Ji, Jiahui Zhao, Jinwang Song, Tingting Cui, Xiaoqing Cheng, Liutao Liutao, and Deyi Xiong. 2024 · 2024
Closest in time.
Exploring collaboration mechanisms for LLM agents: A social psychology view
Jintian Zhang, Xin Xu, Ningyu Zhang, Ruibo Liu, Bryan Hooi, and Shumin Deng. 2024 · 2024
Closest in time.
Large language models are not robust multiple choice selectors
Chujie Zheng, Hao Zhou, Fandong Meng, Jie Zhou, and Minlie Huang. 2024 · 2024
Closest in time.
Can large language models transform computational social science?
Caleb Ziems, William Held, Omar Shaikh, Jiaao Chen, Zhehao Zhang, and Diyi Yang. 2024 · 2024
Closest in time.
Survey of russian elites, moscow, russia, 1993-2020
William Zimmerman, Sharon Werning Rivera, and Kirill Kalinin. 2023 · 2024
Closest in time.
ValueBench: Towards comprehensively evaluating value orientations and understanding of large language models
Yuanyi Ren, Haoran Ye, Hanjun Fang, Xin Zhang, and Guojie Song. 2024 · 2040
Closest in time.