Fetching the paper…
Reading the bibliography…
LLMs are known to exhibit sycophancy: agreeing with and flattering users, even at the cost of correctness.
Sentence-bert: Sentence embeddings using siamese bert-networks
Nils Reimers and Iryna Gurevych · 1908
Earlier work this paper cites.
On face-work: An analysis of ritual elements in social interaction
Erving Goffman · 1955
Earlier work this paper cites.
Politeness: Some universals in language usage
Penelope Brown and Stephen C Levinson · 1987
Earlier work this paper cites.
Cross-cultural pragmatics: The semantics of human interaction, 1991
Eric Pederson · 1991
Earlier work this paper cites.
Culture, face maintenance, and styles of handling interpersonal conflict: A study in five cultures
Stella Ting-Toomey, Ge Gao, Paula Trubisky, Zhizhong Yang, Hak Soo Kim, Sung-Ling Lin, and Tsukasa Nishida · 1991
Earlier work this paper cites.
Varieties of confirmation bias
Joshua Klayman · 1995
Earlier work this paper cites.
Face and interaction
Michael Haugh and Francesca Bargiela-Chiappini · 2009
Earlier work this paper cites.
Framing and face: The relevance of the presentation of self to linguistic discourse analysis
Deborah Tannen · 2009
Earlier work this paper cites.
The python reddit api wrapper
Bryce Boe · 2016
Earlier work this paper cites.
ConvoKit: A toolkit for the analysis of conversations
Jonathan P. Chang, Caleb Chiam, Liye Fu, Andrew Wang, Justine Zhang, and Cristian Danescu-Niculescu-Mizil · 2020
Earlier work this paper cites.
spaCy: Industrial-strength Natural Language Processing in Python
Matthew Honnibal, Ines Montani, Sofie Van Landeghem, and Adriane Boyd · 2020
Earlier work this paper cites.
Can computer based human-likeness endanger humanness?”–a philosophical and ethical perspective on digital assistants expressing feelings they can’t have
Jaana Porra, Mary Lacity, and Michael S Parks · 2020
Earlier work this paper cites.
Why AI alignment could be hard with modern deep learning
Ajeya Cotra · 2021
Earlier work this paper cites.
‘Am I the Bad One’? predicting the moral judgement of the crowd using pre–trained language models
Areej Alhassan, Jinkai Zhang, and Viktor Schlegel · 2022
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al · 2022
Earlier work this paper cites.
Bertopic: Neural topic modeling with a class-based tf-idf procedure
Maarten Grootendorst · 2022
Earlier work this paper cites.
” because AI is 100% right and safe”: User attitudes and sources of AI authority in india
Shivani Kapania, Oliver Siy, Gabe Clapper, Azhagu Meena Sp, and Nithya Sambasivan · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Earlier work this paper cites.
Harms from increasingly agentic algorithmic systems
Alan Chan, Rebecca Salganik, Alva Markelius, Chris Pang, Nitarshan Rajkumar, Dmitrii Krasheninnikov, Lauro Langosco, Zhonghao He, Yawen Duan, Micah Carroll, et al · 2023
Earlier work this paper cites.
Computer says “no”: The case against empathetic conversational AI
Alba Curry and Amanda Cercas Curry · 2023
Earlier work this paper cites.
Alpacafarm: A simulation framework for methods that learn from human feedback
Yann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy S Liang, and Tatsunori B Hashimoto · 2023
Earlier work this paper cites.
Chatgpt outperforms crowd workers for text-annotation tasks
Fabrizio Gilardi, Meysam Alizadeh, and Maël Kubli · 2023
Earlier work this paper cites.
Chatgpt’s advice is perceived as better than that of professional advice columnists
Piers Douglas Lionel Howe, Nicolas Fay, Morgan Saletta, and Eduard Hovy · 2023
Earlier work this paper cites.
Discovering language model behaviors with model-written evaluations
Ethan Perez, Sam Ringer, Kamile Lukosiute, Karina Nguyen, Edwin Chen, Scott Heiner, Craig Pettit, Catherine Olsson, Sandipan Kundu, Saurav Kadavath, Andy Jones, Anna Chen, Benjamin Mann, Brian Israel, Bryan Seethor, Cameron McKinnon, Christopher Olah, Da Yan, Daniela Amodei, Dario Amodei, Dawn Drain, Dustin Li, Eli Tran-Johnson, Guro Khundadze, Jackson Kernion, James Landis, Jamie Kerr, Jared Mueller, Jeeyoon Hyun, Joshua Landau, Kamal Ndousse, Landon Goldberg, Liane Lovitt, Martin Lucas, Michael Sellitto, Miranda Zhang, Neerav Kingsland, Nelson Elhage, Nicholas Joseph, Noemi Mercado, Nova DasSarma, Oliver Rausch, Robin Larson, Sam McCandlish, Scott Johnston, Shauna Kravec, Sheer El Showk, Tamera Lanham, Timothy Telleen-Lawton, Tom Brown, Tom Henighan, Tristan Hume, Yuntao Bai, Zac Hatfield-Dodds, Jack Clark, Samuel R. Bowman, Amanda Askell, Roger Grosse, Danny Hernandez, Deep Ganguli, Evan Hubinger, Nicholas Schiefer, and Jared Kaplan · 2023
Earlier work this paper cites.
Question decomposition improves the faithfulness of model-generated reasoning
Ansh Radhakrishnan, Karina Nguyen, Anna Chen, Carol Chen, Carson Denison, Danny Hernandez, Esin Durmus, Evan Hubinger, Jackson Kernion, Kamile Lukosiute, Newton Cheng, Nicholas Joseph, Nicholas Schiefer, Oliver Rausch, Sam McCandlish, Sheer El Showk, Tamera Lanham, Tim Maxwell, Venkatesa Chandrasekaran, Zac Hatfield-Dodds, Jared Kaplan, Jan Brauner, Samuel R. Bowman, and Ethan Perez · 2023
Earlier work this paper cites.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn · 2023
Earlier work this paper cites.
Simple synthetic data reduces sycophancy in large language models
Jerry Wei, Da Huang, Yifeng Lu, Denny Zhou, and Quoc V Le · 2023
Earlier work this paper cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al · 2023
Cited alongside, same era.
From yes-men to truth-tellers: addressing sycophancy in large language models with pinpoint tuning
Wei Chen, Zhen Huang, Liang Xie, Binbin Lin, Houqiang Li, Le Lu, Xinmei Tian, Deng Cai, Yonggang Zhang, Wenxiao Wan, Xu Shen, and Jieping Ye · 2024
Cited alongside, same era.
AnthroScore: A computational linguistic measure of anthropomorphism
Myra Cheng, Kristina Gligoric, Tiziano Piccardi, and Dan Jurafsky · 2024
Cited alongside, same era.
The illusion of empathy? notes on displays of emotion in human-computer interaction
Andrea Cuadra, Maria Wang, Lynn Andrea Stein, Malte F. Jung, Nicola Dell, Deborah Estrin, and James A. Landay · 2024
Cited alongside, same era.
Ultrafeedback: Boosting language models with scaled AI feedback, 2024
Ganqu Cui, Lifan Yuan, Ning Ding, Guanming Yao, Bingxiang He, Wei Zhu, Yuan Ni, Guotong Xie, Ruobing Xie, Yankai Lin, Zhiyuan Liu, and Maosong Sun · 2024
Explicitly unbiased large language models still form biased associations
Xuechunzi Bai, Angelina Wang, Ilia Sucholutsky, and Thomas L Griffiths · 2025
Closest in time.
From lived experience to insight: Unpacking the psychological risks of using ai conversational agents
Mohit Chandra, Suchismita Naik, Denae Ford, Ebele Okoli, Munmun De Choudhury, Mahsa Ershadi, Gonzalo Ramos, Javier Hernandez, Ananya Bhattacharjee, Shahed Warreth, et al · 2025
Closest in time.
How people use ChatGPT
Aaron Chatterji, Thomas Cunningham, David J. Deming, Zoe Hitzig, Christopher Ong, Carl Yan Shan, and Kevin Wadman · 2025
Closest in time.
Syceval: Evaluating LLM sycophancy
Aaron Fanous, Jacob Goldberg, Ank A Agarwal, Joanna Lin, Anson Zhou, Roxana Daneshjou, and Sanmi Koyejo · 2025
Closest in time.
Gemini 1.5 flash
Google DeepMind · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, et al · 2024
Cited alongside, same era.
Chatgpt giving relationship advice–how reliable is it?
Haonan Hou, Kevin Leach, and Yu Huang · 2024
Cited alongside, same era.
Qwen2. 5-coder technical report
Binyuan Hui, Jian Yang, Zeyu Cui, Jiaxi Yang, Dayiheng Liu, Lei Zhang, Tianyu Liu, Jiajun Zhang, Bowen Yu, Keming Lu, et al · 2024
Cited alongside, same era.
Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, et al · 2024
Cited alongside, same era.
Mitigating sycophancy in large language models via direct preference optimization
Azal Ahmad Khan, Sayan Alam, Xinran Wang, Ahmad Faraz Khan, Debanga Raj Neog, and Ali Anwar · 2024
Cited alongside, same era.
Hannah Rose Kirk, Alexander Whitefield, Paul Röttger, Andrew Bean, Katerina Margatina, Juan Ciro, Rafael Mosquera, Max Bartolo, Adina Williams, He He, et al · 2024
Cited alongside, same era.
Advice from humans and artificial intelligence: Can we distinguish them, and is one better than the other?
Otto JB Kuosmanen · 2024
Cited alongside, same era.
Jiseung Hong, Grace Byun, Seungone Kim, and Kai Shu · 2025
Closest in time.
Thinking beyond the anthropomorphic paradigm benefits LLM research
Lujain Ibrahim and Myra Cheng · 2025
Closest in time.
AdvisorQA: Towards helpful and harmless advice-seeking question answering with collective intelligence
Minbeom Kim, Hwanhee Lee, Joonsuk Park, Hwaran Lee, and Kyomin Jung · 2025
Closest in time.
Darkbench: Benchmarking dark patterns in large language models
Esben Kran, Hieu Minh Nguyen, Akash Kundu, Sami Jawhar, Jinsuk Park, and Mateusz Maria Jurewicz · 2025
Closest in time.
RLHS: Mitigating misalignment in rlhf with hindsight simulation
Kaiqu Liang, Haimin Hu, Ryan Liu, Thomas L Griffiths, and Jaime Fernández Fisac · 2025
Closest in time.
Sycophancy in large language models: Causes and mitigations
Lars Malmqvist · 2025
Closest in time.
Meta llama-3-70b-instruct-turbo
Meta · 2025
Closest in time.
Mistral-7b-instruct-v0.3
Mistral · 2025
Closest in time.
Mistral-small-24b-instruct-2501
Mistral · 2025
Closest in time.
AITA for making this? A public dataset of Reddit posts about moral dilemmas — datachain.ai
Elle O’Brien · 2025
Closest in time.
Expanding on what we missed with sycophancy, May 2025
OpenAI · 2025
Closest in time.
NormAd: A framework for measuring the cultural adaptability of large language models
Abhinav Sukumar Rao, Akhila Yerukola, Vishwa Shah, Katharina Reinecke, and Maarten Sap · 2025
Closest in time.
Giuseppe Russo, Debora Nozza, Paul Röttger, and Dirk Hovy · 2025
Closest in time.
Normative evaluation of large language models with everyday moral dilemmas
Pratik Sachdeva and Tom van Nuenen · 2025
Closest in time.
Navigating rifts in human-LLM grounding: Study and benchmark
Omar Shaikh, Hussein Mozannar, Gagan Bansal, Adam Fourney, and Eric Horvitz · 2025
Closest in time.
AI-LieDar : Examine the trade-off between utility and truthfulness in LLM agents
Zhe Su, Xuhui Zhou, Sanketh Rangreji, Anubha Kabra, Julia Mendelsohn, Faeze Brahman, and Maarten Sap · 2025
Closest in time.
Aligned but blind: Alignment increases implicit bias by reducing awareness of race
Lihao Sun, Chengzhi Mao, Valentin Hofmann, and Xuechunzi Bai · 2025
Closest in time.
When truth is overridden: Uncovering the internal origins of sycophancy in large language models
Keyu Wang, Jin Li, Shu Yang, Zhuoran Zhang, and Di Wang · 2025
Closest in time.
How People Are Really Using Gen AI in 2025 — hbr.org
Marc Zao-Sanders · 2025
Closest in time.
Beyond preferences in ai alignment: T. zhi-xuan et al
Tan Zhi-Xuan, Micah Carroll, Matija Franklin, and Hal Ashton · 2025
Closest in time.
Culture is not trivia: Sociocultural theory for cultural NLP
Naitian Zhou, David Bamman, and Isaac L. Bleaman · 2025
Closest in time.