Fetching the paper…
Reading the bibliography…
Sycophancy refers to the tendency of a large language model to align its outputs with the user's perceived preferences, beliefs, or opinions, in order to look favorable, regardless of whether those statements are factually correct.
Humans and automation: Use, misuse, disuse, abuse
Raja Parasuraman and Victor Riley · 1997
Earlier work this paper cites.
Complacency and automation bias in the use of imperfect automation
Christopher D Wickens, Benjamin A Clegg, Alex Z Vieane, and Angelia L Sebok · 2015
Earlier work this paper cites.
Concrete problems in ai safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Earlier work this paper cites.
Why ai alignment could be hard with modern deep learning
Aleya Cotra · 2021
Earlier work this paper cites.
Discovering language model behaviors with model-written evaluations
Ethan Perez, Sam Ringer, Kamilė Lukošiūtė, Karina Nguyen, Edwin Chen, Scott Heiner, Craig Pettit, Catherine Olsson, Sandipan Kundu, Saurav Kadavath, et al · 2022
Earlier work this paper cites.
A diachronic perspective on user trust in AI under uncertainty
Shehzaad Dhuliawala, Vilém Zouhar, Mennatallah El-Assady, and Mrinmaya Sachan · 2023
Earlier work this paper cites.
Measures for explainable ai: Explanation goodness, user satisfaction, mental models, curiosity, trust, and human-ai performance
Robert R Hoffman, Shane T Mueller, Gary Klein, and Jordan Litman · 2023
Earlier work this paper cites.
A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions
Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, et al · 2023
Cited alongside, same era.
Reducing sycophancy and improving honesty via activation steering
Nina Panickssery · 2023
Cited alongside, same era.
Towards understanding sycophancy in language models
Mrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud, Amanda Askell, Samuel R Bowman, Newton Cheng, Esin Durmus, Zac Hatfield-Dodds, Scott R Johnston, et al · 2023
Cited alongside, same era.
Simple synthetic data reduces sycophancy in large language models
Jerry Wei, Da Huang, Yifeng Lu, Denny Zhou, and Quoc V Le · 2023
Cited alongside, same era.
Germanpartiesqa: Benchmarking commercial large language models for political bias and sycophancy
Users do not trust recommendations from a large language model more than ai-sourced snippets
Melanie J McGrath, Patrick S Cooper, and Andreas Duenser · 2024
Closest in time.
Evaluating the impact of hallucinations on user trust and satisfaction in llm-based systems, 2024
Richard Oelschlager · 2024
Closest in time.
Take it, leave it, or fix it: Measuring productivity and trust in human-ai collaboration
Crystal Qian and James Wexler · 2024
Closest in time.
Aswin RRV, Nemika Tyagi, Md Nayem Uddin, Neeraj Varshney, and Chitta Baral · 2024
Closest in time.
To trust or distrust trust measures: Validating questionnaires for trust in ai
Nicolas Scharowski, Sebastian AC Perrig, Lena Fanya Aeschbach, Nick von Felten, Klaus Opwis, Philipp Wintersberger, and Florian Brühlmann · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jan Batzner, Volker Stocker, Stefan Schmid, and Gjergji Kasneci · 2024
Cited alongside, same era.
From yes-men to truth-tellers: Addressing sycophancy in large language models with pinpoint tuning
Wei Chen, Zhen Huang, Liang Xie, Binbin Lin, Houqiang Li, Le Lu, Xinmei Tian, Deng Cai, Yonggang Zhang, Wenxiao Wan, et al · 2024
Cited alongside, same era.
Ai alignment project ideas
Adam Jones · 2024
Cited alongside, same era.
Closest in time.
Language models don’t always say what they think: unfaithful explanations in chain-of-thought prompting
Miles Turpin, Julian Michael, Ethan Perez, and Samuel Bowman · 2024
Closest in time.