Fetching the paper…
Reading the bibliography…
Large language models (LLMs) are often sycophantic, prioritizing agreement with their users over accurate or objective statements.
Immediate and delayed primacy and recency effects in performance evaluation
Dirk D Steiner and Jeffrey S Rain · 1989
Earlier work this paper cites.
Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales, 2005
Bo Pang and Lillian Lee · 2005
Earlier work this paper cites.
A positivity bias in written and spoken english and its moderation by personality and gender
Adam A Augustine, Matthias R Mehl, and Randy J Larsen · 2011
Earlier work this paper cites.
Interference of the end: Why recency bias in memory determines when a food is consumed again
Emily N Garbinsky, Carey K Morewedge, and Baba Shiv · 2014
Earlier work this paper cites.
Human language reveals a universal positivity bias
Peter Sheridan Dodds, Eric M Clark, Suma Desu, Morgan R Frank, Andrew J Reagan, Jake Ryland Williams, Lewis Mitchell, Kameron Decker Harris, Isabel M Kloumann, James P Bagrow, et al · 2015
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
Understanding intermediate layers using linear classifier probes, 2018
Guillaume Alain and Yoshua Bengio · 2018
Earlier work this paper cites.
On the opportunities and risks of foundation models
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al · 2021
Earlier work this paper cites.
Behavioral biases in the NFL gambling market: Overreaction to news and the recency bias, 2021
Robert B. B. Durand, Fernando Patterson, and Corey A. Shank · 2021
Earlier work this paper cites.
siebert/sentiment-roberta-large-english · hugging face, 2021
Hugging Face · 2021
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al · 2022
Cited alongside, same era.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
Deep Ganguli, Liane Lovitt, Jackson Kernion, Amanda Askell, Yuntao Bai, Saurav Kadavath, Ben Mann, Ethan Perez, Nicholas Schiefer, Kamal Ndousse, et al · 2022
Cited alongside, same era.
More than a feeling: Accuracy and application of sentiment analysis
Jochen Hartmann, Mark Heitmann, Christian Siebert, and Christina Schamp · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Cited alongside, same era.
Discovering language model behaviors with model-written evaluations, 2022
Still no lie detector for language models: Probing empirical and conceptual roadblocks, 2023
B. A. Levinstein and Daniel A. Herrmann · 2023
Later among the works it cites.
Emergent linear representations in world models of self-supervised sequence models, 2023
Neel Nanda, Andrew Lee, and Martin Wattenberg · 2023
Later among the works it cites.
Towards understanding sycophancy in language models, 2023
Mrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud, Amanda Askell, Samuel R. Bowman, Newton Cheng, Esin Durmus, Zac Hatfield-Dodds, Scott R. Johnston, Shauna Kravec, Timothy Maxwell, Sam McCandlish, Kamal Ndousse, Oliver Rausch, Nicholas Schiefer, Da Yan, Miranda Zhang, and Ethan Perez · 2023
Later among the works it cites.
Openchat: Advancing open-source language models with mixed-quality data
Guan Wang, Sijie Cheng, Xianyuan Zhan, Xiangang Li, Sen Song, and Yang Liu · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ethan Perez, Sam Ringer, Kamilė Lukošiūtė, Karina Nguyen, Edwin Chen, Scott Heiner, Craig Pettit, Catherine Olsson, Sandipan Kundu, Saurav Kadavath, Andy Jones, Anna Chen, Ben Mann, Brian Israel, Bryan Seethor, Cameron McKinnon, Christopher Olah, Da Yan, Daniela Amodei, Dario Amodei, Dawn Drain, Dustin Li, Eli Tran-Johnson, Guro Khundadze, Jackson Kernion, James Landis, Jamie Kerr, Jared Mueller, Jeeyoon Hyun, Joshua Landau, Kamal Ndousse, Landon Goldberg, Liane Lovitt, Martin Lucas, Michael Sellitto, Miranda Zhang, Neerav Kingsland, Nelson Elhage, Nicholas Joseph, Noemí Mercado, Nova DasSarma, Oliver Rausch, Robin Larson, Sam McCandlish, Scott Johnston, Shauna Kravec, Sheer El Showk, Tamera Lanham, Timothy Telleen-Lawton, Tom Brown, Tom Henighan, Tristan Hume, Yuntao Bai, Zac Hatfield-Dodds, Jack Clark, Samuel R. Bowman, Amanda Askell, Roger Grosse, Danny Hernandez, Deep Ganguli, Evan Hubinger, Nicholas Schiefer, and Jared Kaplan · 2022
Cited alongside, same era.
Simple synthetic data reduces sycophancy in large language models, 2022
Jerry Wei, Da Huang, Yifeng Lu, Denny Zhou, and Quoc V. Le · 2022
Cited alongside, same era.
Open problems and fundamental limitations of reinforcement learning from human feedback, 2023
Stephen Casper, Xander Davies, Claudia Shi, Thomas Krendl Gilbert, Jérémy Scheurer, Javier Rando, Rachel Freedman, Tomasz Korbak, David Lindner, Pedro Freire, Tony Wang, Samuel Marks, Charbel-Raphaël Segerie, Micah Carroll, Andi Peng, Phillip Christoffersen, Mehul Damani, Stewart Slocum, Usman Anwar, Anand Siththaranjan, Max Nadeau, Eric J. Michaud, Jacob Pfau, Dmitrii Krasheninnikov, Xin Chen, Lauro Langosco, Peter Hase, Erdem Bıyık, Anca Dragan, David Krueger, Dorsa Sadigh, and Dylan Hadfield-Menell · 2023
Cited alongside, same era.
UltraFeedback: Boosting language models with high-quality feedback, 2023
Ganqu Cui, Lifan Yuan, Ning Ding, Guanming Yao, Wei Zhu, Yuan Ni, Guotong Xie, Zhiyuan Liu, and Maosong Sun · 2023
Cited alongside, same era.
Large language models are zero-shot reasoners, 2023
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa · 2023
Cited alongside, same era.
Starling-7b: Improving llm helpfulness & harmlessness with rlaif, 2023
Banghua Zhu, Evan Frick, Tianhao Wu, Hanlin Zhu, and Jiantao Jiao · 2023
Later among the works it cites.
Representation engineering: A top-down approach to ai transparency, 2023
Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, Shashwat Goel, Nathaniel Li, Michael J. Byun, Zifan Wang, Alex Mallen, Steven Basart, Sanmi Koyejo, Dawn Song, Matt Fredrikson, J. Zico Kolter, and Dan Hendrycks · 2023
Later among the works it cites.
Llm agents can autonomously hack websites, 2024
Richard Fang, Rohan Bindu, Akul Gupta, Qiusi Zhan, and Daniel Kang · 2024
Closest in time.
Jaspreet Pannu, Doni Bloomfield, Alex Zhu, Robert MacKnight, Gabe Gomes, Anita Cicero, and Thomas V Inglesby · 2024
Closest in time.
Steering llama 2 via contrastive activation addition, 2024
Nina Rimsky, Nick Gabrieli, Julian Schulz, Meg Tong, Evan Hubinger, and Alexander Matt Turner · 2024
Closest in time.