Fetching the paper…
Reading the bibliography…
As AI grows more powerful, it will increasingly shape how we understand the world.
Thinking, fast and slow
Daniel Kahneman · 2011
Earlier work this paper cites.
Concrete problems in ai safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
Exploring the role of prior beliefs for argument persuasion
Esin Durmus and Claire Cardie · 2018
Earlier work this paper cites.
Geoffrey Irving, Paul Christiano, and Dario Amodei · 2018
Earlier work this paper cites.
Human reasoning: The psychology of deduction
Ruth MJ Byrne, Jonathan St BT Evans, and Stephen E Newstead · 2019
Earlier work this paper cites.
Ai safety needs social scientists
Geoffrey Irving and Amanda Askell · 2019
Earlier work this paper cites.
Climate-fever: A dataset for verification of real-world climate claims
Thomas Diggelmann, Jordan Boyd-Graber, Jannis Bulian, Massimiliano Ciaramita, and Markus Leippold · 2020
Earlier work this paper cites.
Learning to summarize with human feedback
Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano · 2020
Earlier work this paper cites.
CORD-19: The COVID-19 open research dataset
Lucy Lu Wang, Kyle Lo, Yoganand Chandrasekhar, Russell Reas, Jiangjiang Yang, Doug Burdick, Darrin Eide, Kathryn Funk, Yannis Katsis, Rodney Michael Kinney, Yunyao Li, Ziyang Liu, William Merrill, Paul Mooney, Dewey A. Murdick, Devvret Rishi, Jerry Sheehan, Zhihong Shen, Brandon Stilson, Alex D. Wade, Kuansan Wang, Nancy Xin Ru Wang, Christopher Wilhelm, Boya Xie, Douglas M. Raymond, Daniel S. Weld, Oren Etzioni, and Sebastian Kohlmeier · 2020
Earlier work this paper cites.
Patterns of media use, strength of belief in covid-19 conspiracy theories, and the prevention of covid-19 from march to july 2020 in the united states: survey study
Daniel Romer and Kathleen Hall Jamieson · 2021
Earlier work this paper cites.
Political affiliation and risk taking behaviors among adults with elevated chance of severe complications from covid–19
Robert F Schoeni, Emily E Wiemers, Judith A Seltzer, and Kenneth M Langa · 2021
Earlier work this paper cites.
Constitutional ai: Harmlessness from ai feedback
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al · 2022
Earlier work this paper cites.
The impact of cognitive biases on professionals’ decision-making: A review of four occupational areas
Vincent Berthet · 2022
Earlier work this paper cites.
Measuring progress on scalable oversight for large language models
Samuel R. Bowman, Jeeyoon Hyun, and Ethan et al. Perez · 2022
Earlier work this paper cites.
Political ideology and the perceived impact of coronavirus prevention behaviors for the self and others
Aylin Cakanlar, Remi Trudel, and Katherine White · 2022
Earlier work this paper cites.
Misinfo reaction frames: Reasoning about readers’ reactions to news headlines
Saadia Gabriel, Skyler Hallinan, Maarten Sap, Pemi Nguyen, Franziska Roesner, Eunsol Choi, and Yejin Choi · 2022
Earlier work this paper cites.
Improving alignment of dialogue agents via targeted human judgements
Amelia Glaese, Nat McAleese, Maja Trębacz, John Aslanides, Vlad Firoiu, Timo Ewalds, Maribeth Rauh, Laura Weidinger, Martin Chadwick, Phoebe Thacker, et al · 2022
Earlier work this paper cites.
Capturing failures of large language models via human cognitive biases
Erik Jones and Jacob Steinhardt · 2022
Earlier work this paper cites.
Teaching language models to support answers with verified quotes
Jacob Menick, Maja Trebacz, Vladimir Mikulik, John Aslanides, Francis Song, Martin Chadwick, Mia Glaese, Susannah Young, Lucy Campbell-Gillingham, Geoffrey Irving, et al · 2022
Cited alongside, same era.
A general model of cognitive bias in human judgment and systematic review specific to forensic mental health
Tess Neal, Pascal Lienert, Emily Denne, and Jay P Singh · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Cited alongside, same era.
Deciding fast and slow: The role of cognitive biases in ai-assisted decision-making
Charvi Rastogi, Yunfeng Zhang, Dennis Wei, Kush R Varshney, Amit Dhurandhar, and Richard Tomsett · 2022
Cited alongside, same era.
The unreasonable effectiveness of easy training data for hard tasks
Peter Hase, Mohit Bansal, Peter Clark, and Sarah Wiegreffe · 2024
Later among the works it cites.
Large language models as misleading assistants in conversation
Betty Li Hou, Kejian Shi, Jason Phang, James Aung, Steven Adler, and Rosie Campbell · 2024
Later among the works it cites.
Sleeper agents: Training deceptive llms that persist through safety training
Evan Hubinger, Carson Denison, Jesse Mu, Mike Lambert, Meg Tong, Monte MacDiarmid, Tamera Lanham, Daniel M Ziegler, Tim Maxwell, Newton Cheng, et al · 2024
Later among the works it cites.
On scalable oversight with weak llms judging strong llms
Zachary Kenton, Noah Siegel, János Kramár, Jonah Brown-Cohen, Samuel Albanie, Jannis Bulian, Rishabh Agarwal, David Lindner, Yunhao Tang, Noah Goodman, et al · 2024
Later among the works it cites.
Debating with more persuasive llms leads to more truthful answers
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yejin Bang, Samuel Cahyawijaya, Nayeon Lee, Wenliang Dai, Dan Su, Bryan Wilie, Holy Lovenia, Ziwei Ji, Tiezheng Yu, Willy Chung, Quyet V Do, Yan Xu, and Pascale Fung · 2023
Cited alongside, same era.
Weak-to-strong generalization: Eliciting strong capabilities with weak supervision
Collin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker, Leo Gao, Leopold Aschenbrenner, Yining Chen, Adrien Ecoffet, Manas Joglekar, Jan Leike, et al · 2023
Cited alongside, same era.
Benchmarking cognitive biases in large language models as evaluators
Ryan Koo, Minhwa Lee, Vipul Raheja, Jong Inn Park, Zae Myung Kim, and Dongyeop Kang · 2023
Cited alongside, same era.
Debate helps supervise unreliable experts
Julian Michael, Salsabila Mahdi, David Rein, Jackson Petty, Julien Dirani, Vishakh Padmakumar, and Samuel R Bowman · 2023
Cited alongside, same era.
Large language models can strategically deceive their users when put under pressure
Jérémy Scheurer, Mikita Balesni, and Marius Hobbhahn · 2023
Cited alongside, same era.
Towards understanding sycophancy in language models
Mrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud, Amanda Askell, Samuel R Bowman, Newton Cheng, Esin Durmus, Zac Hatfield-Dodds, Scott R Johnston, et al · 2023
Cited alongside, same era.
Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting
Miles Turpin, Julian Michael, Ethan Perez, and Samuel Bowman · 2023
Cited alongside, same era.
Check-covid: Fact-checking covid-19 news claims with scientific evidence
Gengyu Wang, Kate Harwood, Lawrence Chillrud, Amith Ananthram, Melanie Subbiah, and Kathleen McKeown · 2023
Cited alongside, same era.
Akbir Khan, John Hughes, Dan Valentine, Laura Ruis, Kshitij Sachan, Ansh Radhakrishnan, Edward Grefenstette, Samuel R Bowman, Tim Rocktäschel, and Ethan Perez · 2024
Later among the works it cites.
Exploring gender biases in language patterns of human-conversational agent conversations
Weizi Liu · 2024
Later among the works it cites.
Designing human-ai systems: Anthropomorphism and framing bias on human-ai collaboration
Samuel Aleksander Sánchez Olszewski · 2024
Later among the works it cites.
Do language models exhibit the same cognitive biases in problem solving as human learners?
Andreas Opedal, Alessandro Stolfo, Haruki Shirakami, Ying Jiao, Ryan Cotterell, Bernhard Schölkopf, Abulhair Saparov, and Mrinmaya Sachan · 2024
Later among the works it cites.
Spontaneous reward hacking in iterative self-refinement
Jane Pan, He He, Samuel R Bowman, and Shi Feng · 2024
Later among the works it cites.
GPQA: A graduate-level google-proof q&a benchmark
David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R. Bowman · 2024
Later among the works it cites.
On the conversational persuasiveness of large language models: A randomized controlled trial
Francesco Salvi, Manoel Horta Ribeiro, Riccardo Gallotti, and Robert West · 2024
Later among the works it cites.
Don’t be fooled: The misinformation effect of explanations in human-ai collaboration
Philipp Spitzer, Joshua Holstein, Katelyn Morrison, Kenneth Holstein, Gerhard Satzger, and Niklas Kühl · 2024
Later among the works it cites.
Easy-to-hard generalization: Scalable alignment beyond human supervision
Zhiqing Sun, Longhui Yu, Yikang Shen, Weiyang Liu, Yiming Yang, Sean Welleck, and Chuang Gan · 2024
Later among the works it cites.
Language models learn to mislead humans via rlhf
Jiaxin Wen, Ruiqi Zhong, Akbir Khan, Ethan Perez, Jacob Steinhardt, Minlie Huang, Samuel R Bowman, He He, and Shi Feng · 2024
Later among the works it cites.
An alignment safety case sketch based on debate
Marie Davidsen Buhl, Jacob Pfau, Benjamin Hilton, and Geoffrey Irving · 2025
Closest in time.
Just the facts: How dialogues with ai reduce conspiracy beliefs, 2025
Thomas H Costello, Gordon Pennycook, and David Rand · 2025
Closest in time.
Mosaic: Modeling social ai for content dissemination and regulation in multi-agent simulations
Genglin Liu, Salman Rahman, Elisa Kreiss, Marzyeh Ghassemi, and Saadia Gabriel · 2025
Closest in time.
X-teaming: Multi-turn jailbreaks and defenses with adaptive multi-agents
Salman Rahman, Liwei Jiang, James Shiffer, Genglin Liu, Sheriff Issaka, Md Rizwan Parvez, Hamid Palangi, Kai-Wei Chang, Yejin Choi, and Saadia Gabriel · 2025
Closest in time.
A guide to misinformation detection data and evaluation
Camille Thibault, Jacob-Junqi Tian, Gabrielle Péloquin-Skulski, Taylor Lynn Curtis, James Zhou, Florence Laflamme, Luke Yuxiang Guan, Reihaneh Rabbany, Jean-François Godbout, and Kellin Pelrine · 2025
Closest in time.