Fetching the paper…
Reading the bibliography…
This paper critically evaluates the attempts to align Artificial Intelligence (AI) systems, especially Large Language Models (LLMs), with human values and intentions through Reinforcement Learning from Feedback (RLxF) methods, involving either human feedback (RLHF) or AI feedback (RLAIF).
Logic and conversation
Herbert P Grice · 1975
Earlier work this paper cites.
Computer Power and Human Reason: From Judgment to Calculation
Joseph Weizenbaum · 1977
Earlier work this paper cites.
Engineering a Safer World: Systems Thinking Applied to Safety
Nancy G. Leveson · 2012
Earlier work this paper cites.
High-dimensional continuous control using generalized advantage estimation
John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, and Pieter Abbeel · 2015
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
Assessing bert’s syntactic abilities
Yoav Goldberg · 2019
Earlier work this paper cites.
What does bert learn about the structure of language?
Ganesh Jawahar, Benoît Sagot, and Djamé Seddah · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin Ming-Wei Chang Kenton and Lee Kristina Toutanova · 2019
Earlier work this paper cites.
Fairness and abstraction in sociotechnical systems
Andrew D Selbst, Danah Boyd, Sorelle A Friedler, Suresh Venkatasubramanian, and Janet Vertesi · 2019
Earlier work this paper cites.
Concrete problems in ai safety, revisited
ID Raji and R Dobbe · 2020
Earlier work this paper cites.
A general language assistant as a laboratory for alignment
Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain, and Deep et al. Ganguli · 2021
Earlier work this paper cites.
On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?
Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell · 2021
Earlier work this paper cites.
Anticipating safety issues in e2e conversational ai: Framework and tooling
Emily Dinan, Gavin Abercrombie, A Stevie Bergman, Shannon Spruit, Dirk Hovy, Y-Lan Boureau, and Verena Rieser · 2021
Earlier work this paper cites.
Hard choices in artificial intelligence
Roel Dobbe, Thomas Krendl Gilbert, and Yonatan Mintz · 2021
Earlier work this paper cites.
Ethical and social risks of harm from language models
Laura Weidinger, John Mellor, Maribeth Rauh, Conor Griffin, Jonathan Uesato, Po-Sen Huang, Myra Cheng, Mia Glaese, Borja Balle, Atoosa Kasirzadeh, Zac Kenton, Sasha Brown, Will Hawkins, Tom Stepleton, Courtney Biles, Abeba Birhane, Julia Haas, Laura Rimell, Lisa Anne Hendricks, William Isaac, Sean Legassick, Geoffrey Irving, and Iason Gabriel · 2021
Earlier work this paper cites.
Ai anthropomorphism and its effect on users’ self-congruence and self–ai integration: A theoretical framework and research agenda
Amani Alabed, Ana Javornik, and Diana Gregory-Smith · 2022
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al · 2022
Earlier work this paper cites.
Constitutional ai: Harmlessness from ai feedback
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al · 2022
Earlier work this paper cites.
System safety and artificial intelligence
RIJ Dobbe · 2022
Earlier work this paper cites.
CounterFAccTual: How FAccT undermines its organizing principles
Ben Gansky and Sean McDonald · 2022
Earlier work this paper cites.
The data-production dispositif
Milagros Miceli and Julian Posada · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Cited alongside, same era.
AI, opacity, and personal autonomy
Bram Vaassen · 2022
Cited alongside, same era.
Acrocpolis: A descriptive framework for making sense of fairness
Andrea Aler Tubella, Dimitri Coelho Mollo, Adam Dahlgren Lindström, Hannah Devinney, Virginia Dignum, Petter Ericson, Anna Jonsson, Timotheus Kampik, Tom Lenaerts, Julian Alfredo Mendez, and Juan Carlos Nieves · 2023
Cited alongside, same era.
Frontier ai regulation: Managing emerging risks to public safety
Markus Anderljung, Joslyn Barnhart, Anton Korinek, Jade Leung, and Cullen et al. O’Keefe · 2023
Cited alongside, same era.
Model alignment protects against accidental harms, not intentional ones
Arvind Narayanan, Sayash Kapoor, and Lazar Seth · 2023
Later among the works it cites.
Diagnosing and addressing emergent harms in the design process of public AI and algorithmic systems
Sem Nouws, Íñigo Martinez De Rituerto De Troya, Roel Dobbe, and Marijn Janssen · 2023
Later among the works it cites.
Ai deception: A survey of examples, risks, and potential solutions
Peter S Park, Simon Goldstein, Aidan O’Gara, Michael Chen, and Dan Hendrycks · 2023
Later among the works it cites.
Discovering language model behaviors with model-written evaluations
Ethan Perez, Sam Ringer, Kamile Lukosiute, Karina Nguyen, and Edwin et al. Chen · 2023
Later among the works it cites.
Algorithms as Social-Ecological-Technological Systems: an Environmental Justice Lens on Algorithmic Audits
Bogdana Rakova and Roel Dobbe · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mohammad Atari, Mona J Xue, Peter S Park, Damián Blasi, and Joseph Henrich · 2023
Cited alongside, same era.
Open problems and fundamental limitations of reinforcement learning from human feedback
Stephen Casper, Xander Davies, Claudia Shi, Thomas Krendl Gilbert, and Scheurer et al · 2023
Cited alongside, same era.
Ultrafeedback: Boosting language models with high-quality feedback
Ganqu Cui, Lifan Yuan, Ning Ding, Guanming Yao, Wei Zhu, Yuan Ni, Guotong Xie, Zhiyuan Liu, and Maosong Sun · 2023
Cited alongside, same era.
‘Safety Washing’ at the AI Safety Summit, November 2023
Roel Dobbe · 2023
Cited alongside, same era.
Ai is a lot of work
Josh Dzieza · 2023
Cited alongside, same era.
Jailbreak tricks discord’s new chatbot into sharing napalm and meth instructions
Lorenzo Franceschi-Bicchierai · 2023
Cited alongside, same era.
The Past, Present and Better Future of Feedback Learning in Large Language Models for Subjective Human Preferences and Values
Hannah Kirk, Andrew Bean, Bertie Vidgen, Paul Rottger, and Scott Hale · 2023
Cited alongside, same era.
Beyond the ML Model: Applying Safety Engineering Frameworks to Text-to-Image Development
Shalaleh Rismani, Renee Shelby, Andrew Smart, Renelito Delos Santos, AJung Moon, and Negar Rostamzadeh · 2023
Later among the works it cites.
From Plane Crashes to Algorithmic Harm: Applicability of Safety Engineering Frameworks for Responsible ML
Shalaleh Rismani, Renee Shelby, Andrew Smart, Edgar Jatho, Joshua Kroll, AJung Moon, and Negar Rostamzadeh · 2023
Later among the works it cites.
Sociotechnical harms of algorithmic systems: Scoping a taxonomy for harm reduction
Renee Shelby, Shalaleh Rismani, Kathryn Henne, AJung Moon, Negar Rostamzadeh, Paul Nicholas, N’Mah Yilla-Akbari, Jess Gallegos, Andrew Smart, Emilio Garcia, and Gurleen Virk · 2023
Later among the works it cites.
Anthropomorphism in Artificial Intelligence: A Review of Empirical Work Across Domains and Insights for Future Research
Ertugrul Uysal, Sascha Alavi, and Valéry Bezençon · 2023
Later among the works it cites.
Red teaming chatgpt via jailbreaking: Bias, robustness, reliability and toxicity
TY Zhuo, Y Huang, C Chen, and Z Xing · 2023
Later among the works it cites.
Toward Sociotechnical AI and MLOps: Mapping Vulnerabilities for Machine Learning in Context
Roel Dobbe and Anouk Wolters · 2024
Closest in time.
Diversity and language technology: how language modeling bias causes epistemic injustice
Paula Helm, Gábor Bella, Gertraud Koch, and Fausto Giunchiglia · 2024
Closest in time.
Beavertails: Towards improved safety alignment of llm via a human-preference dataset
Jiaming Ji, Mickel Liu, Josef Dai, Xuehai Pan, Chi Zhang, Ce Bian, Boyuan Chen, Ruiyang Sun, Yizhou Wang, and Yaodong Yang · 2024
Closest in time.
The benefits, risks and bounds of personalizing the alignment of large language models to individuals
Hannah Rose Kirk, Bertie Vidgen, Paul Röttger, and Scott A. Hale · 2024
Closest in time.
Understanding the effects of rlhf on llm generalisation and diversity
Robert Kirk, Ishita Mediratta, Christoforos Nalmpantis, Jelena Luketina, Eric Hambro, Edward Grefenstette, and Roberta Raileanu · 2024
Closest in time.
How do large language models navigate conflicts between honesty and helpfulness?
Ryan Liu, Theodore R. Sumers, Ishita Dasgupta, and Thomas L. Griffiths · 2024
Closest in time.
Towards understanding sycophancy in language models
Mrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud, Amanda Askell, Samuel R. Bowman, Esin DURMUS, Zac Hatfield-Dodds, Scott R Johnston, Shauna M Kravec, Timothy Maxwell, Sam McCandlish, Kamal Ndousse, Oliver Rausch, Nicholas Schiefer, Da Yan, Miranda Zhang, and Ethan Perez · 2024
Closest in time.
Preference ranking optimization for human alignment
Feifan Song, Bowen Yu, Minghao Li, Haiyang Yu, Fei Huang, Yongbin Li, and Houfeng Wang · 2024
Closest in time.
Language models don’t always say what they think: unfaithful explanations in chain-of-thought prompting
Miles Turpin, Julian Michael, Ethan Perez, and Samuel Bowman · 2024
Closest in time.