Fetching the paper…
Reading the bibliography…
Aligning AI agents to human intentions and values is a key bottleneck in building safe and deployable AI applications.
“Some Moral and Technical Consequences of Automation: As machines learn they may develop unforeseen strategies at rates that baffle their programmers.”
Norbert Wiener · 1960
Earlier work this paper cites.
“The impossibility of a Paretian liberal”
Amartya Sen · 1970
Earlier work this paper cites.
“Justice as fairness: A restatement”
John Rawls · 2001
Earlier work this paper cites.
“Machine ethics”
Michael Anderson and Susan Anderson · 2011
Earlier work this paper cites.
“Fairness through awareness”
Cynthia Dwork et al · 2012
Earlier work this paper cites.
“Computational social choice”
Felix Brandt, Vincent Conitzer and Ulle Endriss · 2012
Earlier work this paper cites.
“Social choice and individual values”
Kenneth Arrow · 2012
Earlier work this paper cites.
“Programming by Feedback”
Riad Akrour, Marc Schoenauer, Jean-Christophe Souplet and Michele Sebag · 2014
Earlier work this paper cites.
“Human-level control through deep reinforcement learning”
Volodymyr Mnih et al · 2015
Earlier work this paper cites.
“Handbook of computational social choice”
Felix Brandt et al · 2016
Cited alongside, same era.
“Deep reinforcement learning from human preferences”
Paul Christiano et al · 2017
Cited alongside, same era.
“The moral machine experiment”
Edmond Awad et al · 2018
Cited alongside, same era.
“A voting-based system for ethical decision making”
Ritesh Noothigattu et al · 2018
Cited alongside, same era.
“Machine behaviour”
Iyad Rahwan et al · 2019
Cited alongside, same era.
“Fine-tuning language models from human preferences”
Daniel Ziegler et al · 2019
Cited alongside, same era.
“Human-compatible artificial intelligence”
Stuart Russell · 2021
Later among the works it cites.
“A general language assistant as a laboratory for alignment”
Amanda Askell et al · 2021
Later among the works it cites.
“Training a helpful and harmless assistant with reinforcement learning from human feedback”
Yuntao Bai et al · 2022
Later among the works it cites.
“LaMDA: Language Models for Dialog Applications”, 2022
Romal Thoppilan et al · 2022
Later among the works it cites.
“Training language models to follow instructions with human feedback”
Long Ouyang et al · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Margaret Mitchell et al · 2019
Cited alongside, same era.
“The global landscape of AI ethics guidelines”
Anna Jobin, Marcello Ienca and Effy Vayena · 2019
Cited alongside, same era.
“Artificial intelligence, values, and alignment”
Iason Gabriel · 2020
Cited alongside, same era.
“GPT-4 Technical Report”, 2023
OpenAI · 2023
Closest in time.
“Introducing Claude”, 2023
AI Anthropic · 2023
Closest in time.
“Llama 2: Open Foundation and Fine-Tuned Chat Models”, 2023
Hugo Touvron et al · 2023
Closest in time.
“Open problems and fundamental limitations of reinforcement learning from human feedback”
Stephen Casper et al · 2023
Closest in time.