Fetching the paper…
Reading the bibliography…
As AIs rapidly advance and become more agentic, the risk they pose is governed not only by their capabilities but increasingly by their propensities, including goals and values.
Cardinal welfare, individualistic ethics, and interpersonal comparisons of utility
John C Harsanyi · 1955
Earlier work this paper cites.
The structure of utility functions
William M Gorman · 1968
Earlier work this paper cites.
The framing of decisions and the psychology of choice
Amos Tversky and Daniel Kahneman · 1981
Earlier work this paper cites.
An experimental analysis of ultimatum bargaining
Werner Güth, Rolf Schmittberger, and Bernd Schwarze · 1982
Earlier work this paper cites.
Utility functions: from risk theory to finance
Hans U Gerber and Gérard Pafum · 1998
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Andrew Y Ng, Stuart Russell, et al · 2000
Earlier work this paper cites.
On the nature of fair behavior
Armin Falk, Ernst Fehr, and Urs Fischbacher · 2003
Earlier work this paper cites.
Stereoset: Measuring stereotypical bias in pretrained language models, 2020
Moin Nadeem, Anna Bethke, and Siva Reddy · 2004
Earlier work this paper cites.
Uncertainty and hyperbolic discounting
Partha Dasgupta and Eric Maskin · 2005
Earlier work this paper cites.
Does deliberative democracy work?
David M Ryfe · 2005
Earlier work this paper cites.
Designing Deliberative Democracy: The British Columbia Citizens’ Assembly
Mark E. Warren and Hilary Pearse, editors · 2008
Earlier work this paper cites.
Preference reversals and probabilistic decisions
Pavlo R Blavatskyy · 2009
Earlier work this paper cites.
Superintelligence: Paths, dangers, strategies
Nick Bostrom · 2014
Earlier work this paper cites.
Dopamine reward prediction error responses reflect marginal utility
William R Stauffer, Armin Lak, and Wolfram Schultz · 2014
Earlier work this paper cites.
Corrigibility
Nate Soares, Benja Fallenstein, Eliezer Yudkowsky, and Stuart Armstrong · 2015
Earlier work this paper cites.
Cooperative inverse reinforcement learning
Dylan Hadfield-Menell, Stuart J Russell, Pieter Abbeel, and Anca Dragan · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
Understanding intermediate layers using linear classifier probes, 2018
Guillaume Alain and Yoshua Bengio · 2018
Earlier work this paper cites.
Deliberative democracy
André Bächtiger, John S Dryzek, Jane Mansbridge, and Mark Warren · 2018
Earlier work this paper cites.
Decoupled weight decay regularization, 2019
Ilya Loshchilov and Frank Hutter · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Aligning ai with shared human values
Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt · 2020
Earlier work this paper cites.
It’s not just size that matters: Small language models are also few-shot learners
Timo Schick and Hinrich Schütze · 2020
Earlier work this paper cites.
A general language assistant as a laboratory for alignment, 2021
Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain, Deep Ganguli, Tom Henighan, Andy Jones, Nicholas Joseph, Ben Mann, Nova DasSarma, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Jackson Kernion, Kamal Ndousse, Catherine Olsson, Dario Amodei, Tom Brown, Jack Clark, Sam McCandlish, Chris Olah, and Jared Kaplan · 2021
Cited alongside, same era.
Truthful ai: Developing and governing ai that does not lie, 2021
Owain Evans, Owen Cotton-Barratt, Lukas Finnveden, Adam Bales, Avital Balwit, Peter Wills, Luca Righetti, and William Saunders · 2021
Cited alongside, same era.
Are citizen juries and assemblies on climate change driving democratic climate policymaking? an exploration of two case studies in the UK
Rebecca Wells, Candice Howarth, and Lina I Brand-Correa · 2021
Cited alongside, same era.
Training a helpful and harmless assistant with reinforcement learning from human feedback, 2022
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, Jackson Kernion, Tom Conerly, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Tristan Hume, Scott Johnston, Shauna Kravec, Liane Lovitt, Neel Nanda, Catherine Olsson, Dario Amodei, Tom Brown, Jack Clark, Sam McCandlish, Chris Olah, Ben Mann, and Jared Kaplan · 2022
Dailydilemmas: Revealing value preferences of llms with quandaries of daily life, 2024
Yu Ying Chiu, Liwei Jiang, and Yejin Choi · 2024
Later among the works it cites.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al · 2024
Later among the works it cites.
Webvoyager: Building an end-to-end web agent with large multimodal models, 2024
Hongliang He, Wenlin Yao, Kaixin Ma, Wenhao Yu, Yong Dai, Hongming Zhang, Zhenzhong Lan, and Dong Yu · 2024
Later among the works it cites.
Introduction to ai safety, ethics and society, 2024
Dan Hendrycks · 2024
Later among the works it cites.
Learning to be homo economicus: Can an llm learn preferences from choice, 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Discovering latent knowledge in language models without supervision
Collin Burns, Haotian Ye, Dan Klein, and Jacob Steinhardt · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Cited alongside, same era.
Human-compatible artificial intelligence., 2022
Stuart Russell · 2022
Cited alongside, same era.
Goal misgeneralization: Why correct specifications aren’t enough for correct goals
Rohin Shah, Vikrant Varma, Ramana Kumar, Mary Phuong, Victoria Krakovna, Jonathan Uesato, and Zac Kenton · 2022
Cited alongside, same era.
React: Synergizing reasoning and acting in language models
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao · 2022
Cited alongside, same era.
Using large language models to simulate multiple humans and replicate human subject studies, 2023
Gati Aher, Rosa I. Arriaga, and Adam Tauman Kalai · 2023
Cited alongside, same era.
The emergence of economic rationality of gpt, 2023
Yiting Chen, Tracy Xiao Liu, You Shan, and Songfa Zhong · 2023
Cited alongside, same era.
Sortition and its principles: Evaluation of the selection processes of citizens’ assemblies
Adela Gasiorowska · 2023
Cited alongside, same era.
Jeongbin Kim, Matthew Kovach, Kyu-Min Lee, Euncheol Shin, and Hector Tzavellas · 2024
Later among the works it cites.
Are large language models consistent over value-laden questions?
Jared Moore, Tanvi Deshpande, and Diyi Yang · 2024
Later among the works it cites.
Team OLMo, Pete Walsh, Luca Soldaini, Dirk Groeneveld, Kyle Lo, Shane Arora, Akshita Bhagia, Yuling Gu, Shengyi Huang, Matt Jordan, et al · 2024
Later among the works it cites.
Iclr: In-context learning of representations
Core Francisco Park, Andrew Lee, Ekdeep Singh Lubana, Yongyi Yang, Maya Okawa, Kento Nishi, Martin Wattenberg, and Hidenori Tanaka · 2024
Later among the works it cites.
Is temperature the creativity parameter of large language models?
Max Peeperkorn, Tom Kouwenhoven, Dan Brown, and Anna Jordanous · 2024
Later among the works it cites.
Hidden persuaders: Llms’ political leaning and their influence on voters
Yujin Potter, Shiyang Lai, Junsol Kim, James Evans, and Dawn Song · 2024
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn · 2024
Later among the works it cites.
Steer: Assessing the economic rationality of large language models, 2024
Narun Raman, Taylor Lundy, Samuel Amouyal, Yoav Levine, Kevin Leyton-Brown, and Moshe Tennenholtz · 2024
Later among the works it cites.
Assessing political bias in large language models, 2024
Luca Rettenberger, Markus Reischl, and Mark Schutera · 2024
Later among the works it cites.
Do llms have consistent values?, 2024
Naama Rozen, Liat Bezalel, Gal Elidan, Amir Globerson, and Ella Daniel · 2024
Later among the works it cites.
Gemma 2: Improving open language models at a practical size
Gemma Team, Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupatiraju, Léonard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ramé, et al · 2024
Later among the works it cites.
The shutdown problem: An ai engineering puzzle for decision theorists, 2024
Elliott Thornley · 2024
Later among the works it cites.
Grok-2 beta release, August 2024
XAI · 2024
Later among the works it cites.
Magpie: Alignment data synthesis from scratch by prompting aligned llms with nothing, 2024
Zhangchen Xu, Fengqing Jiang, Luyao Niu, Yuntian Deng, Radha Poovendran, Yejin Choi, and Bill Yuchen Lin · 2024
Later among the works it cites.
Focus agent: Llm-powered virtual focus group
Taiyu Zhang, Xuesong Zhang, Robbe Cools, and Adalberto Simeone · 2024
Later among the works it cites.
The claude 3 model family: Opus, sonnet, haiku
Anthropic · 2025
Closest in time.
Gpt-3.5 turbo fine-tuning and api updates
OpenAI · 2025
Closest in time.
Hello gpt-4o
OpenAI · 2025
Closest in time.
Acs 1-year estimates public use microdata sample
U.S. Census Bureau · 2025
Closest in time.