Fetching the paper…
Reading the bibliography…
Social alignment in AI systems aims to ensure that these models behave according to established societal values.
Problems of monetary management: the UK experience
Charles AE Goodhart · 1984
Earlier work this paper cites.
Loss functions for preference levels: Regression with discrete ordered labels
Jason DM Rennie and Nathan Srebro · 2005
Earlier work this paper cites.
Transformative experience
Laurie Ann Paul · 2014
Earlier work this paper cites.
Siamese neural networks for one-shot image recognition
Gregory Koch, Richard Zemel, Ruslan Salakhutdinov, et al · 2015
Earlier work this paper cites.
Facenet: A unified embedding for face recognition and clustering
Florian Schroff, Dmitry Kalenichenko, and James Philbin · 2015
Earlier work this paper cites.
Concrete problems in ai safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Earlier work this paper cites.
Alignment for advanced machine learning systems
Jessica Taylor, Eliezer Yudkowsky, Patrick LaVictoire, and Andrew Critch · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F. Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
Jan Leike, Miljan Martic, Victoria Krakovna, Pedro A Ortega, Tom Everitt, Andrew Lefrancq, Laurent Orseau, and Shane Legg · 2017
Earlier work this paper cites.
Geoffrey Irving, Paul Christiano, and Dario Amodei · 2018
Earlier work this paper cites.
Joel Lehman, Jeff Clune, Dusan Misevic, Christoph Adami, Lee Altenberg, Julie Beaulieu, Peter J Bentley, Samuel Bernard, Guillaume Beslon, David M Bryson, et al · 2018
Earlier work this paper cites.
Scalable agent alignment via reward modeling: a research direction
Jan Leike, David Krueger, Tom Everitt, Miljan Martic, Vishal Maini, and Shane Legg · 2018
Earlier work this paper cites.
Choosing for changing selves
Richard Pettigrew · 2019
Earlier work this paper cites.
Fine-tuning language models from human preferences
Daniel M Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Earlier work this paper cites.
Artificial intelligence, values, and alignment
Iason Gabriel · 2020
Earlier work this paper cites.
RealToxicityPrompts: Evaluating neural toxic degeneration in language models
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith · 2020
Earlier work this paper cites.
Human-centric dialog training via offline reinforcement learning
Natasha Jaques, Judy Hanwen Shen, Asma Ghandeharioun, Craig Ferguson, Agata Lapedriza, Noah Jones, Shixiang Gu, and Rosalind Picard · 2020
Earlier work this paper cites.
Avoiding side effects by considering future tasks
Victoria Krakovna, Laurent Orseau, Richard Ngo, Miljan Martic, and Shane Legg · 2020
Earlier work this paper cites.
A general language assistant as a laboratory for alignment
Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain, Deep Ganguli, Tom Henighan, Andy Jones, Nicholas Joseph, Ben Mann, Nova DasSarma, et al · 2021
Earlier work this paper cites.
On the opportunities and risks of foundation models
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al · 2021
Earlier work this paper cites.
Moral stories: Situated reasoning about norms, intents, actions, and their consequences
Denis Emelin, Ronan Le Bras, Jena D. Hwang, Maxwell Forbes, and Yejin Choi · 2021
Cited alongside, same era.
Reward tampering problems and solutions in reinforcement learning: A causal influence diagram perspective
Tom Everitt, Marcus Hutter, Ramana Kumar, and Victoria Krakovna · 2021
Cited alongside, same era.
SimCSE: Simple contrastive learning of sentence embeddings
Tianyu Gao, Xingcheng Yao, and Danqi Chen · 2021
Cited alongside, same era.
Aligning AI with shared human values
Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt · 2021
Cited alongside, same era.
Surface form competition: Why the highest probability answer isn’t always right
Ari Holtzman, Peter West, Vered Shwartz, Yejin Choi, and Luke Zettlemoyer · 2021
Cited alongside, same era.
The effects of reward misspecification: Mapping and mitigating misaligned models
Alexander Pan, Kush Bhatia, and Jacob Steinhardt · 2022
Later among the works it cites.
Social simulacra: Creating populated prototypes for social computing systems
Joon Sung Park, Lindsay Popowski, Carrie Cai, Meredith Ringel Morris, Percy Liang, and Michael S Bernstein · 2022
Later among the works it cites.
Ml systems will have weird failure modes
Jacob Steinhardt · 2022
Later among the works it cites.
A study of implicit bias in pretrained language models against people with disabilities
Pranav Narayanan Venkit, Mukund Srinath, and Shomir Wilson · 2022
Later among the works it cites.
Self-instruct: Aligning language model with self generated instructions
Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A Smith, Daniel Khashabi, and Hannaneh Hajishirzi · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zachary Kenton, Tom Everitt, Laura Weidinger, Iason Gabriel, Vladimir Mikulik, and Geoffrey Irving · 2021
Cited alongside, same era.
A human blueprint for ai coexistence., 2021
Kai-Fu Lee · 2021
Cited alongside, same era.
Mitigating political bias in language models through reinforced calibration
Ruibo Liu, Chenyan Jia, Jason Wei, Guangxuan Xu, Lili Wang, and Soroush Vosoughi · 2021
Cited alongside, same era.
Scaling language models: Methods, analysis & insights from training gopher
Jack W Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, et al · 2021
Cited alongside, same era.
Understanding the capabilities, limitations, and societal impact of large language models
Alex Tamkin, Miles Brundage, Jack Clark, and Deep Ganguli · 2021
Cited alongside, same era.
Bot-adversarial dialogue for safe conversational agents
Jing Xu, Da Ju, Margaret Li, Y-Lan Boureau, Jason Weston, and Emily Dinan · 2021
Cited alongside, same era.
Language models as agent models
Jacob Andreas · 2022
Cited alongside, same era.
Taxonomy of risks posed by language models
Laura Weidinger, Jonathan Uesato, Maribeth Rauh, Conor Griffin, Po-Sen Huang, John Mellor, Amelia Glaese, Myra Cheng, Borja Balle, Atoosa Kasirzadeh, et al · 2022
Later among the works it cites.
The moral integrity corpus: A benchmark for ethical dialogue systems
Caleb Ziems, Jane Yu, Yi-Chia Wang, Alon Halevy, and Diyi Yang · 2022
Later among the works it cites.
Using large language models to simulate multiple humans and replicate human subject studies, 2023
Gati Aher, Rosa I. Arriaga, and Adam Tauman Kalai · 2023
Closest in time.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E Gonzalez, et al · 2023
Closest in time.
The false promise of imitating proprietary llms
Arnav Gudibande, Eric Wallace, Charlie Snell, Xinyang Geng, Hao Liu, Pieter Abbeel, Sergey Levine, and Dawn Song · 2023
Closest in time.
Chain of hindsight aligns language models with feedback
H Liu, C Sferrazza, and P Abbeel · 2023
Closest in time.
Generative agents: Interactive simulacra of human behavior
Joon Sung Park, Joseph C O’Brien, Carrie J Cai, Meredith Ringel Morris, Percy Liang, and Michael S Bernstein · 2023
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D Manning, and Chelsea Finn · 2023
Closest in time.
The curse of recursion: Training on generated data makes models forget.(may 2023), 2023
Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Yarin Gal, Nicolas Papernot, and Ross Anderson · 2023
Closest in time.
Can large language models change user preference adversarially?
Varshini Subhash · 2023
Closest in time.
Stanford alpaca: An instruction-following llama model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto · 2023
Closest in time.
Transformer reinforcement learning x
Leandro von Werra et al · 2023
Closest in time.
Fundamental limitations of alignment in large language models
Yotam Wolf, Noam Wies, Yoav Levine, and Amnon Shashua · 2023
Closest in time.
Rrhf: Rank responses to align language models with human feedback without tears
Zheng Yuan, Hongyi Yuan, Chuanqi Tan, Wei Wang, Songfang Huang, and Fei Huang · 2023
Closest in time.
Lima: Less is more for alignment
Chunting Zhou, Pengfei Liu, Puxin Xu, Srini Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, Lili Yu, et al · 2023
Closest in time.