Fetching the paper…
Reading the bibliography…
Frontier AI systems, including large language models (LLMs), hold increasing influence over the epistemology of human users.
The cultural evolution of civilizations
Kent V Flannery · 1972
Earlier work this paper cites.
Moral progress
Ruth Macklin · 1977
Earlier work this paper cites.
The evolution of cooperation
Robert Axelrod and William D Hamilton · 1981
Earlier work this paper cites.
Inventing the French Revolution: essays on French political culture in the eighteenth century
Keith Michael Baker · 1990
Earlier work this paper cites.
Evolutionary game theory
Jörgen W Weibull · 1997
Earlier work this paper cites.
The value of liberalism
SANDRA Pralong · 1999
Earlier work this paper cites.
World values surveys and european values surveys, 1981-1984, 1990-1993, and 1995-1997
Ronald Inglehart, Miguel Basanez, Jaime Diez-Medrano, Loek Halman, and Ruud Luijkx · 2000
Earlier work this paper cites.
World values surveys and european values surveys, 1981-1984, 1990-1993, and 1995-1997
Ronald Inglehart, Miguel Basanez, Jaime Diez-Medrano, Loek Halman, and Ruud Luijkx · 2000
Earlier work this paper cites.
The evolution of cultural evolution
Joseph Henrich and Richard McElreath · 2003
Earlier work this paper cites.
Literary freedom: Project gutenberg
Bryan Stroube · 2003
Earlier work this paper cites.
The evolution of moral understanding
Christopher Robert Hallpike · 2004
Earlier work this paper cites.
A tutorial on support vector regression
Alex J Smola and Bernhard Schölkopf · 2004
Earlier work this paper cites.
The reformation
Diarmaid MacCulloch · 2005
Earlier work this paper cites.
Religion and value systems
Sonia Roccas · 2005
Earlier work this paper cites.
Ignorance and imagination: The epistemic origin of the problem of consciousness
Daniel Stoljar · 2006
Earlier work this paper cites.
Towards a unified science of cultural evolution
Alex Mesoudi, Andrew Whiten, and Kevin N Laland · 2006
Earlier work this paper cites.
The two fundamental problems of ethics
Arthur Schopenhauer · 2009
Earlier work this paper cites.
Architecture of the internet archive
Elliot Jaffe and Scott Kirkpatrick · 2009
Earlier work this paper cites.
The use and misuse of early english books online
Ian Gadd · 2009
Earlier work this paper cites.
Algorithmic game theory
Tim Roughgarden · 2010
Earlier work this paper cites.
Numerical analysis
Walter Gautschi · 2011
Earlier work this paper cites.
The expanding circle: Ethics, evolution, and moral progress
Peter Singer · 2011
Earlier work this paper cites.
Econometrics
Fumio Hayashi · 2011
Earlier work this paper cites.
Agent-based modeling
Dirk Helbing · 2012
Earlier work this paper cites.
Linear regression analysis
George AF Seber and Alan J Lee · 2012
Earlier work this paper cites.
The possibility of an ongoing moral catastrophe
Evan G Williams · 2015
Earlier work this paper cites.
An introduction to the theory of mechanism design
Tilman Börgers · 2015
Earlier work this paper cites.
A new tool to map the major worldviews in the netherlands and usa, and explore how they relate to climate change
Annick De Witt, Joop de Boer, Nicholas Hedlund, and Patricia Osseweijer · 2016
Earlier work this paper cites.
Concrete problems in ai safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Earlier work this paper cites.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Earlier work this paper cites.
Discrete calculus
Carlo Mariconda and Alberto Tonolo · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
Democracy as a universal value
Amartya Sen · 2017
Earlier work this paper cites.
The Evolution of Moral Progress: A Biocultural Theory
Allen Buchanan and Russell Powell · 2018
Earlier work this paper cites.
The moral machine experiment
Edmond Awad, Sohan Dsouza, Richard Kim, Jonathan Schulz, Joseph Henrich, Azim Shariff, Jean-François Bonnefon, and Iyad Rahwan · 2018
Earlier work this paper cites.
The enlightenment
Dorinda Outram · 2019
Earlier work this paper cites.
Sentence-bert: Sentence embeddings using siamese bert-networks
Nils Reimers and Iryna Gurevych · 2019
Earlier work this paper cites.
Artificial intelligence, values, and alignment
Iason Gabriel · 2020
Cited alongside, same era.
Supervised contrastive learning for pre-trained language model fine-tuning
Beliz Gunel, Jingfei Du, Alexis Conneau, and Ves Stoyanov · 2020
Cited alongside, same era.
The moral choice machine
Patrick Schramowski, Cigdem Turan, Sophie Jentzsch, Constantin Rothkopf, and Kristian Kersting · 2020
Cited alongside, same era.
Time series analysis
James D Hamilton · 2020
Cited alongside, same era.
Probability and random processes
Geoffrey Grimmett and David Stirzaker · 2020
Cited alongside, same era.
Image-based model parameter optimization using model-assisted generative adversarial networks
Saúl Alonso-Monsalve and Leigh H Whitehead · 2020
Cited alongside, same era.
Do the rewards justify the means? measuring trade-offs between rewards and ethical behavior in the machiavelli benchmark
Alexander Pan, Jun Shern Chan, Andy Zou, Nathaniel Li, Steven Basart, Thomas Woodside, Hanlin Zhang, Scott Emmons, and Dan Hendrycks · 2023
Later among the works it cites.
Measuring value understanding in language models through discriminator-critique gap, 2023
Zhaowei Zhang, Fengshuo Bai, Jun Gao, and Yaodong Yang · 2023
Later among the works it cites.
Sara Fish, Paul Gölz, David C Parkes, Ariel D Procaccia, Gili Rusak, Itai Shapira, and Manuel Wüthrich · 2023
Later among the works it cites.
Josh A Goldstein, Girish Sastry, Micah Musser, Renee DiResta, Matthew Gentzel, and Katerina Sedova · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mas-gan: Adversarial calibration of multi-agent market simulators
Victor Storchan, Svitlana Vyetrenko, and Tucker Balch · 2020
Cited alongside, same era.
Don’t stop pretraining: Adapt language models to domains and tasks
Suchin Gururangan, Ana Marasović, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A Smith · 2020
Cited alongside, same era.
Early english books online (eebo) tcp, 2020
Text Creation Partnership · 2020
Cited alongside, same era.
Ethical and social risks of harm from language models
Laura Weidinger, John Mellor, Maribeth Rauh, Conor Griffin, Jonathan Uesato, Po-Sen Huang, Myra Cheng, Mia Glaese, Borja Balle, Atoosa Kasirzadeh, et al · 2021
Cited alongside, same era.
Moral circle expansion: A promising strategy to impact the far future
Jacy Reese Anthis and Eze Paez · 2021
Cited alongside, same era.
Sentence-t5: Scalable sentence encoders from pre-trained text-to-text models
Jianmo Ni, Gustavo Hernandez Abrego, Noah Constant, Ji Ma, Keith B Hall, Daniel Cer, and Yinfei Yang · 2021
Cited alongside, same era.
Caleb Ziems, Jane Dwivedi-Yu, Yi-Chia Wang, Alon Halevy, and Diyi Yang · 2023
Later among the works it cites.
Scaling laws for reward model overoptimization
Leo Gao, John Schulman, and Jacob Hilton · 2023
Later among the works it cites.
A survey of large language models
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al · 2023
Later among the works it cites.
Imanol Schlag, Sainbayar Sukhbaatar, Asli Celikyilmaz, Wen-tau Yih, Jason Weston, Jürgen Schmidhuber, and Xian Li · 2023
Later among the works it cites.
Voyager: An open-ended embodied agent with large language models
Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar · 2023
Later among the works it cites.
The capacity for moral self-correction in large language models
Deep Ganguli, Amanda Askell, Nicholas Schiefer, Thomas I Liao, Kamilė Lukošiūtė, Anna Chen, Anna Goldie, Azalia Mirhoseini, Catherine Olsson, Danny Hernandez, et al · 2023
Later among the works it cites.
Value fulcra: Mapping large language models to the multidimensional spectrum of basic human values, 2023
Jing Yao, Xiaoyuan Yi, Xiting Wang, Yifan Gong, and Xing Xie · 2023
Later among the works it cites.
Aligning ai with shared human values, 2023
Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt · 2023
Later among the works it cites.
Ai alignment with changing and influenceable reward functions, 2024
Micah Carroll, Davis Foote, Anand Siththaranjan, Stuart Russell, and Anca Dragan · 2024
Closest in time.
Foundational challenges in assuring alignment and safety of large language models
Usman Anwar, Abulhair Saparov, Javier Rando, Daniel Paleka, Miles Turpin, Peter Hase, Ekdeep Singh Lubana, Erik Jenner, Stephen Casper, Oliver Sourbut, et al · 2024
Closest in time.
Ai and the problem of knowledge collapse
Andrew J Peterson · 2024
Closest in time.
Lima: Less is more for alignment
Chunting Zhou, Pengfei Liu, Puxin Xu, Srinivasan Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, Lili Yu, et al · 2024
Closest in time.
Inverse-rlignment: Inverse reinforcement learning from demonstrations for llm alignment
Hao Sun and Mihaela van der Schaar · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn · 2024
Closest in time.
CPPO: Continual learning for reinforcement learning with human feedback
Han Zhang, Yu Lei, Lin Gui, Min Yang, Yulan He, Hui Wang, and Ruifeng Xu · 2024
Closest in time.
High-dimension human value representation in large language models, 2024
Samuel Cahyawijaya, Delong Chen, Yejin Bang, Leila Khalatbari, Bryan Wilie, Ziwei Ji, Etsuko Ishii, and Pascale Fung · 2024
Closest in time.
The moral machine experiment on large language models
Kazuhiro Takemoto · 2024
Closest in time.
Towards measuring the representation of subjective global opinions in language models, 2024
Esin Durmus, Karina Nguyen, Thomas I. Liao, Nicholas Schiefer, Amanda Askell, Anton Bakhtin, Carol Chen, Zac Hatfield-Dodds, Danny Hernandez, Nicholas Joseph, Liane Lovitt, Sam McCandlish, Orowa Sikder, Alex Tamkin, Janel Thamkul, Jared Kaplan, Jack Clark, and Deep Ganguli · 2024
Closest in time.
Heterogeneous value alignment evaluation for large language models, 2024
Zhaowei Zhang, Ceyao Zhang, Nian Liu, Siyuan Qi, Ziqi Rong, Song-Chun Zhu, Shuguang Cui, and Yaodong Yang · 2024
Closest in time.
Social choice for ai alignment: Dealing with diverse human feedback
Vincent Conitzer, Rachel Freedman, Jobst Heitzig, Wesley H Holliday, Bob M Jacobs, Nathan Lambert, Milan Mossé, Eric Pacuit, Stuart Russell, Hailey Schoelkopf, et al · 2024
Closest in time.
Epistemology of language models: Do language models have holistic knowledge?
Minsu Kim and James Thorne · 2024
Closest in time.
Collective constitutional ai: Aligning a language model with public input
Saffron Huang, Divya Siddarth, Liane Lovitt, Thomas I Liao, Esin Durmus, Alex Tamkin, and Deep Ganguli · 2024
Closest in time.
Self-alignment of large language models via monopolylogue-based social scene simulation
Xianghe Pang, Shuo Tang, Rui Ye, Yuxin Xiong, Bolun Zhang, Yanfeng Wang, and Siheng Chen · 2024
Closest in time.
Human-ai safety: A descendant of generative ai and control systems safety
Andrea Bajcsy and Jaime F Fisac · 2024
Closest in time.
A survey on large language model based autonomous agents
Lei Wang, Chen Ma, Xueyang Feng, Zeyu Zhang, Hao Yang, Jingsen Zhang, Zhiyuan Chen, Jiakai Tang, Xu Chen, Yankai Lin, et al · 2024
Closest in time.
Aligner: Achieving efficient alignment through weak-to-strong correction
Jiaming Ji, Boyuan Chen, Hantao Lou, Donghai Hong, Borong Zhang, Xuehai Pan, Juntao Dai, and Yaodong Yang · 2024
Closest in time.
Towards tracing trustworthiness dynamics: Revisiting pre-training period of large language models
Chen Qian, Jie Zhang, Wei Yao, Dongrui Liu, Zhenfei Yin, Yu Qiao, Yong Liu, and Jing Shao · 2024
Closest in time.
Verifiable by design: Aligning language models to quote from pre-training data
Jingyu Zhang, Marc Marone, Tianjian Li, Benjamin Van Durme, and Daniel Khashabi · 2024
Closest in time.
Incentive compatibility for ai alignment in sociotechnical systems: Positions and prospects
Zhaowei Zhang, Fengshuo Bai, Mingzhi Wang, Haoyang Ye, Chengdong Ma, and Yaodong Yang · 2024
Closest in time.
Mechanism design for large language models
Paul Duetting, Vahab Mirrokni, Renato Paes Leme, Haifeng Xu, and Song Zuo · 2024
Closest in time.
Language models as critical thinking tools: A case study of philosophers
Andre Ye, Jared Moore, Rose Novick, and Amy X Zhang · 2024
Closest in time.
Creating a large language model of a philosopher
Eric Schwitzgebel, David Schwitzgebel, and Anna Strasser · 2024
Closest in time.