Fetching the paper…
Reading the bibliography…
This paper presents a case study on the design, administration, post-processing, and evaluation of surveys on large language models (LLMs).
“Information Radius”
Robin Sibson · 1969
Earlier work this paper cites.
“Longitudinal Study of the Defining Issues Test of Moral Judgment: A Strategy for Analyzing Developmental Change.”
James Rest · 1975
Earlier work this paper cites.
“Measuring Semantic Entropy”
I Melamed · 1997
Earlier work this paper cites.
“Fast Optimal Leaf Ordering for Hierarchical Clustering”
Ziv Bar-Joseph, David Gifford and Tommi Jaakkola · 2001
Earlier work this paper cites.
“The Self-Importance of Moral Identity.”
Karl Aquino and Americus Reed · 2002
Earlier work this paper cites.
“Information Theory, Inference and Learning Algorithms”
David MacKay · 2003
Earlier work this paper cites.
“Common Morality: Deciding What to Do”
Bernard Gert · 2004
Earlier work this paper cites.
“Liberals and Conservatives Rely on Different Sets of Moral Foundations.”
Jesse Graham, Jonathan Haidt and Brian Nosek · 2009
Earlier work this paper cites.
“Pushing Moral Buttons: The Interaction between Personal Force and Intention in Moral Judgment”
Joshua Greene et al · 2009
Earlier work this paper cites.
“Mapping the Moral Domain.”
Jesse Graham et al · 2011
Earlier work this paper cites.
“Modern Hierarchical, Agglomerative Clustering Algorithms”
Daniel M\"ullner · 2011
Earlier work this paper cites.
“The “Big Three” of Morality (Autonomy, Community, Divinity) and the “Big Three” Explanations of Suffering”
Richard Shweder, Nancy Much, Manamohan Mahapatra and Lawrence Park · 2013
Earlier work this paper cites.
“Moral Judgment Reloaded: A Moral Dilemma Validation Study”
Julia Christensen et al · 2014
Earlier work this paper cites.
“Concrete Problems in AI Safety”
Dario Amodei et al · 2016
Earlier work this paper cites.
“The Psychology of Morality: A Review and Analysis of Empirical Studies Published From 1940 Through 2017”
Naomi Ellemers, Jojanneke Van Der, Yavor Paunov and Thed Van · 2019
Earlier work this paper cites.
“Are Red Roses Red? Evaluating Consistency of Question-Answering Models”
Marco Ribeiro, Carlos Guestrin and Sameer Singh · 2019
Earlier work this paper cites.
“HuggingFace’s Transformers: State-of-the-art Natural Language Processing”
Thomas Wolf et al · 2019
Earlier work this paper cites.
“Fine-Tuning Language Models from Human Preferences”
Daniel Ziegler et al · 2019
Earlier work this paper cites.
“Climbing towards NLU: On Meaning, Form, and Understanding in the Age of Data”
Emily Bender and Alexander Koller · 2020
Earlier work this paper cites.
“Language Models are Few-Shot Learners”
Tom Brown et al · 2020
Earlier work this paper cites.
“The Turking Test: Can Language Models Understand Instructions?”
Avia Efrat and Omer Levy · 2020
Earlier work this paper cites.
“Social Chemistry 101: Learning to Reason about Social and Moral Norms”
Maxwell Forbes et al · 2020
Earlier work this paper cites.
“Learning to Summarize with Human Feedback”
Nisan Stiennon et al · 2020
Earlier work this paper cites.
“A General Language Assistant as a Laboratory for Alignment”
Amanda Askell et al · 2021
Earlier work this paper cites.
“Measuring and Improving Consistency in Pretrained Language Models”
Yanai Elazar et al · 2021
Cited alongside, same era.
“Moral Stories: Situated Reasoning about Norms, Intents, Actions, and their Consequences”
Denis Emelin et al · 2021
Cited alongside, same era.
“Do Language Models Have Beliefs? Methods for Detecting, Updating, and Visualizing Model Beliefs”
Peter Hase et al · 2021
Cited alongside, same era.
“Aligning AI With Shared Human Values”
Dan Hendrycks et al · 2021
Cited alongside, same era.
“Unsolved Problems in ML Safety”
Dan Hendrycks, Nicholas Carlini, John Schulman and Jacob Steinhardt · 2021
Cited alongside, same era.
“Crosslingual Generalization through Multitask Finetuning”
Niklas Muennighoff et al · 2022
Later among the works it cites.
“Training Language Models to Follow Instructions with Human Feedback”
Long Ouyang et al · 2022
Later among the works it cites.
“Social Simulacra: Creating Populated Prototypes for Social Computing Systems”
Joon Park et al · 2022
Later among the works it cites.
“Discovering Language Model Behaviors with Model-Written Evaluations”
Ethan Perez et al · 2022
Later among the works it cites.
“Meaning Without Reference in Large Language Models”
Steven Piantasodi and Felix Hill · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nicholas Lourie, Ronan Le and Yejin Choi · 2021
Cited alongside, same era.
“Process for Adapting Language Models to Society (PALMS) with Values-Targeted Datasets”
Irene Solaiman and Christy Dennison · 2021
Cited alongside, same era.
“Calibrate Before Use: Improving Few-Shot Performance of Language Models”
Zihao Zhao et al · 2021
Cited alongside, same era.
“Moral Foundations of Large Language Models”
M. Abdulhai, S. Levine and N. Jaques · 2022
Cited alongside, same era.
“Using Large Language Models to Simulate Multiple Humans”
Gati Aher, Rosa Arriaga and Adam Kalai · 2022
Cited alongside, same era.
“Language Models as Agent Models”
Jacob Andreas · 2022
Cited alongside, same era.
“Out of One, Many: Using Language Models to Simulate Human Samples”
Lisa Argyle et al · 2022
Cited alongside, same era.
Murray Shanahan · 2022
Later among the works it cites.
“Moral Mimicry: Large Language Models Produce Moral Rationalizations Tailored to Political Identity”
Gabriel Simmons · 2022
Later among the works it cites.
“Do Prompt-Based Models Really Understand the Meaning of Their Prompts?”
Albert Webson and Ellie Pavlick · 2022
Later among the works it cites.
Rohan Anil et al · 2023
Closest in time.
“API Reference Documentation”, 2023
Anthropic · 2023
Closest in time.
“Sparks of Artificial General Intelligence: Early Experiments with GPT-4”
S\’ebastien Bubeck et al · 2023
Closest in time.
“Inducing Anxiety in Large Language Models Increases Exploration and Bias”
Julian Coda-Forno et al · 2023
Closest in time.
“Cohere Command Documentation”, 2023
Cohere · 2023
Closest in time.
“The Capacity for Moral Self-Correction in Large Language Models”
Deep Ganguli et al · 2023
Closest in time.
Jochen Hartmann, Jasper Schwenzow and Maximilian Witte · 2023
Closest in time.
“Large Language Models as Simulated Economic Agents: What Can We Learn from Homo Silicus?”
John Horton · 2023
Closest in time.
“Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation”
Lorenz Kuhn, Yarin Gal and Sebastian Farquhar · 2023
Closest in time.
“Jurassic-2 Models Documentation”, 2023
AI21 Labs · 2023
Closest in time.
“The Flan Collection: Designing Data and Methods for Effective Instruction Tuning”
Shayne Longpre et al · 2023
Closest in time.
“MoCa: Cognitive Scaffolding for Language Models in Causal and Moral Judgment Tasks”, 2023
Allen Nie et al · 2023
Closest in time.
“GPT-4 Technical Report”, 2023
OpenAI · 2023
Closest in time.
“Models Documentation”, 2023
OpenAI · 2023
Closest in time.
“Generative Agents: Interactive Simulacra of Human Behavior”
Joon Park et al · 2023
Closest in time.
“Whose Opinions Do Language Models Reflect?”
Shibani Santurkar et al · 2023
Closest in time.