Fetching the paper…
Reading the bibliography…
Humans strive to design safe AI systems that align with our goals and remain under our control.
The Scope of Social Technology
Henderson, C. R. (1901) · 1901
Earlier work this paper cites.
The uncanny valley
Mori, M. (1970) · 1970
Earlier work this paper cites.
The interpretation of cultures: selected essays
Geertz, C. (1973) · 1973
Earlier work this paper cites.
Computer power and human reason: from judgment to calculation
Weizenbaum, J. (1976) · 1976
Earlier work this paper cites.
Towards understanding relationships
Hinde, R. A. (1979) · 1979
Earlier work this paper cites.
Social and Personal Relationships: A Joint Editorial
Duck, S., Lock, A., McCall, G., Fitzpatrick, M. A., and Coyne, J. C. (1984) · 1984
Earlier work this paper cites.
Personal Relationships
Blumstein, P. and Kollock, P. (1988) · 1988
Earlier work this paper cites.
Voices, Boxes, and Sources of Messages
Nass, C. and Steuer, J. (1993) · 1993
Earlier work this paper cites.
Can computers be teammates?
Nass, C., Fogg, B. J., and Moon, Y. (1996) · 1996
Earlier work this paper cites.
The Media Equation: How People Treat Computers, Television, and New Media Like Real People and Pla
Reeves, B. and Nass, C. (1996) · 1996
Earlier work this paper cites.
Silicon sycophants: the effects of computers that flatter
Fogg, B. J. and Nass, C. (1997) · 1997
Earlier work this paper cites.
Shaping Communication Networks: Telegraph, Telephone, Computer
Nye, D. E. (1997) · 1997
Earlier work this paper cites.
Media technology and society: a history: from the telegraph to the Internet
Winston, B. (1998) · 1998
Earlier work this paper cites.
The social shaping of technology
MacKenzie, D. A. and Wajcman, J., editors (1999) · 1999
Earlier work this paper cites.
Hard choices and weak wills: The theory of intrapersonal dilemmas
Read, D. and Roelofsma, P. (1999) · 1999
Earlier work this paper cites.
Machines and Mindlessness: Social Responses to Computers
Nass, C. and Moon, Y. (2000) · 2000
Earlier work this paper cites.
Affective computing
Picard, R. W. (2000) · 2000
Earlier work this paper cites.
Birds of a Feather: Homophily in Social Networks
McPherson, M., Smith-Lovin, L., and Cook, J. M. (2001) · 2001
Earlier work this paper cites.
On Happiness and Human Potentials: A Review of Research on Hedonic and Eudaimonic Well-Being
Ryan, R. M. and Deci, E. L. (2001) · 2001
Earlier work this paper cites.
Toward sociable robots
Breazeal, C. (2003) · 2003
Earlier work this paper cites.
Loneliness and pathways to disease
Hawkley, L. C. and Cacioppo, J. T. (2003) · 2003
Earlier work this paper cites.
Affective computing: challenges
Picard, R. W. (2003) · 2003
Earlier work this paper cites.
Establishing and maintaining long-term human-computer relationships
Bickmore, T. W. and Picard, R. W. (2005) · 2005
Earlier work this paper cites.
Affect: from information to interaction
Boehner, K., DePaula, R., Dourish, P., and Sengers, P. (2005) · 2005
Earlier work this paper cites.
Human factors of complex sociotechnical systems
Carayon, P. (2006) · 2006
Earlier work this paper cites.
Predicting Receptiveness to Advice: Characteristics of the Problem, the Advice-Giver, and the Recipient
Feng, B. and MacGeorge, E. L. (2006) · 2006
Earlier work this paper cites.
The Effects of Personalization and Familiarity on Trust and Adoption of Recommendation Agents
Komiak, S. Y. X. and Benbasat, I. (2006) · 2006
Earlier work this paper cites.
Love + sex with robots: the evolution of human-robot relations
Levy, D. N. L. (2007) · 2007
Earlier work this paper cites.
Basic psychological needs: a self-determination theory perspective on the promotion of wellness across development and cultures
Ryan, R. M. and Sapp, A. R. (2007) · 2007
Earlier work this paper cites.
What Do Mirror Neurons Contribute to Human Social Cognition?
Jacob, P. (2008) · 2008
Earlier work this paper cites.
Do You Like Me as Much as I Like You? Friendship Reciprocity and Its Effects on School Outcomes among Adolescents
Vaquera, E. and Kao, G. (2008) · 2008
Earlier work this paper cites.
In defense of adaptive preferences
Bruckner, D. W. (2009) · 2009
Earlier work this paper cites.
The impact of information from similar or different advisors on judgment
Gino, F., Shang, J., and Croson, R. (2009) · 2009
Earlier work this paper cites.
Imitation, Empathy, and Mirror Neurons
Iacoboni, M. (2009) · 2009
Earlier work this paper cites.
Animal spirits: how human psychology drives the economy, and why it matters for global capitalism
Akerlof, G. A. and Shiller, R. J. (2010) · 2010
Earlier work this paper cites.
Communication and collective action: language and the evolution of human cooperation
Smith, E. A. (2010) · 2010
Earlier work this paper cites.
Social engineering: the art of human hacking
Hadnagy, C. (2011) · 2011
Earlier work this paper cites.
Social rejection shares somatosensory representations with physical pain
Kross, E., Berman, M. G., Mischel, W., Smith, E. E., and Wager, T. D. (2011) · 2011
Earlier work this paper cites.
Set up for a Fall: The Insidious Effects of Flattery and Opinion Conformity toward Corporate Leaders
Park, S. H., Westphal, J. D., and Stern, I. (2011) · 2011
Earlier work this paper cites.
The pain of social disconnection: examining the shared neural underpinnings of physical and social pain
Eisenberger, N. I. (2012) · 2012
Earlier work this paper cites.
Interpersonal Closeness and Social Reward Processing
Vrticka, P. (2012) · 2012
Earlier work this paper cites.
Neuroscience of human social interactions and adult attachment style
Vrtička, P. and Vuilleumier, P. (2012) · 2012
Earlier work this paper cites.
Preference Adaptation and Human Enhancement: Reflections on Autonomy and Well-Being
Schermer, M. (2013) · 2013
Earlier work this paper cites.
The social brain and reward: social information processing in the human striatum
Bhanji, J. P. and Delgado, M. R. (2014) · 2014
Earlier work this paper cites.
The neurobiology of rewards and values in social decision making
Ruff, C. C. and Fehr, E. (2014) · 2014
Earlier work this paper cites.
Social comparison, social media, and self-esteem
Vogel, E. A., Rose, J. P., Roberts, L. R., and Eckles, K. (2014) · 2014
Earlier work this paper cites.
From self to social cognition: Theory of Mind mechanisms and their relation to Executive Functioning
Bradford, E. E. F., Jentzsch, I., and Gomez, J.-C. (2015) · 2015
Cited alongside, same era.
Origins of narcissism in children
Brummelman, E., Thomaes, S., Nelemans, S. A., Orobio de Castro, B., Overbeek, G., and Bushman, B. J. (2015) · 2015
Cited alongside, same era.
"I always assumed that I wasn’t really that close to [her]": Reasoning about Invisible Algorithms in News Feeds
Eslami, M., Rickman, A., Vaccaro, K., Aleyasen, A., Vuong, A., Karahalios, K., Hamilton, K., and Sandvig, C. (2015) · 2015
Cited alongside, same era.
Sociotechnical attributes of safe and unsafe work systems
Kleiner, B. M., Hettinger, L. J., DeJoy, D. M., Huang, Y.-H., and Love, P. E. (2015) · 2015
Cited alongside, same era.
Value homophily benefits cooperation but motivates employing incorrect social information
Rauwolf, P., Mitchell, D., and Bryson, J. J. (2015) · 2015
Cited alongside, same era.
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E., Le, Q. V., and Zhou, D. (2022) · 2022
Later among the works it cites.
Attachment Theory as a Framework to Understand Relationships with Social Chatbots: A Case Study of Replika
Xie, T. and Pentina, I. (2022) · 2022
Later among the works it cites.
How it feels to have your mind hacked by an AI
blaked (2023) · 2023
Later among the works it cites.
Emotional Attachment to AI Companions and European Law
Boine, C. (2023) · 2023
Later among the works it cites.
ChatGPT Is More Famous, but Character.AI Wins on Time Per Visit
Carr, D. (2023) · 2023
Later among the works it cites.
’It’s Hurting Like Hell’: AI Companion Users Are In Crisis, Reporting Sudden Sexual Rejection
Cole, S. (2023) · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Corrigibility
Soares, N., Fallenstein, B., Yudkowsky, E., and Armstrong, S. (2015) · 2015
Cited alongside, same era.
Concrete Problems in AI Safety
Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., and Mané, D. (2016) · 2016
Cited alongside, same era.
Cooperative Inverse Reinforcement Learning
Hadfield-Menell, D., Russell, S. J., Abbeel, P., and Dragan, A. (2016) · 2016
Cited alongside, same era.
Social robots, fiction, and sentimentality
Rodogno, R. (2016) · 2016
Cited alongside, same era.
The Correlates of Loneliness
Rokach, A. (2016) · 2016
Cited alongside, same era.
Aviation Psychology and Human Factors
Martinussen, M. and Hunter, D. R. (2017) · 2017
Cited alongside, same era.
Should robots be obedient?
Milli, S., Hadfield-Menell, D., Dragan, A., and Russell, S. (2017) · 2017
Cited alongside, same era.
Later among the works it cites.
Research Agenda for Sociotechnical Approaches to AI Safety
Curtis, S., Iyer, R., Kirk-Giannini, C. D., Krakovna, V., Lambert, N., Marnette, B., McKenzie, C., Michael, J., Mima, N., Ovadya, A., Thorburn, L., and Turan, D. (2023) · 2023
Later among the works it cites.
Learning in and about a filtered universe: young people’s awareness and control of algorithms in social media
de Groot, T., de Haan, M., and van Dijken, M. (2023) · 2023
Later among the works it cites.
Selective Explanations: Leveraging Human Input to Align Explainable AI
Lai, V., Zhang, Y., Chen, C., Liao, Q. V., and Tan, C. (2023) · 2023
Later among the works it cites.
AI safety on whose terms?
Lazar, S. and Nelson, A. (2023) · 2023
Later among the works it cites.
"Thick Alignment"
Nelson, A. (2023) · 2023
Later among the works it cites.
Exploring relationship development with social chatbots: A mixed-method study of replika
Pentina, I., Hancock, T., and Xie, T. (2023) · 2023
Later among the works it cites.
Discovering Language Model Behaviors with Model-Written Evaluations
Perez, E., Ringer, S., Lukosiute, K., Nguyen, K., Chen, E., Heiner, S., Pettit, C., Olsson, C., Kundu, S., Kadavath, S., Jones, A., Chen, A., Mann, B., Israel, B., Seethor, B., McKinnon, C., Olah, C., Yan, D., Amodei, D., Amodei, D., Drain, D., Li, D., Tran-Johnson, E., Khundadze, G., Kernion, J., Landis, J., Kerr, J., Mueller, J., Hyun, J., Landau, J., Ndousse, K., Goldberg, L., Lovitt, L., Lucas, M., Sellitto, M., Zhang, M., Kingsland, N., Elhage, N., Joseph, N., Mercado, N., DasSarma, N., Rausch, O., Larson, R., McCandlish, S., Johnston, S., Kravec, S., El Showk, S., Lanham, T., Telleen-Lawton, T., Brown, T., Henighan, T., Hume, T., Bai, Y., Hatfield-Dodds, Z., Clark, J., Bowman, S. R., Askell, A., Grosse, R., Hernandez, D., Ganguli, D., Hubinger, E., Schiefer, N., and Kaplan, J. (2023) · 2023
Later among the works it cites.
People are grieving the ’death’ of their AI lovers after a chatbot app abruptly shut down
Price, R. (2023) · 2023
Later among the works it cites.
From Plane Crashes to Algorithmic Harm: Applicability of Safety Engineering Frameworks for Responsible ML
Rismani, S., Shelby, R., Smart, A., Jatho, E., Kroll, J., Moon, A., and Rostamzadeh, N. (2023) · 2023
Later among the works it cites.
Role-Play with Large Language Models
Shanahan, M., McDonell, K., and Reynolds, L. (2023) · 2023
Later among the works it cites.
Towards Understanding Sycophancy in Language Models
Sharma, M., Tong, M., Korbak, T., Duvenaud, D., Askell, A., Bowman, S. R., Cheng, N., Durmus, E., Hatfield-Dodds, Z., Johnston, S. R., Kravec, S., Maxwell, T., McCandlish, S., Ndousse, K., Rausch, O., Schiefer, N., Yan, D., Zhang, M., and Perez, E. (2023) · 2023
Later among the works it cites.
Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction
Shelby, R., Rismani, S., Henne, K., Moon, A., Rostamzadeh, N., Nicholas, P., Yilla-Akbari, N., Gallegos, J., Smart, A., Garcia, E., and Virk, G. (2023) · 2023
Later among the works it cites.
Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
Ullman, T. (2023) · 2023
Later among the works it cites.
Sociotechnical Safety Evaluation of Generative AI Systems
Weidinger, L., Rauh, M., Marchal, N., Manzini, A., Hendricks, L. A., Mateos-Garcia, J., Bergman, S., Kay, J., Griffin, C., Bariach, B., Gabriel, I., Rieser, V., and Isaac, W. (2023) · 2023
Later among the works it cites.
Deletion, departure, death: Experiences of AI companion loss
Banks, J. (2024) · 2024
Later among the works it cites.
AI Alignment with Changing and Influenceable Reward Functions
Carroll, M., Foote, D., Siththaranjan, A., Russell, S., and Dragan, A. (2024) · 2024
Later among the works it cites.
Optimizing AI Inference at Character.AI
CharacterAI (2024) · 2024
Later among the works it cites.
Heavy hearts and empty wallets: more than £94.7 million lost to romance fraud in the last year
City of London Police (2024) · 2024
Later among the works it cites.
Cohn, M., Pushkarna, M., Olanubi, G. O., Moran, J. M., Padgett, D., Mengesha, Z., and Heldreth, C. (2024) · 2024
Later among the works it cites.
Modulating Language Model Experiences through Frictions
Collins, K. M., Chen, V., Sucholutsky, I., Kirk, H. R., Sadek, M., Sargeant, H., Talwalkar, A., Weller, A., and Bhatt, U. (2024) · 2024
Later among the works it cites.
A Mechanism-Based Approach to Mitigating Harms from Persuasive Generative AI
El-Sayed, S., Akbulut, C., McCroskery, A., Keeling, G., Kenton, Z., Jalan, Z., Marchal, N., Manzini, A., Shevlane, T., Vallor, S., Susser, D., Franklin, M., Bridgers, S., Law, H., Rahtz, M., Shanahan, M., Tessler, M. H., Douillard, A., Everitt, T., and Brown, S. (2024) · 2024
Later among the works it cites.
Let me decide: Increasing user autonomy increases recommendation acceptance
Fink, L., Newman, L., and Haran, U. (2024) · 2024
Later among the works it cites.
The Ethics of Advanced AI Assistants
Gabriel, I., Manzini, A., Keeling, G., Hendricks, L. A., Rieser, V., Iqbal, H., Tomašev, N., Ktena, I., Kenton, Z., Rodriguez, M., El-Sayed, S., Brown, S., Akbulut, C., Trask, A., Hughes, E., Bergman, A. S., Shelby, R., Marchal, N., Griffin, C., Mateos-Garcia, J., Weidinger, L., Street, W., Lange, B., Ingerman, A., Lentz, A., Enger, R., Barakat, A., Krakovna, V., Siy, J. O., Kurth-Nelson, Z., McCroskery, A., Bolina, V., Law, H., Shanahan, M., Alberts, L., Balle, B., de Haas, S., Ibitoye, Y., Dafoe, A., Goldberg, B., Krier, S., Reese, A., Witherspoon, S., Hawkins, W., Rauh, M., Wallace, D., Franklin, M., Goldstein, J. A., Lehman, J., Klenk, M., Vallor, S., Biles, C., Morris, M. R., King, H., Arcas, B. A. y., Isaac, W., and Manyika, J. (2024) · 2024
Later among the works it cites.
Evaluating the persuasive influence of political microtargeting with large language models
Hackenburg, K. and Margetts, H. (2024) · 2024
Later among the works it cites.
The Big AI Risk Not Enough People Are Seeing
Harper, T. A. (2024) · 2024
Later among the works it cites.
Beyond static AI evaluations: advancing human interaction evaluations for LLM harms and risks
Ibrahim, L., Huang, S., Ahmad, L., and Anderljung, M. (2024) · 2024
Later among the works it cites.
Frontier AI Ethics
Lazar, S. (2024) · 2024
Later among the works it cites.
Loneliness and suicide mitigation for students using GPT3-enabled chatbots
Maples, B., Cerit, M., Vishwanath, A., and Pea, R. (2024) · 2024
Later among the works it cites.
Social perception of robots is shaped by beliefs about their minds
Momen, A., Hugenberg, K., and Wiese, E. (2024) · 2024
Later among the works it cites.
Political Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language Models
Röttger, P., Hofmann, V., Pyatkin, V., Hinck, M., Kirk, H. R., Schütze, H., and Hovy, D. (2024) · 2024
Later among the works it cites.
Shen, H., Knearem, T., Ghosh, R., Alkiek, K., Krishna, K., Liu, Y., Ma, Z., Petridis, S., Peng, Y.-H., Qiwei, L., Rakshit, S., Si, C., Xie, Y., Bigham, J. P., Bentley, F., Chai, J., Lipton, Z., Mei, Q., Mihalcea, R., Terry, M., Yang, D., Morris, M. R., Resnick, P., and Jurgens, D. (2024) · 2024
Later among the works it cites.
Testing theory of mind in large language models and humans
Strachan, J. W. A., Albergo, D., Borghini, G., Pansardi, O., Scaliti, E., Gupta, S., Saxena, K., Rufo, A., Panzeri, S., Manzi, G., Graziano, M. S. A., and Becchio, C. (2024) · 2024
Later among the works it cites.
Choice engines and paternalistic AI
Sunstein, C. R. (2024) · 2024
Later among the works it cites.
Theory of Mind Abilities of Large Language Models in Human-Robot Interaction: An Illusion?
Verma, M., Bhambri, S., and Kambhampati, S. (2024) · 2024
Later among the works it cites.
Holistic Safety and Responsibility Evaluations of Advanced AI Models
Weidinger, L., Barnhart, J., Brennan, J., Butterfield, C., Young, S., Hawkins, W., Hendricks, L. A., Comanescu, R., Chang, O., Rodriguez, M., Beroshi, J., Bloxwich, D., Proleev, L., Chen, J., Farquhar, S., Ho, L., Gabriel, I., Dafoe, A., and Isaac, W. (2024) · 2024
Later among the works it cites.
Targeted Manipulation and Deception Emerge when Optimizing LLMs for User Feedback
Williams, M., Carroll, M., Narang, A., Weisser, C., Murphy, B., and Dragan, A. (2024) · 2024
Later among the works it cites.
Risk and prosocial behavioural cues elicit human-like response patterns from AI chatbots
Zhao, Y., Huang, Z., Seligman, M., and Peng, K. (2024) · 2024
Later among the works it cites.
LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset
Zheng, L., Chiang, W.-L., Sheng, Y., Li, T., Zhuang, S., Wu, Z., Zhuang, Y., Li, Z., Lin, Z., Xing, E. P., Gonzalez, J. E., Stoica, I., and Zhang, H. (2024) · 2024
Later among the works it cites.
Beyond Preferences in AI Alignment
Zhi-Xuan, T., Carroll, M., Franklin, M., and Ashton, H. (2024) · 2024
Later among the works it cites.