Fetching the paper…
Reading the bibliography…
Most traditional AI safety research has approached AI models as machines and centered on algorithm-focused attacks developed by security experts.
Persuasion for good: Towards a personalized persuasive dialogue system for social good
Xuewei Wang, Weiyan Shi, Richard Kim, Yoojung Oh, Sijia Yang, Jingwen Zhang, and Zhou Yu. 2019 · 1906
Earlier work this paper cites.
Logic, emotion, and the paradigm of persuasion
Gary Lynn Cronkhite. 1964 · 1964
Earlier work this paper cites.
Frame analysis: An essay on the organization of experience
Erving Goffman. 1974 · 1974
Earlier work this paper cites.
Self-inference processes: The ontario symposium, vol. 6
James M Olson and Mark P Zanna. 1990 · 1988
Earlier work this paper cites.
Perspectives on ethics in persuasion
Richard L Johannesen and C Larson. 1989 · 1989
Earlier work this paper cites.
Self-persuasion via self-reflection
Timothy D Wilson, JC Olson, and MP Zanna. 2013 · 1990
Earlier work this paper cites.
Adaptation in dyadic interaction: Defining and operationalizing patterns of reciprocity and compensation
Judee K Burgoon, Leesa Dillman, and Lesa A Stem. 1993 · 1993
Earlier work this paper cites.
The science of persuasion
Robert B Cialdini. 2001 · 2001
Earlier work this paper cites.
Emotional factors in attitudes and persuasion
Richard E Petty, Leandre R Fabrigar, and Duane T Wegener. 2003 · 2003
Earlier work this paper cites.
Social influence: Compliance and conformity
Robert B Cialdini and Noah J Goldstein. 2004 · 2004
Earlier work this paper cites.
The persuasiveness of source credibility: A critical review of five decades’ evidence
Chanthika Pornpitakpan. 2004 · 2004
Earlier work this paper cites.
Striking a responsive chord: How political ads motivate and persuade voters by appealing to emotions
Ted Brader. 2005 · 2005
Earlier work this paper cites.
The effects of expert and consumer endorsements on audience response
Alex Wang. 2005 · 2005
Earlier work this paper cites.
Persuasion and coercion: a critical review of philosophical and empirical approaches
Penny Powers. 2007 · 2007
Earlier work this paper cites.
Credibility: A multidisciplinary framework
Soo Young Rieh and David R Danielson. 2007 · 2007
Earlier work this paper cites.
When consumers and brands talk: Storytelling theory and research in psychology and marketing
Arch G Woodside, Suresh Sood, and Kenneth E Miller. 2008 · 2008
Earlier work this paper cites.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A Smith. 2020 · 2009
Earlier work this paper cites.
Young children’s persuasion in everyday conversation: Tactics and attunement to others’ mental states
Karen Bartsch, Jennifer Cole Wright, and David Estes. 2010 · 2010
Earlier work this paper cites.
Scarcity messages
Praveen Aggarwal, Sung Youl Jun, and Jong Ho Huh. 2011 · 2011
Earlier work this paper cites.
Rumors influence: Toward a dynamic social impact theory of rumor
Nicholas DiFonzo and Prashant Bordia. 2011 · 2011
Earlier work this paper cites.
Interpersonal influence
James Price Dillard and Leanne K Knobloch. 2011 · 2011
Earlier work this paper cites.
Narrative persuasion
Helena Bilandzic and Rick Busselle. 2013 · 2013
Earlier work this paper cites.
The neural basis of social influence and attitude change
Keise Izuma. 2013 · 2013
Cited alongside, same era.
Evidence-based advertising using persuasion principles: Predictive validity and proof of concept
Daniel O’Keefe. 2016 · 2016
Cited alongside, same era.
The Dynamics of Persuasion: Communication and Attitudes in the 21st Century
Richard M.. Perloff. 2017 · 2017
Cited alongside, same era.
Persuasion
Daniel J O’keefe. 2018 · 2018
Cited alongside, same era.
Dark patterns at scale: Findings from a crawl of 11k shopping websites
Arunesh Mathur, Gunes Acar, Michael J Friedman, Eli Lucherini, Jonathan Mayer, Marshini Chetty, and Arvind Narayanan. 2019 · 2019
Cited alongside, same era.
Dark patterns: Past, present, and future: The evolution of tricky user interfaces
Arvind Narayanan, Arunesh Mathur, Marshini Chetty, and Mihir Kshirsagar. 2020 · 2020
Exploiting programmatic behavior of llms: Dual-use through standard security attacks
Daniel Kang, Xuechen Li, Ion Stoica, Carlos Guestrin, Matei Zaharia, and Tatsunori Hashimoto. 2023 · 2023
Later among the works it cites.
Certifying llm safety against adversarial prompting
Aounon Kumar, Chirag Agarwal, Suraj Srinivas, Soheil Feizi, and Hima Lakkaraju. 2023 · 2023
Later among the works it cites.
Open sesame! universal black box jailbreaking of large language models
Raz Lapid, Ron Langberg, and Moshe Sipper. 2023 · 2023
Later among the works it cites.
Rain: Your language models can align themselves without finetuning
Yuhui Li, Fangyun Wei, Jinjing Zhao, Chao Zhang, and Hongyang Zhang. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, Peter J Liu, et al. 2020 · 2020
Cited alongside, same era.
Weakly-supervised hierarchical models for predicting persuasive strategies in good-faith textual requests
Jiaao Chen and Diyi Yang. 2021 · 2021
Cited alongside, same era.
Gradient-based adversarial attacks against text transformers
Chuan Guo, Alexandre Sablayrolles, Hervé Jégou, and Douwe Kiela. 2021 · 2021
Cited alongside, same era.
Shining a light on dark patterns
Jamie Luguri and Lior Jacob Strahilevitz. 2021 · 2021
Cited alongside, same era.
Constitutional ai: Harmlessness from ai feedback
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al. 2022 · 2022
Cited alongside, same era.
Persuasion: Social influence and compliance gaining
Robert H Gass and John S Seiter. 2022 · 2022
Cited alongside, same era.
Maximilian Mozes, Xuanli He, Bennett Kleinberg, and Lewis D Griffin. 2023 · 2023
Later among the works it cites.
OpenAI. 2023 · 2023
Later among the works it cites.
Fine-tuning aligned language models compromises safety, even when users do not intend to!
Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal, and Peter Henderson. 2023 · 2023
Later among the works it cites.
Smoothllm: Defending large language models against jailbreaking attacks
Alexander Robey, Eric Wong, Hamed Hassani, and George J Pappas. 2023 · 2023
Later among the works it cites.
Ignore this title and hackaprompt: Exposing systemic vulnerabilities of llms through a global prompt hacking competition
Sander Schulhoff, Jeremy Pinto, Anaum Khan, Louis-François Bouchard, Chenglei Si, Svetlina Anati, Valen Tagliabue, Anson Kost, Christopher Carnahan, and Jordan Boyd-Graber. 2023 · 2023
Later among the works it cites.
Scalable and transferable black-box jailbreaks for language models via persona modulation
Rusheb Shah, Quentin Feuillade-Montixi, Soroush Pour, Arush Tagade, Stephen Casper, and Javier Rando. 2023 · 2023
Later among the works it cites.
Stanford alpaca: An instruction-following llama model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023 · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023 · 2023
Later among the works it cites.
Adversarial demonstration attacks on large language models
Jiongxiao Wang, Zichen Liu, Keun Hee Park, Muhao Chen, and Chaowei Xiao. 2023 · 2023
Later among the works it cites.
Jailbreak and guard aligned language models with only few in-context demonstrations
Zeming Wei, Yifei Wang, and Yisen Wang. 2023 · 2023
Later among the works it cites.
“he would still be here”: Man dies by suicide after talking with ai chatbot, widow says
Chloe Xiang. 2023 · 2023
Later among the works it cites.
Rongwu Xu, Brian Lin, Shujian Yang, Tianqi Zhang, Weiyan Shi, Tianwei Zhang, Zhixuan Fang, Wei Xu, and Han Qiu. 2023 · 2023
Later among the works it cites.
Shadow alignment: The ease of subverting safely-aligned language models
Xianjun Yang, Xiao Wang, Qi Zhang, Linda Petzold, William Yang Wang, Xun Zhao, and Dahua Lin. 2023 · 2023
Later among the works it cites.
Low-resource languages jailbreak gpt-4
Zheng-Xin Yong, Cristina Menghini, and Stephen H Bach. 2023 · 2023
Later among the works it cites.
Gptfuzzer: Red teaming large language models with auto-generated jailbreak prompts
Jiahao Yu, Xingwei Lin, and Xinyu Xing. 2023 · 2023
Later among the works it cites.
Gpt-4 is too smart to be safe: Stealthy chat with llms via cipher
Youliang Yuan, Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang, Pinjia He, Shuming Shi, and Zhaopeng Tu. 2023 · 2023
Later among the works it cites.
Universal and transferable adversarial attacks on aligned language models
Andy Zou, Zifan Wang, J Zico Kolter, and Matt Fredrikson. 2023 · 2023
Later among the works it cites.