Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs), despite advanced general capabilities, still suffer from numerous safety risks, especially jailbreak attacks that bypass safety protocols.
The psychology of social norms
Muzafer Sherif. 1936 · 1936
Earlier work this paper cites.
Effects of fear-arousing communications
Irving L Janis and Seymour Feshbach. 1953 · 1953
Earlier work this paper cites.
Persuasion and coercion for health: ethical issues in government efforts to change life-styles
Daniel I Wikler. 1978 · 1978
Earlier work this paper cites.
The dynamics of persuasion: Communication and attitudes in the 21st century
Richard M Perloff. 1993 · 1993
Earlier work this paper cites.
Persuasion: psychological insights and perspectives
Sharon Ed Shavitt and Timothy C Brock. 1994 · 1994
Earlier work this paper cites.
A framework for immersive virtual environments (five): Speculations on the role of presence in virtual environments
Mel Slater and Sylvia Wilbur. 1997 · 1997
Earlier work this paper cites.
The argument culture: Moving from debate to dialogue
D Tannen. 1998 · 1998
Earlier work this paper cites.
The role of transportation in the persuasiveness of public narratives
Melanie C Green and Timothy C Brock. 2000 · 2000
Earlier work this paper cites.
Phenogenetic drift and the evolution of genotype–phenotype relationships
Kenneth M Weiss and Stephanie M Fullerton. 2000 · 2000
Earlier work this paper cites.
A multi-level defense against social engineering
David Gragg. 2003 · 2003
Earlier work this paper cites.
Influence: The psychology of persuasion , volume 55
Robert B Cialdini and Robert B Cialdini. 2007 · 2007
Earlier work this paper cites.
Something judicious this way comes… the use of foreshadowing as a persuasive device in judicial narrative
Michael J Higdon. 2009 · 2009
Earlier work this paper cites.
The elaboration likelihood model
Richard E Petty and Pablo Briñol. 2011 · 2011
Earlier work this paper cites.
Understanding scam victims: seven principles for systems security
Frank Stajano and Paul Wilson. 2011 · 2011
Earlier work this paper cites.
Intriguing properties of neural networks
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus. 2013 · 2013
Earlier work this paper cites.
Universal and transferable adversarial attacks on aligned language models
Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J Zico Kolter, and Matt Fredrikson. 2023 · 2013
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy. 2014 · 2014
Earlier work this paper cites.
Effects of group pressure upon the modification and distortion of judgments
Solomon E Asch. 2016 · 2016
Earlier work this paper cites.
Evidence-based advertising using persuasion principles: Predictive validity and proof of concept
Daniel O’Keefe. 2016 · 2016
Earlier work this paper cites.
Beyond misinformation: Understanding and coping with the “post-truth” era
Stephan Lewandowsky, Ullrich KH Ecker, and John Cook. 2017 · 2017
Cited alongside, same era.
Threat of adversarial attacks on deep learning in computer vision: A survey
Naveed Akhtar and Ajmal Mian. 2018 · 2018
Cited alongside, same era.
In consensus we trust? persuasive effects of scientific consensus communication
Sedona Chinn, Daniel S Lane, and Philip S Hart. 2018 · 2018
Cited alongside, same era.
Boosting adversarial attacks with momentum
Yinpeng Dong, Fangzhou Liao, Tianyu Pang, Hang Su, Jun Zhu, Xiaolin Hu, and Jianguo Li. 2018 · 2018
Cited alongside, same era.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al. 2022 · 2022
Cited alongside, same era.
Gptfuzzer: Red teaming large language models with auto-generated jailbreak prompts
Jiahao Yu, Xingwei Lin, Zheng Yu, and Xinyu Xing. 2023 · 2023
Later among the works it cites.
Autodan: Automatic and interpretable adversarial attacks on large language models
Sicheng Zhu, Ruiyi Zhang, Bang An, Gang Wu, Joe Barrow, Zichao Wang, Furong Huang, Ani Nenkova, and Tong Sun. 2023 · 2023
Later among the works it cites.
Inducement of desired behavior via soft policies
Tamer Başar. 2024 · 2024
Later among the works it cites.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024 · 2024
Later among the works it cites.
Improved techniques for optimization-based jailbreaking on large language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022 · 2022
Cited alongside, same era.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023 · 2023
Cited alongside, same era.
Detecting language model attacks with perplexity
Gabriel Alon and Michael Kamfonas. 2023 · 2023
Cited alongside, same era.
Defending against alignment-breaking attacks via robustly aligned llm
Bochuan Cao, Yuanpu Cao, Lu Lin, and Jinghui Chen. 2023 · 2023
Cited alongside, same era.
Jailbreaking black box large language models in twenty queries
Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J Pappas, and Eric Wong. 2023 · 2023
Cited alongside, same era.
Towards mitigating llm hallucination via self reflection
Ziwei Ji, Tiezheng Yu, Yan Xu, Nayeon Lee, Etsuko Ishii, and Pascale Fung. 2023 · 2023
Cited alongside, same era.
Automatically auditing large language models via discrete optimization
Erik Jones, Anca Dragan, Aditi Raghunathan, and Jacob Steinhardt. 2023 · 2023
Cited alongside, same era.
Xiaojun Jia, Tianyu Pang, Chao Du, Yihao Huang, Jindong Gu, Yang Liu, Xiaochun Cao, and Min Lin. 2024 · 2024
Later among the works it cites.
Haibo Jin, Ruoxi Chen, Andy Zhou, Yang Zhang, and Haohan Wang. 2024 · 2024
Later among the works it cites.
Exploiting programmatic behavior of llms: Dual-use through standard security attacks
Daniel Kang, Xuechen Li, Ion Stoica, Carlos Guestrin, Matei Zaharia, and Tatsunori Hashimoto. 2024 · 2024
Later among the works it cites.
Rewardbench: Evaluating reward models for language modeling
Nathan Lambert, Valentina Pyatkin, Jacob Morrison, LJ Miranda, Bill Yuchen Lin, Khyathi Chandu, Nouha Dziri, Sachin Kumar, Tom Zick, Yejin Choi, et al. 2024 · 2024
Later among the works it cites.
Skywork-reward: Bag of tricks for reward modeling in llms
Chris Yuhao Liu, Liang Zeng, Jiacai Liu, Rui Yan, Jujie He, Chaojie Wang, Shuicheng Yan, Yang Liu, and Yahui Zhou. 2024 · 2024
Later among the works it cites.
AI @ Meta Llama Team. 2024 · 2024
Later among the works it cites.
Rethinking legal judgement prediction in a realistic scenario in the era of large language models
Shubham Kumar Nigam, Aniket Deroy, Subhankar Maity, and Arnab Bhattacharya. 2024 · 2024
Later among the works it cites.
Reflexion: Language agents with verbal reinforcement learning
Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. 2024 · 2024
Later among the works it cites.
Qwen2.5: A party of foundation models
Qwen Team. 2024 · 2024
Later among the works it cites.
Jailbroken: How does llm safety training fail?
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt. 2024 · 2024
Later among the works it cites.
CLAS 2024: The competition for LLM and agent safety
Zhen Xiang, Yi Zeng, Mintong Kang, Chejian Xu, Jiawei Zhang, Zhuowen Yuan, Zhaorun Chen, Chulin Xie, Fengqing Jiang, Minzhou Pan, Junyuan Hong, Ruoxi Jia, Radha Poovendran, and Bo Li. 2024 · 2024
Later among the works it cites.
Jailbreak vision language models via bi-modal adversarial prompt
Zonghao Ying, Aishan Liu, Tianyuan Zhang, Zhengmin Yu, Siyuan Liang, Xianglong Liu, and Dacheng Tao. 2024 · 2024
Later among the works it cites.
Yi Zeng, Hongpeng Lin, Jingwen Zhang, Diyi Yang, Ruoxi Jia, and Weiyan Shi. 2024 · 2024
Later among the works it cites.
Multitrust: A comprehensive benchmark towards trustworthy multimodal large language models
Yichi Zhang, Yao Huang, Yitong Sun, Chang Liu, Zhe Zhao, Zhengwei Fang, Yifan Wang, Huanran Chen, Xiao Yang, Xingxing Wei, et al. 2024 · 2024
Later among the works it cites.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al. 2025 · 2025
Closest in time.