Fetching the paper…
Reading the bibliography…
As large language models (LLMs) have been deployed in various real-world settings, concerns about the harm they may propagate have grown.
Judgment under uncertainty: Heuristics and biases: Biases in judgments reveal some heuristics of thinking under uncertainty
Amos Tversky and Daniel Kahneman. 1974 · 1974
Earlier work this paper cites.
Undiscovered public knowledge
Don R Swanson. 1986 · 1986
Earlier work this paper cites.
Stereotype accuracy: Toward appreciating group differences
Clark R McCauley, Lee J Jussim, and Yueh-Ting Lee. 1995 · 1995
Earlier work this paper cites.
Advantages of bias and prejudice: An exploration of their neurocognitive templates
A Tobena, I Marks, and R Dar. 1999 · 1999
Earlier work this paper cites.
Stereotype performance boosts: the impact of self-relevance and the manner of stereotype activation
Margaret Shih, Nalini Ambady, Jennifer A Richeson, Kentaro Fujita, and Heather M Gray. 2002 · 2002
Earlier work this paper cites.
Stereoset: Measuring stereotypical bias in pretrained language models
Moin Nadeem, Anna Bethke, and Siva Reddy. 2020 · 2004
Earlier work this paper cites.
Influence: The psychology of persuasion , volume 55
Robert B Cialdini. 2007 · 2007
Earlier work this paper cites.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A Smith. 2020 · 2009
Earlier work this paper cites.
Social psychology of stereotyping and human difference appreciation
Yueh-Ting Lee. 2011 · 2011
Earlier work this paper cites.
Stereotype boost: Positive outcomes from the activation of positive stereotypes
Margaret J Shih, Todd L Pittinsky, and Geoffrey C Ho. 2012 · 2012
Earlier work this paper cites.
Positive stereotypes are pervasive and powerful
Alexander M Czopp, Aaron C Kay, and Sapna Cheryan. 2015 · 2015
Earlier work this paper cites.
Implicit stereotypes and the predictive brain: cognition and culture in “biased” person perception
Perry Hinton. 2017 · 2017
Earlier work this paper cites.
Inducing document structure for aspect-based summarization
Lea Frermann and Alexandre Klementiev. 2019 · 2019
Earlier work this paper cites.
On the dangers of stochastic parrots: Can language models be too big?" "" "
Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021 · 2021
Earlier work this paper cites.
Baco: A background knowledge-and content-based framework for citing sentence generation
Yubin Ge, Ly Dinh, Xiaofeng Liu, Jinsong Su, Ziyao Lu, Ante Wang, and Jana Diesner. 2021 · 2021
Earlier work this paper cites.
Multi-document summarization via deep learning techniques: A survey
Congbo Ma, Wei Emma Zhang, Mingyu Guo, Hu Wang, and Quan Z Sheng. 2022 · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022 · 2022
Cited alongside, same era.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023 · 2023
Cited alongside, same era.
Guardrails for trust, safety, and ethical development and deployment of large language models (llm)
Anjanava Biswas and Wrick Talukdar. 2023 · 2023
Cited alongside, same era.
Jailbreaking black box large language models in twenty queries
Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J Pappas, and Eric Wong. 2023 · 2023
Cited alongside, same era.
Safe rlhf: Safe reinforcement learning from human feedback
Josef Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji, Xinbo Xu, Mickel Liu, Yizhou Wang, and Yaodong Yang. 2023 · 2023
Learning from red teaming: Gender bias provocation and mitigation in large language models
Hsuan Su, Cheng-Chu Cheng, Hua Farn, Shachi H. Kumar, Saurav Sahay, Shang-Tse Chen, and Hung yi Lee. 2023 · 2023
Later among the works it cites.
Gemini: a family of highly capable multimodal models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, Katie Millican, et al. 2023 · 2023
Later among the works it cites.
Shadow alignment: The ease of subverting safely-aligned language models
Xianjun Yang, Xiao Wang, Qi Zhang, Linda Ruth Petzold, William Yang Wang, Xun Zhao, and Dahua Lin. 2023 · 2023
Later among the works it cites.
Universal and transferable adversarial attacks on aligned language models
Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J Zico Kolter, and Matt Fredrikson. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Toxicity in chatgpt: Analyzing persona-assigned language models
A. Deshpande, Vishvak Murahari, Tanmay Rajpurohit, A. Kalyan, and Karthik Narasimhan. 2023 · 2023
Cited alongside, same era.
Robbie: Robust bias evaluation of large generative language models
David Esiobu, Xiaoqing Ellen Tan, Saghar Hosseini, Megan Ung, Yuchen Zhang, Jude Fernandes, Jane Dwivedi-Yu, Eleonora Presani, Adina Williams, and Eric Michael Smith. 2023 · 2023
Cited alongside, same era.
Bias runs deep: Implicit reasoning biases in persona-assigned llms
Shashank Gupta, Vaishnavi Shrivastava, A. Deshpande, A. Kalyan, Peter Clark, Ashish Sabharwal, and Tushar Khot. 2023 · 2023
Cited alongside, same era.
Baseline defenses for adversarial attacks against aligned language models
Neel Jain, Avi Schwarzschild, Yuxin Wen, Gowthami Somepalli, John Kirchenbauer, Ping-yeh Chiang, Micah Goldblum, Aniruddha Saha, Jonas Geiping, and Tom Goldstein. 2023 · 2023
Cited alongside, same era.
StereoMap: Quantifying the awareness of human-like stereotypes in large language models
Sullam Jeoung, Yubin Ge, and Jana Diesner. 2023 · 2023
Cited alongside, same era.
Exploring the landscape of automatic text summarization: a comprehensive survey
Bilal Khan, Zohaib Ali Shah, Muhammad Usman, Inayat Khan, and Badam Niazi. 2023 · 2023
Cited alongside, same era.
Gender bias and stereotypes in large language models
Hadas Kotek, Rikker Dockum, and David Sun. 2023 · 2023
Cited alongside, same era.
Masterkey: Automated jailbreaking of large language model chatbots
Gelei Deng, Yi Liu, Yuekang Li, Kailong Wang, Ying Zhang, Zefeng Li, Haoyu Wang, Tianwei Zhang, and Yang Liu. 2023a · 2024
Later among the works it cites.
Masterkey: Automated jailbreaking of large language model chatbots
Gelei Deng, Yi Liu, Yuekang Li, Kailong Wang, Ying Zhang, Zefeng Li, Haoyu Wang, Tianwei Zhang, and Yang Liu. 2024 · 2024
Later among the works it cites.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024 · 2024
Later among the works it cites.
Stereotype: Cognition and biases
Nitya Ann Eapen. 2024 · 2024
Later among the works it cites.
Controllable citation sentence generation with language models
Nianlong Gu and Richard Hahnloser. 2024 · 2024
Later among the works it cites.
Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, et al. 2024 · 2024
Later among the works it cites.
Decoding biases: Automated methods and llm judges for gender bias detection in language models
Shachi H Kumar, Saurav Sahay, Sahisnu Mazumder, Eda Okur, Ramesh Manuvinakurike, Nicole Beckage, Hsuan Su, Hung-yi Lee, and Lama Nachman. 2024 · 2024
Later among the works it cites.
Great, now write an article about that: The crescendo multi-turn llm jailbreak attack
Mark Russinovich, Ahmed Salem, and Ronen Eldan. 2024 · 2024
Later among the works it cites.
Aclsum: A new dataset for aspect-based summarization of scientific publications
Sotaro Takeshita, Tommaso Green, Ines Reinig, Kai Eckert, and Simone Paolo Ponzetto. 2024 · 2024
Later among the works it cites.
SciMON: Scientific inspiration machines optimized for novelty
Qingyun Wang, Doug Downey, Heng Ji, and Tom Hope. 2024 · 2024
Later among the works it cites.
Chain of attack: a semantic-driven contextual multi-turn attacker for llm
Xikang Yang, Xuehai Tang, Songlin Hu, and Jizhong Han. 2024 · 2024
Later among the works it cites.
Yi Zeng, Hongpeng Lin, Jingwen Zhang, Diyi Yang, Ruoxi Jia, and Weiyan Shi. 2024 · 2024
Later among the works it cites.