Fetching the paper…
Reading the bibliography…
Personality psychologists have analyzed the relationship between personality and safety behaviors in human society.
A coefficient of agreement for nominal scales
Jacob Cohen · 1960
Earlier work this paper cites.
Manual of the Eysenck Personality Questionnaire (junior & adult)
Hans Jurgen Eysenck and Sybil Bianca Giuletta Eysenck · 1975
Earlier work this paper cites.
Myers-Briggs type indicator
Katharine C Briggs · 1976
Earlier work this paper cites.
Joint factors in self-reports and ratings: Neuroticism, extraversion and openness to experience
Robert R. McCrae and Paul T. Costa · 1983
Earlier work this paper cites.
Recent assessments of the myers-briggs type indicator
John G Carlson · 1985
Earlier work this paper cites.
Reinterpreting the myers-briggs type indicator from the perspective of the five-factor model of personality
Robert R McCrae and Paul T Costa Jr · 1989
Earlier work this paper cites.
Personality, affect, and behavior in groups
Jennifer M George · 1990
Earlier work this paper cites.
Big five inventory
Oliver P John, Eileen M Donahue, and Robert L Kentle · 1991
Earlier work this paper cites.
Trait theory as personality theory: Can a part be as great as the whole?
Seymour Epstein · 1994
Earlier work this paper cites.
Evaluation metrics for language models
Stanley F Chen, Douglas Beeferman, and Roni Rosenfeld · 1998
Earlier work this paper cites.
Trait theories of personality
Paul T Costa and Robert R McCrae · 1998
Earlier work this paper cites.
Myers-Briggs Type Indicator : form M
Peter B. Myers, Katharine D. Myers, Isabel Briggs Myers, and Linda K. Kirby · 1998
Earlier work this paper cites.
Is there a relationship between the myers-briggs type indicator and emotional intelligence?
Malcolm Higgs · 2001
Earlier work this paper cites.
Myers-briggs type indicator score reliability across: Studies a meta-analytic reliability generalization study
Robert M Capraro and Mary Margaret Capraro · 2002
Earlier work this paper cites.
MBTI manual: A guide to the development and use of the Myers-Briggs Type Indicator
Isabel Briggs Myers · 2003
Earlier work this paper cites.
Communicator image and Myers-Briggs Type Indicator extraversion-introversion
S. K. Opt and D. A. Loffredo · 2003
Earlier work this paper cites.
Communicator image and myers—briggs type indicator extraversion—introversion
Susan K Opt and Donald A Loffredo · 2003
Earlier work this paper cites.
An empirical test of cognitive style and strategic decision outcomes
Jill R Hough and DT Ogilvie · 2005
Earlier work this paper cites.
Personality and domain-specific risk taking
Nigel Nicholson, Emma Soane, Mark Fenton-O’Creevy, and Paul Willman · 2005
Earlier work this paper cites.
The relations between personality and language use
Chang H Lee, Kyungil Kim, Young Seok Seo, and Cindy K Chung · 2007
Earlier work this paper cites.
Experimental approaches to the study of personality
William Revelle · 2007
Earlier work this paper cites.
The myers-briggs type indicator and transformational leadership
F William Brown and Michael D Reilly · 2009
Earlier work this paper cites.
Performance, personality, and energetics: correlation, causation, and mechanism
Vincent Careau and Theodore Garland Jr · 2012
Earlier work this paper cites.
Correlation and causation in the study of personality
James J Lee · 2012
Earlier work this paper cites.
What has personality and emotional intelligence to do with ‘feeling different’while using a foreign language?
Katarzyna Ożańska-Ponikwia · 2012
Earlier work this paper cites.
Multifaceted personality predictors of workplace safety performance: More than conscientiousness
Joyce Hogan and Jeff Foster · 2013
Earlier work this paper cites.
Is personality modulated by language?
G Marina Veltkamp, Guillermo Recio, Arthur M Jacobs, and Markus Conrad · 2013
Earlier work this paper cites.
Does language affect personality perception? a functional approach to testing the whorfian hypothesis
Sylvia Xiaohua Chen, Verónica Benet-Martínez, and Jacky CK Ng · 2014
Earlier work this paper cites.
A meta-analysis of personality and workplace safety: addressing unanswered questions
Jeremy M Beus, Lindsay Y Dhanani, and Mallory A McCord · 2015
Earlier work this paper cites.
Workplace safety: A review and research synthesis
Jeremy M Beus, Mallory A McCord, and Dov Zohar · 2016
Earlier work this paper cites.
Personality traits across cultures
A Timothy Church · 2016
Earlier work this paper cites.
What is personality?
Otto F Kernberg · 2016
Earlier work this paper cites.
The LAMBADA dataset: Word prediction requiring a broad discourse context
Denis Paperno, Germán Kruszewski, Angeliki Lazaridou, Ngoc Quan Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, and Raquel Fernández · 2016
Earlier work this paper cites.
Personality traits across countries: Support for similarities rather than differences
Petri Kajonius and Erik Mac Giolla · 2017
Earlier work this paper cites.
Inherent trade-offs in the fair determination of risk scores
Jon Kleinberg, Sendhil Mullainathan, and Manish Raghavan · 2017
Earlier work this paper cites.
The Importance of Agreeableness
Ethan Campbell and Matthew P. Kassner · 2018
Earlier work this paper cites.
Culture and personality
Jüri Allik and Anu Realo · 2019
Earlier work this paper cites.
Differential privacy has disparate impact on model accuracy
Eugene Bagdasaryan, Omid Poursaeed, and Vitaly Shmatikov · 2019
Cited alongside, same era.
The policy relevance of personality traits
Wiebke Bleidorn, Patrick L. Hill, Mitja D. Back, Jaap J. A. Denissen, Marie Hennecke, Christopher James Hopwood, Markus Jokela, Christian Kandler, Richard E. Lucas, Maike Luhmann, Ulrich R. Orth, Jenny Wagner, Cornelia Wrzus, Johannes Zimmermann, and Brent W. Roberts · 2019
Cited alongside, same era.
Ethics guidelines for trustworthy AI
European Commission, Content Directorate-General for Communications Networks, and Technology · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Cited alongside, same era.
Privacy risks of securing machine learning models against adversarial examples
Liwei Song, Reza Shokri, and Prateek Mittal · 2019
Cited alongside, same era.
Shengyu Mao, Ningyu Zhang, Xiaohan Wang, Mengru Wang, Yunzhi Yao, Yong Jiang, Pengjun Xie, Fei Huang, and Huajun Chen · 2023
Later among the works it cites.
Can llms keep a secret? testing privacy implications of language models via contextual integrity theory, 2023
Niloofar Mireshghallah, Hyunwoo Kim, Xuhui Zhou, Yulia Tsvetkov, Maarten Sap, Reza Shokri, and Yejin Choi · 2023
Later among the works it cites.
What affects the usage of artificial conversational agents? an agent personality and love theory perspective
Debajyoti Pal, Vajirasak Vanijja, Himanshu Thapliyal, and Xiangmin Zhang · 2023
Later among the works it cites.
Do llms possess a personality? making the mbti test an amazing evaluation for large language models
Keyu Pan and Yawen Zeng · 2023
Later among the works it cites.
Fine-tuning aligned language models compromises safety, even when users do not intend to!
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Daniel M Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving · 2019
Cited alongside, same era.
Personality and safety citizenship: the role of safety motivation and safety knowledge
Julie Laurent, Nik Chmiel, and Isabelle Hansez · 2020
Cited alongside, same era.
A self-refinement strategy for noise reduction in grammatical error correction
Masato Mita, Shun Kiyono, Masahiro Kaneko, Jun Suzuki, and Kentaro Inui · 2020
Cited alongside, same era.
The effects of personality and locus of control on trust in humans versus artificial intelligence. heliyon, 6 (8), e04572, 2020
NN Sharan and DM Romano · 2020
Cited alongside, same era.
A general language assistant as a laboratory for alignment
Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain, Deep Ganguli, Tom Henighan, Andy Jones, Nicholas Joseph, Ben Mann, Nova DasSarma, et al · 2021
Cited alongside, same era.
Proposal for a regulation of the european parliament and of the council laying down harmonised rules on artificial intelligence (artificial intelligence act) and amending certain union legislative acts, pub. l. no. com(2021) 206 final., 2021b
European Commission · 2021
Cited alongside, same era.
StereoSet: Measuring stereotypical bias in pretrained language models
Moin Nadeem, Anna Bethke, and Siva Reddy · 2021
Cited alongside, same era.
Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal, and Peter Henderson · 2023
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn · 2023
Later among the works it cites.
Quantifying language models’ sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt formatting, 2023
Melanie Sclar, Yejin Choi, Yulia Tsvetkov, and Alane Suhr · 2023
Later among the works it cites.
Jailbreaking language models at scale via persona modulation
Rusheb Shah, Quentin Feuillade Montixi, Soroush Pour, Arush Tagade, and Javier Rando · 2023
Later among the works it cites.
Character-llm: A trainable agent for role-playing
Yunfan Shao, Linyang Li, Junqi Dai, and Xipeng Qiu · 2023
Later among the works it cites.
Xiaoyang Song, Akshat Gupta, Kiyan Mohebbizadeh, Shujie Hu, and Anant Singh · 2023
Later among the works it cites.
Artificial intelligence risk management framework (ai rmf 1.0), 2023-01-26 05:01:00 2023
Elham Tabassi · 2023
Later among the works it cites.
The predictors of unsafe behaviors among nuclear power plant workers: An investigation integrating personality, cognitive and attitudinal factors
Da Tao, Xiaofeng Diao, Xingda Qu, Xiaoting Ma, and Tingru Zhang · 2023
Later among the works it cites.
Revisiting the reliability of psychological scales on large language models, 2023
Jen tse Huang, Wenxuan Wang, Man Ho Lam, Eric John Li, Wenxiang Jiao, and Michael R. Lyu · 2023
Later among the works it cites.
Characterchat: Learning towards conversational ai with personalized social support
Quan Tu, Chuanqi Chen, Jinpeng Li, Yanran Li, Shuo Shang, Dongyan Zhao, Ran Wang, and Rui Yan · 2023
Later among the works it cites.
Decodingtrust: A comprehensive assessment of trustworthiness in gpt models
Boxin Wang, Weixin Chen, Hengzhi Pei, Chulin Xie, Mintong Kang, Chenhui Zhang, Chejian Xu, Zidi Xiong, Ritik Dutta, Rylan Schaeffer, et al · 2023
Later among the works it cites.
Haoran Wang and Kai Shu · 2023
Later among the works it cites.
Large language models are not fair evaluators
Peiyi Wang, Lei Li, Liang Chen, Dawei Zhu, Binghuai Lin, Yunbo Cao, Qi Liu, Tianyu Liu, and Zhifang Sui · 2023
Later among the works it cites.
Do changes in personality predict life outcomes?
Amanda Wright and Joshua Jackson · 2023
Later among the works it cites.
Shadow alignment: The ease of subverting safely-aligned language models
Xianjun Yang, Xiao Wang, Qi Zhang, Linda Petzold, William Yang Wang, Xun Zhao, and Dahua Lin · 2023
Later among the works it cites.
Gpt-4 is too smart to be safe: Stealthy chat with llms via cipher
Youliang Yuan, Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang, Pinjia He, Shuming Shi, and Zhaopeng Tu · 2023
Later among the works it cites.
Representation engineering: A top-down approach to ai transparency
Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, et al · 2023
Later among the works it cites.
Is cognition and action consistent or not: Investigating large language model’s personality
Yiming Ai, Zhiwei He, Ziyin Zhang, Wenhong Zhu, Hongkun Hao, Kai Yu, Lingjun Chen, and Rui Wang · 2024
Closest in time.
Afspp: Agent framework for shaping preference and personality with large language models
Zihong He and Changwang Zhang · 2024
Closest in time.
Evaluating and inducing personality in pre-trained language models
Guangyuan Jiang, Manjie Xu, Song-Chun Zhu, Wenjuan Han, Chi Zhang, and Yixin Zhu · 2024
Closest in time.
Can multiple-choice questions really be useful in detecting the abilities of llms?
Wangyue Li, Liangzhi Li, Tong Xiang, Xiao Liu, Wei Deng, and Noa Garcia · 2024
Closest in time.
Codechameleon: Personalized encryption framework for jailbreaking large language models
Huijie Lv, Xiao Wang, Yuansen Zhang, Caishuang Huang, Shihan Dou, Junjie Ye, Tao Gui, Qi Zhang, and Xuanjing Huang · 2024
Closest in time.
Towards tracing trustworthiness dynamics: Revisiting pre-training period of large language models
Chen Qian, Jie Zhang, Wei Yao, Dongrui Liu, Zhenfei Yin, Yu Qiao, Yong Liu, and Jing Shao · 2024
Closest in time.
Identifying multiple personalities in large language models with external evaluation
Xiaoyang Song, Yuta Adachi, Jessie Feng, Mouwei Lin, Linhao Yu, Frank Li, Akshat Gupta, Gopala Anumanchipalli, and Simerjot Kaur · 2024
Closest in time.
Llms simulate big five personality traits: Further evidence
Aleksandra Sorokovikova, Natalia Fedorova, Sharwin Rezagholi, and Ivan P Yamshchikov · 2024
Closest in time.
Trustllm: Trustworthiness in large language models
Lichao Sun, Yue Huang, Haoran Wang, Siyuan Wu, Qihui Zhang, Chujie Gao, Yixin Huang, Wenhan Lyu, Yixuan Zhang, Xiner Li, et al · 2024
Closest in time.
An evolutionary model of personality traits related to cooperative behavior using a large language model
Reiji Suzuki and Takaya Arita · 2024
Closest in time.
Phantom: Personality has an effect on theory-of-mind reasoning in large language models
Fiona Anting Tan, Gerard Christopher Yeo, Fanyou Wu, Weijie Xu, Vinija Jain, Aman Chadha, Kokil Jaidka, Yang Liu, and See-Kiong Ng · 2024
Closest in time.
Incharacter: Evaluating personality fidelity in role-playing agents through psychological interviews, 2024
Xintao Wang, Yunze Xiao, Jen tse Huang, Siyu Yuan, Rui Xu, Haoran Guo, Quan Tu, Yaying Fei, Ziang Leng, Wei Wang, Jiangjie Chen, Cheng Li, and Yanghua Xiao · 2024
Closest in time.
Jailbroken: How does llm safety training fail?
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt · 2024
Closest in time.
Controllm: Crafting diverse personalities for language models
Yixuan Weng, Shizhu He, Kang Liu, Shengping Liu, and Jun Zhao · 2024
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al · 2024
Closest in time.
Easyjailbreak: A unified framework for jailbreaking large language models
Weikang Zhou, Xiao Wang, Limao Xiong, Han Xia, Yingshuang Gu, Mingxu Chai, Fukang Zhu, Caishuang Huang, Shihan Dou, Zhiheng Xi, et al · 2024
Closest in time.