Fetching the paper…
Reading the bibliography…
Recent research has focused on examining Large Language Models' (LLMs) characteristics from a psychological standpoint, acknowledging the necessity of understanding their behavioral characteristics.
A basis for analyzing test-retest reliability
Louis Guttman · 1945
Earlier work this paper cites.
Coefficient alpha and the internal structure of tests
Lee J Cronbach · 1951
Earlier work this paper cites.
The Myers-Briggs Type Indicator: Manual (1962)
Isabel Briggs Myers · 1962
Earlier work this paper cites.
Personality differences across regions of the united states
Samuel E Krug and Raymond W Kulhavy · 1973
Earlier work this paper cites.
Concurrent and predictive validity designs: A critical reanalysis
Gerald V Barrett, James S Phillips, and Ralph A Alexander · 1981
Earlier work this paper cites.
A revised version of the psychoticism scale
Sybil BG Eysenck, Hans J Eysenck, and Paul Barrett · 1985
Earlier work this paper cites.
Test validity: A matter of consequence
Samuel Messick · 1998
Earlier work this paper cites.
The big-five trait taxonomy: History, measurement, and theoretical perspectives
Oliver P John, Sanjay Srivastava, et al · 1999
Earlier work this paper cites.
Development of personality in early and middle adulthood: Set like plaster or persistent change?
Sanjay Srivastava, Oliver P John, Samuel D Gosling, and Jeff Potter · 2003
Earlier work this paper cites.
Regional personality differences in great britain
Peter J Rentfrow, Markus Jokela, and Michael E Lamb · 2015
Earlier work this paper cites.
Constructing validity: New developments in creating objective measuring instruments
Lee Anna Clark and David Watson · 2019
Earlier work this paper cites.
Calibrate before use: Improving few-shot performance of language models
Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh · 2021
Earlier work this paper cites.
Regional personality assessment through social media language
Salvatore Giorgi, Khoa Le Nguyen, Johannes C Eichstaedt, Margaret L Kern, David B Yaden, Michal Kosinski, Martin EP Seligman, Lyle H Ungar, H Andrew Schwartz, and Gregory Park · 2022
Earlier work this paper cites.
Estimating the personality of white-box language models
Saketh Reddy Karra, Son The Nguyen, and Theja Tulabandhula · 2022
Earlier work this paper cites.
Evaluating psychological safety of large language models
Xingxuan Li, Yutong Li, Lin Qiu, Shafiq Joty, and Lidong Bing · 2022
Earlier work this paper cites.
Who is gpt-3? an exploration of personality, values and demographics
Marilù Miotto, Nicola Rossberg, and Bennett Kleinberg · 2022
Earlier work this paper cites.
Introducing chatgpt
OpenAI · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Cited alongside, same era.
Bojana Bodroza, Bojana M Dinic, and Ljubisa Bojic · 2023
Cited alongside, same era.
Sparks of artificial general intelligence: Early experiments with gpt-4
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, Harsha Nori, Hamid Palangi, Marco Tulio Ribeiro, and Yi Zhang · 2023
Cited alongside, same era.
Evaluating the feasibility of chatgpt in healthcare: an analysis of multiple clinical and research scenarios
Marco Cascella, Jonathan Montomoli, Valentina Bellini, and Elena Bignami · 2023
Cited alongside, same era.
Inducing anxiety in large language models increases exploration and bias
Leveraging word guessing games to assess the intelligence of large language models
Tian Liang, Zhiwei He, Jen-tse Huang, Wenxuan Wang, Wenxiang Jiao, Rui Wang, Yujiu Yang, Zhaopeng Tu, Shuming Shi, and Xing Wang · 2023
Closest in time.
OpenAI · 2023
Closest in time.
Introducing gemini: our largest and most capable ai model
Sundar Pichai and Demis Hassabis · 2023
Closest in time.
Do gpt language models suffer from split personality disorder? the advent of substrate-free psychometrics
Peter Romero, Stephen Fitz, and Teruo Nakatsuma · 2023
Closest in time.
Whose opinions do language models reflect?
Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Julian Coda-Forno, Kristin Witte, Akshay K Jagadish, Marcel Binz, Zeynep Akata, and Eric Schulz · 2023
Cited alongside, same era.
Can large language models provide feedback to students? a case study on chatgpt
Wei Dai, Jionghao Lin, Hua Jin, Tongguang Li, Yi-Shan Tsai, Dragan Gašević, and Guanliang Chen · 2023
Cited alongside, same era.
Large language models are zero-shot fuzzers: Fuzzing deep-learning libraries via large language models
Yinlin Deng, Chunqiu Steven Xia, Haoran Peng, Chenyuan Yang, and Lingming Zhang · 2023
Cited alongside, same era.
How ready are pre-trained abstractive models and llms for legal case judgement summarization?
Aniket Deroy, Kripabandhu Ghosh, and Saptarshi Ghosh · 2023
Cited alongside, same era.
Toxicity in chatgpt: Analyzing persona-assigned language models
Ameet Deshpande, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, and Karthik Narasimhan · 2023
Cited alongside, same era.
Can ai language models replace human participants?
Danica Dillion, Niket Tandon, Yuling Gu, and Kurt Gray · 2023
Cited alongside, same era.
Automated repair of programs from large language models
Zhiyu Fan, Xiang Gao, Martin Mirchev, Abhik Roychoudhury, and Shin Hwei Tan · 2023
Cited alongside, same era.
Self-assessment tests are unreliable measures of llm personality
Akshat Gupta, Xiaoyang Song, and Gopala Anumanchipalli · 2023
Cited alongside, same era.
Greg Serapio-García, Mustafa Safdari, Clément Crepy, Luning Sun, Stephen Fitz, Peter Romero, Marwa Abdulhai, Aleksandra Faust, and Maja Matarić · 2023
Closest in time.
Character-llm: A trainable agent for role-playing
Yunfan Shao, Linyang Li, Junqi Dai, and Xipeng Qiu · 2023
Closest in time.
Xiaoyang Song, Akshat Gupta, Kiyan Mohebbizadeh, Shujie Hu, and Anant Singh · 2023
Closest in time.
Challenging the validity of personality tests for large language models
Tom Sühr, Florian E Dorner, Samira Samadi, and Augustin Kelava · 2023
Closest in time.
Chatgpt or grammarly? evaluating chatgpt on grammatical error correction benchmark
Haoran Wu, Wenxuan Wang, Yuxuan Wan, Wenxiang Jiao, and Michael Lyu · 2023
Closest in time.
Red teaming chatgpt via jailbreaking: Bias, robustness, reliability and toxicity
Terry Yue Zhuo, Yujin Huang, Chunyang Chen, and Zhenchang Xing · 2023
Closest in time.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al · 2024
Closest in time.
Personallm: Investigating the ability of large language models to express personality traits
Hang Jiang, Xiajie Zhang, Xubo Cao, Cynthia Breazeal, Deb Roy, and Jad Kabbara · 2024
Closest in time.
The self-perception and political biases of chatgpt
Jérôme Rutinowski, Sven Franke, Jan Endendyk, Ina Dormuth, Moritz Roidl, and Markus Pauly · 2024
Closest in time.
You don’t need a personality test to know these models are unreliable: Assessing the reliability of large language models on psychometric instruments
Bangzhao Shu, Lechen Zhang, Minje Choi, Lavinia Dunagan, Lajanugen Logeswaran, Moontae Lee, Dallas Card, and David Jurgens · 2024
Closest in time.
Rolellm: Benchmarking, eliciting, and enhancing role-playing abilities of large language models
Noah Wang, Zhongyuan Peng, Haoran Que, Jiaheng Liu, Wangchunshu Zhou, Yuhan Wu, Hongcheng Guo, Ruitong Gan, Zehao Ni, Jian Yang, Man Zhang, Zhaoxiang Zhang, Wanli Ouyang, Ke Xu, Wenhao Huang, Jie Fu, and Junran Peng · 2024
Closest in time.
All languages matter: On the multilingual safety of large language models
Wenxuan Wang, Zhaopeng Tu, Chang Chen, Youliang Yuan, Jen-tse Huang, Wenxiang Jiao, and Michael R Lyu · 2024
Closest in time.