Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have significantly advanced dialogue systems and role-playing agents through their ability to generate human-like text.
Behavioral study of obedience
Stanley Milgram. 1963 · 1963
Earlier work this paper cites.
Big five inventory
Oliver P John, Eileen M Donahue, and Robert L Kentle. 1991 · 1991
Earlier work this paper cites.
The origins of attachment theory
John Bowlby, Mary Ainsworth, and I Bretherton. 1992 · 1992
Earlier work this paper cites.
Sparse coding with an overcomplete basis set: A strategy employed by v1?
Bruno A. Olshausen and David J. Field. 1997 · 1997
Earlier work this paper cites.
The influences of family environment on personality traits
Kazuhisa Nakao, Jyo Takaishi, Kenji Tatsuta, Hisanori Katayama, Madoka Iwase, Kazuhiro Yorifuji, and Masatoshi Takeda. 2000 · 2000
Earlier work this paper cites.
Posttraumatic stress disorder and the nature of trauma
B. A. van der Kolk. 2000 · 2000
Earlier work this paper cites.
Technoculture: From Alphabet to Cybersex
L. Green. 2002 · 2002
Earlier work this paper cites.
Cultural influences on personality
Harry Triandis and Eunkook Suh. 2002 · 2002
Earlier work this paper cites.
The relationship between big five personality traits, negative affectivity, type a behavior, and work–family conflict
Carly S Bruck and Tammy D Allen. 2003 · 2003
Earlier work this paper cites.
A six-factor structure of personality-descriptive adjectives: Solutions from psycholexical studies in seven languages
Michael C. Ashton, Kibeom Lee, Marco Perugini, Piotr Szarota, Reinout E. de Vries, Lisa Di Blas, Kathleen Boies, and Boele De Raad. 2004 · 2004
Earlier work this paper cites.
The dark triad and normal personality traits
Sharon Jakobwitz and Vincent Egan. 2006 · 2006
Earlier work this paper cites.
Efficient sparse coding algorithms
Honglak Lee, Alexis Battle, Rajat Raina, and Andrew Ng. 2006 · 2006
Earlier work this paper cites.
Psychological stress and disease
S. Cohen, D. Janicki-Deverts, and G. E. Miller. 2007 · 2007
Earlier work this paper cites.
Ideology: Its resurgence in social, personality, and political psychology
John T Jost, Brian A Nosek, and Samuel D Gosling. 2008 · 2008
Earlier work this paper cites.
Personality trait change in adulthood
B. W. Roberts and D. Mroczek. 2008 · 2008
Earlier work this paper cites.
The multidimensional perfectionism cognitions inventory–english (mpci–e): Reliability, validity, and relationships with positive and negative affect
Joachim Stoeber, Osamu Kobori, and Yoshihiko Tanno. 2010 · 2010
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean. 2013 · 2013
Earlier work this paper cites.
Introducing the short dark triad (sd3): A brief measure of dark personality traits
Daniel N Jones and Delroy L Paulhus. 2014 · 2014
Earlier work this paper cites.
Would you deliver an electric shock in 2015? obedience in the experimental paradigm developed by stanley milgram in the 50 years following the original studies
Dariusz Dolinski, Tomasz Grzyb, Michał Folwarczny, Patrycja Grzybała, Karolina Krzyszycha, Karolina Martynowska, and Jakub Trojanowski. 2017 · 2015
Earlier work this paper cites.
Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav)
Been Kim, Martin Wattenberg, Justin Gilmer, Carrie Cai, James Wexler, Fernanda Viegas, et al. 2018 · 2018
Earlier work this paper cites.
The dark triad and workplace behavior
James M LeBreton, Levi K Shiverdecker, and Elizabeth M Grimaldi. 2018 · 2018
Earlier work this paper cites.
Resilience and big five personality traits: A meta-analysis
Atsushi Oshio, Kanako Taku, Mari Hirano, and Gul Saeed. 2018 · 2018
Earlier work this paper cites.
Connectionism
Cameron Buckner and James Garson. 2019 · 2019
Cited alongside, same era.
Atomic: An atlas of machine commonsense for if-then reasoning
Maarten Sap, Ronan Le Bras, Emily Allaway, Chandra Bhagavatula, Nicholas Lourie, Hannah Rashkin, Brendan Roof, Noah A. Smith, and Yejin Choi. 2019 · 2019
Cited alongside, same era.
A mathematical framework for transformer circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah. 2021 · 2021
Cited alongside, same era.
The dark side of high-fliers: the dark triad, high-flier traits, engagement, and subjective success
Adrian Furnham and Luke Treglown. 2021 · 2021
Cited alongside, same era.
Employment precarity strengthens the relationships between the dark triad and professional commitment
Leah M Kaufmann, Melissa A Wheeler, and Victor E Sojo. 2021 · 2021
Do gpt language models suffer from split personality disorder? the advent of substrate-free psychometrics
Peter Romero, Stephen Fitz, and Teruo Nakatsuma. 2023 · 2023
Later among the works it cites.
Character-LLM: A trainable agent for role-playing
Yunfan Shao, Linyang Li, Junqi Dai, and Xipeng Qiu. 2023 · 2023
Later among the works it cites.
Linear representations of sentiment in large language models
Curt Tigges, Oskar John Hollinsworth, Atticus Geiger, and Neel Nanda. 2023 · 2023
Later among the works it cites.
Activation addition: Steering language models without optimization
Alexander Matt Turner, Lisa Thiergart, Gavin Leech, David Udell, Juan J Vazquez, Ulisse Mini, and Monte MacDiarmid. 2023 · 2023
Later among the works it cites.
Rolellm: Benchmarking, eliciting, and enhancing role-playing abilities of large language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al. 2022 · 2022
Cited alongside, same era.
Xingxuan Li, Yutong Li, Shafiq Joty, Linlin Liu, Fei Huang, Lin Qiu, and Lidong Bing. 2022 · 2022
Cited alongside, same era.
Who is gpt-3? an exploration of personality, values and demographics
Marilù Miotto, Nicola Rossberg, and Bennett Kleinberg. 2022 · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022 · 2022
Cited alongside, same era.
Taking features out of superposition with sparse autoencoders
Lee Sharkey, Dan Braun, and beren. 2022 · 2022
Cited alongside, same era.
Language models can explain neurons in language models
Steven Bills, Nick Cammarata, Dan Mossing, Henk Tillman, Leo Gao, Gabriel Goh, Ilya Sutskever, Jan Leike, Jeff Wu, and William Saunders. 2023 · 2023
Cited alongside, same era.
Towards monosemanticity: Decomposing language models with dictionary learning
Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Tom Conerly, Nicholas L Turner, Cem Anil, Carson Denison, Amanda Askell, et al. 2023 · 2023
Cited alongside, same era.
Zekun Moore Wang, Zhongyuan Peng, Haoran Que, Jiaheng Liu, Wangchunshu Zhou, Yuhan Wu, Hongcheng Guo, Ruitong Gan, Zehao Ni, Man Zhang, et al. 2023 · 2023
Later among the works it cites.
Autogen: Enabling next-gen llm applications via multi-agent conversation framework
Qingyun Wu, Gagan Bansal, Jieyu Zhang, Yiran Wu, Shaokun Zhang, Erkang Zhu, Beibin Li, Li Jiang, Xiaoyun Zhang, and Chi Wang. 2023 · 2023
Later among the works it cites.
Instruction-following evaluation for large language models
Jeffrey Zhou, Tianjian Lu, Swaroop Mishra, Siddhartha Brahma, Sujoy Basu, Yi Luan, Denny Zhou, and Le Hou. 2023 · 2023
Later among the works it cites.
Representation engineering: A top-down approach to AI transparency
Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, et al. 2023 · 2023
Later among the works it cites.
Exploring the psychology of llms’ moral and legal reasoning
Guilherme F.C.F. Almeida, José Luiz Nunes, Neele Engelmann, Alex Wiegmann, and Marcelo de Araújo. 2024 · 2024
Closest in time.
SocialBench: Sociality evaluation of role-playing conversational agents
Hongzhan Chen, Hehong Chen, Ming Yan, Wenshen Xu, Gao Xing, Weizhou Shen, Xiaojun Quan, Chenliang Li, Ji Zhang, and Fei Huang. 2024 · 2024
Closest in time.
Patchscopes: A unifying framework for inspecting hidden representations of language models
Asma Ghandeharioun, Avi Caciularu, Adam Pearce, Lucas Dixon, and Mor Geva. 2024 · 2024
Closest in time.
Causal Determinism
Carl Hoefer. 2024 · 2024
Closest in time.
On the humanity of conversational AI: evaluating the psychological portrayal of llms
Jen-tse Huang, Wenxuan Wang, Eric John Li, Man Ho Lam, Shujie Ren, Youliang Yuan, Wenxiang Jiao, Zhaopeng Tu, and Michael R. Lyu. 2024 · 2024
Closest in time.
Sparse autoencoders find highly interpretable features in language models
Robert Huben, Hoagy Cunningham, Logan Riggs, Aidan Ewart, and Lee Sharkey. 2024 · 2024
Closest in time.
Better zero-shot reasoning with role-play prompting
Aobo Kong, Shiwan Zhao, Hao Chen, Qicheng Li, Yong Qin, Ruiqi Sun, Xin Zhou, Enzhi Wang, and Xiaohang Dong. 2024a · 2024
Closest in time.
Big5-chat: Shaping llm personalities through training on human-grounded data
Wenkai Li, Jiarui Liu, Andy Liu, Xuhui Zhou, Mona Diab, and Maarten Sap. 2024 · 2024
Closest in time.
Llms simulate big five personality traits: Further evidence
Aleksandra Sorokovikova, Natalia Fedorova, AI Toloka, Sharwin Rezagholi, Technikum Wien, and Ivan P Yamshchikov. 2024 · 2024
Closest in time.
Gemma 2: Improving open language models at a practical size
Gemma Team. 2024 · 2024
Closest in time.
A survey on llm-generated text detection: Necessity, methods, and future directions
Junchao Wu, Shu Yang, Runzhe Zhan, Yulin Yuan, Derek F. Wong, and Lidia S. Chao. 2024 · 2024
Closest in time.
Autogen: Enabling next-gen llm applications via multi-agent conversation
Qingyun Wu, Gagan Bansal, Jieyu Zhang, Yiran Wu, Beibin Li, Erkang Zhu, Li Jiang, Xiaoyun Zhang, Shaokun Zhang, Jiale Liu, et al · 2024
Closest in time.
SafetyBench: Evaluating the safety of large language models
Zhexin Zhang, Leqi Lei, Lindong Wu, Rui Sun, Yongkang Huang, Chong Long, Xiao Liu, Xuanyu Lei, Jie Tang, and Minlie Huang. 2024 · 2024
Closest in time.