Fetching the paper…
Reading the bibliography…
Artificial intelligence (AI) developers are increasingly building language models with warm and empathetic personas that millions of people now use for advice, therapy, and companionship.
A coefficient of agreement for nominal scales
Cohen, J · 1960
Earlier work this paper cites.
An application of hierarchical kappa-type statistics in the assessment of majority agreement among multiple observers
Landis, J. R. & Koch, G. G · 1977
Earlier work this paper cites.
White lies in interpersonal communication: A taxonomy and preliminary investigation of social motivations
Camden, C., Motley, M. T. & Wilson, A · 1984
Earlier work this paper cites.
Everyday lies in close and casual relationships
DePaulo, B. M. & Kashy, D. A · 1998
Earlier work this paper cites.
(Im) Politeness, face and perceptions of rapport: unpackaging their bases and interrelationships
Spencer-Oatey, H · 2005
Earlier work this paper cites.
White lies
Erat, S. & Gneezy, U · 2012
Earlier work this paper cites.
Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension
Joshi, M., Choi, E., Weld, D. S. & Zettlemoyer, L · 2017
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D. et al · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K. et al · 2021
Earlier work this paper cites.
What disease does this patient have? a large-scale open domain question answering dataset from medical exams
Jin, D. et al · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Hu, E. J. et al · 2022
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Bai, Y. et al · 2022
Earlier work this paper cites.
Constitutional AI: Harmlessness from AI feedback
Bai, Y. et al · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L. et al · 2022
Earlier work this paper cites.
TruthfulQA: Measuring how models mimic human falsehoods
Lin, S., Hilton, J. & Evans, O · 2022
Earlier work this paper cites.
In conversation with artificial intelligence: aligning language models with human values
Kasirzadeh, A. & Gabriel, I · 2023
Earlier work this paper cites.
Universal and transferable adversarial attacks on aligned language models
Zou, A. et al · 2023
Cited alongside, same era.
Meet my A.I. friends
Roose, K · 2024
Cited alongside, same era.
Towards understanding sycophancy in language models
Sharma, M. et al · 2024
Cited alongside, same era.
Lawsuit claims character.ai is responsible for teen’s suicide
Yang, A · 2024
Cited alongside, same era.
When scaling meets LLM finetuning: The effect of data, model and finetuning method
Zhang, B., Liu, Z., Cherry, C. & Firat, O · 2024
Cited alongside, same era.
Fine-tuning aligned language models compromises safety, even when users do not intend to!
Why AI chatbots lie to us
Mitchell, M · 2025
Closest in time.
Third-party evaluators perceive ai as more compassionate than expert humans
Ovsyannikova, D., de Mello, V. O. & Inzlicht, M · 2025
Closest in time.
AI app Replika accused of deceptive marketing
Chow, A. R · 2025
Closest in time.
They asked chatgpt questions. the answers sent them spiraling
Hill, K · 2025
Closest in time.
Character training: Understanding and crafting a language model’s personality (2025)
Lambert, N · 2025
Closest in time.
A fully open ai foundation model applied to chest radiography
Ma, D., Pang, J., Gotway, M. B. & Liang, J · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Qi, X. et al · 2024
Cited alongside, same era.
How do large language models navigate conflicts between honesty and helpfulness?
Liu, R., Sumers, T. R., Dasgupta, I. & Griffiths, T. L · 2024
Cited alongside, same era.
Gu, J. et al · 2024
Cited alongside, same era.
Who’s asking? user personas and the mechanics of latent misalignment
Ghandeharioun, A. et al · 2024
Cited alongside, same era.
Multi-turn evaluation of anthropomorphic behaviours in large language models
Ibrahim, L. et al · 2025
Cited alongside, same era.
Comparing the value of perceived human versus AI-generated empathy
Rubin, M. et al · 2025
Cited alongside, same era.
OpenAI Model Spec (2025)
OpenAI · 2025
Cited alongside, same era.
Bodnar, C. et al · 2025
Closest in time.
Accurate predictions on small data with a tabular foundation model
Hollmann, N. et al · 2025
Closest in time.
Humt dumt: Measuring and controlling human-like language in llms
Cheng, M., Yu, S. & Jurafsky, D · 2025
Closest in time.
The mask benchmark: Disentangling honesty from accuracy in ai systems
Ren, R. et al · 2025
Closest in time.
Emergent misalignment: Narrow finetuning can produce broadly misaligned llms
Betley, J. et al · 2025
Closest in time.
Why evaluating the impact of AI needs to start now
Hauser, O. P., Light, M., Shelmerdine, L. & Blumenau, J · 2025
Closest in time.
On targeted manipulation and deception when optimizing LLMs for user feedback
Williams, M. et al · 2025
Closest in time.
Persona features control emergent misalignment
Wang, M. et al · 2025
Closest in time.
Social sycophancy: A broader understanding of LLM sycophancy
Cheng, M. et al · 2025
Closest in time.