Fetching the paper…
Reading the bibliography…
Evaluating Large Language Models (LLMs) in the mental health domain poses distinct challenged from other domains, given the subtle and highly subjective nature of symptoms that exhibit significant variability among individuals.
Measuring nominal scale agreement among many raters
Joseph L Fleiss · 1971
Earlier work this paper cites.
Ethical principles of psychologists and code of conduct
American Psychological Association · 2002
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Interrater reliability: the kappa statistic
Mary L McHugh · 2012
Earlier work this paper cites.
Estimating county health statistics with twitter
Aron Culotta · 2014
Earlier work this paper cites.
A diversity-promoting objective function for neural conversation models, 2016
Jiwei Li, Michel Galley, Chris Brockett, Jianfeng Gao, and Bill Dolan · 2016
Earlier work this paper cites.
SMHD: a large-scale resource for exploring online language usage for multiple mental health conditions
Arman Cohan, Bart Desmet, Andrew Yates, Luca Soldaini, Sean MacAvaney, and Nazli Goharian · 2018
Earlier work this paper cites.
Socioeconomic variations in the mental health treatment gap for people with anxiety, mood, and substance use disorders: results from the WHO World Mental Health (WMH) surveys
Sara Evans-Lacko, Sergio Aguilar-Gaxiola, A Al-Hamzawi, et al · 2018
Earlier work this paper cites.
Responding to persons experiencing a mental health crisis
International Association of Chiefs of Police · 2018
Earlier work this paper cites.
Navigating a mental health crisis
National Alliance on Mental Health · 2018
Earlier work this paper cites.
Identifying depression on Reddit: The effect of training data
Inna Pirina and Çağrı Çöltekin · 2018
Earlier work this paper cites.
Knowledge-aware assessment of severity of suicide risk for early intervention
Manas Gaur, Amanuel Alambo, Joy Prakash Sain, Ugur Kursuncu, Krishnaprasad Thirunarayan, Ramakanth Kavuluru, Amit Sheth, Randy Welton, and Jyotishman Pathak · 2019
Earlier work this paper cites.
Dreaddit: A Reddit dataset for stress analysis in social media
Elsbeth Turcan and Kathy McKeown · 2019
Earlier work this paper cites.
Methods in predictive techniques for mental health status on social media: A critical review
Stevie Chancellor and Munmun De Choudhury · 2020
Earlier work this paper cites.
A computational approach to understanding empathy expressed in text-based mental health support
Ashish Sharma, Adam Miner, David Atkins, and Tim Althoff · 2020
Earlier work this paper cites.
Deep learning for suicide and depression identification with unsupervised label correction
Ayaan Haque, Viraaj Reddi, and Tyler Giallanza · 2021
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2021
Earlier work this paper cites.
International declaration of core competencies in professional psychology
International Association of Applied Psychology · 2021
Cited alongside, same era.
What disease does this patient have? a large-scale open domain question answering dataset from medical exams
Dongjin Jin, Eric Pan, Nasim Oufattole, Wei-Hung Weng, Hua Fang, and Peter Szolovits · 2021
Cited alongside, same era.
Charting outcomes in the match: Senior students of u.s. md medical schools, 2020
National Resident Matching Program · 2021
Cited alongside, same era.
Smart conversational agents for the detection of neuropsychiatric disorders: A systematic review
Moisés R. Pacheco-Lorenzo, Sonia M. Valladares-Rodríguez, Luis E. Anido-Rifón, and Manuel J. Fernández-Iglesias · 2021
Cited alongside, same era.
PsyQA: A Chinese dataset for generating long counseling text for mental health support
Hao Sun, Zhenru Lin, Chujie Zheng, Siyang Liu, and Minlie Huang · 2021
Cited alongside, same era.
Efficient and effective text encoding for Chinese LLaMA and Alpaca
Yiming Cui, Ziqing Yang, and Xin Yao · 2023
Closest in time.
MedAlpaca–an open-source collection of medical conversational AI models and training data
Tianyu Han, Lisa C Adams, Jens-Michalis Papaioannou, Paul Grundmann, Tom Oberhauser, Alexander Löser, Daniel Truhn, and Keno K Bressem · 2023
Closest in time.
C-Eval: A multi-level multi-discipline chinese evaluation suite for foundation models
Yuzhen Huang, Yuzhuo Bai, Zhihao Zhu, Junlei Zhang, Jinghan Zhang, Tangjun Su, Junteng Liu, Chuancheng Lv, Yikai Zhang, Jiayi Lei, Yao Fu, Maosong Sun, and Junxian He · 2023
Closest in time.
ChatGPT: Jack of all trades, master of none
Jan Kocoń, Igor Cichecki, Oliwier Kaszyca, Mateusz Kochanek, Dominika Szydło, Joanna Baran, Julita Bielaniewicz, Marcin Gruza, Arkadiusz Janz, Kamil Kanclerz, et al · 2023
Closest in time.
Performance of ChatGPT on USMLE: Potential for AI-assisted medical education using large language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
GLM: General language model pretraining with autoregressive blank infilling
Zhengxiao Du, Yujie Qian, Xiao Liu, Ming Ding, Jiezhong Qiu, Zhilin Yang, and Jie Tang · 2022
Cited alongside, same era.
Holistic evaluation of language models
Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, Benjamin Newman, Binhang Yuan, Bobby Yan, Ce Zhang, Christian Cosgrove, Christopher D. Manning, Christopher Ré, Diana Acosta-Navas, Drew A. Hudson, Eric Zelikman, Esin Durmus, Faisal Ladhak, Frieda Rong, Hongyu Ren, Huaxiu Yao, Jue Wang, Keshav Santhanam, Laurel Orr, Lucia Zheng, Mert Yuksekgonul, Mirac Suzgun, Nathan Kim, Neel Guha, Niladri Chatterji, Omar Khattab, Peter Henderson, Qian Huang, Ryan Chi, Sang Michael Xie, Shibani Santurkar, Surya Ganguli, Tatsunori Hashimoto, Thomas Icard, Tianyi Zhang, Vishrav Chaudhary, William Wang, Xuechen Li, Yifan Mai, Yuhui Zhang, and Yuta Koreeda · 2022
Cited alongside, same era.
Early identification of depression severity levels on Reddit using ordinal classification
Usman Naseem, Adam G. Dunn, Jinman Kim, and Matloob Khushi · 2022
Cited alongside, same era.
ChatGPT: Optimizing language models for dialogue
J. Schulman, B. Zoph, C. Kim, J. Hilton, J. Menick, J. Weng, J. F. C. Uribe, L. Fedus, L. Metz, M. Pokorny, and et al · 2022
Cited alongside, same era.
Putting the ”mental” back in ”mental disorders”: a perspective from research on fear and anxiety
Vincent Taschereau-Dumouchel, Michaël Michel, Hakwan Lau, Stefan G Hofmann, and Joseph E LeDoux · 2022
Cited alongside, same era.
Self-instruct: Aligning language model with self generated instructions
Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A Smith, Daniel Khashabi, and Hannaneh Hajishirzi · 2022
Cited alongside, same era.
D4: a Chinese dialogue dataset for depression-diagnosis-oriented chat
Binwei Yao, Chao Shi, Likai Zou, Lingfeng Dai, Mengyue Wu, Lu Chen, Zhen Wang, and Kai Yu · 2022
Cited alongside, same era.
Tiffany H Kung, Morgan Cheatham, Arielle Medenilla, Czarina Sillos, Lorie De Leon, Camille Elepaño, Maria Madriaga, Rimel Aggabao, Giezel Diaz-Candido, James Maningo, et al · 2023
Closest in time.
Evaluation of ChatGPT for NLP-based mental health applications
Bishal Lamichhane · 2023
Closest in time.
GPT-4 technical report
OpenAI · 2023
Closest in time.
Is ChatGPT a general-purpose natural language processing task solver?
Chengwei Qin, Aston Zhang, Zhuosheng Zhang, Jiaao Chen, Michihiro Yasunaga, and Diyi Yang · 2023
Closest in time.
A benchmark for understanding dialogue safety in mental health support
Huachuan Qiu, Tong Zhao, Anqi Li, Shuai Zhang, Hongliang He, and Zhenzhong Lan · 2023
Closest in time.
Large language models encode clinical knowledge
K. Singhal, S. Azizi, T. Tu, S. S. Mahdavi, J. Wei, H. W. Chung, N. Scales, A. Tanwani, H. Cole-Lewis, S. Pfohl, and et al · 2023
Closest in time.
Scieval: A multi-level large language model evaluation benchmark for scientific research
Liangtai Sun, Yang Han, Zihan Zhao, Da Ma, Zhennan Shen, Baocai Chen, Lu Chen, and Kai Yu · 2023
Closest in time.
Stanford alpaca: An instruction-following llama model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto · 2023
Closest in time.
LLaMA: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Closest in time.
USMLE sample questions
USMLE · 2023
Closest in time.
Depressive disorder (depression)
World Health Organization · 2023
Closest in time.
Leveraging large language models for mental health prediction via online text data
Xuhai Xu, Bingshen Yao, Yuanzhe Dong, Hong Yu, James Hendler, Anind K Dey, and Dakuo Wang · 2023
Closest in time.