Fetching the paper…
Reading the bibliography…
The potential of using Large Language Models (LLMs) themselves to evaluate LLM outputs offers a promising method for assessing model performance across various contexts.
The rating of chessplayers: Past and present
Arpad E Elo and Sam Sloan. 1978 · 1978
Earlier work this paper cites.
Influences on clinical judgement in mental health nursing
Peter J Martin. 1999 · 1999
Earlier work this paper cites.
BLEU: a Method for Automatic Evaluation of Machine Translation. 311–318
Kishore Papineni, Salim Roukos, Todd Ward, and Wei jing Zhu. 2002 · 2002
Earlier work this paper cites.
ROUGE: A Package for Automatic Evaluation of Summaries. In Text Summarization Branches Out . Association for Computational Linguistics, Barcelona, Spain, 74–81
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Design and evaluation guidelines for mental health technologies
Gavin Doherty, David Coyle, and Mark Matthews. 2010 · 2010
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017 · 2017
Earlier work this paper cites.
Reflecting on reflexive thematic analysis
Virginia Braun and Victoria Clarke. 2019 · 2019
Earlier work this paper cites.
Clinical Judgement: an investigation of clinical decision-making
Benjamin Paul Michael. 2019 · 2019
Earlier work this paper cites.
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019 · 2019
Earlier work this paper cites.
BERTScore: Evaluating Text Generation with BERT. In International Conference on Learning Representations
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2020 · 2020
Earlier work this paper cites.
On the opportunities and risks of foundation models
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al · 2021
Earlier work this paper cites.
Precision nutrition: A systematic literature review
Daniel Kirk, Cagatay Catal, and Bedir Tekinerdogan. 2021 · 2021
Earlier work this paper cites.
Journal of Human Nutrition and Dietetics 34, 1 (2021), 124–133
Ruth Vo, M Smith, and N Patton. 2021 · 2021
Earlier work this paper cites.
Bridging the Gap between UX Practitioners’ work practices and AI-enabled design support tools. In CHI Conference on Human Factors in Computing Systems Extended Abstracts . 1–7
Yuwen Lu, Chengzhi Zhang, Iris Zhang, and Toby Jia-Jun Li. 2022 · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Earlier work this paper cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Earlier work this paper cites.
ChainForge: An open-source visual programming environment for prompt engineering. In Adjunct Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology . 1–3
Ian Arawjo, Priyan Vaithilingam, Martin Wattenberg, and Elena Glassman. 2023 · 2023
Earlier work this paper cites.
The now and future of ChatGPT and GPT in psychiatry
Szu-Wei Cheng, Chung-Wen Chang, Wan-Jung Chang, Hao-Wei Wang, Chih-Sung Liang, Taishiro Kishimoto, Jane Pei-Chen Chang, John S Kuo, and Kuan-Pin Su. 2023 · 2023
Earlier work this paper cites.
ChatGPT as a virtual dietitian: Exploring its potential as a tool for improving nutrition knowledge
Manuel B Garcia. 2023 · 2023
Earlier work this paper cites.
Patat: Human-ai collaborative qualitative coding with explainable interactive rule synthesis. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems . 1–19
Simret Araya Gebreegziabher, Zheng Zhang, Xiaohang Tang, Yihao Meng, Elena L Glassman, and Toby Jia-Jun Li. 2023 · 2023
Earlier work this paper cites.
Rethinking large language models in mental health applications
Shaoxiong Ji, Tianlin Zhang, Kailai Yang, Sophia Ananiadou, and Erik Cambria. 2023 · 2023
Earlier work this paper cites.
Evallm: Interactive evaluation of large language model prompts on user-defined criteria
Tae Soo Kim, Yoonjoo Lee, Jamin Shin, Young-Ho Kim, and Juho Kim. 2023 · 2023
Earlier work this paper cites.
An introduction to generative artificial intelligence in mental health care: considerations and guidance
Darlene R King, Guransh Nanda, Joel Stoddard, Allison Dempsey, Sarah Hergert, Jay H Shore, and John Torous. 2023 · 2023
Cited alongside, same era.
Comparison of answers between ChatGPT and human dieticians to common nutrition questions
Daniel Kirk, Elise van Eijnatten, and Guido Camps. 2023 · 2023
Cited alongside, same era.
Chatharuhi: Reviving anime character in reality via large language model
Cheng Li, Ziang Leng, Chenxi Yan, Junyi Shen, Hao Wang, Weishi Mi, Yaying Fei, Xiaoyang Feng, Song Yan, HaoSheng Wang, et al · 2023
Cited alongside, same era.
Chatcounselor: A large language models for mental health support
June M Liu, Donghao Li, He Cao, Tianhe Ren, Zeyi Liao, and Jiamin Wu. 2023 · 2023
Cited alongside, same era.
The credibility of dietary advice formulated by ChatGPT: robo-diets for people with food allergies
Leveraging Variation Theory in Counterfactual Data Augmentation for Optimized Active Learning
Simret Araya Gebreegziabher, Kuangshi Ai, Zheng Zhang, Elena L Glassman, and Toby Jia-Jun Li. 2024a · 2024
Closest in time.
Supporting Co-Adaptive Machine Teaching through Human Concept Learning and Cognitive Theories
Simret Araya Gebreegziabher, Yukun Yang, Elena L Glassman, and Toby Jia-Jun Li. 2024b · 2024
Closest in time.
Overview of Gemini: Large Multimodal Models
Google Research. 2024 · 2024
Closest in time.
Large language models in mental health care: a scoping review
Yining Hua, Fenglin Liu, Kailai Yang, Zehan Li, Yi-han Sheu, Peilin Zhou, Lauren V Moran, Sophia Ananiadou, and Andrew Beam. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Paweł Niszczota and Iga Rybicka. 2023 · 2023
Cited alongside, same era.
Lamp: When large language models meet personalization
Alireza Salemi, Sheshera Mysore, Michael Bendersky, and Hamed Zamani. 2023 · 2023
Cited alongside, same era.
Stanford alpaca: An instruction-following llama model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B Hashimoto. 2023 · 2023
Cited alongside, same era.
Aligning large language models with human: A survey
Yufei Wang, Wanjun Zhong, Liangyou Li, Fei Mi, Xingshan Zeng, Wenyong Huang, Lifeng Shang, Xin Jiang, and Qun Liu. 2023 · 2023
Cited alongside, same era.
Visar: A human-ai argumentative writing assistant with visual programming and rapid draft prototyping. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology . 1–30
Zheng Zhang, Jie Gao, Ranjodh Singh Dhaliwal, and Toby Jia-Jun Li. 2023 · 2023
Cited alongside, same era.
A survey of large language models
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al · 2023
Cited alongside, same era.
Judgelm: Fine-tuned large language models are scalable judges
Lianghui Zhu, Xinggang Wang, and Xinlong Wang. 2023 · 2023
Cited alongside, same era.
Prolific
[n. d.] · 2024
Cited alongside, same era.
Shivani Kapania, Ruiyi Wang, Toby Jia-Jun Li, Tianshi Li, and Hong Shen. 2024 · 2024
Closest in time.
WILDBENCH: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
Bill Yuchen Lin, Yuntian Deng, Khyathi Chandu, Faeze Brahman, Abhilasha Ravichander, Valentina Pyatkin, Nouha Dziri, Ronan Le Bras, and Yejin Choi. 2024 · 2024
Closest in time.
PersonaFlow: Boosting Research Ideation with LLM-Simulated Expert Personas
Yiren Liu, Pranav Sharma, Mehul Jitendra Oswal, Haijun Xia, and Yun Huang. 2024 · 2024
Closest in time.
Ryan Louie, Ananjan Nandi, William Fang, Cheng Chang, Emma Brunskill, and Diyi Yang. 2024 · 2024
Closest in time.
Yuwen Lu, Ziang Tong, Qinyi Zhao, Yewon Oh, Bryan Wang, and Toby Jia-Jun Li. 2024a · 2024
Closest in time.
AI Assistance for UX: A Literature Review Through Human-Centered AI
Yuwen Lu, Yuewen Yang, Qinyi Zhao, Chengzhi Zhang, and Toby Jia-Jun Li. 2024b · 2024
Closest in time.
GPT-4 Research Overview
OpenAI. 2024 · 2024
Closest in time.
Is ChatGPT an Effective Tool for Providing Dietary Advice?
Valentina Ponzo, Ilaria Goitre, Enrica Favaro, Fabio Dario Merlo, Maria Vittoria Mancino, Sergio Riso, and Simona Bo. 2024 · 2024
Closest in time.
Navigating Complexity: Orchestrated Problem Solving with Multi-Agent LLMs
Sumedh Rasal and EJ Hauer. 2024 · 2024
Closest in time.
Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences
Shreya Shankar, JD Zamfirescu-Pereira, Björn Hartmann, Aditya G Parameswaran, and Ian Arawjo. 2024 · 2024
Closest in time.
Luminate: Structured Generation and Exploration of Design Space with Large Language Models for Human-AI Co-Creation. In Proceedings of the CHI Conference on Human Factors in Computing Systems . 1–26
Sangho Suh, Meng Chen, Bryan Min, Toby Jia-Jun Li, and Haijun Xia. 2024 · 2024
Closest in time.
Annalisa Szymanski, Simret Araya Gebreegziabher, Oghenemaro Anuyah, Ronald A Metoyer, and Toby Jia-Jun Li. 2024a · 2024
Closest in time.
Two tales of persona in llms: A survey of role-playing and personalization
Yu-Min Tseng, Yu-Chao Huang, Teng-Yun Hsiao, Yu-Ching Hsu, Jia-Yin Foo, Chao-Wei Huang, and Yun-Nung Chen. 2024 · 2024
Closest in time.
A Comprehensive Survey of LLM Alignment Techniques: RLHF, RLAIF, PPO, DPO and More
Zhichao Wang, Bin Bi, Shiva Kumar Pentyala, Kiran Ramnath, Sougata Chaudhuri, Shubham Mehrotra, Xiang-Bo Mao, Sitaram Asur, et al · 2024
Closest in time.
Mental-llm: Leveraging large language models for mental health prediction via online text data
Xuhai Xu, Bingsheng Yao, Yuanzhe Dong, Saadia Gabriel, Hong Yu, James Hendler, Marzyeh Ghassemi, Anind K Dey, and Dakuo Wang. 2024 · 2024
Closest in time.
ChatDiet: Empowering personalized nutrition-oriented food recommender chatbots through an LLM-augmented framework
Zhongqi Yang, Elahe Khatibi, Nitish Nagesh, Mahyar Abbasian, Iman Azimi, Ramesh Jain, and Amir M Rahmani. 2024 · 2024
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al · 2024
Closest in time.