Repairing the cracked foundation: A survey of obstacles in evaluation practices for generated text
Sebastian Gehrmann, Elizabeth Clark, and Thibault Sellam. 2023 · 2023
Closest in time.
Topical-Chat: Towards Knowledge-Grounded Open-Domain Conversations
Original
Karthik Gopalakrishnan, Behnam Hedayatnia, Qinlang Chen, Anna Gottardi, Sanjeev Kwatra, Anu Venkatesh, Raefer Gabriel, and Dilek Hakkani-Tur. 2023 · 2023
Closest in time.
Cells, Generators, and Lenses: Design Framework for Object-Oriented Interaction with Large Language Models. In The 36th Annual ACM Symposium on User Interface Software and Technology (UIST ’23), October 29-November 1, 2023, San Francisco, CA, USA (San Francisco, CA, USA) (UIST ’23) . Association for Computing Machinery, New York, NY, USA, 18 pages
Tae Soo Kim, Yoonjoo Lee, Minsuk Chang, and Juho Kim. 2023b · 2023
Closest in time.
LongEval: Guidelines for human evaluation of faithfulness in long-form summarization
Original
Kalpesh Krishna, Erin Bransom, Bailey Kuehl, Mohit Iyyer, Pradeep Dasigi, Arman Cohan, and Kyle Lo. 2023 · 2023
Closest in time.
DAPIE: Interactive Step-by-Step Explanatory Dialogues to Answer Children’s Why and How Questions. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI ’23) . Association for Computing Machinery, New York, NY, USA, Article 450, 22 pages
Yoonjoo Lee, Tae Soo Kim, Sungdong Kim, Yohan Yun, and Juho Kim. 2023 · 2023
Closest in time.
PRD: Peer Rank and Discussion Improve Large Language Model based Evaluations
Original
Ruosen Li, Teerth Patel, and Xinya Du. 2023 · 2023
Closest in time.
“What It Wants Me To Say”: Bridging the Abstraction Gap Between End-User Programmers and Code-Generating Large Language Models. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI ’23) . Association for Computing Machinery, New York, NY, USA, Article 598, 31 pages
Michael Xieyang Liu, Advait Sarkar, Carina Negreanu, Benjamin Zorn, Jack Williams, Neil Toronto, and Andrew D. Gordon. 2023b · 2023
Closest in time.
Pre-Train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing
Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2023c · 2023
Closest in time.
Self-refine: Iterative refinement with self-feedback
Original
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, et al · 2023
Closest in time.
PromptAid: Prompt Exploration, Perturbation, Testing and Iteration using Visual Analytics for Large Language Models
Original
Aditi Mishra, Utkarsh Soni, Anjana Arunkumar, Jinbin Huang, Bum Chul Kwon, and Chris Bryan. 2023 · 2023
Closest in time.
GPT-4 Technical Report
Original
OpenAI. 2023 · 2023
Closest in time.
AngleKindling: Supporting Journalistic Angle Ideation with Large Language Models. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI ’23) . Association for Computing Machinery, New York, NY, USA, Article 225, 16 pages
Savvas Petridis, Nicholas Diakopoulos, Kevin Crowston, Mark Hansen, Keren Henderson, Stan Jastrzebski, Jeffrey V Nickerson, and Lydia B Chilton. 2023 · 2023
Closest in time.
Angler: Helping Machine Translation Practitioners Prioritize Model Improvements. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI ’23) . Association for Computing Machinery, New York, NY, USA, Article 832, 20 pages
Samantha Robertson, Zijie J. Wang, Dominik Moritz, Mary Beth Kery, and Fred Hohman. 2023 · 2023
Closest in time.
Whose opinions do language models reflect?
Original
Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto. 2023 · 2023
Closest in time.
Principle-Driven Self-Alignment of Language Models from Scratch with Minimal Human Supervision
Original
Zhiqing Sun, Yikang Shen, Qinhong Zhou, Hongxin Zhang, Zhenfang Chen, David Cox, Yiming Yang, and Chuang Gan. 2023 · 2023
Closest in time.
Large language models are not fair evaluators
Original
Peiyi Wang, Lei Li, Liang Chen, Dawei Zhu, Binghuai Lin, Yunbo Cao, Qi Liu, Tianyu Liu, and Zhifang Sui. 2023a · 2023
Closest in time.
Leveraging Large Language Models to Power Chatbots for Collecting User Self-Reported Data
Original
Jing Wei, Sungdong Kim, Hyunhoon Jung, and Young-Ho Kim. 2023 · 2023
Closest in time.
ScatterShot: Interactive In-Context Example Curation for Text Transformation. In Proceedings of the 28th International Conference on Intelligent User Interfaces (Sydney, NSW, Australia) (IUI ’23) . Association for Computing Machinery, New York, NY, USA, 353–367
Sherry Wu, Hua Shen, Daniel S Weld, Jeffrey Heer, and Marco Tulio Ribeiro. 2023 · 2023
Closest in time.
Evaluating NLG Evaluation Metrics: A Measurement Theory Perspective
Original
Ziang Xiao, Susu Zhang, Vivian Lai, and Q. Vera Liao. 2023 · 2023
Closest in time.
FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets
Original
Seonghyeon Ye, Doyoung Kim, Sungdong Kim, Hyeonbin Hwang, Seungone Kim, Yongrae Jo, James Thorne, Juho Kim, and Minjoon Seo. 2023 · 2023
Closest in time.
Herding AI Cats: Lessons from Designing a Chatbot by Prompting GPT-3. In Proceedings of the 2023 ACM Designing Interactive Systems Conference (Pittsburgh, PA, USA) (DIS ’23) . Association for Computing Machinery, New York, NY, USA, 2206–2220
J.D. Zamfirescu-Pereira, Heather Wei, Amy Xiao, Kitty Gu, Grace Jung, Matthew G Lee, Bjoern Hartmann, and Qian Yang. 2023a · 2023
Closest in time.
Why Johnny Can’t Prompt: How Non-AI Experts Try (and Fail) to Design LLM Prompts. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI ’23) . Association for Computing Machinery, New York, NY, USA, Article 437, 21 pages
J.D. Zamfirescu-Pereira, Richmond Y. Wong, Bjoern Hartmann, and Qian Yang. 2023b · 2023
Closest in time.
LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention
Original
Renrui Zhang, Jiaming Han, Chris Liu, Peng Gao, Aojun Zhou, Xiangfei Hu, Shilin Yan, Pan Lu, Hongsheng Li, and Yu Qiao. 2023 · 2023
Closest in time.
Judging LLM-as-a-judge with MT-Bench and Chatbot Arena
Original
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric. P Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023 · 2023
Closest in time.