2023

EvalLM: Interactive Evaluation of Large Language Model Prompts on User-Defined Criteria

Kim, Tae Soo, Lee, Yoonjoo, Shin, Jamin et al.

Understand

By simply composing prompts, developers can prototype novel generative applications with Large Language Models (LLMs).

  • To refine prototypes into products, however, developers must iteratively revise prompts by evaluating outputs to diagnose weaknesses.
  • Formative interviews (N=8) revealed that developers invest significant effort in manually evaluating outputs as they assess context-specific and subjective criteria.
  • We present EvalLM, an interactive system for iteratively refining prompts by evaluating multiple outputs on user-defined criteria.

Reading the bibliography…