2023

A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity

Bang, Yejin, Cahyawijaya, Samuel, Lee, Nayeon et al.

Understand

This paper proposes a framework for quantitatively evaluating interactive LLMs such as ChatGPT using publicly available data sets.

  • We carry out an extensive technical evaluation of ChatGPT using 23 data sets covering 8 different common NLP application tasks.
  • We evaluate the multitask, multilingual and multi-modal aspects of ChatGPT based on these data sets and a newly designed multimodal dataset.
  • We find that ChatGPT outperforms LLMs with zero-shot learning on most tasks and even outperforms fine-tuned models on some tasks.

Reading the bibliography…