2023

Large Language Models Are Not Robust Multiple Choice Selectors

Zheng, Chujie, Zhou, Hao, Meng, Fandong et al.

Understand

Multiple choice questions (MCQs) serve as a common yet important task format in the evaluation of large language models (LLMs).

  • This work shows that modern LLMs are vulnerable to option position changes in MCQs due to their inherent "selection bias", namely, they prefer to select specific option IDs as answers (like "Option A").
  • Through extensive empirical analyses with 20 LLMs on three benchmarks, we pinpoint that this behavioral bias primarily stems from LLMs' token bias, where the model a priori assigns more probabilistic mass to specific option ID tokens (e.g., A/B/C/D) when predicting answers from the option IDs.
  • To mitigate selection bias, we propose a label-free, inference-time debiasing method, called PriDe, which separates the model's prior bias for option IDs from the overall prediction distribution.

Reading the bibliography…