2022

Learning New Skills after Deployment: Improving open-domain internet-driven dialogue with human feedback

Xu, Jing, Ung, Megan, Komeili, Mojtaba et al.

Understand

Frozen models trained to mimic static datasets can never improve their performance.

  • Models that can employ internet-retrieval for up-to-date information and obtain feedback from humans during deployment provide the promise of both adapting to new information, and improving their performance.
  • In this work we study how to improve internet-driven conversational skills in such a learning framework.
  • We collect deployment data, which we make publicly available, of human interactions, and collect various types of human feedback -- including binary quality measurements, free-form text feedback, and fine-grained reasons for failure.

Reading the bibliography…