Fetching the paper…
Reading the bibliography…
Automatic evaluation is an integral aspect of dialogue system research.
Language Models are Few-Shot Learners
Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; et al. 2020 · 1901
Earlier work this paper cites.
Bleu: a Method for Automatic Evaluation of Machine Translation
Papineni, K.; Roukos, S.; Ward, T.; and Zhu, W.-J. 2002 · 2002
Earlier work this paper cites.
How NOT To Evaluate Your Dialogue System: An Empirical Study of Unsupervised Evaluation Metrics for Dialogue Response Generation
Liu, C.-W.; Lowe, R.; Serban, I.; Noseworthy, M.; Charlin, L.; and Pineau, J. 2016 · 2016
Earlier work this paper cites.
DailyDialog: A Manually Labelled Multi-turn Dialogue Dataset
Li, Y.; Su, H.; Shen, X.; Li, W.; Cao, Z.; and Niu, S. 2017 · 2017
Earlier work this paper cites.
Personalizing Dialogue Agents: I have a dog, do you have pets too?
Zhang, S.; Dinan, E.; Urbanek, J.; Szlam, A.; Kiela, D.; and Weston, J. 2018 · 2018
Earlier work this paper cites.
What makes a good conversation? How controllable attributes affect human judgments
See, A.; Roller, S.; Kiela, D.; and Weston, J. 2019 · 2019
Earlier work this paper cites.
MIME: MIMicking Emotions for Empathetic Response Generation
Majumder, N.; Hong, P.; Peng, S.; Lu, J.; Ghosal, D.; Gelbukh, A.; Mihalcea, R.; and Poria, S. 2020 · 2020
Earlier work this paper cites.
Designing Precise and Robust Dialogue Response Evaluators
Zhao, T.; Lala, D.; and Kawahara, T. 2020 · 2020
Earlier work this paper cites.
Robust Machine Reading Comprehension by Learning Soft labels
Zhao, Z.; Wu, S.; Yang, M.; Chen, K.; and Zhao, T. 2020 · 2020
Earlier work this paper cites.
The Perils of Using Mechanical Turk to Evaluate Open-Ended Text Generation
Karpinska, M.; Akoury, N.; and Iyyer, M. 2021 · 2021
Earlier work this paper cites.
I like fish, especially dolphins: Addressing Contradictions in Dialogue Modeling
Nie, Y.; Williamson, M.; Bansal, M.; Kiela, D.; and Weston, J. 2021 · 2021
Earlier work this paper cites.
Recipes for Building an Open-Domain Chatbot
Roller, S.; Dinan, E.; Goyal, N.; Ju, D.; Williamson, M.; Liu, Y.; Xu, J.; Ott, M.; Smith, E. M.; Boureau, Y.-L.; and Weston, J. 2021 · 2021
Earlier work this paper cites.
Perturbation CheckLists for Evaluating NLG Evaluation Metrics
Sai, A. B.; Dixit, T.; Sheth, D. Y.; Mohan, S.; and Khapra, M. M. 2021 · 2021
Earlier work this paper cites.
A Comprehensive Assessment of Dialog Evaluation Metrics
Yeh, Y.-T.; Eskenazi, M.; and Mehri, S. 2021 · 2021
Earlier work this paper cites.
DynaEval: Unifying Turn and Dialogue Level Evaluation
Zhang, C.; Chen, Y.; D’Haro, L. F.; Zhang, Y.; Friedrichs, T.; Lee, G.; and Li, H. 2021 · 2021
Cited alongside, same era.
Constitutional AI: Harmlessness from AI Feedback
Bai, Y.; Kadavath, S.; Kundu, S.; Askell, A.; Kernion, J.; Jones, A.; et al. 2022 · 2022
Cited alongside, same era.
PaLM: Scaling Language Modeling with Pathways
Chowdhery, A.; Narang, S.; Devlin, J.; Bosma, M.; Mishra, G.; Roberts, A.; et al. 2022 · 2022
Cited alongside, same era.
Scaling Instruction-Finetuned Language Models
Chung, H. W.; Hou, L.; Longpre, S.; Zoph, B.; Tay, Y.; Fedus, W.; et al. 2022 · 2022
Cited alongside, same era.
What is wrong with you?: Leveraging User Sentiment for Automatic Dialog Evaluation
Ghazarian, S.; Hedayatnia, B.; Papangelis, A.; Liu, Y.; and Hakkani-Tur, D. 2022a · 2022
Cited alongside, same era.
Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality
Chiang, W.-L.; Li, Z.; Lin, Z.; Sheng, Y.; Wu, Z.; Zhang, H.; et al. 2023 · 2023
Closest in time.
Free Dolly: Introducing the World’s First Truly Open Instruction-Tuned LLM
Conover, M.; Hayes, M.; Mathur, A.; Xie, J.; Wan, J.; Shah, S.; et al. 2023 · 2023
Closest in time.
GPTScore: Evaluate as You Desire
Fu, J.; Ng, S.-K.; Jiang, Z.; and Liu, P. 2023 · 2023
Closest in time.
OpenLLaMA: An Open Reproduction of LLaMA
Geng, X.; and Liu, H. 2023 · 2023
Closest in time.
ChatGPT Outperforms Crowd-Workers for Text-Annotation Tasks
Gilardi, F.; Alizadeh, M.; and Kubli, M. 2023 · 2023
Closest in time.
Understanding the Effectiveness of Very Large Language Models on Dialog Evaluation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
InstructDial: Improving Zero and Few-shot Generalization in Dialogue through Instruction Tuning
Gupta, P.; Jiao, C.; Yeh, Y.-T.; Mehri, S.; Eskenazi, M.; and Bigham, J. 2022 · 2022
Cited alongside, same era.
Achieving Reliable Human Assessment of Open-Domain Dialogue Systems
Ji, T.; Graham, Y.; Jones, G.; Lyu, C.; and Liu, Q. 2022 · 2022
Cited alongside, same era.
Explaining Dialogue Evaluation Metrics using Adversarial Behavioral Analysis
Khalid, B.; and Lee, S. 2022 · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L.; and et al. 2022 · 2022
Cited alongside, same era.
Human Evaluation of Conversations is an Open Problem: comparing the sensitivity of various methods for evaluating dialogue agents
Smith, E.; Hsu, O.; Qian, R.; Roller, S.; Boureau, Y.-L.; and Weston, J. 2022 · 2022
Cited alongside, same era.
iEval: Interactive Evaluation Framework for Open-Domain Empathetic Chatbots
Svikhnushina, E.; Filippova, A.; and Pu, P. 2022 · 2022
Cited alongside, same era.
FineD-Eval: Fine-grained Automatic Dialogue-Level Evaluation
Zhang, C.; D’Haro, L. F.; Zhang, Q.; Friedrichs, T.; and Li, H. 2022a · 2022
Cited alongside, same era.
Huynh, J.; Jiao, C.; Gupta, P.; Mehri, S.; Bajaj, P.; Chaudhary, V.; and Eskenazi, M. 2023 · 2023
Closest in time.
LLM-Eval: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models
Lin, Y.-T.; and Chen, Y.-N. 2023 · 2023
Closest in time.
G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment
Liu, Y.; Iter, D.; Xu, Y.; Wang, S.; Xu, R.; and Zhu, C. 2023 · 2023
Closest in time.
Long Sequence Modeling with XGen: A 7B LLM Trained on 8K Input Sequence Length
Nijkamp, E.; Xie, T.; Hayashi, H.; Pang, B.; Xia, C.; Xing, C.; et al. 2023 · 2023
Closest in time.
BLOOM: A 176B-Parameter Open-Access Multilingual Language Model
Scao, T. L.; Fan, A.; Akiki, C.; Pavlick, E.; Ilić, S.; Hesslow, D.; et al. 2023 · 2023
Closest in time.
Stanford Alpaca: An Instruction-following LLaMA model
Taori, R.; Gulrajani, I.; Zhang, T.; Dubois, Y.; Li, X.; Guestrin, C.; Liang, P.; and Hashimoto, T. B. 2023 · 2023
Closest in time.
How Far Can Camels Go? Exploring the State of Instruction Tuning on Open Resources
Wang, Y.; Ivison, H.; Dasigi, P.; Hessel, J.; Khot, T.; Chandu, K. R.; Wadden, D.; MacMillan, K.; Smith, N. A.; Beltagy, I.; and Hajishirzi, H. 2023 · 2023
Closest in time.
GLM-130B: An Open Bilingual Pre-trained Model
Zeng, A.; Liu, X.; Du, Z.; Wang, Z.; Lai, H.; Ding, M.; et al. 2023 · 2023
Closest in time.
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Zheng, L.; Chiang, W.-L.; Sheng, Y.; Zhuang, S.; Wu, Z.; Zhuang, Y.; et al. 2023 · 2023
Closest in time.