Fetching the paper…
Reading the bibliography…
Chatbots have been an interesting application of natural language generation since its inception.
1904
Earlier work this paper cites.
Bleu: a Method for Automatic Evaluation of Machine Translation,
K. Papineni, S. Roukos, T. Ward, W.-J. Zhu, · 2002
Earlier work this paper cites.
ROUGE: A Package for Automatic Evaluation of Summaries,
C.-Y. Lin, · 2004
Earlier work this paper cites.
2004
Earlier work this paper cites.
METEOR: An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments,
S. Banerjee, A. Lavie, · 2005
Earlier work this paper cites.
2010
Earlier work this paper cites.
P. Armstrong, Bloom’s Taxonomy, 2010. URL: https://cft.vanderbilt.edu/guides-sub-pages/blooms-taxonomy/
2010
Earlier work this paper cites.
Best practices for the human evaluation of automatically generated text,
C. van der Lee, A. Gatt, E. Van Miltenburg, S. Wubben, E. Krahmer, · 2019
Earlier work this paper cites.
Technical Metrics Used to Evaluate Health Care Chatbots: Scoping Review,
A. Abd-Alrazaq, Z. Safi, M. Alajlani, J. Warren, M. Househ, K. Denecke, · 2020
Earlier work this paper cites.
Human evaluation of automatically generated text: Current trends and best practice guidelines,
C. van der Lee, A. Gatt, E. van Miltenburg, E. Krahmer, · 2020
Earlier work this paper cites.
“This is a Problem, Don’t You Agree?” Framing and Bias in Human Evaluation for Natural Language Generation,
S. Schoch, D. Yang, Y. Ji, · 2020
Cited alongside, same era.
Algorithm Inspection for Chatbot Performance Evaluation,
V. Vijayaraghavan, J. B. Cooper, R. L. J., · 2020
Cited alongside, same era.
2021
Cited alongside, same era.
What happens if you treat ordinal ratings as interval data? Human evaluations in NLP are even more under-powered than you think,
D. M. Howcroft, V. Rieser, · 2021
Cited alongside, same era.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Survey of Hallucination in Natural Language Generation,
Z. Ji, N. Lee, R. Frieske, T. Yu, D. Su, Y. Xu, E. Ishii, Y. J. Bang, A. Madotto, P. Fung, · 2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality | LMSYS Org, 2023. URL: https://lmsys.org/blog/2023-03-30-vicuna
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
Approximating Online Human Evaluation of Social Chatbots with Prompting,
E. Svikhnushina, P. Pu, · 2023
Later among the works it cites.