Fetching the paper…
Reading the bibliography…
The long-standing one-to-many problem of gold standard responses in open-domain dialogue systems presents challenges for automatic evaluation metrics.
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020 · 1904
Earlier work this paper cites.
Sentence-bert: Sentence embeddings using siamese bert-networks
Nils Reimers and Iryna Gurevych. 2019 · 1908
Earlier work this paper cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 1910
Earlier work this paper cites.
A coefficient of agreement for nominal scales
Jacob Cohen. 1960 · 1960
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie. 2005 · 2005
Earlier work this paper cites.
Bootstrapping dialog systems with word embeddings
Gabriel Forgues, Joelle Pineau, Jean-Marie Larchevêque, and Réal Tremblay. 2014 · 2014
Earlier work this paper cites.
Towards universal paraphrastic sentence embeddings
John Wieting, Mohit Bansal, Kevin Gimpel, and Karen Livescu. 2016 · 2016
Earlier work this paper cites.
Ruber: An unsupervised method for automatic evaluation of open-domain dialog systems
Chongyang Tao, Lili Mou, Dongyan Zhao, and Rui Yan. 2018 · 2018
Earlier work this paper cites.
Personalizing dialogue agents: I have a dog, do you have pets too?
Saizheng Zhang, Emily Dinan, Jack Urbanek, Arthur Szlam, Douwe Kiela, and Jason Weston. 2018 · 2018
Earlier work this paper cites.
Topical-chat: Towards knowledge-grounded open-domain conversations
Karthik Gopalakrishnan, Behnam Hedayatnia, Qinlang Chen, Anna Gottardi, Sanjeev Kwatra, Anu Venkatesh, Raefer Gabriel, and Dilek Z. Hakkani-Tür. 2019 · 2019
Cited alongside, same era.
Speaker-aware bert for multi-turn response selection in retrieval-based chatbots
Jia-Chen Gu, Tianda Li, Quan Liu, Xiaodan Zhu, Zhenhua Ling, Zhiming Su, and Si Wei. 2020 · 2020
Cited alongside, same era.
Improving dialog evaluation with a multi-reference adversarial dataset and large scale pretraining
Ananya B. Sai, Akash Kumar Mohankumar, Siddharth Arora, and Mitesh M. Khapra. 2020 · 2020
Cited alongside, same era.
Bridges-2: A platform for rapidly-evolving and data intensive research
Shawn T Brown, Paola Buitrago, Edward Hanna, Sergiu Sanielevici, Robin Scibek, and Nicholas A Nystrom. 2021 · 2021
Cited alongside, same era.
Synthesizing adversarial negative responses for robust response ranking and evaluation
Prakhar Gupta, Yulia Tsvetkov, and Jeffrey P. Bigham. 2021 · 2021
Large language models are state-of-the-art evaluators of translation quality
Tom Kocmi and Christian Federmann. 2023 · 2023
Later among the works it cites.
LLM-eval: Unified multi-dimensional automatic evaluation for open-domain conversations with large language models
Yen-Ting Lin and Yun-Nung Chen. 2023 · 2023
Later among the works it cites.
G-eval: NLG evaluation using gpt-4 with better human alignment
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023 · 2023
Later among the works it cites.
The iron (ic) melting pot: Reviewing human evaluation in humour, irony and sarcasm generation
Tyler Loakman, Aaron Maladry, and Chenghua Lin. 2023 · 2023
Later among the works it cites.
Improving biomedical abstractive summarisation with knowledge aggregation from citation papers
Chen Tang, Shun Wang, Tomas Goldsack, and Chenghua Lin. 2023a · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Fast and scalable dialogue state tracking with explicit modular decomposition
Dingmin Wang, Chenghua Lin, Qi Liu, and Kam-Fai Wong. 2021 · 2021
Cited alongside, same era.
Bartscore: Evaluating generated text as text generation
Weizhe Yuan, Graham Neubig, and Pengfei Liu. 2021 · 2021
Cited alongside, same era.
Affective decoding for empathetic response generation
Chengkun Zeng, Guanyi Chen, Chenghua Lin, Ruizhe Li, and Zhi Chen. 2021 · 2021
Cited alongside, same era.
Mdd-eval: Self-training on augmented data for multi-domain dialogue evaluation
Chen Zhang, L. F. D’Haro, Thomas Friedrichs, and Haizhou Li. 2021 · 2021
Cited alongside, same era.
Can large language models be an alternative to human evaluations?
Cheng-Han Chiang and Hung yi Lee. 2023 · 2023
Cited alongside, same era.
Gptscore: Evaluate as you desire
Jinlan Fu, See-Kiong Ng, Zhengbao Jiang, and Pengfei Liu. 2023 · 2023
Cited alongside, same era.
Enhancing dialogue generation via dynamic graph knowledge aggregation
Chen Tang, Hongbo Zhang, Tyler Loakman, Chenghua Lin, and Frank Guerin. 2023b
Cited in the paper.
Evaluating evaluation metrics: A framework for analyzing nlg evaluation metrics using measurement theory
Ziang Xiao, Susu Zhang, Vivian Lai, and Q Vera Liao. 2023 · 2023
Later among the works it cites.
Evaluating open-domain dialogues in latent space with next sentence prediction and mutual information
Kun Zhao, Bohao Yang, Chenghua Lin, Wenge Rong, Aline Villavicencio, and Xiaohui Cui. 2023 · 2023
Later among the works it cites.
Improving medical dialogue generation with abstract meaning representations
Bohao Yang, Chen Tang, and Chenghua Lin. 2024a · 2024
Closest in time.
Effective distillation of table-based reasoning ability from LLMs
Bohao Yang, Chen Tang, Kun Zhao, Chenghao Xiao, and Chenghua Lin. 2024b · 2024
Closest in time.
Towards a unified multi-dimensional evaluator for text generation
Ming Zhong, Yang Liu, Da Yin, Yuning Mao, Yizhu Jiao, Pengfei Liu, Chenguang Zhu, Heng Ji, and Jiawei Han. 2022 · 2038
Closest in time.