Fetching the paper…
Reading the bibliography…
Recently, amounts of works utilize perplexity~(PPL) to evaluate the quality of the generated text.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
A comparison of greedy and optimal assessment of natural language student input using word-to-word similarity metrics
Vasile Rus and Mihai C. Lintean. 2012 · 2012
Earlier work this paper cites.
Bootstrapping dialog systems with word embeddings
Gabriel Forgues, Joelle Pineau, Jean-Marie Larchevêque, and Réal Tremblay. 2014 · 2014
Earlier work this paper cites.
chrf: character n-gram f-score for automatic MT evaluation
Maja Popovic. 2015 · 2015
Earlier work this paper cites.
Modeling coverage for neural machine translation
Zhaopeng Tu, Zhengdong Lu, Yang Liu, Xiaohua Liu, and Hang Li. 2016 · 2016
Cited alongside, same era.
Question generation for question answering
Nan Duan, Duyu Tang, Peng Chen, and Ming Zhou. 2017 · 2017
Cited alongside, same era.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Cited alongside, same era.
Language model evaluation beyond perplexity
Clara Meister and Ryan Cotterell. 2021b
Cited in the paper.
Language model evaluation beyond perplexity
Clara Meister and Ryan Cotterell. 2021a · 2021
Later among the works it cites.
Neural machine translation with explicit phrase alignment
Jiacheng Zhang, Huanbo Luan, Maosong Sun, Feifei Zhai, Jingfang Xu, and Yang Liu. 2021 · 2021
Later among the works it cites.
MISC: A mixed strategy-aware model integrating COMET for emotional support conversation
Quan Tu, Yanran Li, Jianwei Cui, Bin Wang, Ji-Rong Wen, and Rui Yan. 2022 · 2022
Closest in time.
A survey of evaluation metrics used for NLG systems
Ananya B. Sai, Akash Kumar Mohankumar, and Mitesh M. Khapra. 2023 · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…