Fetching the paper…
Reading the bibliography…
Automatic evaluation metrics are indispensable for evaluating generated text.
Evaluating coherence in dialogue systems using entailment
Nouha Dziri, Ehsan Kamalloo, Kory W Mathewson, and Osmar Zaiane. 2019 · 1904
Earlier work this paper cites.
Language models with transformers
Chenguang Wang, Mu Li, and Alexander J Smola. 2019 · 1904
Earlier work this paper cites.
Albert: A lite bert for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2019 · 1909
Earlier work this paper cites.
Regression Analysis , volume 14
Evan James Williams. 1959 · 1959
Earlier work this paper cites.
Centering theory in discourse
Joshi Prince Walker. 1998 · 1998
Earlier work this paper cites.
Beyond elaboration: The interaction of relations and focus in coherent text
Alistair Knott, Jon Oberlander, Mick O’Donnell, and Chris Mellish. 2001 · 2001
Earlier work this paper cites.
Automatic evaluation of machine translation quality using n-gram co-occurrence statistics
George Doddington. 2002 · 2002
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Automatic evaluation of summaries using n-gram co-occurrence statistics
Chin-Yew Lin and Eduard Hovy. 2003 · 2003
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Evaluating content selection in summarization: The pyramid method
Ani Nenkova and Rebecca Passonneau. 2004 · 2004
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie. 2005 · 2005
Earlier work this paper cites.
Overview of duc 2006
Hoa Trang Dang. 2006 · 2006
Earlier work this paper cites.
Unsupervised multilingual sentence boundary detection
Tibor Kiss and Jan Strunk. 2006 · 2006
Earlier work this paper cites.
A study of translation edit rate with targeted human annotation
Matthew Snover, Bonnie Dorr, Richard Schwartz, Linnea Micciulla, and John Makhoul. 2006 · 2006
Earlier work this paper cites.
Duc in context
Paul Over, Hoa Dang, and Donna Harman. 2007 · 2007
Earlier work this paper cites.
Mind the gap: Dangers of divorcing evaluations of summary content from linguistic quality
John M Conroy and Hoa Trang Dang. 2008 · 2008
Earlier work this paper cites.
The meteor metric for automatic evaluation of machine translation
Alon Lavie and Michael J Denkowski. 2009 · 2009
Earlier work this paper cites.
Learning to predict readability using diverse linguistic features
Rohit J Kate, Xiaoqiang Luo, Siddharth Patwardhan, Martin Franz, Radu Florian, Raymond J Mooney, Salim Roukos, and Chris Welty. 2010 · 2010
Cited alongside, same era.
Phrase-based statistical language generation using graphical models and active learning
François Mairesse, Milica Gašić, Filip Jurčíček, Simon Keizer, Blaise Thomson, Kai Yu, and Steve Young. 2010 · 2010
Cited alongside, same era.
Automatic evaluation of linguistic quality in multi-document summarization
Emily Pitler, Annie Louis, and Ani Nenkova. 2010 · 2010
Cited alongside, same era.
Machine translation evaluation and optimization
Bonnie Dorr, Joseph Olive, John McCary, and Caitlin Christianson. 2011 · 2011
Cited alongside, same era.
Continuous measurement scales in human evaluation of machine translation
Yvette Graham, Timothy Baldwin, Alistair Moffat, and Justin Zobel. 2013 · 2013
Cited alongside, same era.
Why we need new evaluation metrics for nlg
Jekaterina Novikova, Ondřej Dušek, Amanda Cercas Curry, and Verena Rieser. 2017 · 2017
Later among the works it cites.
The price of debiasing automatic metrics in natural language evalaution
Arun Chaganty, Stephen Mussmann, and Percy Liang. 2018 · 2018
Later among the works it cites.
Unsupervised learning of sentence embeddings using compositional n-gram features
Matteo Pagliardini, Prakhar Gupta, and Martin Jaggi. 2018 · 2018
Later among the works it cites.
A graph-theoretic summary evaluation for rouge
Elaheh ShafieiBavani, Mohammad Ebrahimi, Raymond Wong, and Fang Chen. 2018 · 2018
Later among the works it cites.
Ruse: Regressor using sentence embeddings for automatic machine translation evaluation
Hiroki Shimanaka, Tomoyuki Kajiwara, and Mamoru Komachi. 2018 · 2018
Later among the works it cites.
Machine translation quality estimation: Applications and future perspectives
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean. 2013 · 2013
Cited alongside, same era.
Meteor universal: Language specific translation evaluation for any target language
Michael Denkowski and Alon Lavie. 2014 · 2014
Cited alongside, same era.
Testing for significance of increased correlation with human judgment
Yvette Graham and Timothy Baldwin. 2014 · 2014
Cited alongside, same era.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D Manning. 2014 · 2014
Cited alongside, same era.
Teaching machines to read and comprehend
Karl Moritz Hermann, Tomas Kocisky, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom. 2015 · 2015
Cited alongside, same era.
From word embeddings to document distances
Matt Kusner, Yu Sun, Nicholas Kolkin, and Kilian Weinberger. 2015 · 2015
Cited alongside, same era.
Better summarization evaluation with word embeddings for rouge
Jun-Ping Ng and Viktoria Abrecht. 2015 · 2015
Cited alongside, same era.
Lucia Specia and Kashif Shah. 2018 · 2018
Later among the works it cites.
Neural network acceptability judgments
Alex Warstadt, Amanpreet Singh, and Samuel R Bowman. 2018 · 2018
Later among the works it cites.
Quality expectations of machine translation
Andy Way. 2018 · 2018
Later among the works it cites.
Sentence mover’s similarity: Automatic evaluation for multi-sentence texts
Elizabeth Clark, Asli Celikyilmaz, and Noah A Smith. 2019 · 2019
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Meteor++ 2.0: Adopt syntactic level paraphrase knowledge into machine translation evaluation
Yinuo Guo and Junfeng Hu. 2019 · 2019
Later among the works it cites.
Bert has a mouth, and it must speak: Bert as a markov random field language model
Alex Wang and Kyunghyun Cho. 2019 · 2019
Later among the works it cites.
Sum-qe: a bert-based summary quality estimation model
Stratos Xenouleas, Prodromos Malakasiotis, Marianna Apidianaki, and Ion Androutsopoulos. 2019 · 2019
Later among the works it cites.
Moverscore: Text generation evaluating with contextualized embeddings and earth mover distance
Wei Zhao, Maxime Peyrard, Fei Liu, Yang Gao, Christian M Meyer, and Steffen Eger. 2019 · 2019
Later among the works it cites.
Supert: Towards new frontiers in unsupervised evaluation metrics for multi-document summarization
Yang Gao, Wei Zhao, and Steffen Eger. 2020 · 2020
Closest in time.
Facet-aware evaluation for extractive summarization
Yuning Mao, Liyuan Liu, Qi Zhu, Xiang Ren, and Jiawei Han. 2020 · 2020
Closest in time.
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2020 · 2020
Closest in time.