Fetching the paper…
Reading the bibliography…
There is significant interest in developing evaluation metrics which accurately estimate the quality of generated text without the aid of a human-written reference text, which can be time consuming and expensive to collect or entirely unavailable in online applications.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, M, ar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
BLEU: a Method for Automatic Evaluation of Machine Translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Minimum Bayes-Risk Decoding for Statistical Machine Translation
Shankar Kumar and William Byrne. 2004 · 2004
Earlier work this paper cites.
ROUGE: A Package for Automatic Evaluation of Summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
A Smorgasbord of Features for Statistical Machine Translation
Franz Josef Och, Daniel Gildea, Sanjeev Khudanpur, Anoop Sarkar, Kenji Yamada, Alex Fraser, Shankar Kumar, Libin Shen, David Smith, Katherine Eng, Viren Jain, Zhen Jin, and Dragomir Radev. 2004 · 2004
Earlier work this paper cites.
Discriminative Reranking for Machine Translation
Libin Shen, Anoop Sarkar, and Franz Josef Och. 2004 · 2004
Earlier work this paper cites.
A Class of Submodular Functions for Document Summarization
Hui Lin and Jeff Bilmes. 2011 · 2011
Earlier work this paper cites.
Simple-QE: Better Automatic Quality Estimation for Text Simplification
Reno Kriz, Marianna Apidianaki, and Chris Callison-Burch. 2020 · 2012
Earlier work this paper cites.
Automatically Assessing Machine Summary Content Without a Gold Standard
Annie Louis and Ani Nenkova. 2013 · 2013
Earlier work this paper cites.
Re-evaluating Automatic Summarization with BLEU and 192 Shades of ROUGE
Yvette Graham. 2015 · 2015
Earlier work this paper cites.
Abstractive Text Summarization using Sequence-to-sequence RNNs and Beyond
Ramesh Nallapati, Bowen Zhou, Cicero dos Santos, Çağlar Gulçehre, and Bing Xiang. 2016 · 2016
Earlier work this paper cites.
SummaRuNNer: A Recurrent Neural Network Based Sequence Model for Extractive Summarization of Documents
Ramesh Nallapati, Feifei Zhai, and Bowen Zhou. 2017 · 2017
Earlier work this paper cites.
Reference-less Quality Estimation of Text Simplification Systems
Louis Martin, Samuel Humeau, Pierre-Emmanuel Mazaré, Éric de La Clergerie, Antoine Bordes, and Benoît Sagot. 2018 · 2018
Earlier work this paper cites.
A Call for Clarity in Reporting BLEU Scores
Matt Post. 2018 · 2018
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Findings of the WMT 2019 Shared Tasks on Quality Estimation
Erick Fonseca, Lisa Yankovskaya, André F. T. Martins, Mark Fishel, and Christian Federmann. 2019 · 2019
Cited alongside, same era.
Results of the WMT19 Metrics Shared Task: Segment-Level and Strong MT Systems Pose Big Challenges
Qingsong Ma, Johnny Wei, Ondřej Bojar, and Yvette Graham. 2019 · 2019
Cited alongside, same era.
Facebook FAIR’s WMT19 News Translation Task Submission
Nathan Ng, Kyra Yee, Alexei Baevski, Myle Ott, Michael Auli, and Sergey Edunov. 2019 · 2019
Cited alongside, same era.
Answers Unite! Unsupervised Metrics for Reinforced Summarization Models
Thomas Scialom, Sylvain Lamprier, Benjamin Piwowarski, and Jacopo Staiano. 2019 · 2019
Cited alongside, same era.
SUM-QE: a BERT-based Summary Quality Estimation Model
Stratos Xenouleas, Prodromos Malakasiotis, Marianna Apidianaki, and Ion Androutsopoulos. 2019 · 2019
Cited alongside, same era.
Re-evaluating Evaluation in Text Summarization
Fill in the BLANC: Human-free quality estimation of document summaries
Oleg Vasilyev, Vedant Dharnidharka, and John Bohannon. 2020 · 2020
Later among the works it cites.
BERTScore: Evaluating Text Generation with BERT
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020 · 2020
Later among the works it cites.
Assessing Reference-Free Peer Evaluation for Machine Translation
Sweta Agrawal, George Foster, Markus Freitag, and Colin Cherry. 2021 · 2021
Later among the works it cites.
SummEval: Re-evaluating Summarization Evaluation
Alexander Fabbri, Wojciech Kryscinski, Bryan McCann, R. Socher, and Dragomir Radev. 2021 · 2021
Later among the works it cites.
Results of the WMT21 Metrics Shared Task: Evaluating Metrics with Expert-based Human Evaluations on TED and News Domain
Markus Freitag, Ricardo Rei, Nitika Mathur, Chi kiu Lo, Craig Stewart, George Foster, Alon Lavie, and Ondřej Bojar. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Manik Bhandari, Pranav Narayan Gour, Atabak Ashfaq, Pengfei Liu, and Graham Neubig. 2020 · 2020
Cited alongside, same era.
SacreROUGE: An Open-Source Library for Using and Developing Summarization Evaluation Metrics
Daniel Deutsch and Dan Roth. 2020 · 2020
Cited alongside, same era.
SUPERT: Towards New Frontiers in Unsupervised Evaluation Metrics for Multi-Document Summarization
Yang Gao, Wei Zhao, and Steffen Eger. 2020 · 2020
Cited alongside, same era.
BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020 · 2020
Cited alongside, same era.
USR: An Unsupervised and Reference Free Evaluation Metric for Dialog Generation
Shikib Mehri and Maxine Eskenazi. 2020 · 2020
Cited alongside, same era.
COMET: A Neural Framework for MT Evaluation
Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. 2020 · 2020
Cited alongside, same era.
BLEURT: Learning Robust Metrics for Text Generation
Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020 · 2020
Cited alongside, same era.
CLIPScore: A Reference-free Evaluation Metric for Image Captioning
Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi. 2021 · 2021
Later among the works it cites.
Q 2 Q^{2} : Evaluating Factual Consistency in Knowledge-Grounded Dialogues via Question Generation and Question Answering
Or Honovich, Leshem Choshen, Roee Aharoni, Ella Neeman, Idan Szpektor, and Omri Abend. 2021 · 2021
Later among the works it cites.
Are References Really Needed? Unbabel-IST 2021 Submission for the Metrics Shared Task
Ricardo Rei, Ana C Farinha, Chrysoula Zerva, Daan van Stigt, Craig Stewart, Pedro Ramos, Taisiya Glushkova, André FT Martins, and Alon Lavie. 2021 · 2021
Later among the works it cites.
QuestEval: Summarization Asks for Fact-based Evaluation
Thomas Scialom, Paul-Alexis Dray, Sylvain Lamprier, Benjamin Piwowarski, Jacopo Staiano, Alex Wang, and Patrick Gallinari. 2021 · 2021
Later among the works it cites.
Findings of the WMT 2021 Shared Task on Quality Estimation
Lucia Specia, Frédéric Blain, Marina Fomicheva, Chrysoula Zerva, Zhenhao Li, Vishrav Chaudhary, and André F. T. Martins. 2021 · 2021
Later among the works it cites.
Finding a Balanced Degree of Automation for Summary Evaluation
Shiyue Zhang and Mohit Bansal. 2021 · 2021
Later among the works it cites.
Daniel Deutsch and Dan Roth. 2022 · 2022
Closest in time.
Quality-Aware Decoding for Neural Machine Translation
Fern, Patrick es, António Farinhas, Ricardo Rei, José De Souza, Perez Ogayo, Graham Neubig, and Andre Martins. 2022 · 2022
Closest in time.
High Quality Rather than High Model Probability: Minimum Bayes Risk Decoding with Neural Metrics
Markus Freitag, David Grangier, Qijun Tan, and Bowen Liang. 2022 · 2022
Closest in time.