Fetching the paper…
Reading the bibliography…
Automatic evaluation metrics capable of replacing human judgments are critical to allowing fast development of new methods.
Mathematics without numbers
John G Kemeny. 1959 · 1959
Earlier work this paper cites.
A consistent extension of condorcet’s election principle
H Peyton Young and Arthur Levenglick. 1978 · 1978
Earlier work this paper cites.
The computational difficulty of manipulating an election
John J Bartholdi, Craig A Tovey, and Michael A Trick. 1989 · 1989
Earlier work this paper cites.
Rank aggregation methods for the web
Cynthia Dwork, Ravi Kumar, Moni Naor, and Dandapani Sivakumar. 2001 · 2001
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Usr: An unsupervised and reference free evaluation metric for dialog generation
Shikib Mehri and Maxine Eskenazi. 2020 · 2005
Earlier work this paper cites.
Ordering by weighted number of wins gives a good ranking for weighted tournaments
Don Coppersmith, Lisa Fleischer, and Atri Rudra. 2006 · 2006
Earlier work this paper cites.
An information-theoretic approach to automatic evaluation of summaries
Chin-Yew Lin, Guihong Cao, Jianfeng Gao, and Jian-Yun Nie. 2006 · 2006
Earlier work this paper cites.
Fast unfolding of communities in large networks
Vincent D Blondel, Jean-Loup Guillaume, Renaud Lambiotte, and Etienne Lefebvre. 2008 · 2008
Earlier work this paper cites.
Overview of the tac 2008 update summarization task
Hoa Trang Dang and Karolina Owczarzak. 2008 · 2008
Earlier work this paper cites.
Hierarchical pre-training for sequence labelling in spoken dialog
Emile Chapuis, Pierre Colombo, Matteo Manica, Matthieu Labeau, and Chloe Clavel. 2020 · 2009
Earlier work this paper cites.
Overview of the tac 2011 summarization track: Guided task and aesop task
Karolina Owczarzak and Hoa Trang Dang. 2011 · 2011
Earlier work this paper cites.
Experiments with kemeny ranking: What works when?
Alnur Ali and Marina Meilă. 2012 · 2012
Cited alongside, same era.
An Assessment of the Accuracy of Automatic Evaluation in Summarization
Karolina Owczarzak, John M. Conroy, Hoa Trang Dang, and Ani Nenkova. 2012 · 2012
Cited alongside, same era.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier. 2014 · 2014
Cited alongside, same era.
Re-evaluating Automatic Summarization with BLEU and 192 Shades of ROUGE
Yvette Graham. 2015 · 2015
Cited alongside, same era.
Better summarization evaluation with word embeddings for rouge
Jun-Ping Ng and Viktoria Abrecht. 2015 · 2015
Cited alongside, same era.
Moverscore: Text generation evaluating with contextualized embeddings and earth mover distance
Wei Zhao, Maxime Peyrard, Fei Liu, Yang Gao, Christian M Meyer, and Steffen Eger. 2019 · 2019
Later among the works it cites.
Re-evaluating evaluation in text summarization
Manik Bhandari, Pranav Gour, Atabak Ashfaq, Pengfei Liu, and Graham Neubig. 2020 · 2020
Later among the works it cites.
Guiding attention in sequence-to-sequence models for dialogue act prediction
Pierre Colombo, Emile Chapuis, Matteo Manica, Emmanuel Vignon, Giovanna Varni, and Chloe Clavel. 2020 · 2020
Later among the works it cites.
Code-switched inspired losses for generic spoken dialog representations
Emile Chapuis, Pierre Colombo, Matthieu Labeau, and Chloe Clavel. 2021 · 2021
Later among the works it cites.
Learning to represent and generate text using information measures
Pierre Colombo. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jekaterina Novikova, Ondřej Dušek, Amanda Cercas Curry, and Verena Rieser. 2017 · 2017
Cited alongside, same era.
Learning to score system summaries for better content selection evaluation
Maxime Peyrard, Teresa Botschen, and Iryna Gurevych. 2017 · 2017
Cited alongside, same era.
chrf++: words helping character n-grams
Maja Popović. 2017 · 2017
Cited alongside, same era.
RankME: Reliable human ratings for natural language generation
Jekaterina Novikova, Ondřej Dušek, and Verena Rieser. 2018 · 2018
Cited alongside, same era.
A call for clarity in reporting bleu scores
Matt Post. 2018 · 2018
Cited alongside, same era.
Disney at IEST 2018: Predicting Emotions Using an Ensemble
Wojciech Witon, Pierre Colombo, Ashutosh Modi, and Mubbasir Kapadia. 2018 · 2018
Cited alongside, same era.
Affect-Driven Dialog Generation
Pierre Colombo, Wojciech Witon, Ashutosh Modi, James Kennedy, and Mubbasir Kapadia. 2019 · 2019
Cited alongside, same era.
Beam search with bidirectional strategies for neural response generation
Pierre Colombo, Chloé Clavel, Chouchang Yack, and Giovanna Varni. 2021c · 2021
Later among the works it cites.
Summeval: Re-evaluating summarization evaluation
Alexander R Fabbri, Wojciech Kryściński, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir Radev. 2021 · 2021
Later among the works it cites.
Better than average: Paired evaluation of nlp systems
Maxime Peyrard, Wei Zhao, Steffen Eger, and Robert West. 2021 · 2021
Later among the works it cites.
Tharindu Ranasinghe, Constantin Orasan, and Ruslan Mitkov. 2021 · 2021
Later among the works it cites.
Depth-based pseudo-metrics between probability distributions
Guillaume Staerman, Pavlo Mozharovskyi, Stéphan Clémençon, and Florence d’Alché Buc. 2021 · 2021
Later among the works it cites.
Of human criteria and automatic metrics: A benchmark of the evaluation of story generation
Cyril Chhun, Pierre Colombo, Chloé Clavel, and Fabian M Suchanek. 2022 · 2022
Closest in time.
Principal component analysis: a review and recent developments
Ian T Jolliffe and Jorge Cadima. 2016 · 2065
Closest in time.