Fetching the paper…
Reading the bibliography…
A robust evaluation metric has a profound impact on the development of text generation systems.
Linguistic knowledge and transferability of contextual representations
Nelson F Liu, Matt Gardner, Yonatan Belinkov, Matthew Peters, and Noah A Smith. 2019 · 1903
Earlier work this paper cites.
Bertscore: Evaluating text generation with BERT
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2019 · 1904
Earlier work this paper cites.
Ian Tenney, Patrick Xia, Berlin Chen, Alex Wang, Adam Poliak, R Thomas McCoy, Najoung Kim, Benjamin Van Durme, Samuel R Bowman, Dipanjan Das, et al. 2019 · 1905
Earlier work this paper cites.
Better summarization evaluation with word embeddings for rouge
Jun-Ping Ng and Viktoria Abrecht. 2015 · 1930
Earlier work this paper cites.
Kernel smoothing
Matt P Wand and M Chris Jones. 1994 · 1994
Earlier work this paper cites.
The earth mover’s distance as a metric for image retrieval
Yossi Rubner, Carlo Tomasi, and Leonidas J. Guibas. 2000 · 2000
Earlier work this paper cites.
Mean shift: A robust approach toward feature space analysis
Dorin Comaniciu and Peter Meer. 2002 · 2002
Earlier work this paper cites.
BLEU: A Method for Automatic Evaluation of Machine Translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
ROUGE: A Package for Automatic Evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Evaluating content selection in summarization: The pyramid method
Ani Nenkova and Rebecca J. Passonneau. 2004 · 2004
Earlier work this paper cites.
Automated summarization evaluation with basic elements
Eduard Hovy, Chin-Yew Lin, Liang Zhou, and Junichi Fukumoto. 2006 · 2006
Earlier work this paper cites.
Meteor: An Automatic Metric for MT Evaluation with High Levels of Correlation with Human Judgments
Alon Lavie and Abhaya Agarwal. 2007 · 2007
Earlier work this paper cites.
Summarization Evaluation Using Transformed Basic Elements
Stephen Tratz and Eduard H Hovy. 2008 · 2008
Earlier work this paper cites.
Phrase-based statistical language generation using graphical models and active learning
François Mairesse, Milica Gašić, Filip Jurčíček, Simon Keizer, Blaise Thomson, Kai Yu, and Steve Young. 2010 · 2010
Earlier work this paper cites.
Automatically assessing machine summary content without a gold standard
Annie Louis and Ani Nenkova. 2013 · 2013
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. 2014 · 2014
Earlier work this paper cites.
Teaching machines to read and comprehend
Karl Moritz Hermann, Tomas Kocisky, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom. 2015 · 2015
Earlier work this paper cites.
From word embeddings to document distances
Matt J. Kusner, Yu Sun, Nicholas I. Kolkin, and Kilian Q. Weinberger. 2015 · 2015
Cited alongside, same era.
CIDEr: Consensus-based Image Description Evaluation
Ramakrishna Vedantam, C. Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Cited alongside, same era.
Semantically conditioned lstm-based natural language generation for spoken dialogue systems
Tsung-Hsien Wen, Milica Gasic, Nikola Mrksic, Pei-Hao Su, David Vandyke, and Steve Young. 2015 · 2015
Cited alongside, same era.
SPICE: semantic propositional image caption evaluation
Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould. 2016 · 2016
Cited alongside, same era.
A diversity-promoting objective function for neural conversation models
Jiwei Li, Michel Galley, Chris Brockett, Jianfeng Gao, and Bill Dolan. 2016 · 2016
Cited alongside, same era.
Meteor++: Incorporating copy knowledge into machine translation evaluation
Yinuo Guo, Chong Ruan, and Junfeng Hu. 2018 · 2018
Later among the works it cites.
Efficient contextualized representation: Language model pruning for sequence labeling
Liyuan Liu, Xiang Ren, Jingbo Shang, Xiaotao Gu, Jian Peng, and Jiawei Han. 2018 · 2018
Later among the works it cites.
Results of the WMT18 metrics shared task
Qingsong Ma, Ondrej Bojar, and Yvette Graham. 2018 · 2018
Later among the works it cites.
Deep contextualized word representations
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Later among the works it cites.
A call for clarity in reporting bleu scores
Matt Post. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chia-Wei Liu, Ryan Lowe, Iulian V. Serban, Michael Noseworthy, Laurent Charlin, and Joelle Pineau. 2016 · 2016
Cited alongside, same era.
Investigating capsule networks with dynamic routing for text classification
Wei Zhao, Jianbo Ye, Min Yang, Zeyang Lei, Suofei Zhang, and Zhou Zhao. 2018 · 2016
Cited alongside, same era.
Results of the WMT17 metrics shared task
Ondrej Bojar, Yvette Graham, and Amir Kamran. 2017 · 2017
Cited alongside, same era.
MEANT 2.0: Accurate semantic MT evaluation for any output language
Chi-kiu Lo. 2017 · 2017
Cited alongside, same era.
Why We Need New Evaluation Metrics for NLG
Jekaterina Novikova, Ondřej Dušek, Amanda Cercas Curry, and Verena Rieser. 2017 · 2017
Cited alongside, same era.
Learning to score system summaries for better content selection evaluation
Maxime Peyrard, Teresa Botschen, and Iryna Gurevych. 2017 · 2017
Cited alongside, same era.
Get to the point: Summarization with pointer-generator networks
Abigail See, Peter J. Liu, and Christopher D. Manning. 2017 · 2017
Cited alongside, same era.
A structured review of the validity of BLEU
Ehud Reiter. 2018 · 2018
Later among the works it cites.
Concatenated power mean word embeddings as universal cross-lingual sentence representations
Andreas Rücklé, Steffen Eger, Maxime Peyrard, and Iryna Gurevych. 2018 · 2018
Later among the works it cites.
RUSE: Regressor using sentence embeddings for automatic machine translation evaluation
Hiroki Shimanaka, Tomoyuki Kajiwara, and Mamoru Komachi. 2018 · 2018
Later among the works it cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amapreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. 2018 · 2018
Later among the works it cites.
Learning neural templates for text generation
Sam Wiseman, Stuart M. Shieber, and Alexander M. Rush. 2018 · 2018
Later among the works it cites.
Fast dynamic routing based on weighted kernel density estimation
Suofei Zhang, Wei Zhao, Xiaofu Wu, and Quan Zhou. 2018 · 2018
Later among the works it cites.
Better rewards yield better summaries: Learning to summarise without references
Florian Böhm, Yang Gao, Christian M. Meyer, Ori Shapira, Ido Dagan, and Iryna Gurevych. 2019 · 2019
Closest in time.
Sentence mover’s similarity: Automatic evaluation for multi-sentence texts
Elizabeth Clark, Asli Celikyilmaz, and Noah A. Smith. 2019 · 2019
Closest in time.
Text processing like humans do: Visually attacking and shielding NLP systems
Steffen Eger, Gözde Gül Şahin, Andreas Rücklé, Ji-Ung Lee, Claudia Schulz, Mohsen Mesgar, Krishnkant Swarnkar, Edwin Simpson, and Iryna Gurevych. 2019 · 2019
Closest in time.
Putting evaluation in context: Contextual embeddings improve machine translation evaluation
Nitika Mathur, Timothy Baldwin, and Trevor Cohn. 2019 · 2019
Closest in time.
Towards scalable and reliable capsule networks for challenging NLP applications
Wei Zhao, Haiyun Peng, Steffen Eger, Erik Cambria, and Min Yang. 2019 · 2019
Closest in time.