Fetching the paper…
Reading the bibliography…
We propose BERTScore, an automatic evaluation metric for text generation.
Linguistic knowledge and transferability of contextual representations
Nelson F. Liu, Matt Gardner, Yonatan Belinkov, Matthew E. Peters, and Noah A. Smith · 1903
Earlier work this paper cites.
Simple applications of BERT for ad hoc document retrieval
Wei Yang, Haotian Zhang, and Jimmy Lin · 1903
Earlier work this paper cites.
Roberta: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 1907
Earlier work this paper cites.
Regression analysis
Evan James Williams · 1959
Earlier work this paper cites.
Binary Codes Capable of Correcting Deletions, Insertions and Rever sals
Vladimir Iosifovich Levenshtein · 1966
Earlier work this paper cites.
Accelerated dp based search for statistical translation
Christoph Tillmann, Stephan Vogel, Hermann Ney, Arkaitz Zubiaga, and Hassan Sawaf · 1997
Earlier work this paper cites.
A metric for distributions with applications to image databases
Yossi Rubner, Carlo Tomasi, and Leonidas J Guibas · 1998
Earlier work this paper cites.
Automatic evaluation of machine translation quality using n-gram co-occurrence statistics
George Doddington · 2002
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
METEOR: An automatic metric for mt evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie · 2005
Earlier work this paper cites.
Measuring the semantic similarity of texts
Courtney Corley and Rada Mihalcea · 2005
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
William B Dolan and Chris Brockett · 2005
Earlier work this paper cites.
CDER: Efficient MT evaluation using block movements
Gregor Leusch, Nicola Ueffing, and Hermann Ney · 2006
Earlier work this paper cites.
A study of translation edit rate with targeted human annotation
Matthew Snover, Bonnie Dorr, Richard Schwartz, Linnea Micciulla, and John Makhoul · 2006
Earlier work this paper cites.
Moses: Open source toolkit for statistical machine translation
Philipp Koehn, Hieu Hoang, Alexandra Birch, Chris Callison-Burch, Marcello Federico, Nicola Bertoldi, Brooke Cowan, Wade Shen, Christine Moran, Richard Zens, Chris Dyer, Ondrej Bojar, Alexandra Constantin, and Evan Herbst · 2007
Earlier work this paper cites.
Automatic evaluation of translation quality for distant language pairs
Hideki Isozaki, Tsutomu Hirao, Kevin Duh, Katsuhito Sudoh, and Hajime Tsukada · 2010
Earlier work this paper cites.
A comparison of greedy and optimal assessment of natural language student input using word-to-word similarity metrics
Vasile Rus and Mihai Lintean · 2012
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Gregory S. Corrado, and Jeffrey Dean · 2013
Earlier work this paper cites.
Meteor universal: Language specific translation evaluation for any target language
Michael Denkowski and Alon Lavie · 2014
Earlier work this paper cites.
Testing for significance of increased correlation with human judgment
Yvette Graham and Timothy Baldwin · 2014
Earlier work this paper cites.
Microsoft COCO: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge J. Belongie, Lubomir D. Bourdev, Ross B. Girshick, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D. Manning · 2014
Earlier work this paper cites.
Beer: Better evaluation as ranking
Miloš Stanojević and Khalil Sima’an · 2014
Earlier work this paper cites.
deltaBLEU: A discriminative metric for generation tasks with intrinsically diverse targets
Michel Galley, Chris Brockett, Alessandro Sordoni, Yangfeng Ji, Michael Auli, Chris Quirk, Margaret Mitchell, Jianfeng Gao, and William B. Dolan · 2015
Earlier work this paper cites.
From word embeddings to document distances
Matt Kusner, Yu Sun, Nicholas Kolkin, and Kilian Weinberger · 2015
Earlier work this paper cites.
chrf: character n-gram f-score for automatic mt evaluation
Maja Popović · 2015
Earlier work this paper cites.
CIDEr: Consensus-based image description evaluation
Ramakrishna Vedantam, C. Lawrence Zitnick, and Devi Parikh · 2015
Cited alongside, same era.
SPICE: Semantic propositional image caption evaluation
Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould · 2016
Cited alongside, same era.
Results of the WMT16 metrics shared task
Ondřej Bojar, Yvette Graham, Amir Kamran, and Miloš Stanojević · 2016
Cited alongside, same era.
Achieving accurate conclusions in evaluation of automatic machine translation metrics
Yvette Graham and Qun Liu · 2016
Cited alongside, same era.
A decomposable attention model for natural language inference
Ankur Parikh, Oscar Täckström, Dipanjan Das, and Jakob Uszkoreit · 2016
Cited alongside, same era.
A dataset and evaluation metrics for abstractive compression of sentences and short paragraphs
Kristina Toutanova, Chris Brockett, Ke M Tran, and Saleema Amershi · 2016
A call for clarity in reporting BLEU scores
Matt Post · 2018
Later among the works it cites.
Concatenated power mean word embeddings as universal cross-lingual sentence representations
Andreas Rücklé, Steffen Eger, Maxime Peyrard, and Iryna Gurevych · 2018
Later among the works it cites.
Shikhar Sharma, Layla El Asri, Hannes Schulz, and Jeremie Zumer · 2018
Later among the works it cites.
Ruse: Regressor using sentence embeddings for automatic machine translation evaluation
Hiroki Shimanaka, Tomoyuki Kajiwara, and Mamoru Komachi · 2018
Later among the works it cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel Bowman · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Character: Translation edit rate on character level
Weiyue Wang, Jan-Thorsten Peter, Hendrik Rosendahl, and Hermann Ney · 2016
Cited alongside, same era.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V. Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Jeff Klingner, Apurva Shah, Melvin Johnson, Xiaobing Liu, Lukasz Kaiser, Stephan Gouws, Yoshikiyo Kato, Taku Kudo, Hideto Kazawa, Keith Stevens, George Kurian, Nishant Patil, Wei Wang, Cliff Young, Jason Smith, Jason Riesa, Alex Rudnick, Oriol Vinyals, Gregory S. Corrado, Macduff Hughes, and Jeffrey Dean · 2016
Cited alongside, same era.
Results of the WMT17 metrics shared task
Ondřej Bojar, Yvette Graham, and Amir Kamran · 2017
Cited alongside, same era.
Convolutional sequence to sequence learning
Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, and Yann N Dauphin · 2017
Cited alongside, same era.
First quora dataset release: Question pairs
Shankar Iyer, Nikhil Dandekar, and Kornel Csernai · 2017
Cited alongside, same era.
MEANT 2.0: Accurate semantic mt evaluation for any output language
Chi-kiu Lo · 2017
Cited alongside, same era.
SciBERT: A pretrained language model for scientific text
Iz Beltagy, Kyle Lo, and Arman Cohan · 2019
Closest in time.
WMDO: Fluency-based word mover’s distance for machine translation evaluation
Julian Chow, Lucia Specia, and Pranava Madhyastha · 2019
Closest in time.
Sentence mover’s similarity: Automatic evaluation for multi-sentence texts
Elizabeth Clark, Asli Celikyilmaz, and Noah A. Smith · 2019
Closest in time.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Closest in time.
Cooperative generator-discriminator networks for abstractive summarization with narrative flow
Saadia Gabriel, Antoine Bosselut, Ari Holtzman, Kyle Lo, Asli Çelikyilmaz, and Yejin Choi · 2019
Closest in time.
Meteor++ 2.0: Adopt syntactic level paraphrase knowledge into machine translation evaluation
Yinuo Guo and Junfeng Hu · 2019
Closest in time.
Unifying human and statistical evaluation for natural language generation
Tatsu Hashimoto, Hugh Zhang, and Percy Liang · 2019
Closest in time.
Chenyang Huang, Amine Trabelsi, and Osmar R Zaïane · 2019
Closest in time.
Cross-lingual language model pretraining
Guillaume Lample and Alexis Conneau · 2019
Closest in time.
Deep reinforcement learning with distributional semantic rewards for abstractive summarization
Siyao Li, Deren Lei, Pengda Qin, and William Yang Wang · 2019
Closest in time.
Fine-tune BERT for extractive summarization
Yang Liu · 2019
Closest in time.
YiSi - a unified semantic MT quality evaluation and estimation metric for languages with different levels of available resources
Chi-kiu Lo · 2019
Closest in time.
Putting evaluation in context: Contextual embeddings improve machine translation evaluation
Nitika Mathur, Timothy Baldwin, and Trevor Cohn · 2019
Closest in time.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli · 2019
Closest in time.
Alternative weighting schemes for elmo embeddings
Nils Reimers and Iryna Gurevych · 2019
Closest in time.
EED: Extended edit distance measure for machine translation
Peter Stanchev, Weiyue Wang, and Hermann Ney · 2019
Closest in time.
Beyond BLEU:training neural machine translation with semantic similarity
John Wieting, Taylor Berg-Kirkpatrick, Kevin Gimpel, and Graham Neubig · 2019
Closest in time.
Pay less attention with lightweight and dynamic convolutions
Felix Wu, Angela Fan, Alexei Baevski, Yann Dauphin, and Michael Auli · 2019
Closest in time.
Filtering pseudo-references by paraphrasing for automatic evaluation of machine translation
Ryoma Yoshimura, Hiroki Shimanaka, Yukio Matsumura, Hayahide Yamagishi, and Mamoru Komachi · 2019
Closest in time.
PAWS: Paraphrase adversaries from word scrambling
Yuan Zhang, Jason Baldridge, and Luheng He · 2019
Closest in time.
Moverscore: Text generation evaluating with contextualized embeddings and earth mover distance
Wei Zhao, Maxime Peyrard, Fei Liu, Yang Gao, Christian M. Meyer, and Steffen Eger · 2019
Closest in time.