Fetching the paper…
Reading the bibliography…
We review three limitations of BLEU and ROUGE -- the most popular metrics used to assess reference summaries against hypothesis summaries, come up with criteria for what a good metric should behave like and propose concrete ways to use recent Transformers-based Language Models to assess reference summaries against hypothesis summaries.
Bleu: a method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
C.-Y. Lin · 2004
Earlier work this paper cites.
Re-evaluation the role of bleu in machine translation research
C. Callison-Burch, M. Osborne, and P. Koehn · 2006
Earlier work this paper cites.
Recurrent neural network based language model
T. Mikolov, M. Karafiát, L. Burget, J. Černockỳ, and S. Khudanpur · 2010
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
A broad-coverage challenge corpus for sentence understanding through inference
A. Williams, N. Nangia, and S. R. Bowman · 2017
Cited alongside, same era.
A structured review of the validity of bleu
E. Reiter · 2018
Cited alongside, same era.
Ruse: Regressor using sentence embeddings for automatic machine translation evaluation
H. Shimanaka, T. Kajiwara, and M. Komachi · 2018
Cited alongside, same era.
Bleu is not suitable for the evaluation of text simplification
E. Sulem, O. Abend, and A. Rappoport · 2018
Cited alongside, same era.
Sentence mover’s similarity: Automatic evaluation for multi-sentence texts
E. Clark, A. Celikyilmaz, and N. A. Smith · 2019
Cited alongside, same era.
Roberta: A robustly optimized bert pretraining approach
Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov · 2019
Closest in time.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever · 2019
Closest in time.
Bertscore: Evaluating text generation with bert
T. Zhang, V. Kishore, F. Wu, K. Q. Weinberger, and Y. Artzi · 2019
Closest in time.
W. Zhao, L. Wang, K. Shen, R. Jia, and J. Liu · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…