Fetching the paper…
Reading the bibliography…
We present NUBIA, a methodology to build automatic evaluation metrics for text generation using only machine learning models as core components.
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019 · 1904
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Mutt: Metric unit testing for language generation tasks
William Boag, Renan Campos, Kate Saenko, and Anna Rumshisky. 2016 · 1943
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Re-evaluation the role of bleu in machine translation research
Chris Callison-Burch, Miles Osborne, and Philipp Koehn. 2006 · 2006
Earlier work this paper cites.
Recurrent neural network based language model
Tomáš Mikolov, Martin Karafiát, Lukáš Burget, Jan Černockỳ, and Sanjeev Khudanpur. 2010 · 2010
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean. 2013 · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2014 · 2014
Earlier work this paper cites.
Accurate evaluation of segment-level machine translation metrics
Yvette Graham, Timothy Baldwin, and Nitika Mathur. 2015 · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
R. Vedantam, C. L. Zitnick, and D. Parikh. 2015 · 2015
Cited alongside, same era.
Spice: Semantic propositional image caption evaluation
Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould. 2016 · 2016
Cited alongside, same era.
Blend: a novel combined mt metric based on direct assessment—casict-dcu submission to wmt17 metrics task
Qingsong Ma, Yvette Graham, Shugen Wang, and Qun Liu. 2017 · 2017
Cited alongside, same era.
Why we need new evaluation metrics for nlg
Jekaterina Novikova, Ondřej Dušek, Amanda Cercas Curry, and Verena Rieser. 2017 · 2017
Cited alongside, same era.
Results of the wmt17 metrics shared task
MFF UFAL. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
A structured review of the validity of bleu
Ehud Reiter. 2018 · 2018
Later among the works it cites.
Ruse: Regressor using sentence embeddings for automatic machine translation evaluation
Hiroki Shimanaka, Tomoyuki Kajiwara, and Mamoru Komachi. 2018 · 2018
Later among the works it cites.
Bleu is not suitable for the evaluation of text simplification
Elior Sulem, Omri Abend, and Ari Rappoport. 2018 · 2018
Later among the works it cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. 2018 · 2018
Later among the works it cites.
Sentence mover’s similarity: Automatic evaluation for multi-sentence texts
Elizabeth Clark, Asli Celikyilmaz, and Noah A Smith. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Data statements for nlp: Toward mitigating system bias and enabling better science
Emily M Bender and Batya Friedman. 2018 · 2018
Cited alongside, same era.
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumeé III, and Kate Crawford. 2018 · 2018
Cited alongside, same era.
Results of the WMT18 metrics shared task: Both characters and embeddings achieve good performance
Qingsong Ma, Ondřej Bojar, and Yvette Graham. 2018 · 2018
Cited alongside, same era.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018 · 2018
Cited alongside, same era.
Towards neural similarity evaluator
Hassan Kané, Yusuf Kocyigit, Pelkins Ajanoh, Ali Abdalla, and Mohamed Coulibali. 2019 · 2019
Later among the works it cites.
Yisi-a unified semantic mt quality evaluation and estimation metric for languages with different levels of available resources
Chi-kiu Lo. 2019 · 2019
Later among the works it cites.
Model cards for model reporting
Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru. 2019 · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Later among the works it cites.