Fetching the paper…
Reading the bibliography…
Automatic evaluation metrics are crucial to the development of generative systems.
Evaluating the underlying gender bias in contextualized word embeddings
Christine Basta, Marta Ruiz Costa-jussà, and Noe Casas. 2019 · 1904
Earlier work this paper cites.
Distilbert, a distilled version of BERT: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 1910
Earlier work this paper cites.
Redditbias: A real-world resource for bias evaluation and debiasing of conversational language models
Soumya Barikeri, Anne Lauscher, Ivan Vulic, and Goran Glavas. 2021 · 1955
Earlier work this paper cites.
Crows-pairs: A challenge dataset for measuring social biases in masked language models
Nikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel R. Bowman. 2020 · 1967
Earlier work this paper cites.
Catastrophic interference in connectionist networks: Can it be predicted, can it be prevented?
Robert M. French. 1993 · 1993
Earlier work this paper cites.
Automatic evaluation of machine translation quality using n-gram co-occurrence statistics
George Doddington. 2002 · 2002
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
METEOR: an automatic metric for MT evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie. 2005 · 2005
Earlier work this paper cites.
Measuring and reducing gendered correlations in pre-trained models
Kellie Webster, Xuezhi Wang, Ian Tenney, Alex Beutel, Emily Pitler, Ellie Pavlick, Jilin Chen, and Slav Petrov. 2020 · 2010
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey E. Hinton, Oriol Vinyals, and Jeffrey Dean. 2015 · 2015
Earlier work this paper cites.
chrf: character n-gram f-score for automatic MT evaluation
Maja Popovic. 2015 · 2015
Earlier work this paper cites.
Semeval-2017 task 1: Semantic textual similarity - multilingual and cross-lingual focused evaluation
Daniel M. Cer, Mona T. Diab, Eneko Agirre, Iñigo Lopez-Gazpio, and Lucia Specia. 2017 · 2017
Earlier work this paper cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel R. Bowman. 2018 · 2018
Earlier work this paper cites.
Gender bias in coreference resolution: Evaluation and debiasing methods
Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang. 2018a · 2018
Earlier work this paper cites.
Identifying and reducing gender bias in word-level language models
Shikha Bordia and Samuel R. Bowman. 2019 · 2019
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Parameter-efficient transfer learning for NLP
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin de Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019 · 2019
Cited alongside, same era.
Measuring bias in contextualized word representations
Keita Kurita, Nidhi Vyas, Ayush Pareek, Alan W Black, and Yulia Tsvetkov. 2019 · 2019
Cited alongside, same era.
On measuring social biases in sentence encoders
Chandler May, Alex Wang, Shikha Bordia, Samuel R. Bowman, and Rachel Rudinger. 2019 · 2019
Cited alongside, same era.
Assessing social and intersectional biases in contextualized word representations
Yi Chern Tan and L. Elisa Celis. 2019 · 2019
Cited alongside, same era.
Xlnet: Generalized autoregressive pretraining for language understanding
COMET: A neural framework for MT evaluation
Ricardo Rei, Craig Stewart, Ana C. Farinha, and Alon Lavie. 2020 · 2020
Later among the works it cites.
BLEURT: learning robust metrics for text generation
Thibault Sellam, Dipanjan Das, and Ankur P. Parikh. 2020 · 2020
Later among the works it cites.
Automatic machine translation evaluation in many languages via zero-shot paraphrasing
Brian Thompson and Matt Post. 2020 · 2020
Later among the works it cites.
Investigating gender bias in language models using causal mediation analysis
Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Yaron Singer, and Stuart M. Shieber. 2020 · 2020
Later among the works it cites.
Bertscore: Evaluating text generation with BERT
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020 · 2020
Later among the works it cites.
Rethinking embedding coupling in pre-trained language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime G. Carbonell, Ruslan Salakhutdinov, and Quoc V. Le. 2019 · 2019
Cited alongside, same era.
Gender bias in contextualized word embeddings
Jieyu Zhao, Tianlu Wang, Mark Yatskar, Ryan Cotterell, Vicente Ordonez, and Kai-Wei Chang. 2019a · 2019
Cited alongside, same era.
Moverscore: Text generation evaluating with contextualized embeddings and earth mover distance
Wei Zhao, Maxime Peyrard, Fei Liu, Yang Gao, Christian M. Meyer, and Steffen Eger. 2019b · 2019
Cited alongside, same era.
Re-evaluating evaluation in text summarization
Manik Bhandari, Pranav Narayan Gour, Atabak Ashfaq, Pengfei Liu, and Graham Neubig. 2020 · 2020
Cited alongside, same era.
On measuring and mitigating biased inferences of word embeddings
Sunipa Dev, Tao Li, Jeff M. Phillips, and Vivek Srikumar. 2020 · 2020
Cited alongside, same era.
ALBERT: A lite BERT for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2020 · 2020
Cited alongside, same era.
BART: denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020 · 2020
Cited alongside, same era.
Hyung Won Chung, Thibault Févry, Henry Tsai, Melvin Johnson, and Sebastian Ruder. 2021 · 2021
Later among the works it cites.
Paula Czarnowska, Yogarshi Vyas, and Kashif Shah. 2021 · 2021
Later among the works it cites.
Frugalscore: Learning cheaper, lighter and faster evaluation metricsfor automatic text generation
Moussa Kamal Eddine, Guokan Shang, Antoine J.-P. Tixier, and Michalis Vazirgiannis. 2021 · 2021
Later among the works it cites.
A fine-grained analysis of bertscore
Michael Hanna and Ondrej Bojar. 2021 · 2021
Later among the works it cites.
Unmasking the mask - evaluating social biases in masked language models
Masahiro Kaneko and Danushka Bollegala. 2021 · 2021
Later among the works it cites.
Sustainable modular debiasing of language models
Anne Lauscher, Tobias Lüken, and Goran Glavas. 2021 · 2021
Later among the works it cites.
Stereoset: Measuring stereotypical bias in pretrained language models
Moin Nadeem, Anna Bethke, and Siva Reddy. 2021 · 2021
Later among the works it cites.
Adapterfusion: Non-destructive task composition for transfer learning
Jonas Pfeiffer, Aishwarya Kamath, Andreas Rücklé, Kyunghyun Cho, and Iryna Gurevych. 2021 · 2021
Later among the works it cites.
Learning compact metrics for MT
Amy Pu, Hyung Won Chung, Ankur P. Parikh, Sebastian Gehrmann, and Thibault Sellam. 2021 · 2021
Later among the works it cites.
Bartscore: Evaluating generated text as text generation
Weizhe Yuan, Graham Neubig, and Pengfei Liu. 2021 · 2021
Later among the works it cites.