Fetching the paper…
Reading the bibliography…
Quality estimation (QE)-the automatic assessment of translation quality-has recently become crucial across several stages of the translation pipeline, from data curation to training and decoding.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
The web as a parallel corpus
Philip Resnik and Noah A. Smith. 2003 · 2003
Earlier work this paper cites.
Evaluating the output of machine translation systems
Alon Lavie. 2011 · 2011
Earlier work this paper cites.
Parallel corpus refinement as an outlier detection algorithm
Kaveh Taghipour, Shahram Khadivi, and Jia Xu. 2011 · 2011
Earlier work this paper cites.
Findings of the 2012 workshop on statistical machine translation
Chris Callison-Burch, Philipp Koehn, Christof Monz, Matt Post, Radu Soricut, and Lucia Specia. 2012 · 2012
Earlier work this paper cites.
Multidimensional quality metrics: a flexible system for assessing translation quality
Aljoscha Burchardt. 2013 · 2013
Earlier work this paper cites.
Results of the WMT14 metrics shared task
Matouš Macháček and Ondřej Bojar. 2014 · 2014
Earlier work this paper cites.
Results of the WMT15 metrics shared task
Miloš Stanojević, Amir Kamran, Philipp Koehn, and Ondřej Bojar. 2015 · 2015
Earlier work this paper cites.
Is all that glitters in machine translation quality estimation really gold?
Yvette Graham, Timothy Baldwin, Meghan Dowling, Maria Eskevich, Teresa Lynn, and Lamia Tounsi. 2016 · 2016
Earlier work this paper cites.
Can gender-fair language reduce gender stereotyping and discrimination?
Sabine Sczesny, Magda Formanowicz, and Franziska Moser. 2016 · 2016
Earlier work this paper cites.
Semantics derived automatically from language corpora contain human-like biases
Aylin Caliskan, Joanna J Bryson, and Arvind Narayanan. 2017 · 2017
Earlier work this paper cites.
The trouble with bias
Kate Crawford. 2017 · 2017
Earlier work this paper cites.
Measuring and mitigating unintended bias in text classification
Lucas Dixon, John Li, Jeffrey Scott Sorensen, Nithum Thain, and Lucy Vasserman. 2018 · 2018
Earlier work this paper cites.
Getting gender right in neural machine translation
Eva Vanmassenhove, Christian Hardmeier, and Andy Way. 2018 · 2018
Earlier work this paper cites.
Putting fairness principles into practice: Challenges, metrics, and improvements
Alex Beutel, Jilin Chen, Tulsee Doshi, Hai Qian, Allison Woodruff, Christine Luu, Pierre Kreitmann, Jonathan Bischof, and Ed H Chi. 2019 · 2019
Earlier work this paper cites.
On measuring gender bias in translation of gender-neutral pronouns
Won Ik Cho, Ji Won Kim, Seok Min Kim, and Nam Soo Kim. 2019 · 2019
Earlier work this paper cites.
Equalizing gender bias in neural machine translation with word embeddings techniques
Joel Escudé Font and Marta R. Costa-jussà. 2019 · 2019
Earlier work this paper cites.
Taking MT evaluation metrics to extremes: Beyond correlation with human judgments
Marina Fomicheva and Lucia Specia. 2019 · 2019
Earlier work this paper cites.
On measuring social biases in sentence encoders
Chandler May, Alex Wang, Shikha Bordia, Samuel R. Bowman, and Rachel Rudinger. 2019 · 2019
Earlier work this paper cites.
Evaluating gender bias in machine translation
Gabriel Stanovsky, Noah A. Smith, and Luke Zettlemoyer. 2019 · 2019
Earlier work this paper cites.
MoverScore: Text generation evaluating with contextualized embeddings and earth mover distance
Wei Zhao, Maxime Peyrard, Fei Liu, Yang Gao, Christian M. Meyer, and Steffen Eger. 2019 · 2019
Earlier work this paper cites.
Fine-tuning neural machine translation on gender-balanced datasets
Marta R. Costa-jussà and Adrià de Jorge. 2020 · 2020
Earlier work this paper cites.
Towards understanding gender bias in relation extraction
Andrew Gaut, Tony Sun, Shirlyn Tang, Yuxin Huang, Jing Qian, Mai ElSherief, Jieyu Zhao, Diba Mirza, Elizabeth Belding, Kai-Wei Chang, and William Yang Wang. 2020 · 2020
Earlier work this paper cites.
“you sound just like your father” commercial machine translation systems include stylistic biases
Dirk Hovy, Federico Bianchi, and Tommaso Fornaciari. 2020 · 2020
Earlier work this paper cites.
Racial disparities in automated speech recognition
Allison Koenecke, Andrew Nam, Emily Lake, Joe Nudell, Minnie Quartey, Zion Mengesha, Connor Toups, John R Rickford, Dan Jurafsky, and Sharad Goel. 2020 · 2020
Earlier work this paper cites.
Unbabel’s participation in the WMT20 metrics shared task
Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. 2020 · 2020
Earlier work this paper cites.
Reducing gender bias in neural machine translation as a domain adaptation problem
Danielle Saunders and Bill Byrne. 2020 · 2020
Earlier work this paper cites.
Neural machine translation doesn’t translate gender coreference right unless you make it
Danielle Saunders, Rosie Sallis, and Bill Byrne. 2020 · 2020
Earlier work this paper cites.
BLEURT: Learning robust metrics for text generation
Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020 · 2020
Earlier work this paper cites.
Are we estimating or guesstimating translation quality?
Shuo Sun, Francisco Guzmán, and Lucia Specia. 2020 · 2020
Earlier work this paper cites.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020 · 2020
Earlier work this paper cites.
Bertscore: Evaluating text generation with bert
Tianyi Zhang*, Varsha Kishore*, Felix Wu*, Kilian Q. Weinberger, and Yoav Artzi. 2020 · 2020
Earlier work this paper cites.
Quantifying social biases in NLP: A generalization and empirical comparison of extrinsic fairness metrics
Paula Czarnowska, Yogarshi Vyas, and Kashif Shah. 2021 · 2021
Cited alongside, same era.
Multi-task learning for improving gender accuracy in neural machine translation
Carlos Escolano, Graciela Ojeda, Christine Basta, and Marta R. Costa-jussa. 2021 · 2021
Cited alongside, same era.
CLIPScore: A reference-free evaluation metric for image captioning
Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi. 2021 · 2021
Cited alongside, same era.
To ship or not to ship: An extensive evaluation of automatic metrics for machine translation
Tom Kocmi, Christian Federmann, Roman Grundkiewicz, Marcin Junczys-Dowmunt, Hitokazu Matsushita, and Arul Menezes. 2021 · 2021
Cited alongside, same era.
Collecting a large-scale gender bias dataset for coreference resolution and machine translation
Shahar Levy, Koren Lazar, and Gabriel Stanovsky. 2021 · 2021
Cited alongside, same era.
Silvia Alma Piazzolla, Beatrice Savoldi, and Luisa Bentivogli. 2023 · 2023
Later among the works it cites.
Hi guys or hi folks? benchmarking gender-neutral machine translation with the GeNTE corpus
Andrea Piergentili, Beatrice Savoldi, Dennis Fucci, Matteo Negri, and Luisa Bentivogli. 2023 · 2023
Later among the works it cites.
Gender biases in automatic evaluation metrics for image captioning
Haoyi Qiu, Zi-Yi Dou, Tianlu Wang, Asli Celikyilmaz, and Nanyun Peng. 2023 · 2023
Later among the works it cites.
Gate: A challenge set for gender-ambiguous translation examples
Spencer Rarrick, Ranjita Naik, Varun Mathur, Sundar Poudel, and Vishal Chowdhary. 2023 · 2023
Later among the works it cites.
Evaluating metrics for document-context evaluation in machine translation
Vikas Raunak, Tom Kocmi, and Matt Post. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gender and representation bias in GPT-3 generated stories
Li Lucy and David Bamman. 2021 · 2021
Cited alongside, same era.
Gender bias in machine translation
Beatrice Savoldi, Marco Gaido, Luisa Bentivogli, Matteo Negri, and Marco Turchi. 2021 · 2021
Cited alongside, same era.
gENder-IT: An annotated English-Italian parallel challenge set for cross-linguistic natural gender phenomena
Eva Vanmassenhove and Johanna Monti. 2021 · 2021
Cited alongside, same era.
Identifying weaknesses in machine translation metrics through minimum Bayes risk decoding: A case study for COMET
Chantal Amrhein and Rico Sennrich. 2022 · 2022
Cited alongside, same era.
MT-GenEval: A counterfactual and contextual dataset for evaluating gender accuracy in machine translation
Anna Currey, Maria Nadejde, Raghavendra Reddy Pappagari, Mia Mayer, Stanislas Lauly, Xing Niu, Benjamin Hsu, and Georgiana Dinu. 2022 · 2022
Cited alongside, same era.
Quality-aware decoding for neural machine translation
Patrick Fernandes, António Farinhas, Ricardo Rei, José G. C. de Souza, Perez Ogayo, Graham Neubig, and Andre Martins. 2022 · 2022
Cited alongside, same era.
High quality rather than high model probability: Minimum bayes risk decoding with neural metrics
Markus Freitag, David Grangier, Qijun Tan, and Bowen Liang. 2022 · 2022
Cited alongside, same era.
Scaling up CometKiwi: Unbabel-IST 2023 submission for the quality estimation shared task
Ricardo Rei, Nuno M. Guerreiro, José Pombal, Daan van Stigt, Marcos Treviso, Luisa Coheur, José G. C. de Souza, and André Martins. 2023 · 2023
Later among the works it cites.
Gender bias in machine translation: a statistical evaluation of google translate and deepl for english, italian and german
Argentina Rescigno, Johanna Monti, et al. 2023 · 2023
Later among the works it cites.
Language models get a gender makeover: Mitigating gender bias with few-shot data interventions
Himanshu Thakur, Atishay Jain, Praneetha Vaddamanu, Paul Pu Liang, and Louis-Philippe Morency. 2023 · 2023
Later among the works it cites.
INSTRUCTSCORE: Towards explainable text generation evaluation with automatic feedback
Wenda Xu, Danqing Wang, Liangming Pan, Zhenqiao Song, Markus Freitag, William Wang, and Lei Li. 2023 · 2023
Later among the works it cites.
BLEURT has universal translations: An analysis of automatic metrics by minimum risk training
Yiming Yan, Tao Wang, Chengqi Zhao, Shujian Huang, Jiajun Chen, and Mingxuan Wang. 2023 · 2023
Later among the works it cites.
Can automatic metrics assess high-quality translations?
Sweta Agrawal, António Farinhas, Ricardo Rei, and Andre Martins. 2024b · 2024
Closest in time.
Tower: An open multilingual large language model for translation-related tasks
Duarte Miguel Alves, José Pombal, Nuno M Guerreiro, Pedro Henrique Martins, João Alves, Amin Farajian, Ben Peters, Ricardo Rei, Patrick Fernandes, Sweta Agrawal, Pierre Colombo, José G. C. de Souza, and Andre Martins. 2024 · 2024
Closest in time.
Twists, humps, and pebbles: Multilingual speech recognition models exhibit gender performance gaps
Giuseppe Attanasio, Beatrice Savoldi, Dennis Fucci, and Dirk Hovy. 2024 · 2024
Closest in time.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024 · 2024
Closest in time.
Generating gender alternatives in machine translation
Sarthak Garg, Mozhdeh Gheini, Clara Emmanuel, Tatiana Likhomanenko, Qin Gao, and Matthias Paulik. 2024 · 2024
Closest in time.
xcomet: Transparent machine translation evaluation through fine-grained error detection
Nuno M Guerreiro, Ricardo Rei, Daan van Stigt, Luisa Coheur, Pierre Colombo, and André FT Martins. 2024 · 2024
Closest in time.
Improving machine translation with human feedback: An exploration of quality estimation as a reward model
Zhiwei He, Xing Wang, Wenxiang Jiao, Zhuosheng Zhang, Rui Wang, Shuming Shi, and Zhaopeng Tu. 2024 · 2024
Closest in time.
Building bridges: A dataset for evaluating gender-fair machine translation into German
Manuel Lardelli, Giuseppe Attanasio, and Anne Lauscher. 2024 · 2024
Closest in time.
Enhancing gender-inclusive machine translation with neomorphemes and large language models
Andrea Piergentili, Beatrice Savoldi, Matteo Negri, and Luisa Bentivogli. 2024 · 2024
Closest in time.
The power of prompts: Evaluating and mitigating gender bias in MT with LLMs
Aleix Sant, Carlos Escolano, Audrey Mash, Francesca De Luca Fornaciari, and Maite Melero. 2024 · 2024
Closest in time.
What the harm? quantifying the tangible impact of gender bias in machine translation with a human-centered study
Beatrice Savoldi, Sara Papi, Matteo Negri, Ana Guerberof-Arenas, and Luisa Bentivogli. 2024 · 2024
Closest in time.
Whose wife is it anyway? assessing bias against same-gender relationships in machine translation
Ian Stewart and Rada Mihalcea. 2024 · 2024
Closest in time.
Large Language Models are Inconsistent and Biased Evaluators
Rickard Stureborg, Dimitris Alikaniotis, and Yoshi Suhara. 2024 · 2024
Closest in time.
Gemma 2: Improving open language models at a practical size
Gemma Team, Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupatiraju, Léonard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ramé, et al. 2024 · 2024
Closest in time.
Improving statistical significance in human evaluation of automatic metrics via soft pairwise accuracy
Brian Thompson, Nitika Mathur, Daniel Deutsch, and Huda Khayrallah. 2024 · 2024
Closest in time.
Don’t rank, combine! combining machine translation hypotheses using quality estimation
Giorgos Vernikos and Andrei Popescu-Belis. 2024 · 2024
Closest in time.
Analyzing context contributions in LLM-based machine translation
Emmanouil Zaranis, Nuno M Guerreiro, and Andre Martins. 2024 · 2024
Closest in time.
Large language models are not robust multiple choice selectors
Chujie Zheng, Hao Zhou, Fandong Meng, Jie Zhou, and Minlie Huang. 2024 · 2024
Closest in time.
Fine-tuned machine translation metrics struggle in unseen domains
Vilém Zouhar, Shuoyang Ding, Anna Currey, Tatyana Badeka, Jenyuan Wang, and Brian Thompson. 2024 · 2024
Closest in time.
Do llms understand your translations? evaluating paragraph-level mt with question answering
Patrick Fernandes, Sweta Agrawal, Emmanouil Zaranis, André FT Martins, and Graham Neubig. 2025 · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al. 2025 · 2025
Closest in time.
mgente: A multilingual resource for gender-neutral language and translation
Beatrice Savoldi, Eleonora Cupin, Manjinder Thind, Anne Lauscher, Andrea Piergentili, Matteo Negri, and Luisa Bentivogli. 2025 · 2025
Closest in time.