Fetching the paper…
Reading the bibliography…
Recent machine translation (MT) metrics calibrate their effectiveness by correlating with human judgement but without any insights about their behaviour across different error types.
Using test suites in evaluation of machine translation systems
King, Margaret and Kirsten Falkedal. 1990 · 1990
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Papineni, Kishore, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Language models are few-shot learners
Brown, Tom B., Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 2005
Earlier work this paper cites.
Europarl: A parallel corpus for statistical machine translation
Koehn, Philipp. 2005 · 2005
Earlier work this paper cites.
Manual and automatic evaluation of machine translation between European languages
Koehn, Philipp and Christof Monz. 2006 · 2006
Earlier work this paper cites.
Unbounded dependency recovery for parser evaluation
Rimell, Laura, Stephen Clark, and Mark Steedman. 2009 · 2009
Earlier work this paper cites.
Adversarial evaluation for models of natural language
Smith, Noah A. 2012 · 2012
Earlier work this paper cites.
Multidimensional quality metrics (mqm): A framework for declaring and describing translation quality metrics
Lommel, Arle, Aljoscha Burchardt, and Hans Uszkoreit. 2014 · 2014
Earlier work this paper cites.
Neural versus phrase-based machine translation quality: a case study
Bentivogli, Luisa, Arianna Bisazza, Mauro Cettolo, and Marcello Federico. 2016 · 2016
Earlier work this paper cites.
Findings of the 2016 conference on machine translation
Bojar, Ondřej, Rajen Chatterjee, Christian Federmann, Yvette Graham, Barry Haddow, Matthias Huck, Antonio Jimeno Yepes, Philipp Koehn, Varvara Logacheva, Christof Monz, Matteo Negri, Aurélie Névéol, Mariana Neves, Martin Popel, Matt Post, Raphael Rubino, Carolina Scarton, Lucia Specia, Marco Turchi, Karin Verspoor, and Marcos Zampieri. 2016 · 2016
Earlier work this paper cites.
PROTEST: A test suite for evaluating pronouns in machine translation
Guillou, Liane and Christian Hardmeier. 2016 · 2016
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Sennrich, Rico, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
Is neural machine translation the new state of the art?
Castilho, Sheila, Joss Moorkens, Federico Gaspari, Iacer Calixto, John Tinsley, and Andy Way. 2017 · 2017
Earlier work this paper cites.
A challenge set approach to evaluating machine translation
Isabelle, Pierre, Colin Cherry, and George Foster. 2017 · 2017
Earlier work this paper cites.
Improving discourse relation projection to build discourse annotated corpora
Laali, Majid and Leila Kosseim. 2017 · 2017
Earlier work this paper cites.
BIBI system description: Building with CNNs and breaking with deep reinforcement learning
Li, Yitong, Trevor Cohn, and Timothy Baldwin. 2017 · 2017
Earlier work this paper cites.
Breaking NLP: Using morphosyntax, semantics, pragmatics and world knowledge to fool sentiment analysis systems
Mahler, Taylor, Willy Cheung, Micha Elsner, David King, Marie-Catherine de Marneffe, Cory Shain, Symon Stevens-Guille, and Michael White. 2017 · 2017
Earlier work this paper cites.
chrF++: words helping character n-grams
Popović, Maja. 2017 · 2017
Earlier work this paper cites.
Social bias in elicited natural language inferences
Rudinger, Rachel, Chandler May, and Benjamin Van Durme. 2017 · 2017
Earlier work this paper cites.
Breaking sentiment analysis of movie reviews
Staliūnaitė, Ieva and Ben Bonfil. 2017 · 2017
Earlier work this paper cites.
A multifaceted evaluation of neural versus phrase-based machine translation for 9 language directions
Toral, Antonio and Víctor M. Sánchez-Cartagena. 2017 · 2017
Earlier work this paper cites.
Fine-grained evaluation of quality estimation for machine translation based on a linguistically motivated test suite
Avramidis, Eleftherios, Vivien Macketanz, Arle Lommel, and Hans Uszkoreit. 2018 · 2018
Earlier work this paper cites.
XNLI: Evaluating cross-lingual sentence representations
Conneau, Alexis, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel Bowman, Holger Schwenk, and Veselin Stoyanov. 2018 · 2018
Earlier work this paper cites.
A pronoun test suite evaluation of the English–German MT systems at WMT 2018
Guillou, Liane, Christian Hardmeier, Ekaterina Lapshinova-Koltunski, and Sharid Loáiciga. 2018 · 2018
Earlier work this paper cites.
Looking beyond the surface: A challenge set for reading comprehension over multiple sentences
Khashabi, Daniel, Snigdha Chaturvedi, Michael Roth, Shyam Upadhyay, and Dan Roth. 2018 · 2018
Earlier work this paper cites.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Kudo, Taku and John Richardson. 2018 · 2018
Earlier work this paper cites.
ParCorFull: a parallel corpus annotated with full coreference
Lapshinova-Koltunski, Ekaterina, Christian Hardmeier, and Pauline Krielke. 2018 · 2018
Earlier work this paper cites.
The word sense disambiguation test suite at WMT18
Rios, Annette, Mathias Müller, and Rico Sennrich. 2018 · 2018
Earlier work this paper cites.
Gender bias in coreference resolution: Evaluation and debiasing methods
Zhao, Jieyu, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang. 2018 · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, Jacob, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
YiSi - a unified semantic MT quality evaluation and estimation metric for languages with different levels of available resources
Lo, Chi-kiu. 2019 · 2019
Earlier work this paper cites.
Non-entailed subsequences as a challenge for natural language inference
McCoy, Richard T and Tal Linzen. 2019 · 2019
Cited alongside, same era.
Challenge test sets for MT evaluation
Popović, Maja and Sheila Castilho. 2019 · 2019
Cited alongside, same era.
The MuCoW test suite at WMT 2019: Automatically harvested multilingual contrastive word sense disambiguation test sets for machine translation
Raganato, Alessandro, Yves Scherrer, and Jörg Tiedemann. 2019 · 2019
Cited alongside, same era.
Evaluating gender bias in machine translation
Stanovsky, Gabriel, Noah A. Smith, and Luke Zettlemoyer. 2019 · 2019
Cited alongside, same era.
Beto, bentz, becas: The surprising cross-lingual effectiveness of BERT
Wu, Shijie and Mark Dredze. 2019 · 2019
Cited alongside, same era.
PAWS-X: A cross-lingual adversarial dataset for paraphrase identification
DEMETR: Diagnosing evaluation metrics for translation
Karpinska, Marzena, Nishant Raj, Katherine Thai, Yixiao Song, Ankita Gupta, and Mohit Iyyer. 2022 · 2022
Later among the works it cites.
MS-COMET: More and Better Human Judgements Improve Metric Performance
Kocmi, Tom, Hitokazu Matsushita, and Christian Federmann. 2022 · 2022
Later among the works it cites.
Partial Could Be Better Than Whole: HW-TSC 2022 Submission for the Metrics Shared Task
Liu, Yilun, Xiaosong Qiao, Zhanglin Wu, Su Chang, Min Zhang, Yanqing Zhao, shimin tao Song Peng, Hao Yang, Ying Qin, Jiaxin Guo, Minghan Wang, Yinglu Li, Peng Li, and Xiaofeng Zhao. 2022 · 2022
Later among the works it cites.
REUSE: REference-free UnSupervised quality Estimation Metric
Mukherjee, Ananya and Manish Shrivastava. 2022 · 2022
Later among the works it cites.
Is my nlp model working? the answer is harder than you think
Neubig, Graham. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yang, Yinfei, Yuan Zhang, Chris Tar, and Jason Baldridge. 2019 · 2019
Cited alongside, same era.
Extracting training data from large language models
Carlini, Nicholas, Florian Tramèr, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom B. Brown, Dawn Xiaodong Song, Úlfar Erlingsson, Alina Oprea, and Colin Raffel. 2020 · 2020
Cited alongside, same era.
COMET: A neural framework for MT evaluation
Rei, Ricardo, Craig Stewart, Ana C Farinha, and Alon Lavie. 2020 · 2020
Cited alongside, same era.
Learning to evaluate translation beyond English: BLEURT submissions to the WMT metrics 2020 shared task
Sellam, Thibault, Amy Pu, Hyung Won Chung, Sebastian Gehrmann, Qijun Tan, Markus Freitag, Dipanjan Das, and Ankur Parikh. 2020 · 2020
Cited alongside, same era.
Findings of the WMT 2020 shared task on machine translation robustness
Specia, Lucia, Zhenhao Li, Juan Pino, Vishrav Chaudhary, Francisco Guzmán, Graham Neubig, Nadir Durrani, Yonatan Belinkov, Philipp Koehn, Hassan Sajjad, Paul Michel, and Xian Li. 2020 · 2020
Cited alongside, same era.
Bertscore: Evaluating text generation with BERT
Zhang, Tianyi, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020 · 2020
Cited alongside, same era.
Wino-X: Multilingual Winograd schemas for commonsense reasoning and coreference resolution
Emelin, Denis and Rico Sennrich. 2021 · 2021
Cited alongside, same era.
NLLB Team, Marta R. Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Sun, Skyler Wang, Guillaume Wenzek, Al Youngblood, Bapi Akula, Loic Barrault, Gabriel Mejia Gonzalez, Prangthip Hansanti, John Hoffman, Semarley Jarrett, Kaushik Ram Sadagopan, Dirk Rowe, Shannon Spruit, Chau Tran, Pierre Andrews, Necip Fazil Ayan, Shruti Bhosale, Sergey Edunov, Angela Fan, Cynthia Gao, Vedanuj Goswami, Francisco Guzmán, Philipp Koehn, Alexandre Mourachko, Christophe Ropers, Safiyyah Saleem, Holger Schwenk, and Jeff Wang. 2022 · 2022
Later among the works it cites.
COMET-22: Unbabel-IST 2022 Submission for the Metrics Shared Task
Rei, Ricardo, José G. C. de Souza, Duarte Alves, Chrysoula Zerva, Ana C Farinha, Taisiya Glushkova, Alon Lavie, Luisa Coheur, and André F. T. Martins. 2022 · 2022
Later among the works it cites.
BLOOM: A 176b-parameter open-access multilingual language model
Scao, Teven Le, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilic, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, Jonathan Tow, Alexander M. Rush, Stella Biderman, Albert Webson, Pawan Sasanka Ammanamanchi, Thomas Wang, Benoît Sagot, Niklas Muennighoff, Albert Villanova del Moral, Olatunji Ruwase, Rachel Bawden, Stas Bekman, Angelina McMillan-Major, Iz Beltagy, Huu Nguyen, Lucile Saulnier, Samson Tan, Pedro Ortiz Suarez, Victor Sanh, Hugo Laurençon, Yacine Jernite, Julien Launay, Margaret Mitchell, Colin Raffel, Aaron Gokaslan, Adi Simhi, Aitor Soroa, Alham Fikri Aji, Amit Alfassy, Anna Rogers, Ariel Kreisberg Nitzav, Canwen Xu, Chenghao Mou, Chris Emezue, Christopher Klamm, Colin Leong, Daniel van Strien, David Ifeoluwa Adelani, and et al. 2022 · 2022
Later among the works it cites.
CrossQE: HW-TSC 2022 submission for the quality estimation shared task
Tao, Shimin, Su Chang, Ma Miaomiao, Hao Yang, Xiang Geng, Shujian Huang, Min Zhang, Jiaxin Guo, Minghan Wang, and Yinglu Li. 2022 · 2022
Later among the works it cites.
As little as possible, as much as necessary: Detecting over- and undertranslations with contrastive conditioning
Vamvas, Jannis and Rico Sennrich. 2022 · 2022
Later among the works it cites.
Aces: Translation accuracy challenge sets at wmt 2023
Amrhein, Chantal, Nikita Moghe, and Liane Guillou. 2023 · 2023
Later among the works it cites.
Challenging the state-of-the-art machine translation metrics from a linguistic perspective
Avramidis, Eleftherios, Shushen Manakhimova, Vivien Macketanz, and Sebastian Möller. 2023 · 2023
Later among the works it cites.
Instructeval: Towards holistic evaluation of instruction-tuned large language models
Chia, Yew Ken, Pengfei Hong, Lidong Bing, and Soujanya Poria. 2023 · 2023
Later among the works it cites.
Detecting and mitigating hallucinations in machine translation: Model internal workings alone do well, sentence similarity Even better
Dale, David, Elena Voita, Loic Barrault, and Marta R. Costa-jussà. 2023 · 2023
Later among the works it cites.
Faith and fate: Limits of transformers on compositionality
Dziri, Nouha, Ximing Lu, Melanie Sclar, Xiang Lorraine Li, Liwei Jian, Bill Yuchen Lin, Peter West, Chandra Bhagavatula, Ronan Le Bras, Jena D. Hwang, Soumya Sanyal, Sean Welleck, Xiang Ren, Allyson Ettinger, Zaïd Harchaoui, and Yejin Choi. 2023 · 2023
Later among the works it cites.
ebleu: Unexpectedly good machine translation evaluation using simple word embeddings
ElNokrashy, Muhammad and Tom Kocmi. 2023 · 2023
Later among the works it cites.
The devil is in the errors: Leveraging large language models for fine-grained machine translation evaluation
Fernandes, Patrick, Daniel Deutsch, Mara Finkelstein, Parker Riley, André Martins, Graham Neubig, Ankush Garg, Jonathan Clark, Markus Freitag, and Orhan Firat. 2023 · 2023
Later among the works it cites.
Results of WMT23 metrics shared task: Metrics might be guilty but references are not innocent
Freitag, Markus, Nitika Mathur, Chi-kiu Lo, Eleftherios Avramidis, Ricardo Rei, Brian Thompson, Tom Kocmi, Frederic Blain, Daniel Deutsch, Craig Stewart, Chrysoula Zerva, Sheila Castilho, Alon Lavie, and George Foster. 2023 · 2023
Later among the works it cites.
Cometoid: Distilling strong reference-based machine translation metrics into even stronger quality estimation metrics
Gowda, Thamme, Tom Kocmi, and Marcin Junczys-Dowmunt. 2023 · 2023
Later among the works it cites.
xcomet: Transparent machine translation evaluation through fine-grained error detection
Guerreiro, Nuno M., Ricardo Rei, Daan van Stigt, Luisa Coheur, Pierre Colombo, and André F. T. Martins. 2023 · 2023
Later among the works it cites.
Metricx-23: The google submission to the wmt 2023 metrics shared task
Juraska, Juraj, Mara Finkelstein, Daniel Deutsch, Aditya Siddhant, Mehdi Mirzazadeh, and Markus Freitag. 2023 · 2023
Later among the works it cites.
Metric score landscape challenge (mslc23): Understanding metrics’ performance on a wider landscape of translation quality
Lo, Chi-kiu, Samuel Larkin, and Rebecca Knowles. 2023 · 2023
Later among the works it cites.
Error analysis prompting enables human-like translation evaluation in large language models: A case study on chatgpt
Lu, Qingyu, Baopu Qiu, Liang Ding, Liping Xie, and Dacheng Tao. 2023 · 2023
Later among the works it cites.
Extrinsic evaluation of machine translation metrics
Moghe, Nikita, Tom Sherborne, Mark Steedman, and Alexandra Birch. 2023 · 2023
Later among the works it cites.
Mee4 and xlsim : Iiit hyd’s submissions’ for wmt23 metrics shared task
Mukherjee, Ananya and Manish Shrivastava. 2023 · 2023
Later among the works it cites.
The inside story: Towards better understanding of machine translation neural evaluation metrics
Rei, Ricardo, Nuno M. Guerreiro, Marcos Treviso, Luisa Coheur, Alon Lavie, and André Martins. 2023 · 2023
Later among the works it cites.
Stanford alpaca: An instruction-following llama model
Taori, Rohan, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023 · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, Hugo, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023 · 2023
Later among the works it cites.
Empowering a metric with llm-assisted named entity annotation: Hw-tsc’s submission to the wmt23 metrics shared task
Wu, Zhanglin, Yilun Liu, Min Zhang, Xiaofeng Zhao, Junhao Zhu, Ming Zhu, Xiaosong Qiao, Jingfei Zhang, Ma Miaomiao, Zhao Yanqing, Song Peng, shimin tao, Hao Yang, and Yanfei Jiang. 2023 · 2023
Later among the works it cites.
INSTRUCTSCORE: Towards explainable text generation evaluation with automatic feedback
Xu, Wenda, Danqing Wang, Liangming Pan, Zhenqiao Song, Markus Freitag, William Wang, and Lei Li. 2023 · 2023
Later among the works it cites.
Adversarial examples for evaluating reading comprehension systems
Jia, Robin and Percy Liang. 2017 · 2031
Closest in time.