Fetching the paper…
Reading the bibliography…
Despite the steady progress in machine translation evaluation, existing automatic metrics struggle to capture how well meaning is preserved beyond sentence boundaries.
Evaluation of MT systems by TOEFL
Masaru Tomita, Masako Shirai, Junya Tsutsumi, Miki Matsumura, and Yuki · 1993
Earlier work this paper cites.
Evaluation method for determining groups of users who find MT “useful”
M. Fuji, N. Hatanaka, E. Ito, S. Kamei, H. Kumai, T. Sukehiro, T. Yoshimi, and H. Isahara · 2001
Earlier work this paper cites.
Proceedings on the Workshop on Statistical Machine Translation , New York City, June 2006. Association for Computational Linguistics
Philipp Koehn and Christof Monz (eds.) · 2006
Earlier work this paper cites.
Quiz-based evaluation of machine translation
Jan Berkaab, Martin Černỳa, and Ondřej Bojarb · 2011
Earlier work this paper cites.
Lexico-syntactic text simplification and compression with typed dependencies
Mandya Angrosh, Tadashi Nomoto, and Advaith Siddharthan · 2014
Earlier work this paper cites.
chrF: character n-gram F-score for automatic MT evaluation
Maja Popović · 2015
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Earlier work this paper cites.
A reading comprehension corpus for machine translation evaluation
Carolina Scarton and Lucia Specia · 2016
Earlier work this paper cites.
Discourse structure in machine translation evaluation
Shafiq Joty, Francisco Guzmán, Lluís Màrquez, and Preslav Nakov · 2017
Earlier work this paper cites.
Exploring gap filling as a cheaper alternative to reading comprehension questionnaires when evaluating machine translation for gisting
Mikel L. Forcada, Carolina Scarton, Lucia Specia, Barry Haddow, and Alexandra Birch · 2018
Earlier work this paper cites.
Question answering as an automatic evaluation metric for news article summarization
Matan Eyal, Tal Baumel, and Michael Elhadad · 2019
Earlier work this paper cites.
Cross-lingual training for automatic question generation
Vishwajeet Kumar, Nitish Joshi, Arijit Mukherjee, Ganesh Ramakrishnan, and Preethi Jyothi · 2019
Earlier work this paper cites.
On the cross-lingual transferability of monolingual representations
Mikel Artetxe, Sebastian Ruder, and Dani Yogatama · 2020
Earlier work this paper cites.
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov · 2020
Earlier work this paper cites.
FEQA: A question answering evaluation framework for faithfulness assessment in abstractive summarization
Esin Durmus, He He, and Mona Diab · 2020
Earlier work this paper cites.
Tangled up in BLEU: Reevaluating the evaluation of automatic machine translation evaluation metrics
Nitika Mathur, Timothy Baldwin, and Trevor Cohn · 2020
Earlier work this paper cites.
Asking and answering questions to evaluate the factual consistency of summaries
Alex Wang, Kyunghyun Cho, and Mike Lewis · 2020
Earlier work this paper cites.
Controllable open-ended question generation with a new question type ontology
Shuyang Cao and Lu Wang · 2021
Earlier work this paper cites.
InfoXLM: An information-theoretic framework for cross-lingual language model pre-training
Zewen Chi, Li Dong, Furu Wei, Nan Yang, Saksham Singhal, Wenhui Wang, Xia Song, Xian-Ling Mao, Heyan Huang, and Ming Zhou · 2021
Earlier work this paper cites.
Understanding the extent to which content quality metrics measure the information quality of summaries
Daniel Deutsch and Dan Roth · 2021
Earlier work this paper cites.
Towards question-answering as an automatic metric for evaluating the content quality of a summary
Daniel Deutsch, Tania Bedrax-Weiss, and Dan Roth · 2021
Earlier work this paper cites.
q 2 q^{2} : Evaluating factual consistency in knowledge-grounded dialogues via question generation and question answering
Or Honovich, Leshem Choshen, Roee Aharoni, Ella Neeman, Idan Szpektor, and Omri Abend · 2021
Earlier work this paper cites.
Evaluation of review summaries via question-answering
Nannan Huang and Xiuzhen Zhang · 2021
Earlier work this paper cites.
Just ask! evaluating machine translation by asking and answering questions
Mateusz Krubiński, Erfan Ghadery, Marie-Francine Moens, and Pavel Pecina · 2021
Cited alongside, same era.
Improving factual consistency of abstractive summarization via question answering
Feng Nan, Cicero Nogueira dos Santos, Henghui Zhu, Patrick Ng, Kathleen McKeown, Ramesh Nallapati, Dejiao Zhang, Zhiguo Wang, Andrew O. Arnold, and Bing Xiang · 2021
Cited alongside, same era.
QuestEval: Summarization asks for fact-based evaluation
Thomas Scialom, Paul-Alexis Dray, Sylvain Lamprier, Benjamin Piwowarski, Jacopo Staiano, Alex Wang, and Patrick Gallinari · 2021
Cited alongside, same era.
Benchmarking answer verification methods for question answering-based summarization evaluation metrics
Daniel Deutsch and Dan Roth · 2022
Cited alongside, same era.
QAFactEval: Improved QA-based factual consistency evaluation for summarization
Alexander Fabbri, Chien-Sheng Wu, Wenhao Liu, and Caiming Xiong · 2022
Cited alongside, same era.
Large language models are not yet human-level evaluators for abstractive summarization
Chenhui Shen, Liying Cheng, Xuan-Phi Nguyen, Yang You, and Lidong Bing · 2023
Later among the works it cites.
An empirical comparison of LM-based question and answer generation methods
Asahi Ushio, Fernando Alva-Manchego, and Jose Camacho-Collados · 2023
Later among the works it cites.
Evaluating reading comprehension exercises generated by LLMs: A showcase of ChatGPT in education applications
Changrong Xiao, Sean Xin Xu, Kunpeng Zhang, Yufang Wang, and Lei Xia · 2023
Later among the works it cites.
Do text simplification systems preserve meaning? a human evaluation via reading comprehension
Sweta Agrawal and Marine Carpuat · 2024
Later among the works it cites.
Modeling user preferences with automatic metrics: Creating a high-quality preference dataset for machine translation
Sweta Agrawal, José G. C. De Souza, Ricardo Rei, António Farinhas, Gonçalo Faria, Patrick Fernandes, Nuno M Guerreiro, and Andre Martins · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
SimQA: Detecting simultaneous MT errors through word-by-word question answering
HyoJung Han, Marine Carpuat, and Jordan Boyd-Graber · 2022
Cited alongside, same era.
BlonDe: An automatic evaluation metric for document-level machine translation
Yuchen Jiang, Tianyu Liu, Shuming Ma, Dongdong Zhang, Jian Yang, Haoyang Huang, Rico Sennrich, Ryan Cotterell, Mrinmaya Sachan, and Ming Zhou · 2022
Cited alongside, same era.
COMET-22: Unbabel-IST 2022 submission for the metrics shared task
Ricardo Rei, José G. C. de Souza, Duarte Alves, Chrysoula Zerva, Ana C Farinha, Taisiya Glushkova, Alon Lavie, Luisa Coheur, and André F. T. Martins · 2022
Cited alongside, same era.
CometKiwi: IST-unbabel 2022 submission for the quality estimation shared task
Ricardo Rei, Marcos Treviso, Nuno M. Guerreiro, Chrysoula Zerva, Ana C Farinha, Christine Maroti, José G. C. de Souza, Taisiya Glushkova, Duarte Alves, Luisa Coheur, Alon Lavie, and André F. T. Martins · 2022
Cited alongside, same era.
Embarrassingly easy document-level MT metrics: How to convert any pretrained metric into a document-level metric
Giorgos Vernikos, Brian Thompson, Prashant Mathur, and Marcello Federico · 2022
Cited alongside, same era.
Training and meta-evaluating machine translation evaluation metrics at the paragraph level
Daniel Deutsch, Juraj Juraska, Mara Finkelstein, and Markus Freitag · 2023
Cited alongside, same era.
Do multilingual language models think better in english?
Julen Etxaniz, Gorka Azkune, Aitor Soroa, Oier Lopez de Lacalle, and Mikel Artetxe · 2023
Cited alongside, same era.
The claude 3 model family: Opus, sonnet, haiku
AI Anthropic · 2024
Later among the works it cites.
Do multilingual language models think better in English?
Julen Etxaniz, Gorka Azkune, Aitor Soroa, Oier Lopez de Lacalle, and Mikel Artetxe · 2024
Later among the works it cites.
xcomet: Transparent machine translation evaluation through fine-grained error detection
Nuno M. Guerreiro, Ricardo Rei, Daan van Stigt, Luisa Coheur, Pierre Colombo, and André F. T. Martins · 2024
Later among the works it cites.
Are large language model-based evaluators the solution to scaling up multilingual evaluation?
Rishav Hada, Varun Gumma, Adrian Wynter, Harshita Diddee, Mohamed Ahmed, Monojit Choudhury, Kalika Bali, and Sunayana Sitaram · 2024
Later among the works it cites.
Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, et al · 2024
Later among the works it cites.
Prometheus 2: An open source language model specialized in evaluating other language models
Seungone Kim, Juyoung Suk, Shayne Longpre, Bill Yuchen Lin, Jamin Shin, Sean Welleck, Graham Neubig, Moontae Lee, Kyungjae Lee, and Minjoon Seo · 2024
Later among the works it cites.
Findings of the WMT24 general machine translation shared task: The LLM era is here but MT is not solved yet
Tom Kocmi, Eleftherios Avramidis, Rachel Bawden, Ondřej Bojar, Anton Dvorkovich, Christian Federmann, Mark Fishel, Markus Freitag, Thamme Gowda, Roman Grundkiewicz, Barry Haddow, Marzena Karpinska, Philipp Koehn, Benjamin Marie, Christof Monz, Kenton Murray, Masaaki Nagata, Martin Popel, Maja Popović, Mariya Shmatova, Steinthór Steingrímsson, and Vilém Zouhar · 2024
Later among the works it cites.
Leveraging large language models for NLG evaluation: Advances and challenges
Zhen Li, Xiaohan Xu, Tao Shen, Can Xu, Jia-Chen Gu, Yuxuan Lai, Chongyang Tao, and Shuai Ma · 2024
Later among the works it cites.
SummEQuAL: Summarization evaluation via question answering using large language models
Junyuan Liu, Zhengyan Shi, and Aldo Lipani · 2024
Later among the works it cites.
Automatic generation and evaluation of reading comprehension test items with large language models
Andreas Säuberli and Simon Clematide · 2024
Later among the works it cites.
A 2-step framework for automated literary translation evaluation: Its promises and pitfalls
Sheikh Shafayat, Dongkeun Yoon, Woori Jang, Jiwoo Choi, Alice Oh, and Seohyon Jung · 2024
Later among the works it cites.
The good, the bad, and the greedy: Evaluation of llms should not ignore non-determinism
Yifan Song, Guoyin Wang, Sujian Li, and Bill Yuchen Lin · 2024
Later among the works it cites.
ProxyQA: An alternative framework for evaluating long-form text generation with large language models
Haochen Tan, Zhijiang Guo, Zhan Shi, Lu Xu, Zhili Liu, Yunlong Feng, Xiaoguang Li, Yasheng Wang, Lifeng Shang, Qun Liu, and Linqi Song · 2024
Later among the works it cites.
InfoLossQA: Characterizing and recovering information loss in text simplification
Jan Trienes, Sebastian Joseph, Jörg Schlötterer, Christin Seifert, Kyle Lo, Wei Xu, Byron Wallace, and Junyi Jessy Li · 2024
Later among the works it cites.
Foundational autoraters: Taming large language models for better automatic evaluation
Tu Vu, Kalpesh Krishna, Salaheddin Alzubi, Chris Tar, Manaal Faruqui, and Yun-Hsuan Sung · 2024
Later among the works it cites.
Unveiling selection biases: Exploring order and token sensitivity in large language models
Sheng-Lun Wei, Cheng-Kuang Wu, Hen-Hsen Huang, and Hsin-Hsi Chen · 2024
Later among the works it cites.
An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, et al · 2024
Later among the works it cites.