Fetching the paper…
Reading the bibliography…
This study comprehensively evaluates the translation quality of Large Language Models (LLMs), specifically GPT-4, against human translators of varying expertise levels across multiple language pairs and domains.
A coefficient of agreement for nominal scales
Jacob Cohen. 1960 · 1960
Earlier work this paper cites.
Validity in content analysis
Klaus Krippendorff. 1980 · 1980
Earlier work this paper cites.
Clue: A chinese language understanding evaluation benchmark
Liang Xu, Hai Hu, Xuanwei Zhang, Lu Li, Chenjie Cao, Yudong Li, Yechen Xu, Kai Sun, Dian Yu, Cong Yu, et al. 2020 · 2004
Earlier work this paper cites.
Continuous measurement scales in human evaluation of machine translation
Yvette Graham, Timothy Baldwin, Alistair Moffat, and Justin Zobel. 2013 · 2013
Earlier work this paper cites.
Assessing inter-annotator agreement for translation error annotation
Arle Lommel, Maja Popovic, and Aljoscha Burchardt. 2014 · 2014
Earlier work this paper cites.
Achieving human parity on automatic chinese to english news translation
Hany Hassan, Anthony Aue, Chang Chen, Vishal Chowdhary, Jonathan Clark, Christian Federmann, Xuedong Huang, Marcin Junczys-Dowmunt, William Lewis, Mu Li, et al. 2018 · 2018
Earlier work this paper cites.
Quantitative fine-grained human evaluation of machine translation systems: a case study on english to croatian
Filip Klubička, Antonio Toral, and Víctor M Sánchez-Cartagena. 2018 · 2018
Earlier work this paper cites.
doccano: Text annotation tool for human
Hiroki Nakayama, Takahiro Kubo, Junya Kamura, Yasufumi Taniguchi, and Xu Liang. 2018 · 2018
Earlier work this paper cites.
Attaining the unattainable? reassessing claims of human parity in neural machine translation
Antonio Toral, Sheila Castilho, Ke Hu, and Andy Way. 2018 · 2018
Earlier work this paper cites.
What’s the difference between professional human and machine translation? a blind multi-language study on domain-specific MT
Lukas Fischer and Samuel L"̈aubli. 2020 · 2020
Earlier work this paper cites.
Assessing human-parity in machine translation on the segment level
Yvette Graham, Christian Federmann, Maria Eskevich, and Barry Haddow. 2020 · 2020
Earlier work this paper cites.
COMET: A neural framework for MT evaluation
Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. 2020a · 2020
Earlier work this paper cites.
Comet: A neural framework for mt evaluation
Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. 2020b · 2020
Earlier work this paper cites.
Findings of the 2021 conference on machine translation (wmt21)
Akhbardeh Farhad, Arkhangorodsky Arkady, Biesialska Magdalena, Bojar Ondřej, Chatterjee Rajen, Chaudhary Vishrav, Marta R Costa-jussa, España-Bonet Cristina, Fan Angela, Federmann Christian, et al. 2021 · 2021
Earlier work this paper cites.
Experts, errors, and context: A large-scale study of human evaluation for machine translation
Markus Freitag, George Foster, David Grangier, Viresh Ratnakar, Qijun Tan, and Wolfgang Macherey. 2021 · 2021
Cited alongside, same era.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021 · 2021
Cited alongside, same era.
Calibrate before use: Improving few-shot performance of language models
Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2021 · 2021
Cited alongside, same era.
Results of WMT22 metrics shared task: Stop using BLEU – neural metrics are better and more robust
Markus Freitag, Ricardo Rei, Nitika Mathur, Chi-kiu Lo, Craig Stewart, Eleftherios Avramidis, Tom Kocmi, George Foster, Alon Lavie, and André F. T. Martins. 2022 · 2022
Cited alongside, same era.
The suboptimal wmt test sets and its impact on human parity
Ahrii Kim, Yunju Bak, Jimin Sun, Sungwon Lyu, and Changmin Lee. 2023 · 2023
Later among the works it cites.
Findings of the 2023 conference on machine translation (wmt23): Llms are here but not quite there yet
Tom Kocmi, Eleftherios Avramidis, Rachel Bawden, Ondřej Bojar, Anton Dvorkovich, Christian Federmann, Mark Fishel, Markus Freitag, Thamme Gowda, Roman Grundkiewicz, et al. 2023 · 2023
Later among the works it cites.
Deepfake text detection in the wild
Yafu Li, Qintong Li, Leyang Cui, Wei Bi, Longyue Wang, Linyi Yang, Shuming Shi, and Yue Zhang. 2023 · 2023
Later among the works it cites.
Towards making the most of chatgpt for machine translation
Keqin Peng, Liang Ding, Qihuang Zhong, Li Shen, Xuebo Liu, Min Zhang, Yuanxin Ouyang, and Dacheng Tao. 2023 · 2023
Later among the works it cites.
Chatgpt and gpt-4 for professional translators: Exploring the potential of large language models in translation
Sai Cheong Siu. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tanya Goyal, Junyi Jessy Li, and Greg Durrett. 2022 · 2022
Cited alongside, same era.
Findings of the 2022 conference on machine translation (wmt22)
Tom Kocmi, Rachel Bawden, Ondřej Bojar, Anton Dvorkovich, Christian Federmann, Mark Fishel, Thamme Gowda, Yvette Graham, Roman Grundkiewicz, Barry Haddow, et al. 2022 · 2022
Cited alongside, same era.
On" human parity" and" super human performance" in machine translation evaluation
Thierry Poibeau. 2022 · 2022
Cited alongside, same era.
Guangsheng Bao, Yanbin Zhao, Zhiyang Teng, Linyi Yang, and Yue Zhang. 2023 · 2023
Cited alongside, same era.
Gpt-4 surpassing human performance in linguistic pragmatics
Ljubisa Bojic, Predrag Kovacevic, and Milan Cabarkapa. 2023 · 2023
Cited alongside, same era.
Seamlessm4t: Massively multilingual & multimodal machine translation
Seamless Communication, Loïc Barrault, Yu-An Chung, Mariano Cora Meglioli, David Dale, Ning Dong, Paul-Ambroise Duquenne, Hady Elsahar, Hongyu Gong, Kevin Heffernan, John Hoffman, Christopher Klaiber, Pengwei Li, Daniel Licht, Jean Maillard, Alice Rakotoarison, Kaushik Ram Sadagopan, Guillaume Wenzek, Ethan Ye, Bapi Akula, Peng-Jen Chen, Naji El Hachem, Brian Ellis, Gabriel Mejia Gonzalez, Justin Haaheim, Prangthip Hansanti, Russ Howes, Bernie Huang, Min-Jae Hwang, Hirofumi Inaguma, Somya Jain, Elahe Kalbassi, Amanda Kallet, Ilia Kulikov, Janice Lam, Daniel Li, Xutai Ma, Ruslan Mavlyutov, Benjamin Peloquin, Mohamed Ramadan, Abinesh Ramakrishnan, Anna Sun, Kevin Tran, Tuan Tran, Igor Tufanov, Vish Vogeti, Carleigh Wood, Yilin Yang, Bokai Yu, Pierre Andrews, Can Balioglu, Marta R. Costa-jussà, Onur Celebi, Maha Elbayad, Cynthia Gao, Francisco Guzmán, Justine Kao, Ann Lee, Alexandre Mourachko, Juan Pino, Sravya Popuri, Christophe Ropers, Safiyyah Saleem, Holger Schwenk, Paden Tomasello, Changhan Wang, Jeff Wang, and Skyler Wang. 2023 · 2023
Cited alongside, same era.
Results of WMT23 metrics shared task: Metrics might be guilty but references are not innocent
Markus Freitag, Nitika Mathur, Chi-kiu Lo, Eleftherios Avramidis, Ricardo Rei, Brian Thompson, Tom Kocmi, Frederic Blain, Daniel Deutsch, Craig Stewart, Chrysoula Zerva, Sheila Castilho, Alon Lavie, and George Foster. 2023 · 2023
Cited alongside, same era.
Comparative analysis of gpt-4vision, gpt-4 and open source llms in clinical diagnostic accuracy: A benchmark against human expertise
Tianyu Han, Lisa C Adams, Keno Bressem, Felix Busch, Luisa Huck, Sven Nebelung, and Daniel Truhn. 2023 · 2023
Cited alongside, same era.
Later among the works it cites.
A paradigm shift in machine translation: Boosting translation performance of large language models
Haoran Xu, Young Jin Kim, Amr Sharaf, and Hany Hassan Awadalla. 2023 · 2023
Later among the works it cites.
Revisiting out-of-distribution robustness in nlp: Benchmark, analysis, and llms evaluations
Lifan Yuan, Yangyi Chen, Ganqu Cui, Hongcheng Gao, Fangyuan Zou, Xingyi Cheng, Heng Ji, Zhiyuan Liu, and Maosong Sun. 2023 · 2023
Later among the works it cites.
Siren’s song in the ai ocean: a survey on hallucination in large language models
Yue Zhang, Yafu Li, Leyang Cui, Deng Cai, Lemao Liu, Tingchen Fu, Xinting Huang, Enbo Zhao, Yu Zhang, Yulong Chen, et al. 2023 · 2023
Later among the works it cites.
From llm to nmt: Advancing low-resource machine translation with claude
Maxim Enis and Mark Hopkins. 2024 · 2024
Closest in time.
Tear: Improving llm-based machine translation with systematic self-refinement
Zhaopeng Feng, Yan Zhang, Hao Li, Bei Wu, Jiayu Liao, Wenqiang Liu, Jun Lang, Yang Feng, Jian Wu, and Zuozhu Liu. 2024 · 2024
Closest in time.
A comparison of human and gpt-4 use of probabilistic phrases in a coordination game
Laurence T Maloney, Maria F Dal Martello, Vivian Fei, and Valerie Ma. 2024 · 2024
Closest in time.
Using gpt-4 to provide tiered, formative code feedback
Ha Nguyen and Vicki Allan. 2024 · 2024
Closest in time.
Benchmarking llms via uncertainty quantification
Fanghua Ye, Mingming Yang, Jianhui Pang, Longyue Wang, Derek F. Wong, Emine Yilmaz, Shuming Shi, and Zhaopeng Tu. 2024 · 2024
Closest in time.
Jie Zhu, Junhui Li, Yalong Wen, and Lifan Guo. 2024 · 2024
Closest in time.