Fetching the paper…
Reading the bibliography…
In this work, we propose and evaluate the feasibility of a two-stage pipeline to evaluate literary machine translation, in a fine-grained manner, from English to Korean.
What distinguishes major types of translation?
Juan C Sager. 1998 · 1998
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Bleurt: Learning robust metrics for text generation
Thibault Sellam, Dipanjan Das, and Ankur P Parikh. 2020 · 2004
Earlier work this paper cites.
Asking and answering questions to evaluate the factual consistency of summaries
Alex Wang, Kyunghyun Cho, and Mike Lewis. 2020 · 2004
Earlier work this paper cites.
The turns of translation studies
Mary Snell-Hornby. 2006 · 2006
Earlier work this paper cites.
Comet: A neural framework for mt evaluation
Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. 2020 · 2009
Earlier work this paper cites.
The Multidimensional Quality Metric (MQM) framework: A new framework for translation quality assessment
Valerie R Mariana. 2014 · 2014
Earlier work this paper cites.
The translator’s invisibility: A history of translation
Lawrence Venuti. 2017 · 2017
Earlier work this paper cites.
Sociologies of Poetry Translation: Emerging Perspectives
Jacob Blakesley, editor. 2018 · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Bertscore: Evaluating text generation with bert
Tianyi Zhang*, Varsha Kishore*, Felix Wu*, Kilian Q. Weinberger, and Yoav Artzi. 2020 · 2020
Earlier work this paper cites.
Just ask! evaluating machine translation by asking and answering questions
Mateusz Krubiński, Erfan Ghadery, Marie-Francine Moens, and Pavel Pecina. 2021 · 2021
Earlier work this paper cites.
Zero-shot text-to-image generation
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. 2021 · 2021
Earlier work this paper cites.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. 2022 · 2022
Earlier work this paper cites.
Literary translation as a three-stage process: Machine translation, post-editing and revision
Lieve Macken, Bram Vanroy, Luca Desmet, and Arda Tezcan. 2022 · 2022
Earlier work this paper cites.
Introducing translation studies: Theories and applications
Jeremy Munday, Sara Ramos Pinto, and Jacob Blakesley. 2022 · 2022
Earlier work this paper cites.
Comet-22: Unbabel-ist 2022 submission for the metrics shared task
Ricardo Rei, José GC De Souza, Duarte Alves, Chrysoula Zerva, Ana C Farinha, Taisiya Glushkova, Alon Lavie, Luisa Coheur, and André FT Martins. 2022 · 2022
Earlier work this paper cites.
Exploring document-level literary machine translation with parallel paragraphs from world literature
Katherine Thai, Marzena Karpinska, Kalpesh Krishna, Bill Ray, Moira Inghilleri, John Wieting, and Mohit Iyyer. 2022 · 2022
Earlier work this paper cites.
Chateval: Towards better llm-based evaluators through multi-agent debate
Chi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu, Wei Xue, Shanghang Zhang, Jie Fu, and Zhiyuan Liu. 2023 · 2023
Earlier work this paper cites.
Davidsonian scene graph: Improving reliability in fine-grained evaluation for text-image generation
Jaemin Cho, Yushi Hu, Roopal Garg, Peter Anderson, Ranjay Krishna, Jason Baldridge, Mohit Bansal, Jordi Pont-Tuset, and Su Wang. 2023 · 2023
Earlier work this paper cites.
xcomet: Transparent machine translation evaluation through fine-grained error detection
Pierre Colombo, Nuno Guerreiro, Ricardo Rei, Daan Van, Luisa Coheur, and André Martins. 2023 · 2023
Cited alongside, same era.
Alpacafarm: A simulation framework for methods that learn from human feedback
Yann Dubois, Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023 · 2023
Cited alongside, same era.
Results of wmt23 metrics shared task: Metrics might be guilty but references are not innocent
Markus Freitag, Nitika Mathur, Chi-kiu Lo, Eleftherios Avramidis, Ricardo Rei, Brian Thompson, Tom Kocmi, Frédéric Blain, Daniel Deutsch, Craig Stewart, et al. 2023 · 2023
Cited alongside, same era.
Tifa: Accurate and interpretable text-to-image faithfulness evaluation with question answering
Yushi Hu, Benlin Liu, Jungo Kasai, Yizhong Wang, Mari Ostendorf, Ranjay Krishna, and Noah A Smith. 2023 · 2023
Cited alongside, same era.
Gemma 2: Improving open language models at a practical scale
Google DeepMind and Google Research. 2024 · 2024
Closest in time.
Andrew Grattafiori and et al. 2024 · 2024
Closest in time.
Exploring human-like translation strategy with large language models
Zhiwei He, Tian Liang, Wenxiang Jiao, Zhuosheng Zhang, Yujiu Yang, Rui Wang, Zhaopeng Tu, Shuming Shi, and Xing Wang. 2024 · 2024
Closest in time.
Fables: Evaluating faithfulness and content selection in book-length summarization
Yekyung Kim, Yapei Chang, Marzena Karpinska, Aparna Garimella, Varun Manjunatha, Kyle Lo, Tanya Goyal, and Mohit Iyyer. 2024 · 2024
Closest in time.
The last frontier of machine translation
Jeremy Klemin. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang, Xing Wang, Shuming Shi, and Zhaopeng Tu. 2023 · 2023
Cited alongside, same era.
Marzena Karpinska and Mohit Iyyer. 2023 · 2023
Cited alongside, same era.
Prometheus: Inducing fine-grained evaluation capability in language models
Seungone Kim, Jamin Shin, Yejin Cho, Joel Jang, Shayne Longpre, Hwaran Lee, Sangdoo Yun, Seongjin Shin, Sungdong Kim, James Thorne, et al. 2023 · 2023
Cited alongside, same era.
Gemba-mqm: Detecting translation quality error spans with gpt-4
Tom Kocmi and Christian Federmann. 2023 · 2023
Cited alongside, same era.
‘i am a bit surprised’: Literary translation and post-editing processes compared
Waltraud Kolb. 2023 · 2023
Cited alongside, same era.
Generative judge for evaluating alignment
Junlong Li, Shichao Sun, Weizhe Yuan, Run-Ze Fan, Hai Zhao, and Pengfei Liu. 2023 · 2023
Cited alongside, same era.
G-eval: Nlg evaluation using gpt-4 with better human alignment
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023 · 2023
Cited alongside, same era.
Factscore: Fine-grained atomic evaluation of factual precision in long form text generation
Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen-tau Yih, Pang Wei Koh, Mohit Iyyer, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2023 · 2023
Cited alongside, same era.
Hello gpt-4
OpenAI. 2024 · 2024
Closest in time.
SDXL: Improving latent diffusion models for high-resolution image synthesis
Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Müller, Joe Penna, and Robin Rombach. 2024 · 2024
Closest in time.
Mm-eval: A multilingual meta-evaluation benchmark for llm-as-a-judge and reward models
Guijin Son, Dongkeun Yoon, Juyoung Suk, Javier Aula-Blasco, Mano Aslan, Vu Trong Kim, Shayekh Bin Islam, Jaume Prats-Cristià, Lucía Tormo-Bañuelos, and Seungone Kim. 2024 · 2024
Closest in time.
Minghao Wu, Yulin Yuan, Gholamreza Haffari, and Longyue Wang. 2024 · 2024
Closest in time.
Ran Zhang, Wei Zhao, and Steffen Eger. 2024 · 2024
Closest in time.
The claude 3 model family: Opus, sonnet, haiku
Anthropic. 2024a · 2025
Closest in time.
Introducing claude 3.5 sonnet
Anthropic. 2024b · 2025
Closest in time.
Model card addendum: Claude 3.5 haiku and upgraded claude 3.5 sonnet
Anthropic. 2024c · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, et al. 2025 · 2025
Closest in time.
Introducing llama 3.1: Our most capable models to date
Meta AI. 2024 · 2025
Closest in time.
Gpt-4o mini: Advancing cost-efficient intelligence
OpenAI. 2024 · 2025
Closest in time.
Hello gpt-4o
OpenAI. 2024 · 2025
Closest in time.
Introducing gpt-4.1 in the api
OpenAI. 2025a · 2025
Closest in time.
Introducing o3 and o4-mini
OpenAI. 2025b · 2025
Closest in time.