Fetching the paper…
Reading the bibliography…
As LLMs continue to become more powerful and versatile, human evaluation has quickly become intractable at scale and reliance on automatic metrics has become the norm.
Proceedings of the third workshop on statistical machine translation
Callison-Burch, C., Koehn, P., Monz, C., Schroeder, J., and Fordyce, C. S · 2008
Earlier work this paper cites.
Multidimensional quality metrics (mqm): A framework for declaring and describing translation quality metrics
Lommel, A., Uszkoreit, H., and Burchardt, A · 2014
Earlier work this paper cites.
Geva, M., Goldberg, Y., and Berant, J · 2019
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2020
Earlier work this paper cites.
Comet: A neural framework for mt evaluation
Rei, R., Stewart, C., Farinha, A. C., and Lavie, A · 2020
Earlier work this paper cites.
Bleurt: Learning robust metrics for text generation
Sellam, T., Das, D., and Parikh, A. P · 2020
Earlier work this paper cites.
Experts, errors, and context: A large-scale study of human evaluation for machine translation
Freitag, M., Foster, G., Grangier, D., Ratnakar, V., Tan, Q., and Macherey, W · 2021
Earlier work this paper cites.
The perils of using mechanical turk to evaluate open-ended text generation
Karpinska, M., Akoury, N., and Iyyer, M · 2021
Earlier work this paper cites.
Constitutional ai: Harmlessness from ai feedback
Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C., et al · 2022
Earlier work this paper cites.
Holistic evaluation of language models
Liang, P., Bommasani, R., Lee, T., Tsipras, D., Soylu, D., Yasunaga, M., Zhang, Y., Narayanan, D., Wu, Y., Kumar, A., et al · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Earlier work this paper cites.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Cited alongside, same era.
Findings of the wmt 2023 shared task on quality estimation
Blain, F., Zerva, C., Rei, R., Guerreiro, N. M., Kanojia, D., de Souza, J. G., Silva, B., Vaz, T., Jingxuan, Y., Azadi, F., et al · 2023
Cited alongside, same era.
Ties matter: Meta-evaluating modern metrics with pairwise accuracy and tie calibration
Deutsch, D., Foster, G., and Freitag, M · 2023
Cited alongside, same era.
Fernandes, P., Deutsch, D., Finkelstein, M., Riley, P., Martins, A. F., Neubig, G., Garg, A., Clark, J. H., Freitag, M., and Firat, O · 2023
Cited alongside, same era.
Results of wmt23 metrics shared task: Metrics might be guilty but references are not innocent
Chen, B., Wang, X., Peng, S., Litschko, R., Korhonen, A., and Plank, B · 2024
Closest in time.
Are llms breaking mt metrics? results of the wmt24 metrics shared task
Freitag, M., Mathur, N., Deutsch, D., Lo, C.-K., Avramidis, E., Rei, R., Thompson, B., Blain, F., Kocmi, T., Wang, J., et al · 2024
Closest in time.
Gemini: A family of highly capable multimodal models, 2024
Gemini Team · 2024
Closest in time.
Cost-efficient subjective task annotation and modeling through few-shot annotator adaptation
Golazizian, P., Ziabari, A. S., Omrani, A., and Dehghani, M · 2024
Closest in time.
Metricx-24: The google submission to the wmt 2024 metrics shared task
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Freitag, M., Mathur, N., Lo, C.-k., Avramidis, E., Rei, R., Thompson, B., Kocmi, T., Blain, F., Deutsch, D., Stewart, C., et al · 2023
Cited alongside, same era.
xcomet: Transparent machine translation evaluation through fine-grained error detection
Guerreiro, N. M., Rei, R., van Stigt, D., Coheur, L., Colombo, P., and Martins, A. F · 2023
Cited alongside, same era.
Prometheus: Inducing fine-grained evaluation capability in language models
Kim, S., Shin, J., Cho, Y., Jang, J., Longpre, S., Lee, H., Yun, S., Shin, S., Kim, S., Thorne, J., et al · 2023
Cited alongside, same era.
Longeval: Guidelines for human evaluation of faithfulness in long-form summarization
Krishna, K., Bransom, E., Kuehl, B., Iyyer, M., Dasigi, P., Cohan, A., and Lo, K · 2023
Cited alongside, same era.
Generative judge for evaluating alignment
Li, J., Sun, S., Yuan, W., Fan, R.-Z., Zhao, H., and Liu, P · 2023
Cited alongside, same era.
A benchmark for learning to translate a new language from one grammar book
Tanzer, G., Suzgun, M., Visser, E., Jurafsky, D., and Melas-Kyriazi, L · 2023
Cited alongside, same era.
Understanding in-context learning from repetitions
Yan, J., Xu, J., Song, C., Wu, C., Li, Y., and Zhang, Y · 2023
Cited alongside, same era.
Gemba-mqm: Detecting translation quality error spans with gpt-4
Kocmi, T. and Federmann, C
Cited in the paper.
Juraska, J., Deutsch, D., Finkelstein, M., and Freitag, M · 2024
Closest in time.
Evaluating llms at detecting errors in llm responses
Kamoi, R., Das, S. S. S., Lou, R., Ahn, J. J., Zhao, Y., Lu, X., Zhang, N., Zhang, Y., Zhang, R. H., Vummanthala, S. R., et al · 2024
Closest in time.
Prometheus 2: An open source language model specialized in evaluating other language models
Kim, S., Suk, J., Longpre, S., Lin, B. Y., Shin, J., Welleck, S., Neubig, G., Lee, M., Lee, K., and Seo, M · 2024
Closest in time.
Finding replicable human evaluations via stable ranking probability
Riley, P., Deutsch, D., Foster, G., Ratnakar, V., Dabirmoghaddam, A., and Freitag, M · 2024
Closest in time.
Foundational autoraters: Taming large language models for better automatic evaluation
Vu, T., Krishna, K., Alzubi, S., Tar, C., Faruqui, M., and Sung, Y.-H · 2024
Closest in time.
LLMRefine: Pinpointing and refining large language models via fine-grained actionable feedback
Xu, W., Deutsch, D., Finkelstein, M., Juraska, J., Zhang, B., Liu, Z., Wang, W. Y., Li, L., and Freitag, M · 2024
Closest in time.