Fetching the paper…
Reading the bibliography…
With an increasing number of parameters and pre-training data, generative large language models (LLMs) have shown remarkable capabilities to solve tasks with minimal or no task-related examples.
THE TREATMENT OF TIES IN RANKING PROBLEMS
M. G. KENDALL. 1945 · 1945
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Evaluating content selection in summarization: The pyramid method
Ani Nenkova and Rebecca Passonneau. 2004 · 2004
Earlier work this paper cites.
Overview of duc 2005
Hoa Trang Dang. 2005 · 2005
Earlier work this paper cites.
Multidimensional quality metrics (mqm): A framework for declaring and describing translation quality metrics
Arle Lommel, Aljoscha Burchardt, and Hans Uszkoreit. 2014 · 2014
Earlier work this paper cites.
XGBoost
Tianqi Chen and Carlos Guestrin. 2016 · 2016
Earlier work this paper cites.
A structured review of the validity of BLEU
Ehud Reiter. 2018 · 2018
Earlier work this paper cites.
A study of reinforcement learning for neural machine translation
Lijun Wu, Fei Tian, Tao Qin, Jianhuang Lai, and Tie-Yan Liu. 2018 · 2018
Earlier work this paper cites.
Sentence-BERT: Sentence embeddings using Siamese BERT-networks
Nils Reimers and Iryna Gurevych. 2019 · 2019
Earlier work this paper cites.
MoverScore: Text generation evaluating with contextualized embeddings and earth mover distance
Wei Zhao, Maxime Peyrard, Fei Liu, Yang Gao, Christian M. Meyer, and Steffen Eger. 2019 · 2019
Earlier work this paper cites.
SUPERT: Towards new frontiers in unsupervised evaluation metrics for multi-document summarization
Yang Gao, Wei Zhao, and Steffen Eger. 2020 · 2020
Earlier work this paper cites.
Results of the WMT20 metrics shared task
Nitika Mathur, Johnny Wei, Markus Freitag, Qingsong Ma, and Ondřej Bojar. 2020 · 2020
Earlier work this paper cites.
COMET: A neural framework for MT evaluation
Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. 2020 · 2020
Earlier work this paper cites.
BLEURT: Learning robust metrics for text generation
Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020 · 2020
Earlier work this paper cites.
Pre-trained summarization distillation
Sam Shleifer and Alexander M. Rush. 2020 · 2020
Earlier work this paper cites.
Findings of the WMT 2020 shared task on quality estimation
Lucia Specia, Frédéric Blain, Marina Fomicheva, Erick Fonseca, Vishrav Chaudhary, Francisco Guzmán, and André F. T. Martins. 2020 · 2020
Earlier work this paper cites.
Automatic machine translation evaluation in many languages via zero-shot paraphrasing
Brian Thompson and Matt Post. 2020 · 2020
Earlier work this paper cites.
Big bird: Transformers for longer sequences
Manzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, and Amr Ahmed. 2020 · 2020
Earlier work this paper cites.
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020 · 2020
Earlier work this paper cites.
Evaluation of text generation: A survey
Asli Celikyilmaz, Elizabeth Clark, and Jianfeng Gao. 2021 · 2021
Earlier work this paper cites.
A statistical analysis of summarization evaluation metrics using resampling methods
Daniel Deutsch, Rotem Dror, and Dan Roth. 2021 · 2021
Earlier work this paper cites.
SummEval: Re-evaluating summarization evaluation
Alexander R. Fabbri, Wojciech Kryściński, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir Radev. 2021 · 2021
Earlier work this paper cites.
Beyond english-centric multilingual machine translation
Angela Fan, Shruti Bhosale, Holger Schwenk, Zhiyi Ma, Ahmed El-Kishky, Siddharth Goyal, Mandeep Baines, Onur Celebi, Guillaume Wenzek, Vishrav Chaudhary, Naman Goyal, Tom Birch, Vitaliy Liptchinsky, Sergey Edunov, Michael Auli, and Armand Joulin. 2021 · 2021
Earlier work this paper cites.
The Eval4NLP shared task on explainable quality estimation: Overview and results
Marina Fomicheva, Piyawat Lertvittayakumjorn, Wei Zhao, Steffen Eger, and Yang Gao. 2021 · 2021
Earlier work this paper cites.
XL-sum: Large-scale multilingual abstractive summarization for 44 languages
Tahmid Hasan, Abhik Bhattacharjee, Md. Saiful Islam, Kazi Mubasshir, Yuan-Fang Li, Yong-Bin Kang, M. Sohel Rahman, and Rifat Shahriyar. 2021 · 2021
Earlier work this paper cites.
Global explainability of BERT-based evaluation metrics by disentangling along linguistic factors
Marvin Kaster, Wei Zhao, and Steffen Eger. 2021 · 2021
Earlier work this paper cites.
Findings of the WMT 2021 shared task on quality estimation
Lucia Specia, Frédéric Blain, Marina Fomicheva, Chrysoula Zerva, Zhenhao Li, Vishrav Chaudhary, and André F. T. Martins. 2021 · 2021
Cited alongside, same era.
Multilingual translation from denoising pre-training
Yuqing Tang, Chau Tran, Xian Li, Peng-Jen Chen, Naman Goyal, Vishrav Chaudhary, Jiatao Gu, and Angela Fan. 2021 · 2021
Cited alongside, same era.
Bartscore: Evaluating generated text as text generation
Weizhe Yuan, Graham Neubig, and Pengfei Liu. 2021 · 2021
Cited alongside, same era.
Language-agnostic BERT sentence embedding
Fangxiaoyu Feng, Yinfei Yang, Daniel Cer, Naveen Arivazhagan, and Wei Wang. 2022 · 2022
Cited alongside, same era.
Quality-aware decoding for neural machine translation
Patrick Fernandes, António Farinhas, Ricardo Rei, José G. C. de Souza, Perez Ogayo, Graham Neubig, and Andre Martins. 2022 · 2022
Cited alongside, same era.
Results of WMT22 metrics shared task: Stop using BLEU – neural metrics are better and more robust
xcomet: Transparent machine translation evaluation through fine-grained error detection
Nuno M. Guerreiro, Ricardo Rei, Daan van Stigt, Luisa Coheur, Pierre Colombo, and André F. T. Martins. 2023 · 2023
Closest in time.
Chatgpt in education: Strategies for responsible implementation
Mohanad Halaweh. 2023 · 2023
Closest in time.
On the blind spots of model-based evaluation metrics for text generation
Tianxing He, Jingyu Zhang, Tianle Wang, Sachin Kumar, Kyunghyun Cho, James Glass, and Yulia Tsvetkov. 2023 · 2023
Closest in time.
Yunjie Ji, Yan Gong, Yiping Peng, Chao Ni, Peiyan Sun, Dongyu Pan, Baochang Ma, and Xiangang Li. 2023 · 2023
Closest in time.
Which is better? exploring prompting strategy for llm-based metrics
JoongHoon Kim, Sangmin Lee, Seung Hun, Saeran Park, Jiyoon Lee, Kiyoon Jeong, and Pilsung Kang. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Markus Freitag, Ricardo Rei, Nitika Mathur, Chi-kiu Lo, Craig Stewart, Eleftherios Avramidis, Tom Kocmi, George Foster, Alon Lavie, and André F. T. Martins. 2022 · 2022
Cited alongside, same era.
Can we do that simpler? simple, efficient, high-quality evaluation metrics for nlg
Jens Grünwald, Christoph Leiter, and Steffen Eger. 2022 · 2022
Cited alongside, same era.
FrugalScore: Learning cheaper, lighter and faster evaluation metrics for automatic text generation
Moussa Kamal Eddine, Guokan Shang, Antoine Tixier, and Michalis Vazirgiannis. 2022 · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F Christiano, Jan Leike, and Ryan Lowe. 2022 · 2022
Cited alongside, same era.
Searching for COMETINHO: The little metric that could
Ricardo Rei, Ana C Farinha, José G.C. de Souza, Pedro G. Ramos, André F.T. Martins, Luisa Coheur, and Alon Lavie. 2022 · 2022
Cited alongside, same era.
A survey of evaluation metrics used for nlg systems
Ananya B. Sai, Akash Kumar Mohankumar, and Mitesh M. Khapra. 2022 · 2022
Cited alongside, same era.
Layer or representation space: What makes BERT-based evaluation metrics robust?
Doan Nam Long Vu, Nafise Sadat Moosavi, and Steffen Eger. 2022 · 2022
Cited alongside, same era.
Closest in time.
Large language models are state-of-the-art evaluators of translation quality
Tom Kocmi and Christian Federmann. 2023 · 2023
Closest in time.
Little giants: Exploring the potential of small llms as evaluation metrics in summarization in the eval4nlp 2023 shared task
Neema Kotonya, Saran Krishnasamy, Joel R. Tetreault, and Alejandro Jaimes. 2023 · 2023
Closest in time.
Team nllg submission for eval4nlp 2023 shared task: Retrieval-augmented in-context learning for nlg evaluation
Daniil Larionov, Vasiliy Viskov, George Kokush, Alexander Panchenko, and Steffen Eger. 2023 · 2023
Closest in time.
Qingyu Lu, Baopu Qiu, Liang Ding, Kanjian Zhang, Tom Kocmi, and Dacheng Tao. 2023 · 2023
Closest in time.
Characterised llms affect its evaluation of summary and translation
Yuan Lu and Lin Yu-Ting. 2023 · 2023
Closest in time.
Exploring prompting large language models as explainable metrics
Ghazaleh Mahmoudi. 2023 · 2023
Closest in time.
orca_mini_v3_7b: An explain tuned llama2-7b model
Pankaj Mathur. 2023 · 2023
Closest in time.
Orca: Progressive learning from complex explanation traces of gpt-4
Subhabrata Mukherjee, Arindam Mitra, Ganesh Jawahar, Sahaj Agarwal, Hamid Palangi, and Ahmed Awadallah. 2023 · 2023
Closest in time.
OpenAI. 2023 · 2023
Closest in time.
Codalab competitions: An open source platform to organize scientific challenges
Adrien Pavao, Isabelle Guyon, Anne-Catherine Letournel, Dinh-Tuan Tran, Xavier Baro, Hugo Jair Escalante, Sergio Escalera, Tyler Thomas, and Zhen Xu. 2023 · 2023
Closest in time.
Understanding large language model based metrics for text summarization
Abhishek Pradhan and Ketan Kumar Todi. 2023 · 2023
Closest in time.
Scaling up cometkiwi: Unbabel-ist 2023 submission for the quality estimation shared task
Ricardo Rei, Nuno M. Guerreiro, José Pombal, Daan van Stigt, Marcos Treviso, Luisa Coheur, José G. C. de Souza, and André F. T. Martins. 2023a · 2023
Closest in time.
Code llama: Open foundation models for code
Baptiste Rozière, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Tal Remez, Jérémy Rapin, Artyom Kozhevnikov, Ivan Evtimov, Joanna Bitton, Manish Bhatt, Cristian Canton Ferrer, Aaron Grattafiori, Wenhan Xiong, Alexandre Défossez, Jade Copet, Faisal Azhar, Hugo Touvron, Louis Martin, Nicolas Usunier, Thomas Scialom, and Gabriel Synnaeve. 2023 · 2023
Closest in time.
Are large language models good evaluators for abstractive summarization?
Chenhui Shen, Liying Cheng, Yang You, and Lidong Bing. 2023 · 2023
Closest in time.
Miles Turpin, Julian Michael, Ethan Perez, and Samuel R. Bowman. 2023 · 2023
Closest in time.
Self-consistency improves chain of thought reasoning in language models
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2023 · 2023
Closest in time.
Knowledge-prompted estimator: A novel approach to explainable machine translation assessment
Hao Yang, Min Zhang, Shimin Tao, Minghan Wang, Daimeng Wei, and Yanfei Jiang. 2023 · 2023
Closest in time.
Tree of thoughts: Deliberate problem solving with large language models
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, and Karthik Narasimhan. 2023 · 2023
Closest in time.
Hit-mi&t lab’s submission to eval4nlp 2023 shared task
Rui Zhang, Fuhai Song, Hui Huang, Jinghao Yuan, Muyun Yang, and Tiejun Zhao. 2023 · 2023
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric. P Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023 · 2023
Closest in time.
Large language models are human-level prompt engineers
Yongchao Zhou, Andrei Ioan Muresanu, Ziwen Han, Keiran Paster, Silviu Pitis, Harris Chan, and Jimmy Ba. 2023 · 2023
Closest in time.