Fetching the paper…
Reading the bibliography…
The zero-shot capability of Large Language Models (LLMs) has enabled highly flexible, reference-free metrics for various tasks, making LLM evaluators common tools in NLP.
Attitudinal effects of mere exposure
Robert B Zajonc. 1968 · 1968
Earlier work this paper cites.
Judgment under uncertainty: Heuristics and biases
Amos Tversky and Daniel Kahneman. 1974 · 1974
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie. 2005 · 2005
Earlier work this paper cites.
Heuristics in numerical cognition: Implications for pricing
Manoj Thomas and Vicki Morwitz. 2009 · 2009
Earlier work this paper cites.
How frequent are numbers?
Nikolas Coupland. 2011 · 2011
Earlier work this paper cites.
Automatically assessing machine summary content without a gold standard
Annie Louis and Ani Nenkova. 2013 · 2013
Earlier work this paper cites.
Likert-scale questionnaires
Tomoko Nemoto and David Beglar. 2014 · 2013
Earlier work this paper cites.
Abstractive text summarization using sequence-to-sequence rnns and beyond
Ramesh Nallapati, Bowen Zhou, Cicero Nogueira dos santos, Caglar Gulcehre, and Bing Xiang. 2016 · 2016
Earlier work this paper cites.
Shashi Narayan, Shay B. Cohen, and Mirella Lapata. 2018 · 2018
Earlier work this paper cites.
Findings of the WMT 2019 shared tasks on quality estimation
Erick Fonseca, Lisa Yankovskaya, André F. T. Martins, Mark Fishel, and Christian Federmann. 2019 · 2019
Earlier work this paper cites.
SAMSum corpus: A human-annotated dialogue dataset for abstractive summarization
Bogdan Gliwa, Iwona Mochol, Maciej Biesek, and Aleksander Wawer. 2019 · 2019
Earlier work this paper cites.
Answers unite! unsupervised metrics for reinforced summarization models
Thomas Scialom, Sylvain Lamprier, Benjamin Piwowarski, and Jacopo Staiano. 2019 · 2019
Earlier work this paper cites.
MoverScore: Text generation evaluating with contextualized embeddings and earth mover distance
Wei Zhao, Maxime Peyrard, Fei Liu, Yang Gao, Christian M. Meyer, and Steffen Eger. 2019 · 2019
Earlier work this paper cites.
Re-evaluating evaluation in text summarization
Manik Bhandari, Pranav Narayan Gour, Atabak Ashfaq, Pengfei Liu, and Graham Neubig. 2020 · 2020
Cited alongside, same era.
BLEURT: Learning robust metrics for text generation
Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020 · 2020
Cited alongside, same era.
Fill in the BLANC: Human-free quality estimation of document summaries
Oleg Vasilyev, Vedant Dharnidharka, and John Bohannon. 2020 · 2020
Cited alongside, same era.
BertScore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2020 · 2020
Cited alongside, same era.
SummEval: Re-evaluating summarization evaluation
Alexander R. Fabbri, Wojciech Kryściński, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir Radev. 2021 · 2021
Cited alongside, same era.
Are references really needed? unbabel-IST 2021 submission for the metrics shared task
Human-like summarization evaluation with chatgpt
Mingqi Gao, Jie Ruan, Renliang Sun, Xunjian Yin, Shiping Yang, and Xiaojun Wan. 2023 · 2023
Later among the works it cites.
Large language models are state-of-the-art evaluators of translation quality
Tom Kocmi and Christian Federmann. 2023 · 2023
Later among the works it cites.
Benchmarking cognitive biases in large language models as evaluators
Ryan Koo, Minhwa Lee, Vipul Raheja, Jong Inn Park, Zae Myung Kim, and Dongyeop Kang. 2023 · 2023
Later among the works it cites.
Yen-Ting Lin and Yun-Nung Chen. 2023 · 2023
Later among the works it cites.
Chatgpt as a factual inconsistency evaluator for text summarization
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ricardo Rei, Ana C Farinha, Chrysoula Zerva, Daan van Stigt, Craig Stewart, Pedro Ramos, Taisiya Glushkova, André F. T. Martins, and Alon Lavie. 2021 · 2021
Cited alongside, same era.
QuestEval: Summarization asks for fact-based evaluation
Thomas Scialom, Paul-Alexis Dray, Sylvain Lamprier, Benjamin Piwowarski, Jacopo Staiano, Alex Wang, and Patrick Gallinari. 2021 · 2021
Cited alongside, same era.
On the limitations of reference-free evaluations of generated text
Daniel Deutsch, Rotem Dror, and Dan Roth. 2022 · 2022
Cited alongside, same era.
On the round number bias and wisdom of crowds in different response formats for numerical estimation
Hidehito Honda, Rina Kagawa, and Masaru Shirasuna. 2022 · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed Chi, Quoc V Le, and Denny Zhou. 2022 · 2022
Cited alongside, same era.
Chateval: Towards better llm-based evaluators through multi-agent debate
Chi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu, Wei Xue, Shanghang Zhang, Jie Fu, and Zhiyuan Liu. 2023 · 2023
Cited alongside, same era.
Yi Chen, Rui Wang, Haiyun Jiang, Shuming Shi, and Ruifeng Xu. 2023 · 2023
Cited alongside, same era.
Zheheng Luo, Qianqian Xie, and Sophia Ananiadou. 2023 · 2023
Later among the works it cites.
Large language models are not yet human-level evaluators for abstractive summarization
Chenhui Shen, Liying Cheng, Xuan-Phi Nguyen, Yang You, and Lidong Bing. 2023 · 2023
Later among the works it cites.
Style over substance: Evaluation biases for large language models
Minghao Wu and Alham Fikri Aji. 2023 · 2023
Later among the works it cites.
Less is more for long document summary evaluation by llms
Yunshu Wu, Hayate Iso, Pouya Pezeshkpour, Nikita Bhutani, and Estevam Hruschka. 2023 · 2023
Later among the works it cites.
Hui Huang, Yingqi Qu, Jing Liu, Muyun Yang, and Tiejun Zhao. 2024 · 2024
Closest in time.
Beyond reference-based metrics: Analyzing behaviors of open llms on data-to-text generation
Zdeněk Kasner and Ondřej Dušek. 2024 · 2024
Closest in time.
Leveraging large language models for nlg evaluation: A survey
Zhen Li, Xiaohan Xu, Tao Shen, Can Xu, Jia-Chen Gu, and Chongyang Tao. 2024 · 2024
Closest in time.
Likelihood-based mitigation of evaluation bias in large language models
Masanari Ohi, Masahiro Kaneko, Ryuto Koike, Mengsay Loem, and Naoaki Okazaki. 2024 · 2024
Closest in time.
Llm evaluators recognize and favor their own generations
Arjun Panickssery, Samuel R. Bowman, and Shi Feng. 2024 · 2024
Closest in time.
Characterizing the confidence of large language model-based automatic evaluation metrics
Rickard Stureborg, Dimitris Alikaniotis, and Yoshi Suhara. 2024 · 2024
Closest in time.
Hierarchical multi-label classification of online vaccine concerns
Chloe Qinyu Zhu, Rickard Stureborg, and Bhuwan Dhingra. 2024 · 2024
Closest in time.