Fetching the paper…
Reading the bibliography…
Significant progress has been made in automatic text evaluation with the introduction of large language models (LLMs) as evaluators.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
The meteor metric for automatic evaluation of machine translation
Alon Lavie and Michael J. Denkowski. 2009 · 2009
Earlier work this paper cites.
Ensemble methods: foundations and algorithms
Zhi-Hua Zhou. 2012 · 2012
Earlier work this paper cites.
Bootstrapping dialog systems with word embeddings
Gabriel Forgues, Joelle Pineau, Jean-Marie Larchevêque, and Réal Tremblay. 2014 · 2014
Earlier work this paper cites.
Teaching machines to read and comprehend
Karl Moritz Hermann, Tomas Kocisky, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom. 2015 · 2015
Earlier work this paper cites.
Towards an automatic turing test: Learning to evaluate dialogue responses
Ryan Lowe, Michael D. Noseworthy, Iulian Vlad Serban, Nicolas Angelard-Gontier, Yoshua Bengio, and Joelle Pineau. 2017 · 2017
Earlier work this paper cites.
Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization
Shashi Narayan, Shay Cohen, and Maria Lapata. 2018 · 2018
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 2020
Earlier work this paper cites.
Unsupervised evaluation of interactive dialog with dialogpt
Shikib Mehri and Maxine Eskenazi. 2020 · 2020
Cited alongside, same era.
USR: an unsupervised and reference free evaluation metric for dialog generation
Shikib Mehri and Maxine Eskénazi. 2020 · 2020
Cited alongside, same era.
Asking and answering questions to evaluate the factual consistency of summaries
Alex Wang, Kyunghyun Cho, and Mike Lewis. 2020 · 2020
Cited alongside, same era.
Bertscore: Evaluating text generation with BERT
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020 · 2020
Cited alongside, same era.
Annotating and modeling fine-grained factuality in summarization
Tanya Goyal and Greg Durrett. 2021 · 2021
Cited alongside, same era.
Dynaeval: Unifying turn and dialogue level evaluation
Chen Zhang, Yiming Chen, Luis Fernando D’Haro, Yan Zhang, Thomas Friedrichs, Grandee Lee, and Haizhou Li. 2021 · 2021
Can large language models be an alternative to human evaluations?
David Cheng-Han Chiang and Hung-yi Lee. 2023a · 2023
Closest in time.
A closer look into using large language models for automatic evaluation
David Cheng-Han Chiang and Hung-yi Lee. 2023b · 2023
Closest in time.
Alpacafarm: A simulation framework for methods that learn from human feedback
Yann Dubois, Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy Liang, and Tatsunori B Hashimoto. 2023 · 2023
Closest in time.
Yen-Ting Lin and Yun-Nung Chen. 2023 · 2023
Closest in time.
G-eval: NLG evaluation using gpt-4 with better human alignment
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023 · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Of human criteria and automatic metrics: A benchmark of the evaluation of story generation
Cyril Chhun, Pierre Colombo, Fabian M Suchanek, and Chloé Clavel. 2022 · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou. 2022 · 2022
Cited alongside, same era.
MME-CRS: multi-metric evaluation based on correlation re-scaling for evaluating open-domain dialogue
Pengfei Zhang, Xiaohui Hu, Kaidong Yu, Jian Wang, Song Han, Cao Liu, and Chunyang Yuan. 2022 · 2022
Cited alongside, same era.
Evaluating large language models: A comprehensive survey
Zishan Guo, Renren Jin, Chuang Liu, Yufei Huang, Dan Shi, Supryadi, Linhao Yu, Yan Liu, Jiaxuan Li, Bojian Xiong, and Deyi Xiong. 2023a
Cited in the paper.
Evaluating large language models: A comprehensive survey
Zishan Guo, Renren Jin, Chuang Liu, Yufei Huang, Dan Shi, Supryadi, Linhao Yu, Yan Liu, Jiaxuan Li, Bojian Xiong, and Deyi Xiong. 2023b
Cited in the paper.
OpenAI. 2023 · 2023
Closest in time.
A survey of evaluation metrics used for NLG systems
Ananya B. Sai, Akash Kumar Mohankumar, and Mitesh M. Khapra. 2023 · 2023
Closest in time.
Better correlation and robustness: A distribution-balanced self-supervised learning framework for automatic dialogue evaluation
Peiwen Yuan, Xinglin Wang, Jiayi Shi, Bin Sun, Yiwei Li, and Kan Li. 2023 · 2023
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P Xing, et al. 2023 · 2023
Closest in time.