Fetching the paper…
Reading the bibliography…
Recent advances in large language models (LLMs) show the potential of using LLMs as evaluators for assessing the quality of text generations from LLMs.
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2019 · 1904
Earlier work this paper cites.
Maximum likelihood estimation of observer error-rates using the em algorithm
Alexander Philip Dawid and Allan M Skene. 1979 · 1979
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
A cautionary note on the robustness of latent class models for estimating diagnostic error without a gold standard
Paul S Albert and Lori E Dodd. 2004 · 2004
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Multilevel bayesian models of categorical data annotation
Bob Carpenter. 2008 · 2008
Earlier work this paper cites.
Whose vote should count more: Optimal integration of labels from labelers of unknown expertise
Jacob Whitehill, Ting-fan Wu, Jacob Bergsma, Javier Movellan, and Paul Ruvolo. 2009 · 2009
Earlier work this paper cites.
The no-u-turn sampler: Adaptively setting path lengths in hamiltonian monte carlo
Matthew D. Hoffman and Andrew Gelman. 2011 · 2011
Earlier work this paper cites.
Bayesian classifier combination
Hyun-Chul Kim and Zoubin Ghahramani. 2012 · 2012
Earlier work this paper cites.
Learning whom to trust with MACE
Dirk Hovy, Taylor Berg-Kirkpatrick, Ashish Vaswani, and Eduard Hovy. 2013 · 2013
Earlier work this paper cites.
The benefits of a model of annotation
Rebecca J. Passonneau and Bob Carpenter. 2014 · 2014
Earlier work this paper cites.
Teaching machines to read and comprehend
Karl Moritz Hermann, Tomáš Kočiský, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom. 2015 · 2015
Earlier work this paper cites.
Spectral methods meet em: A provably optimal algorithm for crowdsourcing
Yuchen Zhang, Xi Chen, Dengyong Zhou, and Michael I Jordan. 2016 · 2016
Earlier work this paper cites.
Hierarchical neural story generation
Angela Fan, Mike Lewis, and Yann Dauphin. 2018 · 2018
Earlier work this paper cites.
Comparing Bayesian models of annotation
Silviu Paun, Bob Carpenter, Jon Chamberlain, Dirk Hovy, Udo Kruschwitz, and Massimo Poesio. 2018 · 2018
Earlier work this paper cites.
SummEval: Re-evaluating Summarization Evaluation
Alexander R. Fabbri, Wojciech Kryściński, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir Radev. 2021 · 2021
Cited alongside, same era.
OpenMEVA: A benchmark for evaluating open-ended story generation metrics
Jian Guan, Zhexin Zhang, Zhuoer Feng, Zitao Liu, Wenbiao Ding, Xiaoxi Mao, Changjie Fan, and Minlie Huang. 2021 · 2021
Cited alongside, same era.
Mauve: Measuring the gap between neural text and human text using divergence frontiers
Krishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun, Sean Welleck, Yejin Choi, and Zaid Harchaoui. 2021 · 2021
Cited alongside, same era.
Of human criteria and automatic metrics: A benchmark of the evaluation of story generation
Cyril Chhun, Pierre Colombo, Fabian M. Suchanek, and Chloé Clavel. 2022 · 2022
Cited alongside, same era.
QAFactEval: Improved QA-based factual consistency evaluation for summarization
Alexander Fabbri, Chien-Sheng Wu, Wenhao Liu, and Caiming Xiong. 2022 · 2022
Cited alongside, same era.
AlignScore: Evaluating factual consistency with a unified alignment function
Yuheng Zha, Yichi Yang, Ruichen Li, and Zhiting Hu. 2023 · 2023
Later among the works it cites.
Wider and deeper llm networks are fairer llm evaluators
Xinghua Zhang, Bowen Yu, Haiyang Yu, Yangyu Lv, Tingwen Liu, Fei Huang, Hongbo Xu, and Yongbin Li. 2023 · 2023
Later among the works it cites.
Alpacafarm: A simulation framework for methods that learn from human feedback
Yann Dubois, Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2024 · 2024
Closest in time.
Llm-ensemble: Optimal large language model ensemble method for e-commerce product attribute value extraction
Chenhao Fang, Xiaohan Li, Zezhong Fan, Jianpeng Xu, Kaushiki Nag, Evren Korpeoglu, Sushant Kumar, and Kannan Achan. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023 · 2022
Cited alongside, same era.
Chateval: Towards better llm-based evaluators through multi-agent debate
Chi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu, Wei Xue, Shanghang Zhang, Jie Fu, and Zhiyuan Liu. 2023 · 2023
Cited alongside, same era.
Menli: Robust evaluation metrics from natural language inference
Yanran Chen and Steffen Eger. 2023 · 2023
Cited alongside, same era.
A closer look into using large language models for automatic evaluation
Cheng-Han Chiang and Hung-yi Lee. 2023b · 2023
Cited alongside, same era.
Benchmarking cognitive biases in large language models as evaluators
Ryan Koo, Minhwa Lee, Vipul Raheja, Jong Inn Park, Zae Myung Kim, and Dongyeop Kang. 2023 · 2023
Cited alongside, same era.
Prd: Peer rank and discussion improve large language model based evaluations
Ruosen Li, Teerth Patel, and Xinya Du. 2023 · 2023
Cited alongside, same era.
G-eval: NLG evaluation using gpt-4 with better human alignment
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023a · 2023
Cited alongside, same era.
Yixin Liu, Kejian Shi, Alexander R. Fabbri, Yilun Zhao, Peifeng Wang, Chien-Sheng Wu, Shafiq Joty, and Arman Cohan. 2024 · 2024
Closest in time.
LLM comparative assessment: Zero-shot NLG evaluation through pairwise comparisons using large language models
Adian Liusie, Potsawee Manakul, and Mark Gales. 2024 · 2024
Closest in time.
Pushing the boundaries of legal information processing with integration of large language models
Chau Nguyen, Thanh Tran, Khang Le, Hien Nguyen, Truong Do, Trang Pham, Son T. Luu, Trung Vo, and Le-Minh Nguyen. 2024 · 2024
Closest in time.
Gemini: A family of highly capable multimodal models
Gemini Team. 2024 · 2024
Closest in time.
Judging the judges: Evaluating alignment and vulnerabilities in llms-as-judges
Aman Singh Thakur, Kartik Choudhary, Venkat Srinik Ramayapally, Sankaran Vaidyanathan, and Dieuwke Hupkes. 2024 · 2024
Closest in time.
Pandalm: An automatic evaluation benchmark for llm instruction tuning optimization
Yidong Wang, Zhuohao Yu, Zhengran Zeng, Linyi Yang, Cunxiang Wang, Hao Chen, Chaoya Jiang, Rui Xie, Jindong Wang, Xing Xie, Wei Ye, Shikun Zhang, and Yue Zhang. 2024 · 2024
Closest in time.
Large language models are active critics in nlg evaluation
Shuying Xu, Junjie Hu, and Ming Jiang. 2024 · 2024
Closest in time.
A bayesian approach towards crowdsourcing the truths from LLMs
Peiran Yao, Jerin George Mathew, Shehraj Singh, Donatella Firmani, and Denilson Barbosa. 2024 · 2024
Closest in time.
Evaluating large language models at evaluating instruction following
Zhiyuan Zeng, Jiatong Yu, Tianyu Gao, Yu Meng, Tanya Goyal, and Danqi Chen. 2024 · 2024
Closest in time.
Auto arena of llms: Automating llm evaluations with agent peer-battles and committee discussions
Ruochen Zhao, Wenxuan Zhang, Yew Ken Chia, Deli Zhao, and Lidong Bing. 2024 · 2024
Closest in time.