Fetching the paper…
Reading the bibliography…
Assessing the quality of outputs generated by generative models, such as large language models and vision language models, presents notable challenges.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie · 2005
Earlier work this paper cites.
Usr: An unsupervised and reference free evaluation metric for dialog generation, 2020
Shikib Mehri and Maxine Eskenazi · 2005
Earlier work this paper cites.
Teaching machines to read and comprehend, 2015
Karl Moritz Hermann, Tomáš Kočiský, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom · 2015
Earlier work this paper cites.
Program induction by rationale generation: Learning to solve and explain algebraic word problems
Wang Ling, Dani Yogatama, Chris Dyer, and Phil Blunsom · 2017
Earlier work this paper cites.
Blind image quality assessment using a deep bilinear convolutional neural network
Weixia Zhang, Kede Ma, Jia Yan, Dexiang Deng, and Zhou Wang · 2018
Earlier work this paper cites.
Topical-Chat: Towards Knowledge-Grounded Open-Domain Conversations
Karthik Gopalakrishnan, Behnam Hedayatnia, Qinlang Chen, Anna Gottardi, Sanjeev Kwatra, Anu Venkatesh, Raefer Gabriel, and Dilek Hakkani-Tür · 2019
Earlier work this paper cites.
Quality estimation for image captions based on large-scale human evaluations
T. Levinboim, A. Thapliyal, P. Sharma, and R. Soricut · 2019
Earlier work this paper cites.
Blindly assess image quality in the wild guided by a self-adaptive hyper network
Shaolin Su, Qingsen Yan, Yu Zhu, Cheng Zhang, Xin Ge, Jinqiu Sun, and Yanning Zhang · 2020
Earlier work this paper cites.
Explanations for CommonsenseQA: New Dataset and Models
Shourya Aggarwal, Divyanshu Mandowara, Vishwajeet Agrawal, Dinesh Khandelwal, Parag Singla, and Dinesh Garg · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems, 2021
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman · 2021
Earlier work this paper cites.
Summeval: Re-evaluating summarization evaluation, 2021
Alexander R. Fabbri, Wojciech Kryściński, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir Radev · 2021
Earlier work this paper cites.
Clipscore: A reference-free evaluation metric for image captioning, 2022
Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi · 2022
Earlier work this paper cites.
Towards a unified multi-dimensional evaluator for text generation, 2022
Ming Zhong, Yang Liu, Da Yin, Yuning Mao, Yizhu Jiao, Pengfei Liu, Chenguang Zhu, Heng Ji, and Jiawei Han · 2022
Cited alongside, same era.
Chateval: Towards better llm-based evaluators through multi-agent debate, 2023
Chi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu, Wei Xue, Shanghang Zhang, Jie Fu, and Zhiyuan Liu · 2023
Cited alongside, same era.
Exploring the use of large language models for reference-free text quality evaluation: An empirical study
Yi Chen, Rui Wang, Haiyun Jiang, Shuming Shi, and Ruifeng Xu · 2023
Cited alongside, same era.
A closer look into automatic evaluation using large language models, 2023
Cheng-Han Chiang and Hung-yi Lee · 2023
Cited alongside, same era.
Improving factuality and reasoning in language models through multiagent debate, 2023
Yilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum, and Igor Mordatch · 2023
Large language models are human-level prompt engineers, 2023
Yongchao Zhou, Andrei Ioan Muresanu, Ziwen Han, Keiran Paster, Silviu Pitis, Harris Chan, and Jimmy Ba · 2023
Later among the works it cites.
Reconcile: Round-table conference improves reasoning via consensus among diverse llms, 2024
Justin Chih-Yao Chen, Swarnadeep Saha, and Mohit Bansal · 2024
Closest in time.
Gemini 1.5-flash
Google DeepMind · 2024
Closest in time.
Gemma-2-9b
Google-Research · 2024
Closest in time.
Large language models cannot self-correct reasoning yet, 2024
Jie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng, Adams Wei Yu, Xinying Song, and Denny Zhou · 2024
Closest in time.
Open llm leaderboard
HuggingFace · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Gptscore: Evaluate as you desire, 2023
Jinlan Fu, See-Kiong Ng, Zhengbao Jiang, and Pengfei Liu · 2023
Cited alongside, same era.
Tigerscore: Towards building explainable metric for all text generation tasks, 2023
Dongfu Jiang, Yishan Li, Ge Zhang, Wenhao Huang, Bill Yuchen Lin, and Wenhu Chen · 2023
Cited alongside, same era.
Pick-a-pic: An open dataset of user preferences for text-to-image generation, 2023
Yuval Kirstain, Adam Polyak, Uriel Singer, Shahbuland Matiana, Joe Penna, and Omer Levy · 2023
Cited alongside, same era.
Large language models are zero-shot reasoners, 2023
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa · 2023
Cited alongside, same era.
Let models speak ciphers: Multiagent debate through embeddings, 2023
Chau Pham, Boyi Liu, Yingxiang Yang, Zhengyu Chen, Tianyi Liu, Jianbo Yuan, Bryan A. Plummer, Zhaoran Wang, and Hongxia Yang · 2023
Cited alongside, same era.
Automatic prompt optimization with “gradient descent” and beam search
Reid Pryzant, Dan Iter, Jerry Li, Yin Lee, Chenguang Zhu, and Michael Zeng · 2023
Cited alongside, same era.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models, 2023
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R. Brown, and et el · 2023
Cited alongside, same era.
Closest in time.
Llama-3.1-8b-instruct
Meta-AI · 2024
Closest in time.
Mistral-nemo-instruct-2407
Mistral-AI · 2024
Closest in time.
Bringing textual prompt to ai-generated image quality assessment, 2024
Bowen Qu, Haohui Li, and Wei Gao · 2024
Closest in time.
Fusion-eval: Integrating evaluators with llms, 2024
Lei Shu, Nevan Wichers, Liangchen Luo, Yun Zhu, Yinxiao Liu, Jindong Chen, and Lei Meng · 2024
Closest in time.
Unleashing the emergent cognitive synergy in large language models: A task-solving agent through multi-persona self-collaboration, 2024
Zhenhailong Wang, Shaoguang Mao, Wenshan Wu, Tao Ge, Furu Wei, and Heng Ji · 2024
Closest in time.
Large language models as optimizers, 2024
Chengrun Yang, Xuezhi Wang, Yifeng Lu, Hanxiao Liu, Quoc V. Le, Denny Zhou, and Xinyun Chen · 2024
Closest in time.
Quality assessment in the era of large models: A survey, 2024
Zicheng Zhang, Yingjie Zhou, Chunyi Li, Baixuan Zhao, Xiaohong Liu, and Guangtao Zhai · 2024
Closest in time.