Fetching the paper…
Reading the bibliography…
Assessment and evaluation have long been critical challenges in artificial intelligence (AI) and natural language processing (NLP).
Large language models can accurately predict searcher preferences
Paul Thomas, Seth Spielman, Nick Craswell, and Bhaskar Mitra. 2024 · 1940
Earlier work this paper cites.
Are large language models good at utility judgments?
Hengran Zhang, Ruqing Zhang, Jiafeng Guo, Maarten de Rijke, Yixing Fan, and Xueqi Cheng. 2024c · 1951
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
How NOT to evaluate your dialogue system: An empirical study of unsupervised evaluation metrics for dialogue response generation
Chia-Wei Liu, Ryan Lowe, Iulian Serban, Mike Noseworthy, Laurent Charlin, and Joelle Pineau. 2016 · 2016
Earlier work this paper cites.
A call for clarity in reporting BLEU scores
Matt Post. 2018 · 2018
Earlier work this paper cites.
A structured review of the validity of BLEU
Ehud Reiter. 2018 · 2018
Earlier work this paper cites.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, and 12 others. 2020 · 2020
Earlier work this paper cites.
Bertscore: Evaluating text generation with BERT
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020 · 2020
Earlier work this paper cites.
Bartscore: Evaluating generated text as text generation
Weizhe Yuan, Graham Neubig, and Pengfei Liu. 2021 · 2021
Earlier work this paper cites.
Constitutional ai: Harmlessness from ai feedback
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, and 1 others. 2022 · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F. Christiano, Jan Leike, and Ryan Lowe. 2022 · 2022
Earlier work this paper cites.
A survey of evaluation metrics used for nlg systems
Ananya B Sai, Akash Kumar Mohankumar, and Mitesh M Khapra. 2022 · 2022
Earlier work this paper cites.
Bertscore is unfair: On social bias in language model-based metrics for text generation
Tianxiang Sun, Junliang He, Xipeng Qiu, and Xuan-Jing Huang. 2022 · 2022
Earlier work this paper cites.
Supervised graph contrastive learning for few-shot node classification
Zhen Tan, Kaize Ding, Ruocheng Guo, and Huan Liu. 2022 · 2022
Earlier work this paper cites.
Finetuned language models are zero-shot learners
Jason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V. Le. 2022a · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou. 2022b · 2022
Earlier work this paper cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, and 1 others. 2023 · 2023
Earlier work this paper cites.
Benchmarking foundation models with language-model-as-an-examiner
Yushi Bai, Jiahao Ying, Yixin Cao, Xin Lv, Yuze He, Xiaozhi Wang, Jifan Yu, Kaisheng Zeng, Yijia Xiao, Haozhe Lyu, Jiayin Zhang, Juanzi Li, and Lei Hou. 2023a · 2023
Earlier work this paper cites.
Benchmarking foundation models with language-model-as-an-examiner
Yushi Bai, Jiahao Ying, Yixin Cao, Xin Lv, Yuze He, Xiaozhi Wang, Jifan Yu, Kaisheng Zeng, Yijia Xiao, Haozhe Lyu, Jiayin Zhang, Juanzi Li, and Lei Hou. 2023b · 2023
Earlier work this paper cites.
Oceangpt: A large language model for ocean science tasks
Zhen Bi, Ningyu Zhang, Yida Xue, Yixin Ou, Daxiong Ji, Guozhou Zheng, and Huajun Chen. 2023 · 2023
Earlier work this paper cites.
Chateval: Towards better llm-based evaluators through multi-agent debate
Chi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu, Wei Xue, Shanghang Zhang, Jie Fu, and Zhiyuan Liu. 2023 · 2023
Earlier work this paper cites.
Exploring the use of large language models for reference-free text quality evaluation: An empirical study
Yi Chen, Rui Wang, Haiyun Jiang, Shuming Shi, and Ruifeng Xu. 2023 · 2023
Earlier work this paper cites.
Evaluating hallucinations in chinese large language models
Qinyuan Cheng, Tianxiang Sun, Wenwei Zhang, Siyin Wang, Xiangyang Liu, Mozhi Zhang, Junliang He, Mianqiu Huang, Zhangyue Yin, Kai Chen, and 1 others. 2023 · 2023
Earlier work this paper cites.
Selection-inference: Exploiting large language models for interpretable logical reasoning
Antonia Creswell, Murray Shanahan, and Irina Higgins. 2023 · 2023
Earlier work this paper cites.
A survey on in-context learning
Qingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Zhiyong Wu, Baobao Chang, Xu Sun, Jingjing Xu, and Zhifang Sui. 2023 · 2023
Earlier work this paper cites.
Perspectives on large language models for relevance judgment
Guglielmo Faggioli, Laura Dietz, Charles LA Clarke, Gianluca Demartini, Matthias Hagen, Claudia Hauff, Noriko Kando, Evangelos Kanoulas, Martin Potthast, Benno Stein, and 1 others. 2023 · 2023
Earlier work this paper cites.
Lawbench: Benchmarking legal knowledge of large language models
Zhiwei Fei, Xiaoyu Shen, Dawei Zhu, Fengzhe Zhou, Zhuo Han, Songyang Zhang, Kai Chen, Zongwen Shen, and Jidong Ge. 2023 · 2023
Earlier work this paper cites.
Human-like summarization evaluation with chatgpt
Mingqi Gao, Jie Ruan, Renliang Sun, Xunjian Yin, Shiping Yang, and Xiaojun Wan. 2023 · 2023
Earlier work this paper cites.
Reasoning with language model is planning with world model
Shibo Hao, Yi Gu, Haodi Ma, Joshua Hong, Zhen Wang, Daisy Wang, and Zhiting Hu. 2023 · 2023
Earlier work this paper cites.
Allure: auditing and improving llm-based evaluation of text using iterative in-context-learning
Hosein Hasanbeig, Hiteshi Sharma, Leo Betthauser, Felipe Vieira Frujeri, and Ida Momennejad. 2023 · 2023
Earlier work this paper cites.
Socreval: Large language models with the socratic method for reference-free reasoning evaluation
Hangfeng He, Hongming Zhang, and Dan Roth. 2023 · 2023
Earlier work this paper cites.
Large language models can self-improve
Jiaxin Huang, Shixiang Gu, Le Hou, Yuexin Wu, Xuezhi Wang, Hongkun Yu, and Jiawei Han. 2023a · 2023
Earlier work this paper cites.
Llama guard: Llm-based input-output safeguard for human-ai conversations
Hakan Inan, Kartikeya Upasani, Jianfeng Chi, Rashi Rungta, Krithika Iyer, Yuning Mao, Michael Tontchev, Qing Hu, Brian Fuller, Davide Testuggine, and 1 others. 2023 · 2023
Earlier work this paper cites.
Multi-dimensional evaluation of text summarization with in-context learning
Sameer Jain, Vaishakh Keshava, Swarnashree Mysore Sathyendra, Patrick Fernandes, Pengfei Liu, Graham Neubig, and Chunting Zhou. 2023a · 2023
Earlier work this paper cites.
Multi-dimensional evaluation of text summarization with in-context learning
Sameer Jain, Vaishakh Keshava, Swarnashree Mysore Sathyendra, Patrick Fernandes, Pengfei Liu, Graham Neubig, and Chunting Zhou. 2023b · 2023
Earlier work this paper cites.
Towards mitigating llm hallucination via self reflection
Ziwei Ji, Tiezheng Yu, Yan Xu, Nayeon Lee, Etsuko Ishii, and Pascale Fung. 2023 · 2023
Earlier work this paper cites.
Large language models are state-of-the-art evaluators of translation quality
Tom Kocmi and Christian Federmann. 2023 · 2023
Earlier work this paper cites.
Benchmarking cognitive biases in large language models as evaluators
Ryan Koo, Minhwa Lee, Vipul Raheja, Jong Inn Park, Zae Myung Kim, and Dongyeop Kang. 2023 · 2023
Earlier work this paper cites.
Little giants: Exploring the potential of small LLMs as evaluation metrics in summarization in the Eval4NLP 2023 shared task
Neema Kotonya, Saran Krishnasamy, Joel Tetreault, and Alejandro Jaimes. 2023 · 2023
Earlier work this paper cites.
Improving diversity of demographic representation in large language models via collective-critiques and self-voting
Preethi Lahoti, Nicholas Blumm, Xiao Ma, Raghavendra Kotikalapudi, Sahitya Potluri, Qijun Tan, Hansa Srinivasan, Ben Packer, Ahmad Beirami, Alex Beutel, and Jilin Chen. 2023 · 2023
Earlier work this paper cites.
Rlaif: Scaling reinforcement learning from human feedback with ai feedback
Harrison Lee, Samrat Phatale, Hassan Mansoor, Thomas Mesnard, Johan Ferret, Kellie Lu, Colton Bishop, Ethan Hall, Victor Carbune, Abhinav Rastogi, and 1 others. 2023 · 2023
Earlier work this paper cites.
MoT: Memory-of-thought enables ChatGPT to self-improve
Xiaonan Li and Xipeng Qiu. 2023a · 2023
Earlier work this paper cites.
Mot: Memory-of-thought enables chatgpt to self-improve
Xiaonan Li and Xipeng Qiu. 2023b · 2023
Earlier work this paper cites.
Encouraging divergent thinking in large language models through multi-agent debate
Tian Liang, Zhiwei He, Wenxiang Jiao, Xing Wang, Yan Wang, Rui Wang, Yujiu Yang, Zhaopeng Tu, and Shuming Shi. 2023 · 2023
Earlier work this paper cites.
Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. 2023 · 2023
Earlier work this paper cites.
The unlocking spell on base llms: Rethinking alignment via in-context learning
Bill Yuchen Lin, Abhilasha Ravichander, Ximing Lu, Nouha Dziri, Melanie Sclar, Khyathi Chandu, Chandra Bhagavatula, and Yejin Choi. 2023 · 2023
Earlier work this paper cites.
LLM-eval: Unified multi-dimensional automatic evaluation for open-domain conversations with large language models
Yen-Ting Lin and Yun-Nung Chen. 2023a · 2023
Earlier work this paper cites.
LLM-eval: Unified multi-dimensional automatic evaluation for open-domain conversations with large language models
Yen-Ting Lin and Yun-Nung Chen. 2023b · 2023
Earlier work this paper cites.
G-eval: NLG evaluation using gpt-4 with better human alignment
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023b · 2023
Earlier work this paper cites.
Zero-shot nlg evaluation through pairware comparisons with llms
Adian Liusie, Potsawee Manakul, and Mark JF Gales. 2023 · 2023
Earlier work this paper cites.
Zero-shot listwise document reranking with a large language model
Xueguang Ma, Xinyu Zhang, Ronak Pradeep, and Jimmy Lin. 2023 · 2023
Earlier work this paper cites.
FActScore: Fine-grained atomic evaluation of factual precision in long form text generation
Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen-tau Yih, Pang Koh, Mohit Iyyer, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2023 · 2023
Earlier work this paper cites.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning, Stefano Ermon, and Chelsea Finn. 2023 · 2023
Earlier work this paper cites.
Evaluating the moral beliefs encoded in llms
Nino Scherrer, Claudia Shi, Amir Feder, and David M. Blei. 2023 · 2023
Earlier work this paper cites.
Languagempc: Large language models as decision makers for autonomous driving
Hao Sha, Yao Mu, Yuxuan Jiang, Li Chen, Chenfeng Xu, Ping Luo, Shengbo Eben Li, Masayoshi Tomizuka, Wei Zhan, and Mingyu Ding. 2023 · 2023
Earlier work this paper cites.
Is ChatGPT good at search? investigating large language models as re-ranking agents
Weiwei Sun, Lingyong Yan, Xinyu Ma, Shuaiqiang Wang, Pengjie Ren, Zhumin Chen, Dawei Yin, and Zhaochun Ren. 2023 · 2023
Earlier work this paper cites.
Large language models can accurately predict searcher preferences, 2023
Paul Thomas, Seth Spielman, Nick Craswell, and Bhaskar Mitra. 2023 · 2023
Earlier work this paper cites.
Can ChatGPT defend its belief in truth? evaluating LLM reasoning via debate
Boshi Wang, Xiang Yue, and Huan Sun. 2023a · 2023
Earlier work this paper cites.
Large language models are diverse role-players for summarization evaluation
Ning Wu, Ming Gong, Linjun Shou, Shining Liang, and Daxin Jiang. 2023 · 2023
Earlier work this paper cites.
The rise and potential of large language model based agents: A survey
Zhiheng Xi, Wenxiang Chen, Xin Guo, Wei He, Yiwen Ding, Boyang Hong, Ming Zhang, Junzhe Wang, Senjie Jin, Enyu Zhou, and 1 others. 2023 · 2023
Earlier work this paper cites.
Data selection for language models via importance resampling
Sang Michael Xie, Shibani Santurkar, Tengyu Ma, and Percy Liang. 2023 · 2023
Earlier work this paper cites.
INSTRUCTSCORE: Towards explainable text generation evaluation with automatic feedback
Wenda Xu, Danqing Wang, Liangming Pan, Zhenqiao Song, Markus Freitag, William Wang, and Lei Li. 2023a · 2023
Earlier work this paper cites.
Auto-gpt for online decision making: Benchmarks and additional opinions
Hui Yang, Sifu Yue, and Yunzhong He. 2023 · 2023
Earlier work this paper cites.
Tree of thoughts: Deliberate problem solving with large language models
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan. 2023a · 2023
Earlier work this paper cites.
React: Synergizing reasoning and acting in language models
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R. Narasimhan, and Yuan Cao. 2023b · 2023
Earlier work this paper cites.
Automatic evaluation of attribution by large language models
Xiang Yue, Boshi Wang, Ziru Chen, Kai Zhang, Yu Su, and Huan Sun. 2023 · 2023
Earlier work this paper cites.
Wider and deeper llm networks are fairer llm evaluators
Xinghua Zhang, Bowen Yu, Haiyang Yu, Yangyu Lv, Tingwen Liu, Fei Huang, Hongbo Xu, and Yongbin Li. 2023 · 2023
Earlier work this paper cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023 · 2023
Earlier work this paper cites.
Sotopia: Interactive evaluation for social intelligence in language agents
Xuhui Zhou, Hao Zhu, Leena Mathur, Ruohong Zhang, Haofei Yu, Zhengyang Qi, Louis-Philippe Morency, Yonatan Bisk, Daniel Fried, Graham Neubig, and 1 others. 2023 · 2023
Earlier work this paper cites.
Judgelm: Fine-tuned large language models are scalable judges
Lianghui Zhu, Xinggang Wang, and Xinlong Wang. 2023 · 2023
Earlier work this paper cites.
Can we use large language models to fill relevance judgment holes?
Zahra Abbasiantaeb, Chuan Meng, Leif Azzopardi, and Mohammad Aliannejadi. 2024 · 2024
Earlier work this paper cites.
Many-shot in-context learning
Rishabh Agarwal, Avi Singh, Lei M Zhang, Bernd Bohnet, Luis Rosias, Stephanie CY Chan, Biao Zhang, Aleksandra Faust, and Hugo Larochelle · 2024
Earlier work this paper cites.
i-srt: Aligning large multimodal models for videos by iterative self-retrospective judgment
Daechul Ahn, Yura Choi, San Kim, Youngjae Yu, Dongyeop Kang, and Jonghyun Choi. 2024 · 2024
Earlier work this paper cites.
Generative information retrieval evaluation
Marwah Alaofi, Negar Arabzadeh, Charles LA Clarke, and Mark Sanderson. 2024 · 2024
Earlier work this paper cites.
A survey on data selection for language models
Alon Albalak, Yanai Elazar, Sang Michael Xie, Shayne Longpre, Nathan Lambert, Xinyi Wang, Niklas Muennighoff, Bairu Hou, Liangming Pan, Haewon Jeong, and 1 others. 2024 · 2024
Earlier work this paper cites.
Critique-out-loud reward models
Zachary Ankner, Mansheej Paul, Brandon Cui, Jonathan D Chang, and Prithviraj Ammanabrolu. 2024 · 2024
Earlier work this paper cites.
Samee Arif, Sualeha Farid, Abdul Hameed Azeemi, Awais Athar, and Agha Ali Raza. 2024 · 2024
Earlier work this paper cites.
Self-RAG: Learning to retrieve, generate, and critique through self-reflection
Akari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil, and Hannaneh Hajishirzi. 2024 · 2024
Earlier work this paper cites.
Reference-guided verdict: Llms-as-judges in automatic evaluation of free-form text
Sher Badshah and Hassan Sajjad. 2024 · 2024
Earlier work this paper cites.
Longbench v2: Towards deeper understanding and reasoning on realistic long-context multitasks
Yushi Bai, Shangqing Tu, Jiajie Zhang, Hao Peng, Xiaozhi Wang, Xin Lv, Shulin Cao, Jiazheng Xu, Lei Hou, Yuxiao Dong, and 1 others. 2024 · 2024
Earlier work this paper cites.
Adversarial multi-agent evaluation of large language models through iterative debates
Chaithanya Bandi and Abir Harrasse. 2024 · 2024
Earlier work this paper cites.
The vulnerability of language model benchmarks: Do they accurately reflect true llm performance?
Sourav Banerjee, Ayushi Agarwal, and Eishkaran Singh. 2024 · 2024
Earlier work this paper cites.
Lrq-fact: Llm-generated relevant questions for multimodal fact-checking
Alimohammad Beigi, Bohan Jiang, Dawei Li, Tharindu Kumarage, Zhen Tan, Pouya Shaeri, and Huan Liu. 2024 · 2024
Cited alongside, same era.
Graph of thoughts: Solving elaborate problems with large language models
Maciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger, Michal Podstawski, Lukas Gianinazzi, Joanna Gajda, Tomasz Lehmann, Hubert Niewiadomski, Piotr Nyczyk, and Torsten Hoefler. 2024 · 2024
Cited alongside, same era.
Comparing two model designs for clinical note generation; is an llm a useful evaluator of consistency?
Nathan Brake and Thomas Schaaf. 2024 · 2024
Cited alongside, same era.
A survey on evaluation of large language models
Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, and 1 others. 2024 · 2024
Cited alongside, same era.
Conqret: Benchmarking fine-grained evaluation of retrieval augmented argumentation with llm judges
Kaustubh D. Dhole, Kai Shu, and Eugene Agichtein. 2024 · 2024
Alma: Alignment with minimal annotation
Michihiro Yasunaga, Leonid Shamis, Chunting Zhou, Andrew Cohen, Jason Weston, Luke Zettlemoyer, and Marjan Ghazvininejad. 2024 · 2024
Closest in time.
FLASK: Fine-grained language model evaluation based on alignment skill sets
Seonghyeon Ye, Doyoung Kim, Sungdong Kim, Hyeonbin Hwang, Seungone Kim, Yongrae Jo, James Thorne, Juho Kim, and Minjoon Seo. 2024b · 2024
Closest in time.
Seungjun Yi, Jaeyoung Lim, and Juyong Yoon. 2024 · 2024
Closest in time.
Advancing llm reasoning generalists with preference trees
Lifan Yuan, Ganqu Cui, Hanbin Wang, Ning Ding, Xingyao Wang, Jia Deng, Boji Shan, Huimin Chen, Ruobing Xie, Yankai Lin, and 1 others · 2024
Closest in time.
Learning reward for robot skills using large language models via self-alignment
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Can llm be a personalized judge?
Yijiang River Dong, Tiancheng Hu, and Nigel Collier. 2024 · 2024
Cited alongside, same era.
Length-controlled alpacaeval: A simple way to debias automatic evaluators
Yann Dubois, Balázs Galambosi, Percy Liang, and Tatsunori B Hashimoto. 2024 · 2024
Cited alongside, same era.
Aparna Elangovan, Lei Xu, Jongwoo Ko, Mahsa Elyasi, Ling Liu, Sravan Bodapati, and Dan Roth. 2024 · 2024
Cited alongside, same era.
Test of time: A benchmark for evaluating llms on temporal reasoning
Bahare Fatemi, Mehran Kazemi, Anton Tsitsulin, Karishma Malkan, Jinyeong Yim, John Palowitch, Sungyong Seo, Jonathan Halcrow, and Bryan Perozzi. 2024 · 2024
Cited alongside, same era.
M-mad: Multidimensional multi-agent debate for advanced machine translation evaluation
Zhaopeng Feng, Jiayuan Su, Jiamei Zheng, Jiahan Ren, Yan Zhang, Jian Wu, Hongwei Wang, and Zuozhu Liu. 2024 · 2024
Cited alongside, same era.
Gptscore: Evaluate as you desire
Jinlan Fu, See Kiong Ng, Zhengbao Jiang, and Pengfei Liu. 2024 · 2024
Cited alongside, same era.
Llm-based nlg evaluation: Current status and challenges
Mingqi Gao, Xinyu Hu, Jie Ruan, Xiao Pu, and Xiaojun Wan. 2024 · 2024
Cited alongside, same era.
Yuwei Zeng, Yao Mu, and Lin Shao. 2024 · 2024
Closest in time.
Online self-preferring language models
Yuanzhao Zhai, Zhuo Zhang, Kele Xu, Hanyang Peng, Yue Yu, Dawei Feng, Cheng Yang, Bo Ding, and Huaimin Wang. 2024 · 2024
Closest in time.
A comprehensive analysis of the effectiveness of large language models as automatic dialogue evaluators
Chen Zhang, Luis Fernando D’Haro, Yiming Chen, Malu Zhang, and Haizhou Li. 2024a · 2024
Closest in time.
An llm feature-based framework for dialogue constructiveness assessment
Lexin Zhou, Youmna Farag, and Andreas Vlachos. 2024c · 2024
Closest in time.
Is llm a reliable reviewer? a comprehensive evaluation of llm on automatic paper reviewing tasks
Ruiyang Zhou, Lu Chen, and Kai Yu. 2024e · 2024
Closest in time.
Beyond yes and no: Improving zero-shot LLM rankers via scoring fine-grained relevance labels
Honglei Zhuang, Zhen Qin, Kai Hui, Junru Wu, Le Yan, Xuanhui Wang, and Michael Bendersky. 2024a · 2024
Closest in time.
Agent-as-a-judge: Evaluate agents with agents
Mingchen Zhuge, Changsheng Zhao, Dylan Ashley, Wenyi Wang, Dmitrii Khizbullin, Yunyang Xiong, Zechun Liu, Ernie Chang, Raghuraman Krishnamoorthi, Yuandong Tian, and 1 others. 2024 · 2024
Closest in time.
Ice-score: Instructing large language models to evaluate code
Terry Yue Zhuo. 2024 · 2024
Closest in time.
Argument summarization and its evaluation in the era of large language models
Moritz Altemeyer, Steffen Eger, Johannes Daxenberger, Tim Altendorf, Philipp Cimiano, and Benjamin Schiller. 2025 · 2025
Closest in time.
Defense against the dark prompts: Mitigating best-of-n jailbreaking with prompt evaluation
Stuart Armstrong, Matija Franklin, Connor Stevens, and Rebecca Gorman. 2025 · 2025
Closest in time.
Dafe: Llm-based evaluation through dynamic arbitration for free-form question-answering
Sher Badshah and Hassan Sajjad. 2025 · 2025
Closest in time.
Krisztian Balog, Donald Metzler, and Zhen Qin. 2025 · 2025
Closest in time.
Jeremy Barnes, Naiara Perez, Alba Bonet-Jover, and Begoña Altuna. 2025 · 2025
Closest in time.
Nitay Calderon, Roi Reichart, and Rotem Dror. 2025 · 2025
Closest in time.
Riccardo Cantini, Alessio Orsino, Massimo Ruggiero, and Domenico Talia. 2025 · 2025
Closest in time.
Exploring the multilingual nlg evaluation abilities of llm-based evaluators
Jiayi Chang, Mingqi Gao, Xinyu Hu, and Xiaojun Wan. 2025 · 2025
Closest in time.
Safer or luckier? llms as safety evaluators are not robust to artifacts
Hongyu Chen and Seraphina Goldfarb-Tarrant. 2025 · 2025
Closest in time.
Copilot arena: A platform for code llm evaluation in the wild
Wayne Chi, Valerie Chen, Anastasios Nikolas Angelopoulos, Wei-Lin Chiang, Aditya Mittal, Naman Jain, Tianjun Zhang, Ion Stoica, Chris Donahue, and Ameet Talwalkar. 2025 · 2025
Closest in time.
Tract: Regression-aware fine-tuning meets chain-of-thought reasoning for llm-as-a-judge
Cheng-Han Chiang, Hung-yi Lee, and Michal Lukasik. 2025 · 2025
Closest in time.
Marianne Chuang, Gabriel Chuang, Cheryl Chuang, and John Chuang. 2025 · 2025
Closest in time.
Dynamic-kgqa: A scalable framework for generating adaptive question answering datasets
Preetam Prabhu Srikar Dammu, Himanshu Naidu, and Chirag Shah. 2025 · 2025
Closest in time.
To judge or not to judge: Using llm judgements for advertiser keyphrase relevance at ebay
Soumik Dey, Hansi Wu, and Binbin Li. 2025 · 2025
Closest in time.
Llm-evaluation tropes: Perspectives on the validity of llm-evaluations
Laura Dietz, Oleg Zendel, Peter Bailey, Charles Clarke, Ellese Cotterill, Jeff Dalton, Faegheh Hasibi, Mark Sanderson, and Nick Craswell. 2025 · 2025
Closest in time.
Tobias Domhan and Dawei Zhu. 2025 · 2025
Closest in time.
Know thy judge: On the robustness meta-evaluation of llm safety judges
Francisco Eiras, Eliott Zemour, Eric Lin, and Vaikkunth Mugunthan. 2025 · 2025
Closest in time.
Sedareval: Automated evaluation using self-adaptive rubrics
Zhiyuan Fan, Weinong Wang, Xing Wu, and Debing Zhang. 2025 · 2025
Closest in time.
Grapheval: A lightweight graph-based llm framework for idea evaluation
Tao Feng, Yihang Sun, and Jiaxuan You. 2025 · 2025
Closest in time.
How reliable is multilingual llm-as-a-judge?
Xiyan Fu and Wei Liu. 2025 · 2025
Closest in time.
Great models think alike and this undermines ai oversight
Shashwat Goel, Joschka Struber, Ilze Amanda Auzina, Karuna K Chandra, Ponnurangam Kumaraguru, Douwe Kiela, Ameya Prabhu, Matthias Bethge, and Jonas Geiping. 2025 · 2025
Closest in time.
Validating llm-as-a-judge systems in the absence of gold labels
Luke Guerdan, Solon Barocas, Kenneth Holstein, Hanna Wallach, Zhiwei Steven Wu, and Alexandra Chouldechova. 2025 · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, and 1 others. 2025 · 2025
Closest in time.
From code to courtroom: Llms as the new software judges
Junda He, Jieke Shi, Terry Yue Zhuo, Christoph Treude, Jiamou Sun, Zhenchang Xing, Xiaoning Du, and David Lo. 2025 · 2025
Closest in time.
Amey Hengle, Aswini Kumar, Anil Bandhakavi, and Tanmoy Chakraborty. 2025 · 2025
Closest in time.
Llm-as-a-judge: Reassessing the performance of llms in extractive qa
Xanh Ho, Jiahao Huang, Florian Boudin, and Akiko Aizawa. 2025 · 2025
Closest in time.
Agent-as-judge for factual summarization of long narratives
Yeonseok Jeong, Minsoo Kim, Seung-won Hwang, and Byung-Hak Kim. 2025 · 2025
Closest in time.
Safechain: Safety of language models with long chain-of-thought reasoning capabilities
Fengqing Jiang, Zhangchen Xu, Yuetai Li, Luyao Niu, Zhen Xiang, Bo Li, Bill Yuchen Lin, and Radha Poovendran. 2025 · 2025
Closest in time.
Verdict: A library for scaling judge-time compute
Nimit Kalra and Leonard Tang. 2025 · 2025
Closest in time.
Openlgauge: An explainable metric for nlg evaluation with open-weights llms
Ivan Kartáč, Mateusz Lango, and Ondřej Dušek. 2025 · 2025
Closest in time.
Boshra Khalili and Andrew W Smyth. 2025 · 2025
Closest in time.
Heegyu Kim, Taeyang Jeon, Seungtaek Choi, Ji Hoon Hong, Dong Won Jeon, Ga-Yeon Baek, Gyeong-Won Kwak, Dong-Hee Lee, Jisu Bae, Chihoon Lee, and 1 others. 2025 · 2025
Closest in time.
Revieweval: An evaluation framework for ai-generated reviews
Chhavi Kirtani, Madhav Krishan Garg, Tejash Prasad, Tanmay Singhal, Murari Mandal, and Dhruv Kumar. 2025 · 2025
Closest in time.
No free labels: Limitations of llm-as-a-judge without human grounding
Michael Krumdick, Charles Lovering, Varshini Reddy, Seth Ebner, and Chris Tanner. 2025 · 2025
Closest in time.
Evaluating text-to-visual generation with image-to-text generation
Zhiqiu Lin, Deepak Pathak, Baiqi Li, Jiayao Li, Xide Xia, Graham Neubig, Pengchuan Zhang, and Deva Ramanan. 2025 · 2025
Closest in time.
Decoding ai judgment: How llms assess news credibility and bias
Edoardo Loru, Jacopo Nudo, Niccolò Di Marco, Matteo Cinelli, and Walter Quattrociocchi. 2025 · 2025
Closest in time.
Agentrewardbench: Evaluating automatic evaluations of web agent trajectories
Xing Han Lù, Amirhossein Kazemnejad, Nicholas Meade, Arkil Patel, Dongchan Shin, Alejandra Zambrano, Karolina Stańczak, Peter Shaw, Christopher J Pal, and Siva Reddy. 2025 · 2025
Closest in time.
Wenhan Mu, Ling Xu, Shuren Pei, Le Mi, and Huichi Zhou. 2025 · 2025
Closest in time.
Synthetic data can mislead evaluations: Membership inference as machine text detection
Ali Naseh and Niloofar Mireshghallah. 2025 · 2025
Closest in time.
How to get your llm to generate challenging problems for evaluation
Arkil Patel, Siva Reddy, and Dzmitry Bahdanau. 2025 · 2025
Closest in time.
An llm-as-a-judge approach for scalable gender-neutral translation evaluation
Andrea Piergentili, Beatrice Savoldi, Matteo Negri, and Luisa Bentivogli. 2025 · 2025
Closest in time.
The great nugget recall: Automating fact extraction and rag evaluation with large language models
Ronak Pradeep, Nandan Thakur, Shivani Upadhyay, Daniel Campos, Nick Craswell, and Jimmy Lin. 2025 · 2025
Closest in time.
Learning to generate unit tests for automated debugging
Archiki Prasad, Elias Stengel-Eskin, Justin Chih-Yao Chen, Zaid Khan, and Mohit Bansal. 2025 · 2025
Closest in time.
Judge anything: Mllm as a judge across any modality
Shu Pu, Yaochen Wang, Dongping Chen, Yuhang Chen, Guohao Wang, Qi Qin, Zhongyi Zhang, Zhiyuan Zhang, Zetong Zhou, Shuang Gong, and 1 others. 2025 · 2025
Closest in time.
Evaluating llms’ assessment of mixed-context hallucination through the lens of summarization
Siya Qi, Rui Cao, Yulan He, and Zheng Yuan. 2025 · 2025
Closest in time.
Efficient map estimation of llm judgment performance with prior transfer
Huaizhi Qu, Inyoung Choi, Zhen Tan, Song Wang, Sukwon Yun, Qi Long, Faizan Siddiqui, Kwonjoon Lee, and Tianlong Chen. 2025 · 2025
Closest in time.
Melissa Kazemi Rad, Huy Nghiem, Andy Luo, Sahil Wadhwa, Mohammad Sorower, and Stephen Rawls. 2025 · 2025
Closest in time.
Towards safer chatbots: A framework for policy compliance evaluation of custom gpts
David Rodriguez, William Seymour, Jose M Del Alamo, and Jose Such. 2025 · 2025
Closest in time.
Learning to plan & reason for evaluation with thinking-llm-as-a-judge
Swarnadeep Saha, Xian Li, Marjan Ghazvininejad, Jason Weston, and Tianlu Wang. 2025 · 2025
Closest in time.
Tuning llm judge design decisions for 1/1000 of the cost
David Salinas, Omar Swelam, and Frank Hutter. 2025 · 2025
Closest in time.
Piotr Sawicki, Marek Grześ, Dan Brown, and Fabrício Góes. 2025 · 2025
Closest in time.
Validating llm-generated relevance labels for educational resource search
Ratan J Sebastian and Anett Hoppe. 2025 · 2025
Closest in time.
Kwangwook Seo, Donguk Kwon, and Dongha Lee. 2025 · 2025
Closest in time.
Heimdall: test-time scaling on the generative verification
Wenlei Shi and Xing Jin. 2025 · 2025
Closest in time.
Lces: Zero-shot automated essay scoring via pairwise comparisons using large language models
Takumi Shibata and Yuichi Miyamura. 2025 · 2025
Closest in time.
Grp: Goal-reversed prompting for zero-shot evaluation with llms
Mingyang Song, Mao Zheng, and Xuan Luo. 2025 · 2025
Closest in time.
Stop overthinking: A survey on efficient reasoning for large language models
Yang Sui, Yu-Neng Chuang, Guanchu Wang, Jiamu Zhang, Tianyi Zhang, Jiayi Yuan, Hongyi Liu, Andrew Wen, Shaochen Zhong, Hanjie Chen, and 1 others. 2025 · 2025
Closest in time.
Badjudge: Backdoor vulnerabilities of llm-as-a-judge
Terry Tong, Fei Wang, Zhe Zhao, and Muhao Chen. 2025 · 2025
Closest in time.
Pairwise or pointwise? evaluating feedback protocols for bias in llm-based evaluation
Tuhina Tripathi, Manya Wadhwa, Greg Durrett, and Scott Niekum. 2025 · 2025
Closest in time.
Aligning black-box language models with human judgments
Gerrit JJ van den Burg, Gen Suzuki, Wei Liu, and Murat Sensoy. 2025 · 2025
Closest in time.
Rocketeval: Efficient automated llm evaluation via grading checklist
Tianjun Wei, Wei Wen, Ruizhi Qiao, Xing Sun, and Jianghong Ma. 2025 · 2025
Closest in time.
Hpss: Heuristic prompting strategy search for llm evaluators
Bosi Wen, Pei Ke, Yufei Sun, Cunxiang Wang, Xiaotao Gu, Jinfeng Zhou, Jie Tang, Hongning Wang, and Minlie Huang. 2025 · 2025
Closest in time.
J1: Incentivizing thinking in llm-as-a-judge via reinforcement learning
Chenxi Whitehouse, Tianlu Wang, Ping Yu, Xian Li, Jason Weston, Ilia Kulikov, and Swarnadeep Saha. 2025 · 2025
Closest in time.
Longeval: A comprehensive analysis of long-text generation through a plan-based paradigm
Siwei Wu, Yizhi Li, Xingwei Qu, Rishi Ravikumar, Yucheng Li, Tyler Loakman, Shanghaoran Quan, Xiaoyong Wei, Riza Batista-Navarro, and Chenghua Lin. 2025 · 2025
Closest in time.
Uncertainty-aware step-wise verification with generative reward models
Zihuiwen Ye, Luckeciano Carvalho Melo, Younesse Kaddar, Phil Blunsom, Sam Staton, and Yarin Gal. 2025 · 2025
Closest in time.
Improve llm-as-a-judge ability as a general ability
Jiachen Yu, Shaoning Sun, Xiaohui Hu, Jiaxu Yan, Kaidong Yu, and Xuelong Li. 2025 · 2025
Closest in time.
Hierarchical divide-and-conquer for fine-grained alignment in llm-based medical evaluation
Shunfan Zheng, Xiechi Zhang, Gerard de Melo, Xiaoling Wang, and Linlin Wang. 2025 · 2025
Closest in time.
Yilun Zhou, Austin Xu, Peifeng Wang, Caiming Xiong, and Shafiq Joty. 2025 · 2025
Closest in time.
Deepreview: Improving llm-based paper review with human-like deep thinking process
Minjun Zhu, Yixuan Weng, Linyi Yang, and Yue Zhang. 2025 · 2025
Closest in time.
Trueteacher: Learning factual consistency evaluation with large language models
Zorik Gekhman, Jonathan Herzig, Roee Aharoni, Chen Elkind, and Idan Szpektor. 2023 · 2070
Closest in time.