Fetching the paper…
Reading the bibliography…
The large language model (LLM)-as-judge paradigm has been used to meet the demand for a cheap, reliable, and fast evaluation of model outputs during AI system development and post-deployment monitoring.
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020 · 1904
Earlier work this paper cites.
Moverscore: Text generation evaluating with contextualized embeddings and earth mover distance
Wei Zhao, Maxime Peyrard, Fei Liu, Yang Gao, Christian M Meyer, and Steffen Eger. 2019 · 1909
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie. 2005 · 2005
Earlier work this paper cites.
Computing krippendorff’s alpha-reliability
Klaus Krippendorff. 2011 · 2011
Earlier work this paper cites.
MRQA 2019 shared task: Evaluating generalization in reading comprehension
Adam Fisch, Alon Talmor, Robin Jia, Minjoon Seo, Eunsol Choi, and Danqi Chen. 2019 · 2019
Earlier work this paper cites.
Evaluating the factual consistency of abstractive text summarization
Wojciech Kryscinski, Bryan McCann, Caiming Xiong, and Richard Socher. 2020 · 2020
Earlier work this paper cites.
CLIFF: Contrastive learning for improving faithfulness and factuality in abstractive summarization
Shuyang Cao and Lu Wang. 2021 · 2021
Earlier work this paper cites.
Annotating and modeling fine-grained factuality in summarization
Tanya Goyal and Greg Durrett. 2021 · 2021
Earlier work this paper cites.
Bartscore: Evaluating generated text as text generation
Weizhe Yuan, Graham Neubig, and Pengfei Liu. 2021 · 2021
Earlier work this paper cites.
SummaC: Re-visiting NLI-based models for inconsistency detection in summarization
Philippe Laban, Tobias Schnabel, Paul N. Bennett, and Marti A. Hearst. 2022 · 2022
Earlier work this paper cites.
Self-consistency improves chain of thought reasoning in language models
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2022 · 2022
Earlier work this paper cites.
Ragas: Automated evaluation of retrieval augmented generation
Shahul Es, Jithin James, Luis Espinosa-Anke, and Steven Schockaert. 2023 · 2023
Earlier work this paper cites.
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. 2023 · 2023
Earlier work this paper cites.
Pressure testing gpt-4-128k with long context recall
Greg Kamradt. 2023 · 2023
Earlier work this paper cites.
Prometheus: Inducing fine-grained evaluation capability in language models
Seungone Kim, Jamin Shin, Yejin Cho, Joel Jang, Shayne Longpre, Hwaran Lee, Sangdoo Yun, Seongjin Shin, Sungdong Kim, James Thorne, et al. 2023 · 2023
Earlier work this paper cites.
SummEdits: Measuring LLM ability at factual reasoning through the lens of summarization
Philippe Laban, Wojciech Kryscinski, Divyansh Agarwal, Alexander Fabbri, Caiming Xiong, Shafiq Joty, and Chien-Sheng Wu. 2023 · 2023
Earlier work this paper cites.
Yixin Liu, Alexander R Fabbri, Jiawen Chen, Yilun Zhao, Simeng Han, Shafiq Joty, Pengfei Liu, Dragomir Radev, Chien-Sheng Wu, and Arman Cohan. 2023 · 2023
Earlier work this paper cites.
Ares: An automated evaluation framework for retrieval-augmented generation systems
Jon Saad-Falcon, Omar Khattab, Christopher Potts, and Matei Zaharia. 2023 · 2023
Earlier work this paper cites.
Delucionqa: Detecting hallucinations in domain-specific question answering
Mobashir Sadat, Zhengyu Zhou, Lukas Lange, Jun Araki, Arsalan Gundroo, Bingqing Wang, Rakesh R Menon, Md Rizwan Parvez, and Zhe Feng. 2023 · 2023
Earlier work this paper cites.
Large language models are not fair evaluators
Peiyi Wang, Lei Li, Liang Chen, Zefan Cai, Dawei Zhu, Binghuai Lin, Yunbo Cao, Qi Liu, Tianyu Liu, and Zhifang Sui. 2023 · 2023
Earlier work this paper cites.
Fine-grained human feedback gives better rewards for language model training
Zeqiu Wu, Yushi Hu, Weijia Shi, Nouha Dziri, Alane Suhr, Prithviraj Ammanabrolu, Noah A. Smith, Mari Ostendorf, and Hannaneh Hajishirzi. 2023 · 2023
Earlier work this paper cites.
Efficient continual pre-training for building domain specific large language models
Yong Xie, Karan Aggarwal, and Aitzaz Ahmad. 2023 · 2023
Earlier work this paper cites.
A critical evaluation of evaluations for long-form question answering
Fangyuan Xu, Yixiao Song, Mohit Iyyer, and Eunsol Choi. 2023 · 2023
Earlier work this paper cites.
Evaluating large language models at evaluating instruction following
Zhiyuan Zeng, Jiatong Yu, Tianyu Gao, Yu Meng, Tanya Goyal, and Danqi Chen. 2023 · 2023
Earlier work this paper cites.
Phi-3 technical report: A highly capable language model locally on your phone
Marah Abdin, Jyoti Aneja, Hany Awadalla, Ahmed Awadallah, Ammar Ahmad Awan, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Jianmin Bao, Harkirat Behl, et al. 2024 · 2024
Cited alongside, same era.
Command r: Retrieval-augmented generation at production scale
Cohere Team. 2024 · 2024
Cited alongside, same era.
Saullm-54b & saullm-141b: Scaling up domain adaptation for the legal domain
Pierre Colombo, Telmo Pires, Malik Boudiaf, Rui Filipe Coimbra Pereira de Melo, Gabriel Hautreux, Etienne Malaboeuf, Johanne Charpentier, Dominic Culver, and Michael Desa. 2024 · 2024
Cited alongside, same era.
Introducing rag 2.0
Contextual AI Team. 2024 · 2024
Cited alongside, same era.
Glider: Grading llm interactions and decisions using explainable ranking
Darshan Deshpande, Selvan Sunitha Ravi, Sky CH-Wang, Bartosz Mielczarek, Anand Kannappan, and Rebecca Qian. 2024 · 2024
Unanswerability evaluation for retreival augmented generation
Xiangyu Peng, Prafulla Kumar Choubey, Caiming Xiong, and Chien-Sheng Wu. 2024 · 2024
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. 2024 · 2024
Later among the works it cites.
Veritas: A unified approach to reliability evaluation
Rajkumar Ramamurthy, Meghana Arakkal Rajeev, Oliver Molenschot, James Zou, and Nazneen Rajani. 2024 · 2024
Later among the works it cites.
Lynx: An open source hallucination evaluation model
Selvan Sunitha Ravi, Bartosz Mielczarek, Anand Kannappan, Douwe Kiela, and Rebecca Qian. 2024 · 2024
Later among the works it cites.
Lmunit: Fine-grained evaluation with natural language unit tests
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Karel D’Oosterlinck, Winnie Xu, Chris Develder, Thomas Demeester, Amanpreet Singh, Christopher Potts, Douwe Kiela, and Shikib Mehri. 2024 · 2024
Cited alongside, same era.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024 · 2024
Cited alongside, same era.
How to evaluate reward models for rlhf
Evan Frick, Tianle Li, Connor Chen, Wei-Lin Chiang, Anastasios N Angelopoulos, Jiantao Jiao, Banghua Zhu, Joseph E Gonzalez, and Ion Stoica. 2024 · 2024
Cited alongside, same era.
Ragbench: Explainable benchmark for retrieval-augmented generation systems
Robert Friel, Masha Belyi, and Atindriyo Sanyal. 2024 · 2024
Cited alongside, same era.
M-rewardbench: Evaluating reward models in multilingual settings
Srishti Gureja, Lester James V Miranda, Shayekh Bin Islam, Rishabh Maheshwary, Drishti Sharma, Gusti Winata, Nathan Lambert, Sebastian Ruder, Sara Hooker, and Marzieh Fadaee. 2024 · 2024
Cited alongside, same era.
RAG-QA arena: Evaluating domain robustness for long-form retrieval augmented question answering
Rujun Han, Yuhao Zhang, Peng Qi, Yumo Xu, Jenyuan Wang, Lan Liu, William Yang Wang, Bonan Min, and Vittorio Castelli. 2024 · 2024
Cited alongside, same era.
Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, et al. 2024 · 2024
Cited alongside, same era.
Jon Saad-Falcon, Rajan Vivek, William Berrios, Nandita Shankar Naik, Matija Franklin, Bertie Vidgen, Amanpreet Singh, Douwe Kiela, and Shikib Mehri. 2024 · 2024
Later among the works it cites.
Skywork critic model series
Tu Shiwen, Zhao Liang, Chris Yuhao Liu, Liang Zeng, and Yang Liu. 2024 · 2024
Later among the works it cites.
Scaling llm test-time compute optimally can be more effective than scaling model parameters
Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar. 2024 · 2024
Later among the works it cites.
FineSurE: Fine-grained summarization evaluation using LLMs
Hwanjun Song, Hang Su, Igor Shalyminov, Jason Cai, and Saab Mansour. 2024 · 2024
Later among the works it cites.
Judgebench: A benchmark for evaluating llm-based judges
Sijun Tan, Siyuan Zhuang, Kyle Montgomery, William Y Tang, Alejandro Cuadron, Chenguang Wang, Raluca Ada Popa, and Ion Stoica. 2024 · 2024
Later among the works it cites.
Minicheck: Efficient fact-checking of llms on grounding documents
Liyan Tang, Philippe Laban, and Greg Durrett. 2024 · 2024
Later among the works it cites.
Mistral NeMo
The Mistral AI Team. 2024 · 2024
Later among the works it cites.
Replacing judges with juries: Evaluating llm generations with a panel of diverse models
Pat Verga, Sebastian Hofstatter, Sophia Althammer, Yixuan Su, Aleksandra Piktus, Arkady Arkhangorodsky, Minjie Xu, Naomi White, and Patrick Lewis. 2024 · 2024
Later among the works it cites.
Foundational autoraters: Taming large language models for better automatic evaluation
Tu Vu, Kalpesh Krishna, Salaheddin Alzubi, Chris Tar, Manaal Faruqui, and Yun-Hsuan Sung. 2024 · 2024
Later among the works it cites.
On positional bias of faithfulness for long-form summarization
David Wan, Jesse Vig, Mohit Bansal, and Shafiq Joty. 2024 · 2024
Later among the works it cites.
Measuring short-form factuality in large language models
Jason Wei, Nguyen Karina, Hyung Won Chung, Yunxin Joy Jiao, Spencer Papay, Amelia Glaese, John Schulman, and William Fedus. 2024 · 2024
Later among the works it cites.
Cbr-rag: case-based reasoning for retrieval augmented generation in llms for legal question answering
Nirmalie Wiratunga, Ramitha Abeyratne, Lasal Jayawardena, Kyle Martin, Stewart Massie, Ikechukwu Nkisi-Orji, Ruvan Weerasinghe, Anne Liret, and Bruno Fleisch. 2024 · 2024
Later among the works it cites.
Benchmarking retrieval-augmented generation for medicine
Guangzhi Xiong, Qiao Jin, Zhiyong Lu, and Aidong Zhang. 2024 · 2024
Later among the works it cites.
Beyond scalar reward model: Learning generative judge from preference data
Ziyi Ye, Xiangsheng Li, Qiuchi Li, Qingyao Ai, Yujia Zhou, Wei Shen, Dong Yan, and Yiqun Liu. 2024 · 2024
Later among the works it cites.
Atla selene mini: A general purpose evaluation model
Andrei Alexandru, Antonia Calvi, Henry Broomfield, Jackson Golden, Kyle Dai, Mathias Leys, Maurice Burger, Max Bartolo, Roman Engeler, Sashank Pisupati, et al. 2025 · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al. 2025 · 2025
Closest in time.
The facts grounding leaderboard: Benchmarking llms’ ability to ground responses to long-form input
Alon Jacovi, Andrew Wang, Chris Alberti, Connie Tao, Jon Lipovetz, Kate Olszewska, Lukas Haas, Michelle Liu, Nate Keating, Adam Bloniarz, et al. 2025 · 2025
Closest in time.
Demystifying domain-adaptive post-training for financial llms
Zixuan Ke, Yifei Ming, Xuan-Phi Nguyen, Caiming Xiong, and Shafiq Joty. 2025 · 2025
Closest in time.
Structured chain-of-thought prompting for code generation
Jia Li, Ge Li, Yongmin Li, and Zhi Jin. 2025 · 2025
Closest in time.
Learning to verify summary facts with fine-grained LLM feedback
Jihwan Oh, Jeonghwan Choi, Nicole Hee-Yoen Kim, Taewon Yun, and Hwanjun Song. 2025 · 2025
Closest in time.
Learning to plan & reason for evaluation with thinking-llm-as-a-judge
Swarnadeep Saha, Xian Li, Marjan Ghazvininejad, Jason Weston, and Tianlu Wang. 2025 · 2025
Closest in time.