Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) are increasingly employed as automated evaluators to assess the safety of generated content, yet their reliability in this role remains uncertain.
The price of debiasing automatic metrics in natural language evalaution
Arun Chaganty, Stephen Mussmann, and Percy Liang. 2018 · 2018
Earlier work this paper cites.
Shortcut learning in deep neural networks
Robert Geirhos, Jörn-Henrik Jacobsen, Claudio Michaelis, Richard Zemel, Wieland Brendel, Matthias Bethge, and Felix A Wichmann. 2020 · 2020
Earlier work this paper cites.
The halo effect: A longitudinal approach
Juan Luis Nicolau, Juan Pedro Mellinas, and Eva Martín-Fuentes. 2020 · 2020
Earlier work this paper cites.
A survey of reinforcement learning from human feedback
Timo Kaufmann, Paul Weng, Viktor Bengs, and Eyke Hüllermeier. 2023 · 2023
Earlier work this paper cites.
HuaSLIM: Human attention motivated shortcut learning identification and mitigation for large language models
Yuqi Ren and Deyi Xiong. 2023 · 2023
Earlier work this paper cites.
Large language models can be lazy learners: Analyze shortcuts in in-context learning
Ruixiang Tang, Dehan Kong, Longtao Huang, and Hui Xue. 2023 · 2023
Earlier work this paper cites.
Style over substance: Evaluation biases for large language models
Minghao Wu and Alham Fikri Aji. 2023 · 2023
Earlier work this paper cites.
Judging LLM-as-a-judge with MT-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023 · 2023
Earlier work this paper cites.
The multilingual alignment prism: Aligning global and local preferences to reduce harm
Aakanksha, Arash Ahmadian, Beyza Ermis, Seraphina Goldfarb-Tarrant, Julia Kreutzer, Marzieh Fadaee, and Sara Hooker. 2024 · 2024
Earlier work this paper cites.
Bhashithe Abeysinghe and Ruhan Circi. 2024 · 2024
Earlier work this paper cites.
Llama 3 model card
AI@Meta. 2024 · 2024
Cited alongside, same era.
Humans or LLMs as the judge? a study on judgement bias
Guiming Hardy Chen, Shunian Chen, Ziche Liu, Feng Jiang, and Benyou Wang. 2024 · 2024
Cited alongside, same era.
Chatbot arena: An open platform for evaluating llms by human preference
Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Banghua Zhu, Hao Zhang, Michael Jordan, Joseph E Gonzalez, et al. 2024 · 2024
Cited alongside, same era.
ConSiDERS-the-human evaluation framework: Rethinking human evaluation for generative large language models
Aparna Elangovan, Ling Liu, Lei Xu, Sravan Babu Bodapati, and Dan Roth. 2024 · 2024
Cited alongside, same era.
Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, Lélio Renard Lavaud, Lucile Saulnier, Marie-Anne Lachaux, Pierre Stock, Sandeep Subramanian, Sophia Yang, Szymon Antoniak, Teven Le Scao, Théophile Gervet, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2024 · 2024
Shortcut learning in in-context learning: A survey
Rui Song, Yingji Li, Lida Shi, Fausto Giunchiglia, and Hao Xu. 2024 · 2024
Later among the works it cites.
Judging the judges: Evaluating alignment and vulnerabilities in llms-as-judges
Aman Singh Thakur, Kartik Choudhary, Venkat Srinik Ramayapally, Sankaran Vaidyanathan, and Dieuwke Hupkes. 2024 · 2024
Later among the works it cites.
Replacing judges with juries: Evaluating llm generations with a panel of diverse models
Pat Verga, Sebastian Hofstatter, Sophia Althammer, Yixuan Su, Aleksandra Piktus, Arkady Arkhangorodsky, Minjie Xu, Naomi White, and Patrick Lewis. 2024 · 2024
Later among the works it cites.
Large language models are not fair evaluators
Peiyi Wang, Lei Li, Liang Chen, Zefan Cai, Dawei Zhu, Binghuai Lin, Yunbo Cao, Lingpeng Kong, Qi Liu, Tianyu Liu, and Zhifang Sui. 2024 · 2024
Later among the works it cites.
Sorry-bench: Systematically evaluating large language model safety refusal behaviors
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Benchmarking cognitive biases in large language models as evaluators
Ryan Koo, Minhwa Lee, Vipul Raheja, Jong Inn Park, Zae Myung Kim, and Dongyeop Kang. 2024 · 2024
Cited alongside, same era.
Prd: Peer rank and discussion improve large language model based evaluations
Ruosen Li, Teerth Patel, and Xinya Du. 2024 · 2024
Cited alongside, same era.
LLM comparative assessment: Zero-shot NLG evaluation through pairwise comparisons using large language models
Adian Liusie, Potsawee Manakul, and Mark Gales. 2024 · 2024
Cited alongside, same era.
Au large
Mistral AI. 2024 · 2024
Cited alongside, same era.
Is LLM-as-a-judge robust? investigating universal adversarial attacks on zero-shot LLM assessment
Vyas Raina, Adian Liusie, and Mark Gales. 2024 · 2024
Cited alongside, same era.
Protecting Children from Online Grooming
ActiveFence
Cited in the paper.
The claude 3 model family: Opus, sonnet, haiku
Anthropic
Cited in the paper.
Tinghao Xie, Xiangyu Qi, Yi Zeng, Yangsibo Huang, Udari Madhushani Sehwag, Kaixuan Huang, Luxi He, Boyi Wei, Dacheng Li, Ying Sheng, Ruoxi Jia, Bo Li, Kai Li, Danqi Chen, Peter Henderson, and Prateek Mittal. 2024 · 2024
Later among the works it cites.
Fantastic llms for preference data annotation and how to (not) find them
Guangxuan Xu, Kai Xu, Shivchander Sudalairaj, Hao Wang, and Akash Srivastava. 2024 · 2024
Later among the works it cites.
Air-bench 2024: A safety benchmark based on risk categories from regulations and policies
Yi Zeng, Yu Yang, Andy Zhou, Jeffrey Ziwei Tan, Yuheng Tu, Yifan Mai, Kevin Klyman, Minzhou Pan, Ruoxi Jia, Dawn Song, Percy Liang, and Bo Li. 2024a · 2024
Later among the works it cites.
Command a: An enterprise-ready large language model
Team Cohere, Arash Ahmadian, Marwan Ahmed, Jay Alammar, Milad Alizadeh, Yazeed Alnumay, Sophia Althammer, Arkady Arkhangorodsky, Viraat Aryabumi, Dennis Aumiller, et al. 2025 · 2025
Closest in time.
RewardBench: Evaluating reward models for language modeling
Nathan Lambert, Valentina Pyatkin, Jacob Morrison, LJ Miranda, Bill Yuchen Lin, Khyathi Chandu, Nouha Dziri, Sachin Kumar, Tom Zick, Yejin Choi, Noah A. Smith, and Hannaneh Hajishirzi. 2025 · 2025
Closest in time.