Fetching the paper…
Reading the bibliography…
The "LLM-as-an-annotator" and "LLM-as-a-judge" paradigms employ Large Language Models (LLMs) as annotators, judges, and evaluators in tasks traditionally performed by humans.
The control of the false discovery rate in multiple testing under dependency
Yoav Benjamini and Daniel Yekutieli. 2001 · 2001
Earlier work this paper cites.
Crowdtruth: Machine-human computation framework for harnessing disagreement in gathering annotated data
Oana Inel, Khalid Khamkham, Tatiana Cristea, Anca Dumitrache, Arne Rutjes, Jelle van der Ploeg, Lukasz Romaszko, Lora Aroyo, and Robert-Jan Sips. 2014 · 2014
Earlier work this paper cites.
Replicability analysis for natural language processing: Testing significance with multiple datasets
Rotem Dror, Gili Baumer, Marina Bogomolov, and Roi Reichart. 2017 · 2017
Earlier work this paper cites.
Crowd disagreement about medical images is informative
Veronika Cheplygina and Josien P. W. Pluim. 2018 · 2018
Earlier work this paper cites.
Skin lesion analysis toward melanoma detection: A challenge at the 2017 international symposium on biomedical imaging (isbi), hosted by the international skin imaging collaboration (ISIC)
Noel C. F. Codella, David A. Gutman, M. Emre Celebi, Brian Helba, Michael A. Marchetti, Stephen W. Dusza, Aadi Kalloo, Konstantinos Liopyris, Nabin K. Mishra, Harald Kittler, and Allan Halpern. 2018 · 2018
Earlier work this paper cites.
The hitchhiker’s guide to testing statistical significance in natural language processing
Rotem Dror, Gili Baumer, Segev Shlomov, and Roi Reichart. 2018 · 2018
Earlier work this paper cites.
We need to consider disagreement in evaluation
Valerio Basile, Michael Fell, Tommaso Fornaciari, Dirk Hovy, Silviu Paun, Barbara Plank, Massimo Poesio, Alexandra Uma, et al. 2021 · 2021
Earlier work this paper cites.
Summeval: Re-evaluating summarization evaluation
Alexander R Fabbri, Wojciech Kryściński, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir Radev. 2021 · 2021
Earlier work this paper cites.
Learning from disagreement: A survey
Alexandra Uma, Tommaso Fornaciari, Dirk Hovy, Silviu Paun, Barbara Plank, and Massimo Poesio. 2021 · 2021
Earlier work this paper cites.
Cebab: Estimating the causal effects of real-world concepts on nlp model behavior
Eldar D Abraham, Karel D’Oosterlinck, Amir Feder, Yair Gat, Atticus Geiger, Christopher Potts, Roi Reichart, and Zhengxuan Wu. 2022 · 2022
Earlier work this paper cites.
Abstract visual reasoning with tangram shapes
Anya Ji, Noriyuki Kojima, Noah Rush, Alane Suhr, Wai Keen Vong, Robert Hawkins, and Yoav Artzi. 2022 · 2022
Earlier work this paper cites.
Large language models are zero-shot reasoners
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022 · 2022
Earlier work this paper cites.
Wax: A new dataset for word association explanations
Chunhua Liu, Trevor Cohn, Simon De Deyne, and Lea Frermann. 2022 · 2022
Earlier work this paper cites.
The “problem” of human label variation: On ground truth in data, modeling and evaluation
Barbara Plank. 2022 · 2022
Earlier work this paper cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023 · 2023
Earlier work this paper cites.
Using large language models to simulate multiple humans and replicate human subject studies
Gati V. Aher, Rosa I. Arriaga, and Adam Tauman Kalai. 2023 · 2023
Earlier work this paper cites.
Benchmarking foundation models with language-model-as-an-examiner
Yushi Bai, Jiahao Ying, Yixin Cao, Xin Lv, Yuze He, Xiaozhi Wang, Jifan Yu, Kaisheng Zeng, Yijia Xiao, Haozhe Lyu, Jiayin Zhang, Juanzi Li, and Lei Hou. 2023 · 2023
Earlier work this paper cites.
Self-consistency of large language models under ambiguity
Henning Bartsch, Ole Jorgensen, Domenic Rosati, Jason Hoelscher-Obermaier, and Jacob Pfau. 2023 · 2023
Earlier work this paper cites.
Can large language models be an alternative to human evaluations?
Cheng-Han Chiang and Hung-yi Lee. 2023 · 2023
Earlier work this paper cites.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E Gonzalez, et al. 2023 · 2023
Earlier work this paper cites.
Using imperfect surrogates for downstream inference: Design-based supervised learning for social science applications of large language models
Naoki Egami, Musashi Hinck, Brandon M. Stewart, and Hanying Wei. 2023 · 2023
Earlier work this paper cites.
Conflicts, villains, resolutions: Towards models of narrative media framing
Lea Frermann, Jiatong Li, Shima Khanehzar, and Gosia Mikolajczak. 2023 · 2023
Earlier work this paper cites.
Chatgpt outperforms crowd-workers for text-annotation tasks
Fabrizio Gilardi, Meysam Alizadeh, and Maël Kubli. 2023 · 2023
Earlier work this paper cites.
A systematic study and comprehensive evaluation of ChatGPT on benchmark datasets
Md Tahmid Rahman Laskar, M Saiful Bari, Mizanur Rahman, Md Amran Hossen Bhuiyan, Shafiq Joty, and Jimmy Huang. 2023 · 2023
Cited alongside, same era.
140 characters of justice? the promise and perils of using social media to reveal lay punishment perspectives
Itay Ravid and Rotem Dror. 2023 · 2023
Cited alongside, same era.
Performance of large language models on a neurology board–style examination
Marc Cicero Schubert, Wolfgang Wick, and Varun Venkataramani. 2023 · 2023
Cited alongside, same era.
Navigating cultural chasms: Exploring and unlocking the cultural POV of text-to-image models
Mor Ventura, Eyal Ben-David, Anna Korhonen, and Roi Reichart. 2023 · 2023
Cited alongside, same era.
Style over substance: Evaluation biases for large language models
Minghao Wu and Alham Fikri Aji. 2023 · 2023
Rewardbench: Evaluating reward models for language modeling
Nathan Lambert, Valentina Pyatkin, Jacob Morrison, LJ Miranda, Bill Yuchen Lin, Khyathi Raghavi Chandu, Nouha Dziri, Sachin Kumar, Tom Zick, Yejin Choi, Noah A. Smith, and Hannaneh Hajishirzi. 2024 · 2024
Later among the works it cites.
Large language models: An applied econometric framework
Jens Ludwig, Sendhil Mullainathan, and Ashesh Rambachan. 2024 · 2024
Later among the works it cites.
Large language models surpass human experts in predicting neuroscience results
Xiaoliang Luo, Akilles Rechardt, Guangzhi Sun, Kevin K Nejad, Felipe Yáñez, Bati Yilmaz, Kangjoo Lee, Alexandra O Cohen, Valentina Borghesani, Anton Pashkov, et al. 2024 · 2024
Later among the works it cites.
Evaluating the performance of large language models via debates
Behrad Moniri, Hamed Hassani, and Edgar Dobriban. 2024 · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Automatic evaluation of attribution by large language models
Xiang Yue, Boshi Wang, Ziru Chen, Kai Zhang, Yu Su, and Huan Sun. 2023 · 2023
Cited alongside, same era.
Judgelm: Fine-tuned large language models are scalable judges
Lianghui Zhu, Xinggang Wang, and Xinlong Wang. 2023 · 2023
Cited alongside, same era.
Can llms replace manual annotation of software engineering artifacts?
Toufique Ahmed, Premkumar T. Devanbu, Christoph Treude, and Michael Pradel. 2024 · 2024
Cited alongside, same era.
Zahra Ashktorab, Michael Desmond, Qian Pan, James M. Johnson, Martin Santillan Cooper, Elizabeth M. Daly, Rahul Nair, Tejaswini Pedapati, Swapnaja Achintalwar, and Werner Geyer. 2024 · 2024
Cited alongside, same era.
Llms instead of human judges? A large scale empirical study across 20 NLP evaluation tasks
Anna Bavaresco, Raffaella Bernardi, Leonardo Bertolazzi, Desmond Elliott, Raquel Fernández, Albert Gatt, Esam Ghaleb, Mario Giulianelli, Michael Hanna, Alexander Koller, André F. T. Martins, Philipp Mondorf, Vera Neplenbroek, Sandro Pezzelle, Barbara Plank, David Schlangen, Alessandro Suglia, Aditya K. Surikuchi, Ece Takmaz, and Alberto Testoni. 2024 · 2024
Cited alongside, same era.
On behalf of the stakeholders: Trends in nlp model interpretability in the era of llms
Nitay Calderon and Roi Reichart. 2024 · 2024
Cited alongside, same era.
Prediction-powered ranking of large language models
Ivi Chatzi, Eleni Straitouri, Suhas Thejaswi, and Manuel Gomez Rodriguez. 2024 · 2024
Cited alongside, same era.
Omer Nahum, Nitay Calderon, Orgad Keller, Idan Szpektor, and Roi Reichart. 2024 · 2024
Later among the works it cites.
CARE: extracting experimental findings from clinical literature
Aakanksha Naik, Bailey Kuehl, Erin Bransom, Doug Downey, and Tom Hope. 2024 · 2024
Later among the works it cites.
Chatgpt label: Comparing the quality of human-generated and llm-generated annotations in low-resource language NLP tasks
Arbi Haza Nasution and Aytug Onan. 2024 · 2024
Later among the works it cites.
Maja Pavlovic and Massimo Poesio. 2024 · 2024
Later among the works it cites.
Can large language models replace economic choice prediction labs?
Eilam Shapira, Omer Madmon, Roi Reichart, and Moshe Tennenholtz. 2024 · 2024
Later among the works it cites.
Can many-shot in-context learning help long-context LLM judges? see more, judge better!
Mingyang Song, Mao Zheng, and Xuan Luo. 2024 · 2024
Later among the works it cites.
Limitations of the llm-as-a-judge approach for evaluating LLM outputs in expert knowledge tasks
Annalisa Szymanski, Noah Ziems, Heather A. Eicher-Miller, Toby Jia-Jun Li, Meng Jiang, and Ronald A. Metoyer. 2024 · 2024
Later among the works it cites.
Large language models for data annotation and synthesis: A survey
Zhen Tan, Dawei Li, Song Wang, Alimohammad Beigi, Bohan Jiang, Amrita Bhattacharjee, Mansooreh Karami, Jundong Li, Lu Cheng, and Huan Liu. 2024b · 2024
Later among the works it cites.
Large language models and the wisdom of small crowds
Sean Trott. 2024 · 2024
Later among the works it cites.
Replacing judges with juries: Evaluating llm generations with a panel of diverse models
Pat Verga, Sebastian Hofstatter, Sophia Althammer, Yixuan Su, Aleksandra Piktus, Arkady Arkhangorodsky, Minjie Xu, Naomi White, and Patrick Lewis. 2024 · 2024
Later among the works it cites.
Large language models are not fair evaluators
Peiyi Wang, Lei Li, Liang Chen, Zefan Cai, Dawei Zhu, Binghuai Lin, Yunbo Cao, Lingpeng Kong, Qi Liu, Tianyu Liu, and Zhifang Sui. 2024 · 2024
Later among the works it cites.
Pride and prejudice: LLM amplifies self-bias in self-refinement
Wenda Xu, Guanglei Zhu, Xuandong Zhao, Liangming Pan, Lei Li, and William Wang. 2024 · 2024
Later among the works it cites.
Harnessing the power of llms in practice: A survey on chatgpt and beyond
Jingfeng Yang, Hongye Jin, Ruixiang Tang, Xiaotian Han, Qizhang Feng, Haoming Jiang, Shaochen Zhong, Bing Yin, and Xia Ben Hu. 2024 · 2024
Later among the works it cites.
Justice or prejudice? quantifying biases in llm-as-a-judge
Jiayi Ye, Yanbo Wang, Yue Huang, Dongping Chen, Qihui Zhang, Nuno Moniz, Tian Gao, Werner Geyer, Chao Huang, Pin-Yu Chen, Nitesh V. Chawla, and Xiangliang Zhang. 2024 · 2024
Later among the works it cites.
Can large language models transform computational social science?
Caleb Ziems, William Held, Omar Shaikh, Jiaao Chen, Zhehao Zhang, and Diyi Yang. 2024 · 2024
Later among the works it cites.
Chatgpt out-scores medical students on complex clinical care exam questions
Adam Hadhazy. 2023 · 2025
Closest in time.
Trueteacher: Learning factual consistency evaluation with large language models
Zorik Gekhman, Jonathan Herzig, Roee Aharoni, Chen Elkind, and Idan Szpektor. 2023 · 2070
Closest in time.
The colorful future of LLMs: Evaluating and improving LLMs as emotional supporters for queer youth
Shir Lissak, Nitay Calderon, Geva Shenkman, Yaakov Ophir, Eyal Fruchter, Anat Brunstein Klomek, and Roi Reichart. 2024 · 2079
Closest in time.