Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) are increasingly used to assess NLP tasks due to their ability to generate human-like judgments.
Maximum likelihood estimation of observer error-rates using the em algorithm
Alexander Philip Dawid and Allan M Skene. 1979 · 1979
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Comparing automatic and human evaluation of NLG systems
Anja Belz and Ehud Reiter. 2006 · 2006
Earlier work this paper cites.
Learning from crowdsourced labeled data: a survey
Jing Zhang, Xindong Wu, and Victor S Sheng. 2016 · 2016
Earlier work this paper cites.
Truth inference in crowdsourcing: Is the problem solved?
Yudian Zheng, Guoliang Li, Yuanbing Li, Caihua Shan, and Reynold Cheng. 2017 · 2017
Earlier work this paper cites.
Discourse coherence in the wild: A dataset, evaluation and methods
Alice Lai and Joel Tetreault. 2018 · 2018
Earlier work this paper cites.
Deep learning from crowds
Filipe Rodrigues and Francisco Pereira. 2018 · 2018
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019 · 2019
Earlier work this paper cites.
A technical survey on statistical modelling and design methods for crowdsourcing quality control
Yuan Jin, Mark Carman, Ye Zhu, and Yong Xiang. 2020 · 2020
Earlier work this paper cites.
SummEval: Re-evaluating Summarization Evaluation
Alexander R. Fabbri, Wojciech Kryściński, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir Radev. 2021 · 2021
Earlier work this paper cites.
TruthfulQA: Measuring how models mimic human falsehoods
Stephanie Lin, Jacob Hilton, and Owain Evans. 2022 · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F Christiano, Jan Leike, and Ryan Lowe. 2022 · 2022
Cited alongside, same era.
Finetuned language models are zero-shot learners
Jason Wei, Maarten Bosma, Vincent Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V Le. 2022 · 2022
Cited alongside, same era.
Knowledge learning with crowdsourcing: A brief review and systematic perspective
Jing Zhang. 2022 · 2022
Cited alongside, same era.
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2023 · 2023
Cited alongside, same era.
Rlaif vs. rlhf: Scaling reinforcement learning from human feedback with ai feedback
Lmsys - chatbot arena human preference predictions
Weilin Chiang, Lianmin Zheng, Lisa Dunlap, Joseph E. Gonzalez, Ion Stoica, Paul Mooney, Sohier Dane, Addison Howard, and Nate Keating. 2024 · 2024
Closest in time.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, et al. 2024 · 2024
Closest in time.
Athene-70b: Redefining the boundaries of post-training for open models
Evan Frick, Peter Jin, Tianle Li, Karthik Ganesan, Jian Zhang, Jiantao Jiao, and Banghua Zhu. 2024 · 2024
Closest in time.
Mixtral of experts
Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, et al. 2024 · 2024
Closest in time.
LLM comparative assessment: Zero-shot NLG evaluation through pairwise comparisons using large language models
Adian Liusie, Potsawee Manakul, and Mark Gales. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Harrison Lee, Samrat Phatale, Hassan Mansoor, Kellie Lu, Thomas Mesnard, Colton Bishop, Victor Carbune, and Abhinav Rastogi. 2023 · 2023
Cited alongside, same era.
HaluEval: A large-scale hallucination evaluation benchmark for large language models
Junyi Li, Xiaoxue Cheng, Xin Zhao, Jian-Yun Nie, and Ji-Rong Wen. 2023 · 2023
Cited alongside, same era.
Orca: Progressive learning from complex explanation traces of gpt-4
Subhabrata Mukherjee, Arindam Mitra, Ganesh Jawahar, Sahaj Agarwal, Hamid Palangi, and Ahmed Awadallah. 2023 · 2023
Cited alongside, same era.
Openhermes 2.5: An open dataset of synthetic data for generalist llm assistants
Teknium. 2023 · 2023
Cited alongside, same era.
Zephyr: Direct distillation of lm alignment
Lewis Tunstall, Edward Beeching, Nathan Lambert, Nazneen Rajani, Kashif Rasul, Younes Belkada, Shengyi Huang, Leandro von Werra, Clémentine Fourrier, Nathan Habib, Nathan Sarrazin, Omar Sanseviero, Alexander M. Rush, and Thomas Wolf. 2023 · 2023
Cited alongside, same era.
Judging LLM-as-a-judge with MT-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023 · 2023
Cited alongside, same era.
Starling-7b: Improving llm helpfulness & harmlessness with rlaif
Banghua Zhu, Evan Frick, Tianhao Wu, Hanlin Zhu, and Jiantao Jiao. 2023 · 2023
Cited alongside, same era.
Reference-guided verdict: Llms-as-judges in automatic evaluation of free-form text
Sher Badshah and Hassan Sajjad. 2024 · 2024
Cited alongside, same era.
Rickard Stureborg, Dimitris Alikaniotis, and Yoshi Suhara. 2024 · 2024
Closest in time.
Crosscheckgpt: Universal hallucination ranking for multimodal foundation models
Guangzhi Sun, Potsawee Manakul, Adian Liusie, Kunat Pipatanakul, Chao Zhang, Phil Woodland, and Mark Gales. 2024 · 2024
Closest in time.
Ryan Teknium, Jeffrey Quesnelle, and Chen Guang. 2024 · 2024
Closest in time.
Replacing judges with juries: Evaluating llm generations with a panel of diverse models
Pat Verga, Sebastian Hofstatter, Sophia Althammer, Yixuan Su, Aleksandra Piktus, Arkady Arkhangorodsky, Minjie Xu, Naomi White, and Patrick Lewis. 2024 · 2024
Closest in time.
Measuring and reducing llm hallucination without gold-standard answers via expertise-weighting
Jiaheng Wei, Yuanshun Yao, Jean-Francois Ton, Hongyi Guo, Andrew Estornell, and Yang Liu. 2024 · 2024
Closest in time.
Qwen2 technical report
An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, et al. 2024 · 2024
Closest in time.