Fetching the paper…
Reading the bibliography…
Existing large language models (LLMs) evaluation methods typically focus on testing the performance on some closed-environment and domain-specific benchmarks with human annotations.
A new measure of rank correlation
Maurice G Kendall · 1938
Earlier work this paper cites.
Mjrty—a fast majority vote algorithm
Robert S Boyer and J Strother Moore · 1991
Earlier work this paper cites.
Introduction to algorithms , volume 3
Charles Eric Leiserson, Ronald L Rivest, Thomas H Cormen, and Clifford Stein · 1994
Earlier work this paper cites.
Validity problems comparing values across cultures and possible solutions
Kaiping Peng, Richard E Nisbett, and Nancy YC Wong · 1997
Earlier work this paper cites.
Permutation entropy: a natural complexity measure for time series
Christoph Bandt and Bernd Pompe · 2002
Earlier work this paper cites.
The wisdom of crowds
James Surowiecki · 2005
Earlier work this paper cites.
Majority voting
Allan M. Feldman · 2006
Earlier work this paper cites.
Cultural consensus theory: Applications and frequently asked questions
Susan C Weller · 2007
Earlier work this paper cites.
Rating through voting: An iterative method for robust rating
Mohammad Allahbakhsh and Aleksandar Ignjatovic · 2012
Earlier work this paper cites.
Pearson’s correlation coefficient
Philip Sedgwick · 2012
Earlier work this paper cites.
JMP for basic univariate and multivariate statistics: methods for researchers and social scientists
Ann Lehman, Norm O’Rourke, Larry Hatcher, and Edward Stepanski · 2013
Earlier work this paper cites.
Identifying expertise to extract the wisdom of crowds
David V Budescu and Eva Chen · 2015
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman · 2018
Earlier work this paper cites.
Superglue: A stickier benchmark for general-purpose language understanding systems
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al · 2021
Earlier work this paper cites.
Webgpt: Browser-assisted question-answering with human feedback
Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, et al · 2021
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al · 2022
Earlier work this paper cites.
Holistic evaluation of language models
Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al · 2022
Earlier work this paper cites.
Active learning for open-set annotation
Kun-Peng Ning, Xun Zhao, Yu Li, and Sheng-Jun Huang · 2022
Cited alongside, same era.
Introducing chatgpt
OpenAI · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Cited alongside, same era.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, et al · 2022
Cited alongside, same era.
Self-instruct: Aligning language model with self generated instructions
Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A Smith, Daniel Khashabi, and Hannaneh Hajishirzi · 2022
Cited alongside, same era.
Pandalm: An automatic evaluation benchmark for llm instruction tuning optimization
Yidong Wang, Zhuohao Yu, Zhengran Zeng, Linyi Yang, Cunxiang Wang, Hao Chen, Chaoya Jiang, Rui Xie, Jindong Wang, Xing Xie, et al · 2023
Later among the works it cites.
Skywork: A more open bilingual foundation model
Tianwen Wei, Liang Zhao, Lichang Zhang, Bo Zhu, Lijie Wang, Haihua Yang, Biye Li, Cheng Cheng, Weiwei Lü, Rui Hu, et al · 2023
Later among the works it cites.
Wizardlm: Empowering large language models to follow complex instructions, 2023
Can Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng, Pu Zhao, Jiazhan Feng, Chongyang Tao, and Daxin Jiang · 2023
Later among the works it cites.
Llm lies: Hallucinations are not bugs, but features as adversarial examples
Jia-Yu Yao, Kun-Peng Ning, Zhen-Hui Liu, Mu-Nan Ning, and Li Yuan · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Glm-130b: An open bilingual pre-trained model
Aohan Zeng, Xiao Liu, Zhengxiao Du, Zihan Wang, Hanyu Lai, Ming Ding, Zhuoyi Yang, Yifan Xu, Wendi Zheng, Xiao Xia, et al · 2022
Cited alongside, same era.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Cited alongside, same era.
Can we trust the evaluation on chatgpt?, 2023
Rachith Aiyappa, Jisun An, Haewoon Kwak, and Yong-Yeol Ahn · 2023
Cited alongside, same era.
Gpt4all: Training an assistant-style chatbot with large scale data distillation from gpt-3.5-turbo
Yuvanesh Anand, Zach Nussbaum, Brandon Duderstadt, Benjamin Schmidt, and Andriy Mulyar · 2023
Cited alongside, same era.
Yejin Bang, Samuel Cahyawijaya, Nayeon Lee, Wenliang Dai, Dan Su, Bryan Wilie, Holy Lovenia, Ziwei Ji, Tiezheng Yu, Willy Chung, et al · 2023
Cited alongside, same era.
Sparks of artificial general intelligence: Early experiments with gpt-4
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al · 2023
Cited alongside, same era.
Chateval: Towards better llm-based evaluators through multi-agent debate
Chi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu, Wei Xue, Shanghang Zhang, Jie Fu, and Zhiyuan Liu · 2023
Cited alongside, same era.
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al · 2023
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena, 2023
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric. P Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica · 2023
Later among the works it cites.
Don’t make your llm an evaluation benchmark cheater
Kun Zhou, Yutao Zhu, Zhipeng Chen, Wentong Chen, Wayne Xin Zhao, Xu Chen, Yankai Lin, Ji-Rong Wen, and Jiawei Han · 2023
Later among the works it cites.
Can large language models transform computational social science?
Caleb Ziems, William Held, Omar Shaikh, Jiaao Chen, Zhehao Zhang, and Diyi Yang · 2023
Later among the works it cites.
https://guanaco-model.github.io/ , 2023
Guanaco - generative universal assistant for natural-language adaptive context-aware omnilingual outputs · 2024
Closest in time.
Stablelm-tuned-alpha-7b: A fine-tuned language model for diverse applications
Stability AI · 2024
Closest in time.
Vicuna: An open-source chatbot impressing gpt-4 with 90% chatgpt quality
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E Gonzalez, et al · 2024
Closest in time.
Pre: A peer review based large language model evaluator
Zhumin Chu, Qingyao Ai, Yiteng Tu, Haitao Li, and Yiqun Liu · 2024
Closest in time.
Oasst-sft-4-pythia-12b: A supervised fine-tuning model for language understanding
Open-Assistant Contributors · 2024
Closest in time.
Qlora: Efficient finetuning of quantized llms
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer · 2024
Closest in time.
Koala-13b: Dialogue model for effective human-ai interaction
Xinyang Geng, Arnav Gudibande, Hao Liu, Eric Wallace, Pieter Abbeel, Sergey Levine, and Dawn Song · 2024
Closest in time.
Sparse orthogonal parameters tuning for continual learning
Kun-Peng Ning, Hai-Jian Ke, Yu-Yang Liu, Jia-Yu Yao, Yong-Hong Tian, and Li Yuan · 2024
Closest in time.
How far can camels go? exploring the state of instruction tuning on open resources
Yizhong Wang, Hamish Ivison, Pradeep Dasigi, Jack Hessel, Tushar Khot, Khyathi Chandu, David Wadden, Kelsey MacMillan, Noah A Smith, Iz Beltagy, et al · 2024
Closest in time.
Is parameter collision hindering continual learning in llms?
Shuo Yang, Kun-Peng Ning, Yu-Yang Liu, Jia-Yu Yao, Yong-Hong Tian, Yi-Bing Song, and Li Yuan · 2024
Closest in time.
Gpt as a monte carlo language tree: A probabilistic perspective
Kun-Peng Ning, Jia-Yu Yao, Yu-Yang Liu, Mu-Nan Ning, and Li Yuan · 2025
Closest in time.