Fetching the paper…
Reading the bibliography…
The rapid evolution of large language models (LLMs) has ushered in the need for comprehensive assessments of their performance across various dimensions.
MCTest: A challenge dataset for the open-domain machine comprehension of text
Matthew Richardson, Christopher J. C. Burges, and Erin Renshaw. 2013 · 2013
Earlier work this paper cites.
MS MARCO: A human generated machine reading comprehension dataset
Tri Nguyen, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder, and Li Deng. 2016 · 2016
Earlier work this paper cites.
SQuAD: 100, 000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
RACE: large-scale reading comprehension dataset from examinations
Guokun Lai, Qizhe Xie, Hanxiao Liu, Yiming Yang, and Eduard H. Hovy. 2017 · 2017
Earlier work this paper cites.
The NarrativeQA Reading Comprehension Challenge
Tomás Kociský, Jonathan Schwarz, Phil Blunsom, Chris Dyer, Karl Moritz Hermann, Gábor Melis, and Edward Grefenstette. 2018 · 2018
Earlier work this paper cites.
MCScript: A novel dataset for assessing machine comprehension using script knowledge
Simon Ostermann, Ashutosh Modi, Michael Roth, Stefan Thater, and Manfred Pinkal. 2018 · 2018
Earlier work this paper cites.
DRCD: a chinese machine reading comprehension dataset
Chih-Chieh Shao, Trois Liu, Yuting Lai, Yiying Tseng, and Sam Tsai. 2018 · 2018
Earlier work this paper cites.
A span-extraction dataset for chinese machine reading comprehension
Yiming Cui, Ting Liu, Wanxiang Che, Li Xiao, Zhipeng Chen, Wentao Ma, Shijin Wang, and Guoping Hu. 2019 · 2019
Cited alongside, same era.
CJRC: A reliable human-annotated benchmark dataset for chinese judicial reading comprehension
Xingyi Duan, Baoxin Wang, Ziyue Wang, Wentao Ma, Yiming Cui, Dayong Wu, Shijin Wang, Ting Liu, Tianxiang Huo, Zhen Hu, Heng Wang, and Zhiyuan Liu. 2019 · 2019
Cited alongside, same era.
BiPaR: A bilingual parallel dataset for multilingual and cross-lingual reading comprehension on novels
Yimin Jing, Deyi Xiong, and Yan Zhen. 2019 · 2019
Cited alongside, same era.
CoQA: A conversational question answering challenge
Siva Reddy, Danqi Chen, and Christopher D. Manning. 2019 · 2019
Cited alongside, same era.
Native chinese reader: A dataset towards native-level chinese machine reading comprehension
Shusheng Xu, Yichen Liu, Xiaoyu Yi, Siyuan Zhou, Huizi Li, and Yi Wu. 2021 · 2021
Cited alongside, same era.
Evaluating large language models: A comprehensive survey
Zishan Guo, Renren Jin, Chuang Liu, Yufei Huang, Dan Shi, Supryadi, Linhao Yu, Yan Liu, Jiaxuan Li, Bojian Xiong, and Deyi Xiong. 2023 · 2023
Later among the works it cites.
CBBQ: A chinese bias benchmark dataset curated with human-ai collaboration for large language models
Yufei Huang and Deyi Xiong. 2023 · 2023
Later among the works it cites.
Chuang Liu, Renren Jin, Yuqi Ren, Linhao Yu, Tianyu Dong, Xiaohan Peng, Shuting Zhang, Jianxiang Peng, Peiyi Zhang, Qingqing Lyu, Xiaowen Su, Qun Liu, and Deyi Xiong. 2023 · 2023
Later among the works it cites.
Roleeval: A bilingual role evaluation benchmark for large language models
Tianhao Shen, Sun Li, and Deyi Xiong. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
GLM: general language model pretraining with autoregressive blank infilling
Zhengxiao Du, Yujie Qian, Xiao Liu, Ming Ding, Jiezhong Qiu, Zhilin Yang, and Jie Tang. 2022 · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F. Christiano, Jan Leike, and Ryan Lowe. 2022 · 2022
Cited alongside, same era.
WYWEB: A NLP evaluation benchmark for classical chinese
Bo Zhou, Qianglong Chen, Xiaomi Zhong, and Yin Zhang. 2023 · 2023
Later among the works it cites.
Chuang Liu, Renren Jin, Yuqi Ren, and Deyi Xiong. 2024 · 2024
Closest in time.