Fetching the paper…
Reading the bibliography…
We evaluate questions generated by large language models (LLMs) from context, comparing them to human-authored questions across six dimensions: question type, question length, context coverage, answerability, uncommonness, and required answer length.
Recent advances in neural question generation, 2019
Liangming Pan, Wenqiang Lei, Tat-Seng Chua, and Min-Yen Kan · 1905
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
Automatic generation of short answer questions for reading comprehension assessment
YAN HUANG and LIANZHEN HE · 2016
Earlier work this paper cites.
Pointer sentinel mixture models, 2016
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Earlier work this paper cites.
Generating natural questions about an image
Nasrin Mostafazadeh, Ishan Misra, Jacob Devlin, Margaret Mitchell, Xiaodong He, and Lucy Vanderwende · 2016
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text, 2016
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Earlier work this paper cites.
Iulian Vlad Serban, Alberto García-Durán, Caglar Gulcehre, Sungjin Ahn, Sarath Chandar, Aaron Courville, and Yoshua Bengio · 2016
Earlier work this paper cites.
Towards Topic-to-Question Generation
Yllias Chali and Sadid A. Hasan · 2017
Earlier work this paper cites.
Question generation for question answering
Nan Duan, Duyu Tang, Peng Chen, and Ming Zhou · 2017
Earlier work this paper cites.
Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension
Mandar Joshi, Eunsol Choi, Daniel S. Weld, and Luke Zettlemoyer · 2017
Earlier work this paper cites.
Generating natural language question-answer pairs from a knowledge graph using a RNN based question generation model
Sathish Reddy, Dinesh Raghu, Mitesh M. Khapra, and Sachindra Joshi · 2017
Earlier work this paper cites.
Visual question generation as dual task of visual question answering
Yikang Li, Nan Duan, Bolei Zhou, Xiao Chu, Wanli Ouyang, Xiaogang Wang, and Ming Zhou · 2018
Cited alongside, same era.
Hotpotqa: A dataset for diverse, explainable multi-hop question answering
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W Cohen, Ruslan Salakhutdinov, and Christopher D Manning · 2018
Cited alongside, same era.
Difficulty-controllable multi-hop question generation from knowledge graphs
Vishwajeet Kumar, Yuncheng Hua, Ganesh Ramakrishnan, Guilin Qi, Lianli Gao, and Yuan-Fang Li · 2019
Cited alongside, same era.
Natural questions: a benchmark for question answering research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al · 2019
Cited alongside, same era.
ToxiGen: A large-scale machine-generated dataset for adversarial and implicit hate speech detection
Thomas Hartvigsen, Saadia Gabriel, Hamid Palangi, Maarten Sap, Dipankar Ray, and Ece Kamar · 2022
Hagrid: A human-llm collaborative dataset for generative information-seeking with attribution, 2023
Ehsan Kamalloo, Aref Jafari, Xinyu Zhang, Nandan Thakur, and Jimmy Lin · 2023
Later among the works it cites.
Dynosaur: A dynamic growth paradigm for instruction-tuning data curation, 2023
Da Yin, Xiao Liu, Fan Yin, Ming Zhong, Hritik Bansal, Jiawei Han, and Kai-Wei Chang · 2023
Later among the works it cites.
Toolqa: A dataset for llm question answering with external tools
Yuchen Zhuang, Yue Yu, Kuan Wang, Haotian Sun, and Chao Zhang · 2023
Later among the works it cites.
Cosmopedia, 2024
Loubna Ben Allal, Anton Lozhkov, Guilherme Penedo, Thomas Wolf, and Leandro von Werra · 2024
Later among the works it cites.
Under the surface: Tracking the artifactuality of llm-generated data, 2024
Debarati Das, Karin De Langis, Anna Martin, Jaehyung Kim, Minhwa Lee, Zae Myung Kim, Shirley Hayati, Risako Owan, Bin Hu, Ritik Parkar, Ryan Koo, Jonginn Park, Aahan Tyagi, Libby Ferland, Sanjali Roy, Vincent Liu, and Dongyeop Kang · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Knowledge-based visual question generation
Jiayuan Xie, Wenhao Fang, Yi Cai, Qingbao Huang, and Qing Li · 2022
Cited alongside, same era.
Autoqgs: Auto-prompt for low-resource knowledge-based question generation from sparql
Guanming Xiong, Junwei Bao, Wen Zhao, Youzheng Wu, and Xiaodong He · 2022
Cited alongside, same era.
Towards a unified multi-dimensional evaluator for text generation
Ming Zhong, Yang Liu, Da Yin, Yuning Mao, Yizhu Jiao, Pengfei Liu, Chenguang Zhu, Heng Ji, and Jiawei Han · 2022
Cited alongside, same era.
Tinystories: How small can language models be and still speak coherent english?, 2023
Ronen Eldan and Yuanzhi Li · 2023
Cited alongside, same era.
Ragas: Automated evaluation of retrieval augmented generation
Shahul Es, Jithin James, Luis Espinosa-Anke, and Steven Schockaert · 2023
Cited alongside, same era.
Gptscore: Evaluate as you desire
Jinlan Fu, See-Kiong Ng, Zhengbao Jiang, and Pengfei Liu · 2023
Cited alongside, same era.
Beavertails: Towards improved safety alignment of LLM via a human-preference dataset
Jiaming Ji, Mickel Liu, Juntao Dai, Xuehai Pan, Chi Zhang, Ce Bian, Boyuan Chen, Ruiyang Sun, Yizhou Wang, and Yaodong Yang · 2023
Cited alongside, same era.
Later among the works it cites.
Deepseek-v3 technical report, 2024
DeepSeek-AI · 2024
Later among the works it cites.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al · 2024
Later among the works it cites.
Qgeval: A benchmark for question generation evaluation, 2024
Weiping Fu, Bifan Wei, Jianxiang Hu, Zhongmin Cai, and Jun Liu · 2024
Later among the works it cites.
Exploring quality criteria and evaluation methods in automated question generation: A comprehensive survey
Guher Gorgun and Okan Bulut · 2024
Later among the works it cites.
A survey on neural question generation: Methods, applications, and prospects, 2024
Shasha Guo, Lizi Liao, Cuiping Li, and Tat-Seng Chua · 2024
Later among the works it cites.
Where is the answer? investigating positional bias in language model knowledge extraction, 2024
Kuniaki Saito, Kihyuk Sohn, Chen-Yu Lee, and Yoshitaka Ushiku · 2024
Later among the works it cites.
Evaluating open-qa evaluation
Cunxiang Wang, Sirui Cheng, Qipeng Guo, Yuanhao Yue, Bowen Ding, Zhikun Xu, Yidong Wang, Xiangkun Hu, Zheng Zhang, and Yue Zhang · 2024
Later among the works it cites.