Fetching the paper…
Reading the bibliography…
Question generation (QGen) models are often evaluated with standardized NLG metrics that are based on n-gram overlap.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 1910
Earlier work this paper cites.
Response time in man-computer conversational transactions
Robert B Miller. 1968 · 1968
Earlier work this paper cites.
The instruction of reading comprehension
P David Pearson and Margaret C Gallagher. 1983 · 1983
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie. 2005 · 2005
Earlier work this paper cites.
Pearson correlation coefficient
Jacob Benesty, Jingdong Chen, Yiteng Huang, and Israel Cohen. 2009 · 2009
Earlier work this paper cites.
Multidimensional quality metrics (mqm): A framework for declaring and describing translation quality metrics
Arle Lommel, Hans Uszkoreit, and Aljoscha Burchardt. 2014 · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
Generating questions and multiple-choice answers using semantic analysis of texts
Jun Araki, Dheeraj Rajagopal, Sreecharan Sankaranarayanan, Susan Holm, Yukari Yamakawa, and Teruko Mitamura. 2016 · 2016
Earlier work this paper cites.
Reading comprehension: Core components and processes
Panayiota Kendeou, Kristen L McMaster, and Theodore J Christ. 2016 · 2016
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019 · 2016
Cited alongside, same era.
What do you mean exactly? analyzing clarification questions in cqa
Pavel Braslavski, Denis Savenkov, Eugene Agichtein, and Alina Dubatovka. 2017 · 2017
Cited alongside, same era.
Learning to ask: Neural question generation for reading comprehension
Xinya Du, Junru Shao, and Claire Cardie. 2017 · 2017
Cited alongside, same era.
Generating natural language question-answer pairs from a knowledge graph using a rnn based question generation model
Sathish Reddy Indurthi, Dinesh Raghu, Mitesh M Khapra, and Sachindra Joshi. 2017 · 2017
Cited alongside, same era.
Why we need new evaluation metrics for nlg
Jekaterina Novikova, Ondřej Dušek, Amanda Cercas Curry, and Verena Rieser. 2017 · 2017
Cited alongside, same era.
Newsqa: A machine comprehension dataset
Correlation coefficients: appropriate use and interpretation
Patrick Schober, Christa Boer, and Lothar A Schwarte. 2018 · 2018
Later among the works it cites.
Answer-focused and position-aware neural question generation
Xingwu Sun, Jing Liu, Yajuan Lyu, Wei He, Yanjun Ma, and Shi Wang. 2018 · 2018
Later among the works it cites.
Evaluating rewards for question generation models
Tom Hosking and Sebastian Riedel. 2019 · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Later among the works it cites.
A systematic review of automatic question generation for educational purposes
Ghader Kurdi, Jared Leo, Bijan Parsia, Uli Sattler, and Salam Al-Emari. 2020 · 2020
Later among the works it cites.
What’s the latest? a question-driven news chatbot
Philippe Laban, John Canny, and Marti A Hearst. 2020 · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Adam Trischler, Tong Wang, Xingdi Yuan, Justin Harris, Alessandro Sordoni, Philip Bachman, and Kaheer Suleman. 2017 · 2017
Cited alongside, same era.
Evaluation methodologies in automatic question generation 2013-2018
Jacopo Amidei, Paul Piwek, and Alistair Willis. 2018 · 2018
Cited alongside, same era.
The price of debiasing automatic metrics in natural language evalaution
Arun Chaganty, Stephen Mussmann, and Percy Liang. 2018 · 2018
Cited alongside, same era.
Survey of the state of the art in natural language generation: Core tasks, applications and evaluation
Albert Gatt and Emiel Krahmer. 2018 · 2018
Cited alongside, same era.
Towards a better metric for evaluating question generation systems
Preksha Nema and Mitesh M Khapra. 2018 · 2018
Cited alongside, same era.
Learning to ask good questions: Ranking clarification questions using neural expected value of perfect information
Sudha Rao and Hal Daumé III. 2018 · 2018
Cited alongside, same era.
Later among the works it cites.
Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020 · 2020
Later among the works it cites.
Prophetnet: Predicting future n-gram for sequence-to-sequence pre-training
Weizhen Qi, Yu Yan, Yeyun Gong, Dayiheng Liu, Nan Duan, Jiusheng Chen, Ruofei Zhang, and Ming Zhou. 2020 · 2020
Later among the works it cites.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Julien Chaumond, Lysandre Debut, Victor Sanh, Clement Delangue, Anthony Moi, Pierric Cistac, Morgan Funtowicz, Joe Davison, Sam Shleifer, et al. 2020 · 2020
Later among the works it cites.
Tim Steuer, Anna Filighera, Tobias Meuser, and Christoph Rensing. 2021 · 2021
Later among the works it cites.
Mixqg: Neural question generation with mixed answer types
Lidiya Murakhovs’ka, Chien-Sheng Wu, Philippe Laban, Tong Niu, Wenhao Liu, and Caiming Xiong. 2022 · 2022
Closest in time.