Fetching the paper…
Reading the bibliography…
Existing metrics for evaluating the quality of automatically generated questions such as BLEU, ROUGE, BERTScore, and BLEURT compare the reference and predicted questions, providing a high score when there is a considerable lexical overlap or semantic similarity between the candidate and the reference questions.
Reasoning over paragraph effects in situations
Kevin Lin, Oyvind Tafjord, Peter Clark, and Matt Gardner. 2019 · 1908
Earlier work this paper cites.
Answers unite! unsupervised metrics for reinforced summarization models
Thomas Scialom, Sylvain Lamprier, Benjamin Piwowarski, and Jacopo Staiano. 2019a · 1909
Earlier work this paper cites.
Qasc: A dataset for question answering via sentence composition
Tushar Khot, Peter Clark, Michal Guerquin, Peter Alexander Jansen, and Ashish Sabharwal. 2020 · 1910
Earlier work this paper cites.
Harvesting paragraph-level question-answer pairs from Wikipedia
Xinya Du and Claire Cardie. 2018 · 1917
Earlier work this paper cites.
Wordnet: A lexical database for english
George A. Miller. 1995 · 1995
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Supervised automatic evaluation for summarization with voted regression model
Tsutomu Hirao, Manabu Okumura, Norihito Yasuda, and Hideki Isozaki. 2007 · 2007
Earlier work this paper cites.
Natural language processing with Python: analyzing text with the natural language toolkit
Steven Bird, Ewan Klein, and Edward Loper. 2009 · 2009
Earlier work this paper cites.
Extending the METEOR machine translation evaluation metric to the phrase level
Michael Denkowski and Alon Lavie. 2010 · 2010
Earlier work this paper cites.
The first question generation shared task evaluation challenge
Vasile Rus, Brendan Wyse, Paul Piwek, Mihai Lintean, Svetlana Stoyanchev, and Christian Moldovan. 2010 · 2010
Earlier work this paper cites.
MCTest: A challenge dataset for the open-domain machine comprehension of text
Matthew Richardson, Christopher J.C. Burges, and Erin Renshaw. 2013 · 2013
Earlier work this paper cites.
Modeling biological processes for reading comprehension
Jonathan Berant, Vivek Srikumar, Pei-Chun Chen, Abby Vander Linden, Brittany Harding, Brad Huang, Peter Clark, and Christopher D. Manning. 2014 · 2014
Earlier work this paper cites.
BEER: BEtter evaluation as ranking
Miloš Stanojević and Khalil Sima’an. 2014 · 2014
Earlier work this paper cites.
chrF: character n-gram F-score for automatic MT evaluation
Maja Popović. 2015 · 2015
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
Newsqa: A machine comprehension dataset
Adam Trischler, Tong Wang, Xingdi Yuan, Justin Harris, Alessandro Sordoni, Philip Bachman, and Kaheer Suleman. 2016 · 2016
Earlier work this paper cites.
Learning to ask: Neural question generation for reading comprehension
Xinya Du, Junru Shao, and Claire Cardie. 2017 · 2017
Earlier work this paper cites.
Question generation for question answering
Nan Duan, Duyu Tang, Peng Chen, and Ming Zhou. 2017 · 2017
Earlier work this paper cites.
RACE: Large-scale ReAding comprehension dataset from examinations
Guokun Lai, Qizhe Xie, Hanxiao Liu, Yiming Yang, and Eduard Hovy. 2017 · 2017
Earlier work this paper cites.
Learning to score system summaries for better content selection evaluation
Maxime Peyrard, Teresa Botschen, and Iryna Gurevych. 2017 · 2017
Earlier work this paper cites.
Learningq: A large-scale dataset for educational question generation
Guanliang Chen, Jie Yang, Claudia Hauff, and Geert-Jan Houben. 2018 · 2018
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018 · 2018
Earlier work this paper cites.
Marian: Fast neural machine translation in C++
Marcin Junczys-Dowmunt, Roman Grundkiewicz, Tomasz Dwojak, Hieu Hoang, Kenneth Heafield, Tom Neckermann, Frank Seide, Ulrich Germann, Alham Fikri Aji, Nikolay Bogoychev, André F. T. Martins, and Alexandra Birch. 2018 · 2018
Earlier work this paper cites.
Looking beyond the surface: A challenge set for reading comprehension over multiple sentences
Daniel Khashabi, Snigdha Chaturvedi, Michael Roth, Shyam Upadhyay, and Dan Roth. 2018 · 2018
Earlier work this paper cites.
The NarrativeQA reading comprehension challenge
Tomáš Kočiský, Jonathan Schwarz, Phil Blunsom, Chris Dyer, Karl Moritz Hermann, Gábor Melis, and Edward Grefenstette. 2018 · 2018
Earlier work this paper cites.
Can a suit of armor conduct electricity? a new dataset for open book question answering
Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. 2018 · 2018
Earlier work this paper cites.
Towards a better metric for evaluating question generation systems
Preksha Nema and Mitesh M. Khapra. 2018 · 2018
Earlier work this paper cites.
MCScript: A Novel Dataset for Assessing Machine Comprehension Using Script Knowledge
Simon Ostermann, Ashutosh Modi, Michael Roth, Stefan Thater, and Manfred Pinkal. 2018 · 2018
Earlier work this paper cites.
Know what you don’t know: Unanswerable questions for SQuAD
Pranav Rajpurkar, Robin Jia, and Percy Liang. 2018 · 2018
Earlier work this paper cites.
Record: Bridging the gap between human and machine commonsense reading comprehension
Sheng Zhang, Xiaodong Liu, Jingjing Liu, Jianfeng Gao, Kevin Duh, and Benjamin Van Durme. 2018 · 2018
Earlier work this paper cites.
Piqa: Reasoning about physical commonsense in natural language
Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi. 2019 · 2019
Cited alongside, same era.
BoolQ: Exploring the surprising difficulty of natural yes/no questions
Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Quoref: A reading comprehension dataset with questions requiring coreferential reasoning
Pradeep Dasigi, Nelson F. Liu, Ana Marasović, Noah A. Smith, and Matt Gardner. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs
Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. 2019 · 2019
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020 · 2020
Later among the works it cites.
COMET: A neural framework for MT evaluation
Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. 2020 · 2020
Later among the works it cites.
Getting closer to ai complete question answering: A set of prerequisite real tasks
Anna Rogers, Olga Kovaleva, Matthew Downey, and Anna Rumshisky. 2020 · 2020
Later among the works it cites.
BLEURT: Learning robust metrics for text generation
Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020 · 2020
Later among the works it cites.
Automatic machine translation evaluation in many languages via zero-shot paraphrasing
Brian Thompson and Matt Post. 2020 · 2020
Later among the works it cites.
Reclor: A reading comprehension dataset requiring logical reasoning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Evaluating rewards for question generation models
Tom Hosking and Sebastian Riedel. 2019 · 2019
Cited alongside, same era.
Large-scale, diverse, paraphrastic bitexts via sampling and clustering
J. Edward Hu, Abhinav Singh, Nils Holzenberger, Matt Post, and Benjamin Van Durme. 2019 · 2019
Cited alongside, same era.
Cosmos QA: Machine reading comprehension with contextual commonsense reasoning
Lifu Huang, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2019 · 2019
Cited alongside, same era.
PubMedQA: A dataset for biomedical research question answering
Qiao Jin, Bhuwan Dhingra, Zhengping Liu, William Cohen, and Xinghua Lu. 2019 · 2019
Cited alongside, same era.
Natural questions: A benchmark for question answering research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. 2019 · 2019
Cited alongside, same era.
A new multi-choice reading comprehension dataset for curriculum learning
Yichan Liang, Jianheng Li, and Jian Yin. 2019 · 2019
Cited alongside, same era.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. 2019 · 2019
Cited alongside, same era.
Weihao Yu, Zihang Jiang, Yanfei Dong, and Jiashi Feng. 2020 · 2020
Later among the works it cites.
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020 · 2020
Later among the works it cites.
PROST: Physical reasoning about objects through space and time
Stéphane Aroca-Ouellette, Cory Paik, Alessandro Roncone, and Katharina Kann. 2021 · 2021
Later among the works it cites.
mmarco: A multilingual version of the ms marco passage ranking dataset
Luiz Bonifacio, Vitor Jeronymo, Hugo Queiroz Abonizio, Israel Campiotti, Marzieh Fadaee, Roberto Lotufo, and Rodrigo Nogueira. 2021 · 2021
Later among the works it cites.
Guiding the growth: Difficulty-controllable question generation through step-by-step rewriting
Yi Cheng, Siyao Li, Bang Liu, Ruihui Zhao, Sujian Li, Chenghua Lin, and Yefeng Zheng. 2021 · 2021
Later among the works it cites.
Onestop qamaker: Extract question-answer pairs from text in a one-stop approach
Shaobo Cui, Xintong Bao, Xinxing Zu, Yangyang Guo, Zhongzhou Zhao, Ji Zhang, and Haiqing Chen. 2021 · 2021
Later among the works it cites.
Compression, transduction, and creation: A unified framework for evaluating natural language generation
Mingkai Deng, Bowen Tan, Zhengzhong Liu, Eric Xing, and Zhiting Hu. 2021 · 2021
Later among the works it cites.
Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies
Mor Geva, Daniel Khashabi, Elad Segal, Tushar Khot, Dan Roth, and Jonathan Berant. 2021 · 2021
Later among the works it cites.
Deberta: Decoding-enhanced bert with disentangled attention
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Wei Chen. 2021 · 2021
Later among the works it cites.
Qace: Asking questions to evaluate an image caption
Hwanhee Lee, Thomas Scialom, Seunghyun Yoon, Franck Dernoncourt, and Kyomin Jung. 2021 · 2021
Later among the works it cites.
Data-QuestEval: A referenceless metric for data-to-text semantic evaluation
Clement Rebuffel, Thomas Scialom, Laure Soulier, Benjamin Piwowarski, Sylvain Lamprier, Jacopo Staiano, Geoffrey Scoutheeten, and Patrick Gallinari. 2021 · 2021
Later among the works it cites.
Winogrande: An adversarial winograd schema challenge at scale
Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2021 · 2021
Later among the works it cites.
Questeval: Summarization asks for fact-based evaluation
Thomas Scialom, Paul-Alexis Dray, Patrick Gallinari, Sylvain Lamprier, Benjamin Piwowarski, Jacopo Staiano, and Alex Wang. 2021 · 2021
Later among the works it cites.
Bartscore: Evaluating generated text as text generation
Weizhe Yuan, Graham Neubig, and Pengfei Liu. 2021 · 2021
Later among the works it cites.
QAFactEval: Improved QA-based factual consistency evaluation for summarization
Alexander Fabbri, Chien-Sheng Wu, Wenhao Liu, and Caiming Xiong. 2022 · 2022
Closest in time.
TRUE: Re-evaluating factual consistency evaluation
Or Honovich, Roee Aharoni, Jonathan Herzig, Hagai Taitelbaum, Doron Kukliansy, Vered Cohen, Thomas Scialom, Idan Szpektor, Avinatan Hassidim, and Yossi Matias. 2022 · 2022
Closest in time.
Unifiedqa-v2: Stronger generalization via broader cross-format training
Daniel Khashabi, Yeganeh Kordi, and Hannaneh Hajishirzi. 2022 · 2022
Closest in time.
Quiz design task: Helping teachers create quizzes with automated question generation
Philippe Laban, Chien-Sheng Wu, Lidiya Murakhovs’ka, Wenhao Liu, and Caiming Xiong. 2022 · 2022
Closest in time.
SMaLL-100: Introducing shallow multilingual machine translation model for low-resource languages
Alireza Mohammadshahi, Vassilina Nikoulina, Alexandre Berard, Caroline Brun, James Henderson, and Laurent Besacier. 2022a · 2022
Closest in time.
What do compressed multilingual machine translation models forget?
Alireza Mohammadshahi, Vassilina Nikoulina, Alexandre Berard, Caroline Brun, James Henderson, and Laurent Besacier. 2022b · 2022
Closest in time.
MixQG: Neural question generation with mixed answer types
Lidiya Murakhovs’ka, Chien-Sheng Wu, Philippe Laban, Tong Niu, Wenhao Liu, and Caiming Xiong. 2022 · 2022
Closest in time.
Generative language models for paragraph-level question generation
Asahi Ushio, Fernando Alva-Manchego, and Jose Camacho-Collados. 2022 · 2022
Closest in time.
A practical toolkit for multilingual question and answer generation, acl 2022, system demonstration
Asahi Ushio, Fernando Alva-Manchego, and Jose Camacho-Collados. 2023b · 2022
Closest in time.
Xiaoqiang Wang, Bang Liu, Siliang Tang, and Lingfei Wu. 2022 · 2022
Closest in time.
QAConv: Question answering on informative conversations
Chien-Sheng Wu, Andrea Madotto, Wenhao Liu, Pascale Fung, and Caiming Xiong. 2022 · 2022
Closest in time.
Are factuality checkers reliable? adversarial meta-evaluation of factuality in summarization
Yiran Chen, Pengfei Liu, and Xipeng Qiu. 2021 · 2095
Closest in time.