Fetching the paper…
Reading the bibliography…
The evaluation of question answering models compares ground-truth annotations with model predictions.
Wordnet: a lexical database for english
George A Miller. 1995 · 1995
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Automatic evaluation of summaries using n-gram co-occurrence statistics
Chin-Yew Lin and Eduard Hovy. 2003 · 2003
Earlier work this paper cites.
Automatic evaluation of machine translation quality using longest common subsequence and skip-bigram statistics
Chin-Yew Lin and Franz Josef Och. 2004 · 2004
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie. 2005 · 2005
Earlier work this paper cites.
Squibs: What is a paraphrase?
Rahul Bhagat and Eduard Hovy. 2013 · 2013
Earlier work this paper cites.
Semeval-2017 task 1: Semantic textual similarity-multilingual and cross-lingual focused evaluation
Daniel Cer, Mona Diab, Eneko Agirre, Inigo Lopez-Gazpio, and Lucia Specia. 2017 · 2017
Earlier work this paper cites.
Learning to ask: Neural question generation for reading comprehension
Xinya Du, Junru Shao, and Claire Cardie. 2017 · 2017
Earlier work this paper cites.
Towards a better metric for evaluating question generation systems
Preksha Nema and Mitesh M. Khapra. 2018 · 2018
Earlier work this paper cites.
Know what you don’t know: Unanswerable questions for SQuAD
Pranav Rajpurkar, Robin Jia, and Percy Liang. 2018 · 2018
Earlier work this paper cites.
Systematic error analysis of the Stanford question answering dataset
Marc-Antoine Rondeau and T. J. Hazen. 2018 · 2018
Earlier work this paper cites.
Comparative analysis of neural QA models on SQuAD
Soumya Wadhwa, Khyathi Chandu, and Eric Nyberg. 2018 · 2018
Earlier work this paper cites.
Adaptations of ROUGE and BLEU to better evaluate machine reading comprehension task
An Yang, Kai Liu, Jing Liu, Yajuan Lyu, and Sujian Li. 2018 · 2018
Cited alongside, same era.
Reqa: An evaluation for end-to-end answer retrieval models
Amin Ahmad, Noah Constant, Yinfei Yang, and Daniel Cer. 2019 · 2019
Cited alongside, same era.
Evaluating question answering evaluation
Anthony Chen, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Are we modeling the task or the annotator? an investigation of annotator bias in natural language understanding datasets
Mor Geva, Yoav Goldberg, and Jonathan Berant. 2019 · 2019
Cited alongside, same era.
Natural questions: A benchmark for question answering research
More bang for your buck: Natural perturbation for robust question answering
Daniel Khashabi, Tushar Khot, and Ashish Sabharwal. 2020 · 2020
Later among the works it cites.
Look at the first sentence: Position bias in question answering
Miyoung Ko, Jinhyuk Lee, Hyunjae Kim, Gangwoo Kim, and Jaewoo Kang. 2020 · 2020
Later among the works it cites.
NeurIPS 2020 EfficientQA competition: Systems, analyses and lessons learned
Sewon Min, Jordan Boyd-Graber, Chris Alberti, Danqi Chen, Eunsol Choi, Michael Collins, Kelvin Guu, Hannaneh Hajishirzi, Kenton Lee, Jennimaria Palomaki, Colin Raffel, Adam Roberts, Tom Kwiatkowski, Patrick Lewis, Yuxiang Wu, Heinrich Küttler, Linqing Liu, Pasquale Minervini, Pontus Stenetorp, Sebastian Riedel, Sohee Yang, Minjoon Seo, Gautier Izacard, Fabio Petroni, Lucas Hosseini, Nicola De Cao, Edouard Grave, Ikuya Yamada, Sonse Shimaoka, Masatoshi Suzuki, Shumpei Miyawaki, Shun Sato, Ryo Takahashi, Jun Suzuki, Martin Fajcik, Martin Docekal, Karel Ondrej, Pavel Smrz, Hao Cheng, Yelong Shen, Xiaodong Liu, Pengcheng He, Weizhu Chen, Jianfeng Gao, Barlas Oguz, Xilun Chen, Vladimir Karpukhin, Stan Peshterliev, Dmytro Okhonko, Michael Schlichtkrull, Sonal Gupta, Yashar Mehdad, and Wen tau Yih. 2021 · 2020
Later among the works it cites.
BLEURT: Learning robust metrics for text generation
Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. 2019 · 2019
Cited alongside, same era.
Latent retrieval for weakly supervised open domain question answering
Kenton Lee, Ming-Wei Chang, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Linguistic knowledge and transferability of contextual representations
Nelson F. Liu, Matt Gardner, Yonatan Belinkov, Matthew E. Peters, and Noah A. Smith. 2019 · 2019
Cited alongside, same era.
Bend but don’t break? multi-challenge stress test for qa models
Hemant Pugaliya, James Route, Kaixin Ma, Yixuan Geng, and Eric Nyberg. 2019 · 2019
Cited alongside, same era.
Sentence-BERT: Sentence embeddings using Siamese BERT-networks
Nils Reimers and Iryna Gurevych. 2019 · 2019
Cited alongside, same era.
Are red roses red? evaluating consistency of question-answering models
Marco Tulio Ribeiro, Carlos Guestrin, and Sameer Singh. 2019 · 2019
Cited alongside, same era.
Errudite: Scalable, reproducible, and testable error analysis
Tongshuang Wu, Marco Tulio Ribeiro, Jeffrey Heer, and Daniel S Weld. 2019 · 2019
Cited alongside, same era.
What do we expect from multiple-choice QA systems?
Krunal Shah, Nitish Gupta, and Dan Roth. 2020 · 2020
Later among the works it cites.
BERTScore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020 · 2020
Later among the works it cites.
Can NLI models verify QA systems’ predictions?
Jifan Chen, Eunsol Choi, and Greg Durrett. 2021 · 2021
Closest in time.
The GEM benchmark: Natural language generation, its evaluation and metrics
Sebastian Gehrmann, Tosin Adewumi, Karmanya Aggarwal, Pawan Sasanka Ammanamanchi, Anuoluwapo Aremu, Antoine Bosselut, Khyathi Raghavi Chandu, Miruna-Adriana Clinciu, Dipanjan Das, Kaustubh Dhole, Wanyu Du, Esin Durmus, Ondřej Dušek, Chris Chinenye Emezue, Varun Gangal, Cristina Garbacea, Tatsunori Hashimoto, Yufang Hou, Yacine Jernite, Harsh Jhamtani, Yangfeng Ji, Shailza Jolly, Mihir Kale, Dhruv Kumar, Faisal Ladhak, Aman Madaan, Mounica Maddela, Khyati Mahajan, Saad Mahamood, Bodhisattwa Prasad Majumder, Pedro Henrique Martins, Angelina McMillan-Major, Simon Mille, Emiel van Miltenburg, Moin Nadeem, Shashi Narayan, Vitaly Nikolaev, Andre Niyongabo Rubungo, Salomey Osei, Ankur Parikh, Laura Perez-Beltrachini, Niranjan Ramesh Rao, Vikas Raunak, Juan Diego Rodriguez, Sashank Santhanam, João Sedoc, Thibault Sellam, Samira Shaikh, Anastasia Shimorina, Marco Antonio Sobrevilla Cabezudo, Hendrik Strobelt, Nishant Subramani, Wei Xu, Diyi Yang, Akhila Yerukola, and Jiawei Zhou. 2021 · 2021
Closest in time.
A tutorial on evaluation metrics used in natural language generation
Mitesh M. Khapra and Ananya B. Sai. 2021 · 2021
Closest in time.
GermanQuAD and GermanDPR: Improving non-english question answering and passage retrieval
Timo Möller, Julian Risch, and Malte Pietsch. 2021 · 2021
Closest in time.
Towards a more robust evaluation for conversational question answering
Wissam Siblini, Baris Sayil, and Yacine Kessaci. 2021 · 2021
Closest in time.