Fetching the paper…
Reading the bibliography…
The predictions of question answering (QA)systems are typically evaluated against manually annotated finite sets of one or more answers.
ALBERT: A lite BERT for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2019 · 1909
Earlier work this paper cites.
A sys called qanda
Eric Breck, John Burger, Lisa Ferro, David House, Marc Light, and Inderjeet Mani. 1999 · 1999
Earlier work this paper cites.
Pay enough or don’t pay at all
Uri Gneezy and Aldo Rustichini. 2000 · 2000
Earlier work this paper cites.
Building a question answering test collection
Ellen M. Voorhees and Dawn M. Tice. 2000 · 2000
Earlier work this paper cites.
Bleu: A method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
A comparison of string distance metrics for name-matching tasks
William W. Cohen, Pradeep Ravikumar, and Stephen E. Fienberg. 2003 · 2003
Earlier work this paper cites.
Content Analysis: An Introduction to Its Methodology (second edition)
Klaus Krippendorff. 2004 · 2004
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Ambigqa: Answering ambiguous open-domain questions
Sewon Min, Julian Michael, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2020b · 2004
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie. 2005 · 2005
Earlier work this paper cites.
Algorithmic Learning in a Random World
Vladimir Vovk, Alex Gammerman, and Glenn Shafer. 2005 · 2005
Earlier work this paper cites.
A tutorial on conformal prediction
Glenn Shafer and Vladimir Vovk. 2007 · 2007
Earlier work this paper cites.
Freebase: A collaboratively created graph database for structuring human knowledge
Kurt Bollacker, Colin Evans, Praveen Paritosh, Tim Sturge, and Jamie Taylor. 2008 · 2008
Earlier work this paper cites.
From word embeddings to document distances
Matt J Kusner, Yu Sun, Nicholas I Kolkin, and Kilian Q Weinberger. 2015 · 2015
Earlier work this paper cites.
Towards ai-complete question answering: A set of prerequisite toy tasks
Jason Weston, Antoine Bordes, Sumit Chopra, Alexander M. Rush, Bart van Merriënboer, Armand Joulin, and Tomas Mikolov. 2015 · 2015
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
Bidirectional attention flow for machine comprehension
Min Joon Seo, Aniruddha Kembhavi, Ali Farhadi, and Hannaneh Hajishirzi. 2016 · 2016
Cited alongside, same era.
Adaptations of ROUGE and BLEU to better evaluate machine reading comprehension task
An Yang, Kai Liu, Jing Liu, Yajuan Lyu, and Sujian Li. 2018 · 2018
Cited alongside, same era.
Evaluating question answering evaluation
Anthony Chen, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. 2019 · 2019
Cited alongside, same era.
Sentence mover’s similarity: Automatic evaluation for multi-sentence texts
Elizabeth Clark, Asli Celikyilmaz, and Noah A. Smith. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
A gentle introduction to conformal prediction and distribution-free uncertainty quantification
Anastasios N. Angelopoulos and Stephen Bates. 2021 · 2021
Later among the works it cites.
Distribution-free, risk-controlling prediction sets
Stephen Bates, Anastasios Angelopoulos, Lihua Lei, Jitendra Malik, and Michael Jordan. 2021 · 2021
Later among the works it cites.
Can NLI models verify QA systems’ predictions?
Jifan Chen, Eunsol Choi, and Greg Durrett. 2021 · 2021
Later among the works it cites.
Qafacteval: Improved qa-based factual consistency evaluation for summarization
Alexander R. Fabbri, Chien-Sheng Wu, Wenhao Liu, and Caiming Xiong. 2021 · 2021
Later among the works it cites.
Efficient conformal prediction via cascaded inference with expanded admission
Adam Fisch, Tal Schuster, Tommi S. Jaakkola, and Regina Barzilay. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Where’s My Head? Definition, Data Set, and Models for Numeric Fused-Head Identification and Resolution
Yanai Elazar and Yoav Goldberg. 2019 · 2019
Cited alongside, same era.
Question answering as an automatic evaluation metric for news article summarization
Matan Eyal, Tal Baumel, and Michael Elhadad. 2019 · 2019
Cited alongside, same era.
Latent retrieval for weakly supervised open domain question answering
Kenton Lee, Ming-Wei Chang, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le. 2019 · 2019
Cited alongside, same era.
MOCHA: A dataset for training and evaluating generative reading comprehension metrics
Anthony Chen, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. 2020 · 2020
Cited alongside, same era.
Neurips 2020 efficientqa competition: Systems, analyses and lessons learned
Sewon Min, Jordan L. Boyd-Graber, Chris Alberti, Danqi Chen, Eunsol Choi, Michael Collins, Kelvin Guu, Hannaneh Hajishirzi, Kenton Lee, Jennimaria Palomaki, Colin Raffel, Adam Roberts, Tom Kwiatkowski, Patrick S. H. Lewis, Yuxiang Wu, Heinrich Küttler, Linqing Liu, Pasquale Minervini, Pontus Stenetorp, Sebastian Riedel, Sohee Yang, Minjoon Seo, Gautier Izacard, Fabio Petroni, Lucas Hosseini, Nicola De Cao, Edouard Grave, Ikuya Yamada, Sonse Shimaoka, Masatoshi Suzuki, Shumpei Miyawaki, Shun Sato, Ryo Takahashi, Jun Suzuki, Martin Fajcik, Martin Docekal, Karel Ondrej, Pavel Smrz, Hao Cheng, Yelong Shen, Xiaodong Liu, Pengcheng He, Weizhu Chen, Jianfeng Gao, Barlas Oguz, Xilun Chen, Vladimir Karpukhin, Stan Peshterliev, Dmytro Okhonko, Michael Sejr Schlichtkrull, Sonal Gupta, Yashar Mehdad, and Wen-tau Yih. 2020a · 2020
Cited alongside, same era.
COMET: A neural framework for MT evaluation
Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. 2020 · 2020
Cited alongside, same era.
The gem benchmark: Natural language generation, its evaluation and metrics
Sebastian Gehrmann, Tosin Adewumi, Karmanya Aggarwal, Pawan Sasanka Ammanamanchi, Aremu Anuoluwapo, Antoine Bosselut, Khyathi Raghavi Chandu, Miruna Clinciu, Dipanjan Das, Kaustubh D. Dhole, Wanyu Du, Esin Durmus, Ondřej Dušek, Chris Emezue, Varun Gangal, Cristina Garbacea, Tatsunori Hashimoto, Yufang Hou, Yacine Jernite, Harsh Jhamtani, Yangfeng Ji, Shailza Jolly, Mihir Kale, Dhruv Kumar, Faisal Ladhak, Aman Madaan, Mounica Maddela, Khyati Mahajan, Saad Mahamood, Bodhisattwa Prasad Majumder, Pedro Henrique Martins, Angelina McMillan-Major, Simon Mille, Emiel van Miltenburg, Moin Nadeem, Shashi Narayan, Vitaly Nikolaev, Rubungo Andre Niyongabo, Salomey Osei, Ankur Parikh, Laura Perez-Beltrachini, Niranjan Ramesh Rao, Vikas Raunak, Juan Diego Rodriguez, Sashank Santhanam, João Sedoc, Thibault Sellam, Samira Shaikh, Anastasia Shimorina, Marco Antonio Sobrevilla Cabezudo, Hendrik Strobelt, Nishant Subramani, Wei Xu, Diyi Yang, Akhila Yerukola, and Jiawei Zhou. 2021 · 2021
Later among the works it cites.
q 2 q^{2} : Evaluating factual consistency in knowledge-grounded dialogues via question generation and question answering
Or Honovich, Leshem Choshen, Roee Aharoni, Ella Neeman, Idan Szpektor, and Omri Abend. 2021 · 2021
Later among the works it cites.
To ship or not to ship: An extensive evaluation of automatic metrics for machine translation
Tom Kocmi, Christian Federmann, Roman Grundkiewicz, Marcin Junczys-Dowmunt, Hitokazu Matsushita, and Arul Menezes. 2021 · 2021
Later among the works it cites.
Semantic answer similarity for evaluating question answering models
Julian Risch, Timo Möller, Julian Gutsch, and Malte Pietsch. 2021 · 2021
Later among the works it cites.
Get your vitamin C! robust fact verification with contrastive evidence
Tal Schuster, Adam Fisch, and Regina Barzilay. 2021a · 2021
Later among the works it cites.
Consistent accelerated inference via confident adaptive transformers
Tal Schuster, Adam Fisch, Tommi Jaakkola, and Regina Barzilay. 2021b · 2021
Later among the works it cites.
What’s in a name? answer equivalence for open-domain question answering
Chenglei Si, Chen Zhao, and Jordan Boyd-Graber. 2021 · 2021
Later among the works it cites.
Uncertainty estimation for natural language processing
Adam Fisch, Robin Jia, and Tal Schuster. 2022 · 2022
Closest in time.
TRUE: Re-evaluating factual consistency evaluation
Or Honovich, Roee Aharoni, Jonathan Herzig, Hagai Taitelbaum, Doron Kukliansy, Vered Cohen, Thomas Scialom, Idan Szpektor, Avinatan Hassidim, and Yossi Matias. 2022 · 2022
Closest in time.
Confident adaptive language modeling
Tal Schuster, Adam Fisch, Jai Gupta, Mostafa Dehghani, Dara Bahri, Vinh Quang Tran, Yi Tay, and Donald Metzler. 2022 · 2022
Closest in time.