Fetching the paper…
Reading the bibliography…
To avoid giving wrong answers, question answering (QA) models need to know when to abstain from answering.
Learning and evaluating general linguistic intelligence
D. Yogatama, C. de M. d’Autume, J. Connor, T. Kocisky, M. Chrzanowski, L. Kong, A. Lazaridou, W. Ling, L. Yu, C. Dyer, et al. 2019 · 1901
Earlier work this paper cites.
Quizbowl: The case for incremental question answering
P. Rodriguez, S. Feng, M. Iyyer, H. He, and J. Boyd-Graber. 2019 · 1904
Earlier work this paper cites.
Selective prediction-set models with coverage guarantees
J. Feng, A. Sondhi, J. Perry, and N. Simon. 2019 · 1906
Earlier work this paper cites.
An optimum character recognition system using decision functions
C. K. Chow. 1957 · 1957
Earlier work this paper cites.
Understanding Natural Language
T. Winograd. 1972 · 1972
Earlier work this paper cites.
The Process of Question Answering
W. Lehnert. 1977 · 1977
Earlier work this paper cites.
Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods
J. Platt. 1999 · 1999
Earlier work this paper cites.
Support vector method for novelty detection
B. Schölkopf, R. Williamson, A. Smola, J. Shawe-Taylor, and J. Platt. 1999 · 1999
Earlier work this paper cites.
Is it the right answer? exploiting web redundancy for answer validation
B. Magnini, M. Negri, R. Prevete, and H. Tanev. 2002 · 2002
Earlier work this paper cites.
Domain adaptation with structural correspondence learning
J. Blitzer, R. McDonald, and F. Pereira. 2006 · 2006
Earlier work this paper cites.
Frustratingly easy domain adaptation
H. Daume III. 2007 · 2007
Earlier work this paper cites.
A probabilistic framework for answer selection in question answering
J. Ko, L. Si, and E. Nyberg. 2007 · 2007
Earlier work this paper cites.
What is the jeopardy model? a quasi-synchronous grammar for QA
M. Wang, N. A. Smith, and T. Mitamura. 2007 · 2007
Earlier work this paper cites.
Overview of ResPubliQA 2009: Question answering evaluation over european legislation
A. Peñas, P. Forner, R. Sutcliffe, Álvaro Rodrigo, C. Forăscu, I. Alegria, D. Giampiccolo, N. Moreau, and P. Osenova. 2009 · 2009
Earlier work this paper cites.
On the foundations of noise-free selective classification
R. El-Yaniv and Y. Wiener. 2010 · 2010
Earlier work this paper cites.
Scikit-learn: Machine learning in Python
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay. 2011 · 2011
Earlier work this paper cites.
A framework for merging and ranking of answers in DeepQA
D. C. Gondek, A. Lally, A. Kalyanpur, J. W. Murdock, P. A. Duboue, L. Zhang, Y. Pan, Z. M. Qiu, and C. Welty. 2012 · 2012
Earlier work this paper cites.
QA4MRE 2011-2013: Overview of question answering for machine reading evaluation
A. Peñas, E. Hovy, P. Forner, Álvaro Rodrigo, R. Sutcliffe, and R. Morante. 2013 · 2013
Cited alongside, same era.
Assessment of machine learning reliability methods for quantifying the applicability domain of QSAR regression models
M. Toplak, R. Močnik, M. Polajnar, Z. Bosnić, L. Carlsson, C. Hasselgren, J. Demšar, S. Boyer, B. Zupan, and J. Stålring. 2014 · 2014
Cited alongside, same era.
A large annotated corpus for learning natural language inference
S. Bowman, G. Angeli, C. Potts, and C. D. Manning. 2015 · 2015
Cited alongside, same era.
Dropout as a Bayesian approximation: Representing model uncertainty in deep learning
Y. Gal and Z. Ghahramani. 2016 · 2016
Cited alongside, same era.
SQuAD: 100,000+ questions for machine comprehension of text
P. Rajpurkar, J. Zhang, K. Lopyrev, and P. Liang. 2016 · 2016
Cited alongside, same era.
Know what you don’t know: Unanswerable questions for SQuAD
P. Rajpurkar, R. Jia, and P. Liang. 2018 · 2018
Later among the works it cites.
Understanding measures of uncertainty for adversarial example detection
L. Smith and Y. Gal. 2018 · 2018
Later among the works it cites.
Fever: a large-scale dataset for fact extraction and verification
J. Thorne, A. Vlachos, C. Christodoulopoulos, and A. Mittal. 2018 · 2018
Later among the works it cites.
WikiQA: A challenge dataset for open-domain question answering
Y. Yang, W. Yih, and C. Meek. 2015 · 2018
Later among the works it cites.
HotpotQA: A dataset for diverse, explainable multi-hop question answering
Z. Yang, P. Qi, S. Zhang, Y. Bengio, W. W. Cohen, R. Salakhutdinov, and C. D. Manning. 2018 · 2018
Later among the works it cites.
Evaluating question answering evaluation
A. Chen, G. Stanovsky, S. Singh, and M. Gardner. 2019 · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Chen, A. Fisch, J. Weston, and A. Bordes. 2017 · 2017
Cited alongside, same era.
SearchQA: A new Q&A dataset augmented with context from a search engine
M. Dunn, , L. Sagun, M. Higgins, U. Guney, V. Cirik, and K. Cho. 2017 · 2017
Cited alongside, same era.
Detecting adversarial samples from artifacts
R. Feinman, R. R. Curtin, S. Shintre, and A. B. Gardner. 2017 · 2017
Cited alongside, same era.
Selective classification for deep neural networks
Y. Geifman and R. El-Yaniv. 2017 · 2017
Cited alongside, same era.
A baseline for detecting misclassified and out-of-distribution examples in neural networks
D. Hendrycks and K. Gimpel. 2017 · 2017
Cited alongside, same era.
Adversarial examples for evaluating reading comprehension systems
R. Jia and P. Liang. 2017 · 2017
Cited alongside, same era.
TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension
M. Joshi, E. Choi, D. Weld, and L. Zettlemoyer. 2017 · 2017
Cited alongside, same era.
Later among the works it cites.
Understanding dataset design choices for multi-hop reasoning
J. Chen and G. Durrett. 2019 · 2019
Later among the works it cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M. Chang, K. Lee, and K. Toutanova. 2019 · 2019
Later among the works it cites.
MRQA 2019 shared task: Evaluating generalization in reading comprehension
A. Fisch, A. Talmor, R. Jia, M. Seo, E. Choi, and D. Chen. 2019 · 2019
Later among the works it cites.
Posing fair generalization tasks for natural language inference
A. Geiger, I. Cases, L. Karttunen, and C. Potts. 2019 · 2019
Later among the works it cites.
Natural questions: a benchmark for question answering research
T. Kwiatkowski, J. Palomaki, O. Redfield, M. Collins, A. Parikh, C. Alberti, D. Epstein, I. Polosukhin, M. Kelcey, J. Devlin, K. Lee, K. N. Toutanova, L. Jones, M. Chang, A. Dai, J. Uszkoreit, Q. Le, and S. Petrov. 2019 · 2019
Later among the works it cites.
Compositional questions do not necessitate multi-hop reasoning
S. Min, E. Wallace, S. Singh, M. Gardner, H. Hajishirzi, and L. Zettlemoyer. 2019 · 2019
Later among the works it cites.
Can you trust your model’s uncertainty? evaluating predictive uncertainty under dataset shift
Y. Ovadia, E. Fertig, J. Ren, Z. Nado, D. Sculley, S. Nowozin, J. V. Dillon, B. Lakshminarayanan, and J. Snoek. 2019 · 2019
Later among the works it cites.
MultiQA: An empirical investigation of generalization and transfer in reading comprehension
A. Talmor and J. Berant. 2019 · 2019
Later among the works it cites.
Universal adversarial triggers for attacking and analyzing NLP
E. Wallace, S. Feng, N. Kandpal, M. Gardner, and S. Singh. 2019 · 2019
Later among the works it cites.
Transforming classifier scores into accurate multiclass probability estimates
B. Zadrozny and C. Elkan. 2002 · 2019
Later among the works it cites.
Seq2Sick: Evaluating the robustness of sequence-to-sequence models with adversarial examples
M. Cheng, J. Yi, H. Zhang, P. Chen, and C. Hsieh. 2020 · 2020
Closest in time.