Fetching the paper…
Reading the bibliography…
We study calibration in question answering, estimating whether model correctly predicts answer for each question.
A BERT baseline for the natural questions
Chris Alberti, Kenton Lee, and Michael Collins. 2019 · 1901
Earlier work this paper cites.
Calibration of encoder decoder models for neural machine translation
A. Kumar and Sunita Sarawagi. 2019 · 1903
Earlier work this paper cites.
Xlda: Cross-lingual data augmentation for natural language inference and question answering
Jasdeep Singh, Bryan McCann, Nitish Shirish Keskar, Caiming Xiong, and Richard Socher. 2019 · 1905
Earlier work this paper cites.
Latent retrieval for weakly supervised open domain question answering
Kenton Lee, Ming-Wei Chang, and Kristina Toutanova. 2019 · 1906
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Mrqa 2019 shared task: Evaluating generalization in reading comprehension
Adam Fisch, Alon Talmor, Robin Jia, Minjoon Seo, Eunsol Choi, and Danqi Chen. 2019 · 1910
Earlier work this paper cites.
Unsupervised out-of-distribution detection with batch normalization
Jiaming Song, Yang Song, and Stefano Ermon. 2019 · 1910
Earlier work this paper cites.
Verification of forecasts expressed in terms of probability
G. W. Brier. 1950 · 1950
Earlier work this paper cites.
The meaning and use of the area under a receiver operating characteristic (roc) curve
James A Hanley and Barbara J McNeil. 1982 · 1982
Earlier work this paper cites.
Building a question answering test collection
Ellen M. Voorhees and Dawn M. Tice. 2000 · 2000
Earlier work this paper cites.
Ensembling neural networks: Many could be better than all
Z. Zhou, Jianxin Wu, and Wei Tang. 2002 · 2002
Earlier work this paper cites.
Calibration of pre-trained transformers
Shrey Desai and Greg Durrett. 2020 · 2003
Earlier work this paper cites.
Using confidence scores to improve hands-free speech based navigation in continuous dictation systems
Jinjuan Feng and Andrew Sears. 2004 · 2004
Earlier work this paper cites.
Discriminative reranking for machine translation
Libin Shen, Anoop Sarkar, and F. Och. 2004 · 2004
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Patrick Lewis, Ethan Perez, Aleksandara Piktus, F. Petroni, V. Karpukhin, Naman Goyal, Heinrich Kuttler, M. Lewis, Wen tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020 · 2005
Earlier work this paper cites.
Using bayesian model averaging to calibrate forecast ensembles
A. Raftery, T. Gneiting, F. Balabdaoui, and M. Polakowski. 2005 · 2005
Earlier work this paper cites.
Uci machine learning repository
Arthur Asuncion and David Newman. 2007 · 2007
Earlier work this paper cites.
Diverse ensembles improve calibration
Asa Cooper Stickland and Iain Murray. 2020 · 2007
Earlier work this paper cites.
It’s better to say "i can’t answer" than answering incorrectly: Towards safety critical nlp systems
N. Varshney, Swaroop Mishra, and Chitta Baral. 2020 · 2008
Earlier work this paper cites.
Natural language processing with Python: analyzing text with the natural language toolkit
Steven Bird, Ewan Klein, and Edward Loper. 2009 · 2009
Earlier work this paper cites.
Self-training improves pre-training for natural language understanding
Jingfei Du, Edouard Grave, Beliz Gunel, Vishrav Chaudhary, Onur Celebi, Michael Auli, Ves Stoyanov, and Alexis Conneau. 2020 · 2010
Cited alongside, same era.
Scikit-learn: Machine learning in python
Fabian Pedregosa, Gaël Varoquaux, Alexandre Gramfort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel, Peter Prettenhofer, Ron Weiss, Vincent Dubourg, et al. 2011 · 2011
Cited alongside, same era.
Sabrina J. Mielke, Arthur Szlam, Y.-Lan Boureau, and Emily Dinan. 2020 · 2012
Cited alongside, same era.
Posterior calibration and exploratory analysis for natural language processing models
Khanh Nguyen and Brendan T. O’Connor. 2015 · 2015
Cited alongside, same era.
Xgboost: A scalable tree boosting system
Tianqi Chen and Carlos Guestrin. 2016 · 2016
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Later among the works it cites.
Confidence modeling for neural semantic parsing
Li Dong, Chris Quirk, and Mirella Lapata. 2018 · 2018
Later among the works it cites.
On calibration of modern neural networks
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger. 2018 · 2018
Later among the works it cites.
Marian: Fast neural machine translation in C++
Marcin Junczys-Dowmunt, Roman Grundkiewicz, Tomasz Dwojak, Hieu Hoang, Kenneth Heafield, Tom Neckermann, Frank Seide, Ulrich Germann, Alham Fikri Aji, Nikolay Bogoychev, André F. T. Martins, and Alexandra Birch. 2018 · 2018
Later among the works it cites.
Know what you don’t know: Unanswerable questions for squad
Pranav Rajpurkar, Robin Jia, and Percy Liang. 2018 · 2018
Later among the works it cites.
Understanding measures of uncertainty for adversarial example detection
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Yarin Gal and Zoubin Ghahramani. 2016 · 2016
Cited alongside, same era.
Squad: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Cited alongside, same era.
Improving neural machine translation models with monolingual data
Rico Sennrich, B. Haddow, and Alexandra Birch. 2016 · 2016
Cited alongside, same era.
Reading Wikipedia to answer open-domain questions
Danqi Chen, Adam Fisch, Jason Weston, and Antoine Bordes. 2017 · 2017
Cited alongside, same era.
Supervised learning of universal sentence representations from natural language inference data
Alexis Conneau, Douwe Kiela, Holger Schwenk, Loïc Barrault, and Antoine Bordes. 2017 · 2017
Cited alongside, same era.
Learning to paraphrase for question answering
Li Dong, Jonathan Mallinson, Siva Reddy, and Mirella Lapata. 2017 · 2017
Cited alongside, same era.
Searchqa: A new q&a dataset augmented with context from a search engine
Matthew Dunn, Levent Sagun, Mike Higgins, V Ugur Guney, Volkan Cirik, and Kyunghyun Cho. 2017 · 2017
Cited alongside, same era.
Lewis Smith and Yarin Gal. 2018 · 2018
Later among the works it cites.
I know there is no answer: modeling answer validation for machine reading comprehension
Chuanqi Tan, Furu Wei, Qingyu Zhou, Nan Yang, Weifeng Lv, and Ming Zhou. 2018 · 2018
Later among the works it cites.
Hotpotqa: A dataset for diverse, explainable multi-hop question answering
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D Manning. 2018 · 2018
Later among the works it cites.
Qanet: Combining local convolution with global self-attention for reading comprehension
Adams Wei Yu, David Dohan, Minh-Thang Luong, R. Zhao, Kai Chen, Mohammad Norouzi, and Quoc V. Le. 2018 · 2018
Later among the works it cites.
Diverse ensemble evolution: Curriculum data-model marriage
Tianyi Zhou, S. Wang, and J. Bilmes. 2018 · 2018
Later among the works it cites.
Using pre-training can improve model robustness and uncertainty
Dan Hendrycks, Kimin Lee, and Mantas Mazeika. 2019 · 2019
Later among the works it cites.
Read+ verify: Machine reading comprehension with unanswerable questions
Minghao Hu, Furu Wei, Yuxing Peng, Zhen Huang, Nan Yang, and Dongsheng Li. 2019 · 2019
Later among the works it cites.
Natural questions: a benchmark for question answering research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al. 2019 · 2019
Later among the works it cites.
Do not have enough data? deep learning to the rescue!
Ateret Anaby-Tavor, Boaz Carmeli, Esther Goldbraich, Amir Kantor, George Kour, Segev Shlomov, Naama Tepper, and Naama Zwerdling. 2020 · 2020
Later among the works it cites.
Calibrating structured output predictors for natural language processing
Abhyuday N. Jagannatha and Hong Yu. 2020 · 2020
Later among the works it cites.
Selective question answering under domain shift
Amita Kamath, Robin Jia, and Percy Liang. 2020 · 2020
Later among the works it cites.
Dense passage retrieval for open-domain question answering
Vladimir Karpukhin, Barlas Oğuz, Sewon Min, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020 · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, M. Matena, Yanqi Zhou, W. Li, and Peter J. Liu. 2020 · 2020
Later among the works it cites.
No answer is better than wrong answer: A reflection model for document level machine reading comprehension
Xuguang Wang, Linjun Shou, Ming Gong, Nan Duan, and Daxin Jiang. 2020 · 2020
Later among the works it cites.
Contextual dropout: An efficient sample-dependent dropout module
XINJIE FAN, Shujian Zhang, Korawat Tanwisuth, Xiaoning Qian, and Mingyuan Zhou. 2021 · 2021
Closest in time.